Feedback Requests
Complete

Storage half of the connection-hang incident documented

The incident write-up for the connection-hang outage covered the web-entry-point side thoroughly and stopped at the kernel fix. The same limit had also starved the distributed storage platform, and that damage outlived the fix by two days.

Added to the write-up: the storage failure chain; the repairs the kernel fix does not undo (two filesystem checks, a database write-ahead-log rebuild, a database cluster stuck in an unrecoverable state, and the cache log above); why a warm restart preserves the faulty hardware table while a cold power-cycle rebuilds it; the two measurements that misled the diagnosis; and two destructive actions to avoid while repairing it.

Also records that the fix survived a warm fleet restart — which is what proves it is applied persistently at boot rather than only to the running system.

0 Comments

Sign in to comment

No comments yet. Be the first to share your thoughts!