Storage redundancy could not rebuild itself after the outage
Once the stuck machines were restarted, the platform had to rebuild the redundant copies of a great deal of data at once. It never finished - the number of under-protected volumes stalled around sixty instead of falling to zero.
Cause
The platform was allowed to run up to six rebuilds per machine at the same time, which across nine machines meant more than fifty at once. Under that much simultaneous work, the storage components missed their own health checks. Each missed check aborted a rebuild, marked that copy as failed, and scheduled a fresh one - so the work undid itself as fast as it was done.
Fix
The concurrency limit was lowered so rebuilds finish instead of colliding. Failures immediately fell from roughly six a minute to one, and the number of fully-protected volumes began climbing again.
This is a durable setting rather than a one-off intervention, and it is now recorded with the reasoning behind it.
0 Comments
Sign in to comment
No comments yet. Be the first to share your thoughts!
