Bug Reports
Complete

Storage redundancy could not rebuild itself after the outage

Once the stuck machines were restarted, the platform had to rebuild the redundant copies of a great deal of data at once. It never finished - the number of under-protected volumes stalled around sixty instead of falling to zero.

Cause

The platform was allowed to run up to six rebuilds per machine at the same time, which across nine machines meant more than fifty at once. Under that much simultaneous work, the storage components missed their own health checks. Each missed check aborted a rebuild, marked that copy as failed, and scheduled a fresh one - so the work undid itself as fast as it was done.

Fix

The concurrency limit was lowered so rebuilds finish instead of colliding. Failures immediately fell from roughly six a minute to one, and the number of fully-protected volumes began climbing again.

This is a durable setting rather than a one-off intervention, and it is now recorded with the reasoning behind it.

0 Comments

Sign in to comment

No comments yet. Be the first to share your thoughts!