Bug Reports
Complete

Four services lost access to their data during the storage fault

Four services were left without usable data when the storage fault ran its course, and could not start.

Recovery

Two were recovered in place: the underlying copies were intact and only marked bad, so clearing that marking brought them back to the moment of failure with nothing lost. One of those had been reported by the platform as having no recoverable data at all, so attempting recovery before falling back to a backup is always worth it.

One database was rebuilt directly from its own live replica rather than from a backup, recovering it to the current moment instead of losing the day's changes.

The remainder were restored from that morning's backups, into their original storage so that nothing else had to be reconfigured to find them.

Prevention

The two faults behind this are now documented with their symptoms, and the maintenance procedure that triggered them has gained a check that would have caught the condition within minutes rather than hours.

0 Comments

Sign in to comment

No comments yet. Be the first to share your thoughts!