Two production machines stopped when their storage silently filled
Two machines stopped accepting writes and shut down during the nightly backup. The backups themselves were fine; the storage underneath them was not.
Space on this platform is allocated on demand rather than reserved up front, which means deleted data has to be actively handed back or the space is never recovered. That handback was never switched on. The result is not a leak that shows up gradually - the usage figure looks reasonable right up until it reaches the limit, and then writes stop.
Both affected machines were on pools that had reached their limit. Machines on a healthy pool on the same hardware were entirely unaffected, which is what identified the cause.
The handback is now enabled on every disk of every workload, and each was told to return the space it was already holding unnecessarily. The pools that were completely full are now between a seventh and a half used.
The detail worth keeping: the piece that reclaims space inside each machine had been running all along, on a weekly schedule, achieving nothing - because the layer beneath it was discarding the request. It looked correctly configured from the inside for the entire life of the platform.
0 Comments
Sign in to comment
No comments yet. Be the first to share your thoughts!
