Bug Reports
Complete

Services would not start for hours after an overnight maintenance restart

After a planned restart of every machine in the platform, a large number of services could not start again. They stayed stuck waiting for their storage, for several hours, overnight.

What was wrong

Two thirds of the machines came back carrying stale storage devices left over from before the restart. Anything touching one of those devices became stuck in a state the system cannot interrupt or kill, which in turn stopped each machine from tidying up its own storage processes. Storage could then never be handed to the services that needed it.

The fault was unusually well hidden: every machine reported itself healthy, every component reported itself running, and nothing raised an alert. One machine showed a load figure five times its normal level while its processors were 70% idle - a signature of waiting on storage rather than doing work.

Because the same machine also handles inbound traffic, that stall made the whole platform feel slow from outside. The slowness and the stuck services were the same fault, not two.

Fix

The affected machines were restarted one at a time, each drained of work beforehand and checked afterwards. There is no software way out of the stuck state - a restart is the only exit.

A standing post-maintenance check has been added to the runbook, because the condition is invisible to ordinary health reporting and was found only by reading a machine's kernel log by hand.

0 Comments

Sign in to comment

No comments yet. Be the first to share your thoughts!