Everything starts itself again after a power cut
When the platform was taken apart for the hardware move, automatic start-up was switched off on every workload so nothing would come back in the wrong place partway through. It was never switched back on, and no step existed to do it.
That gap only shows itself during a real outage. The machines now shut themselves down cleanly when utility power is lost, but nothing would have started again once power returned - including the directory and name service that the rest of the platform depends on. The estate would have survived the power cut and then stayed down until somebody noticed.
Automatic start-up is on again, in a deliberate order: the directory and name service first, then the cluster control plane, then the remaining workloads, with no artificial wait between them.
One limitation is worth stating plainly: start-up order is honoured per machine rather than across the cluster, so if all three machines power on at the same moment, a control plane on one machine can come up before the directory service on another. Within a single machine the ordering is exact.
The workloads belonging to the offline tooling cluster were deliberately left switched off - they are waiting on their replacement storage, and starting them into missing disks would be worse than leaving them down.
0 Comments
Sign in to comment
No comments yet. Be the first to share your thoughts!
