The management cluster now survives losing any single machine
The three coordinators that keep the management cluster running are now spread one per machine.
They had ended up two-on-one-machine, which meant that losing that single machine would have taken two of the three at once — and a cluster that loses more than half its coordinators stops accepting changes until someone intervenes. Nothing was broken day to day; the risk only showed up on the day a machine failed.
The reason they were bunched up was deliberate: the third machine had a faulty write-cache module, and putting a coordinator on unprotected storage is exactly what caused an earlier outage. That hardware fault is now repaired, so the coordinator has been moved back where the design intended.
The move was done live, with no downtime and no interruption to anything running on the cluster.
0 Comments
Sign in to comment
No comments yet. Be the first to share your thoughts!
