Bug Reports
Planned

Management cluster cannot survive losing one server

Found during a full audit of the migration plan against the live estate, 2026-09-07.

The platform runs two clusters. The one that serves you is correctly spread so that any single server can fail without interrupting anything � that was verified and is working as designed.

The second cluster, which runs the platform's own tooling rather than the services you use, is not spread that way right now. Two of its three coordinating members ended up on the same physical server during the migration. If that one server goes down, that cluster stops until it comes back.

What this does and does not affect. No service you use depends on this. The risk is to administration and automation � deployments, backups tooling and configuration management would pause until the server returned.

Why it happened, and why it was reasonable. The third server still has a disabled controller cache from a hardware fault. A previous outage was traced to exactly that combination � an unbuffered write path under a coordinating member � so one was deliberately kept off it. That was the right call at the time; the cost is the concentration this ticket describes.

Two honest ways out, no third: repair the controller cache and move the member back, or record a deliberate decision to accept that this cluster tolerates no loss of that one server. Leaving it unstated is the only bad option, because the written plan still says the members are spread one-per-server when they are not.

Tracked in the migration plan as P4.15c.

0 Comments

Sign in to comment

No comments yet. Be the first to share your thoughts!