Bug Reports
Complete

Production control-plane repeatedly unstable

Control-plane nodes on the production cluster flapped for several days: API servers restarted repeatedly and nodes cycled between Ready and NotReady.

Root cause was storage replica data sharing one filesystem with the cluster database's write-ahead log, so replica rebuild traffic starved database commits. Storage was moved off all five control-plane nodes; commit latency fell from 18.02 ms/op to 1.64 ms/op and the flapping stopped.

0 Comments

Sign in to comment

No comments yet. Be the first to share your thoughts!