redis-operator restarts on rke-prod caused by leader-election lease timeout
The redis-operator pod in namespace redis-operator on the rke-prod cluster has been restarting intermittently (223 restarts over 18 days as of ~2026-08-19, down to ~17 restarts over 11 days / ~1.5 per day as of 2026-09-13, stable for the last 3 days).
Root cause (confirmed via --previous pod logs): a leader-election lease-renewal timeout against the Kubernetes API server (context deadline exceeded on the coordination.k8s.io Lease API call), causing the controller-runtime manager to self-exit(1) with "leader election lost". Not an OOMKill (exit code 1, not 137; memory well within its 500Mi Guaranteed limit) and not a liveness/readiness probe kill (the health-probe server was up; the shutdown was self-triggered and orderly).
Correlates with node-level API-connectivity flakiness on worker node lprkeprod08: recurring client-go "watch ended: short buffer" errors in the operator's own log, several unrelated pods on that same node showing similarly elevated restart counts (rke2-canal CNI, authentik-rke-prod-server, yamtrack-rke-prod, node-exporter), and a rke2-cert-monitor CertificateExpirationWarning event firing repeatedly (84 times over 4d2h) about RKE2 system certs on that node. Circumstantial, not proven as the direct trigger, but consistent with the kind of issue that produces intermittent apiserver TLS/API-call timeouts.
Recommended follow-up: run a cert check against node lprkeprod08 specifically (a worker-node cert check, distinct from the already-tracked lprkegest04 control-plane cert expiry on 2026-09-19). Not urgent - the restart rate has already dropped ~8x and the pod has been stable for 3 days - but the underlying trigger is unaddressed and should be expected to recur.
Separate from the ThinkRedis application-instance hardening work (auth/NetworkPolicy/TLS) - this is about the operator itself, not a consumer.
0 Comments
Sign in to comment
No comments yet. Be the first to share your thoughts!
