Bug Reports
Complete
Fixed an out-of-memory condition that brought down the production cluster and the services on it
What happens
The production cluster ran out of memory. When that happens the system starts terminating whatever is using the most memory, which on a busy node means killing the components that keep the cluster itself running — so the failure cascades rather than staying contained to one service.
Impact
Services across the platform went down on 2025-10-12 and did not come back on their own.
Resolution
The cluster was recovered node by node, memory limits were set on the workloads that had no ceiling, and reservations were put in place so the cluster's own components can no longer be starved by an application.
0 Comments
Sign in to comment
No comments yet. Be the first to share your thoughts!
