Bug Reports
Complete

Fixed an out-of-memory condition that brought down the production cluster and the services on it

What happens

The production cluster ran out of memory. When that happens the system starts terminating whatever is using the most memory, which on a busy node means killing the components that keep the cluster itself running — so the failure cascades rather than staying contained to one service.

Impact

Services across the platform went down on 2025-10-12 and did not come back on their own.

Resolution

The cluster was recovered node by node, memory limits were set on the workloads that had no ceiling, and reservations were put in place so the cluster's own components can no longer be starved by an application.

0 Comments

Sign in to comment

No comments yet. Be the first to share your thoughts!