Name resolution failed across the platform
Services could not look each other up by name, and several could not reach their databases. The feedback portal itself was among them.
What happened. The component that answers internal name lookups was capped at a fraction of a CPU core. Under normal load that was enough; once it fell behind, the systems that depend on it retried, the retries added load, and it fell further behind — a loop that does not recover on its own.
Why it was hard to spot. Lookups for outside addresses kept working the whole time, because those take a different path. So the symptom looked like individual machines misbehaving rather than one shared component running out of headroom.
Fix. The cap was raised with room to absorb a surge instead of amplifying one. Name resolution recovered immediately on every machine, and the component now idles at a small fraction of what it is allowed. The same limit was raised on the management cluster before it could cause the same outage there.
0 Comments
Sign in to comment
No comments yet. Be the first to share your thoughts!
