Name resolution no longer depends on a single machine
Every service on the platform depends on name resolution, and until now that ran on one small machine with one network connection on one switch port. When it stopped, nothing announced it: the router had also been configured to answer, and quietly did so, so queries kept being answered while filtering, internal name lookups and logging were all silently missing. It went unnoticed for two days.
There is now a second, independent resolver on different hardware, connected to two separate switches rather than one. Both answer on their own, with the same filtering rules and the same internal lookups. If one stops, the other simply continues.
The two are kept identical automatically. The original remains the single place where settings are changed, and its configuration is copied to the second every fifteen minutes, so the two cannot quietly drift into different behaviour. Their individual network addresses are deliberately excluded from that copy, which is what lets two otherwise-identical copies run side by side.
Every device now receives both addresses - those that configure themselves automatically, and the servers and virtual machines that are set by hand and would otherwise never have been told the second one exists.
What this does not do: both sit in the same building on the same power, so this removes the single-machine failure, not a site-wide one.
0 Comments
Sign in to comment
No comments yet. Be the first to share your thoughts!
