Bug Reports
Complete

Sign-in does not work on the backup name resolver

The second, independent name resolver was brought online with its own protected web interface, but signing in to it never worked.

What was happening

The access rules for both resolvers are declared as code and handed to the sign-on platform as a single document. That document is applied all-or-nothing. One entry inside it referred to something defined further down the same file, which does not exist yet at the moment it is read, so the entry was rejected and the entire document was discarded — every time it was applied, for as long as the second resolver had existed.

Why it was not obvious

Nothing reported an error anywhere a person would look. The document was correct in version control, correct on the cluster, and correctly delivered to the sign-on platform. The only visible signs were things that looked unrelated: the first resolver kept an address it should have moved away from, the second resolver's access rules were simply absent, and the sign-on gateway went on answering for the old shared address. It read as though the work had never been applied at all.

Fix

The document was reordered so nothing refers to anything defined later in it. The correction was proven against the live platform before release using a mode that performs the real application inside a transaction and then rolls it back, so both the old and new orderings could be tested safely: the old order is rejected, the new order succeeds and creates the missing access rules.

The house templates and the internal procedure have been updated so the next service published on two addresses cannot repeat this, including a direct way to check whether the platform is actually applying a document rather than assuming a correct-looking file means a correct result.

0 Comments

Sign in to comment

No comments yet. Be the first to share your thoughts!