Fixed ThinkWatch's database major-version upgrade repeatedly failing on a storage attachment race
What happened
ThinkWatch's database was being moved up several major PostgreSQL versions. Each attempt restarted the database pods faster than the storage layer could release the underlying volume, so the new pod tried to attach a disk the old pod had not finished letting go of. The upgrade rolled back, and the next attempt hit the same race.
Impact
Six failed attempts across two sessions, 2026-07-31 to 2026-08-01, with ThinkWatch's database unavailable during each one. This was the one service out of six in that upgrade round that would not complete on the first pass.
Resolution
The remaining version steps were run with the node temporarily fenced off so the pods could not be rescheduled mid-upgrade, which removes the race entirely. Both remaining steps then succeeded on the first attempt. ThinkWatch now runs PostgreSQL 18.4 with all three replicas healthy.
0 Comments
Sign in to comment
No comments yet. Be the first to share your thoughts!
