Feedback Requests
Complete

A weekly all-clear from the backup platform, not just failure alarms

Alerting until now only spoke up when something broke. That sounds sufficient and is not: a backup that succeeded and a backup that never started look exactly the same from a phone that stays quiet.

That is not a hypothetical. One of the three nightly backup jobs silently missed its very first trigger, by two minutes, and nothing said so. It was found by hand a day later.

Every machine in the platform now sends a short weekly summary: how many backups ran, whether any failed, whether every workload is actually covered, how much room is left on the backup store, and - on the backup server itself - whether its housekeeping, pruning and integrity checks are healthy.

Each machine reports for itself rather than one machine reporting for all of them. That is deliberate. If a machine is down you lose exactly one report, and the one that is missing names the machine that went quiet. A single combined report would simply fail to arrive and tell you nothing about why.

Reports also catch up after an outage rather than being skipped, so a report that never turns up means the machine is still down.

0 Comments

Sign in to comment

No comments yet. Be the first to share your thoughts!