On this page5 sections
In a hypothetical backup review, the cleanup change is merged and backups finish within their window in the tested configuration. An older data format still has no verification result. Is the postmortem action done? Is the risk closed? Asking only one of those questions forces the team to hide either the progress or the uncertainty.
A risk registry is useful when it preserves both. It records failure scenarios that still matter, their treatments and the evidence for what has changed. Delivery and risk disposition need separate decisions: completing the agreed work does not by itself establish which exposure that work removed. Equally, an unresolved format should not make the delivered improvement vanish from the record.
Why stale records delay backups
In a hypothetical backup service, stale records accumulate until enumeration takes too long and jobs miss their completion window. Enumeration lists the records the backup needs to process; extra stale entries make that step longer. The mechanism gives a vague action such as “improve backup monitoring” a concrete problem to address.
It also exposes several choices. Repairing record lifecycle can prevent accumulation. Reducing enumeration work can shorten the job. Earlier detection can give a responder more notice, while a recovery path may tolerate delayed backups. These treatments change different parts of the problem. A faster warning does not necessarily shorten backup duration, and cleanup can introduce harm if it removes records an active operation still needs.
Google’s postmortem guidance emphasizes learning and effective follow-up. The registry carries unresolved exposure from that review into engineering priorities and acceptance decisions. Linking the source incident and delivery work preserves the evidence without copying the whole postmortem into another document.
What the cleanup test establishes
The desired claim should determine the test. “Enumeration stays within the backup window for the tested record count” identifies a property and scope that representative evidence can support. “Backup reliability is fixed” may also imply formats, regions, scale and recovery paths that nobody tested. The broader wording creates obligations the result may not fulfill.
Suppose the hypothetical team implements bounded cleanup and demonstrates backup and restore for one supported configuration. That is meaningful evidence. The untested older format then becomes its own remaining decision: test it, contain it, retire it or accept the exposure through an authorized owner and review trigger. Closing everything would overstate the result; leaving everything open would discard useful progress.
The separate restore check matters even when backup is now faster. Backup duration helps establish freshness of saved data. It does not establish the time needed to restore that data. A later recovery plan may rely on the registry, so a recovery-time claim needs evidence from recovery itself rather than a nearby measurement that sounds related.
Read diagram description
Delivery and verification are separate decisions. Evidence may close only part of a failure scenario; remaining exposure still needs an owner and a treatment or time-limited acceptance decision. Diagram labels: Failure scenario: Evidence, affected scope and consequence; Owner + treatment decision: Mitigate, investigate or accept with an expiry; Delivery work: Implement the selected change; Verification evidence: Test the mechanism and the stated scope; Scope demonstrated: Close that portion; Exposure remains: Retain an owned decision.
Who can accept the untested exposure?
Treatment may exceed available capacity or introduce a more serious competing risk. Acceptance can be a legitimate outcome when the actual tradeoff is recorded. It needs defined scope, an authorized decision maker, a rationale, a review or expiry date and an event that would reopen the choice earlier.
The engineer assigned the cleanup may not have authority to accept customer or business exposure from the old format. Likewise, an expired acceptance should return for review rather than renew itself through inattention. These distinctions give the record an operational purpose: it changes who must decide, what gets funded or where the response boundary lies.
Priority labels can organize that discussion without pretending to measure a calibrated probability of loss. A destructive scenario with sparse evidence may deserve attention before a frequent inconvenience. Saving the reason for that choice lets a later reviewer reconsider it when evidence, capacity or consequences change, rather than arguing over a color whose meaning has been lost.
Task counts and risk reduction measure different things
Task and risk counts do not share a natural denominator. Several tasks can treat one risk, and one task can affect several risks. A percentage of tasks closed therefore cannot be translated into the same percentage of risk removed. Reporting both levels separately keeps the arithmetic from making a claim the underlying work did not establish.
Useful progress statements say what became true: a recovery path was exercised, an alert reaches a staffed owner, a failure mechanism is contained for a defined population, or an acceptance expired and requires review. Those statements remain informative even when new findings increase the total backlog.
Durable risk tracking is warranted for material exposure, incomplete verification and decisions crossing ownership or planning boundaries; small local corrections can stay in the delivery workflow when their evidence is easy to find. Maintaining a second permanent record for every fix would add work without preserving an otherwise lost decision.
The companion guide to turning postmortem lessons into verifiable changes covers acceptance evidence for an individual action. The registry’s broader job is to retain the exposure and choices that must survive across those actions.
A treatment record for the backup failure
| Field | Hypothetical backup example |
|---|---|
| Failure mechanism | Stale records increase enumeration time until the backup window is missed. |
| Exposed population | Record the specific controller version, configuration, and tested scale. |
| Treatment and alternative | Use bounded cleanup plus lifecycle repair; reject unbounded deletion because active operations may depend on records. |
| Claim to demonstrate | Enumeration and backup complete within the required window for the stated population. |
| Evidence | Link representative test conditions, observed duration, successful backup, and a separate restore check. |
| Residual exposure | Older formats and untested populations remain explicit. |
| Disposition | Close only the demonstrated scope; assign further treatment or authorized acceptance to the remainder. |
| Reopen trigger | Revisit after population growth, a relevant controller change, failed verification, or acceptance expiry. |
The backup entry can close the cleanup work and the configuration whose backup and restore tests passed, while retaining the older format as an unresolved case. That remaining entry tells planning what is still needed: a representative test, containment, retirement or authorized acceptance. Progress stays visible without extending a successful result to data nobody exercised.
Put this into practice
Check what your corrective actions demonstrated
Bring: The failure scenario, proposed treatment, owner, due date, verification test and available evidence.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- Google’s postmortem guidancesre.google
