Guide

Writing a post-incident review the next responder can use

In brief

A delayed rollback connects the incident timeline to the information responders had, the compatibility question they faced and the follow-up that remains.

5 min read

Sources
An unfolded procedure leads to separate application and data inputs beneath a raised, unlatched checking arm.
Conceptual illustration of the hypothetical rollback review: clearer instructions do not establish compatibility with migrated data. The raised arm represents a pending exercise, not a failed or successful recovery.
On this page4 sections

In a hypothetical post-incident review, everyone agrees that rollback took too long. The team updates the procedure, closes the action and moves on. A later responder can now find the command quickly, but still cannot tell whether the old application version will work with data that has already been migrated. The review improved the instructions while leaving the difficult decision unresolved.

A useful review makes that distinction visible. It explains what happened, reconstructs why the available choices looked the way they did, and leaves the next person enough context to continue the improvement. That requires more than a timeline and a list of tasks, but it does not require an elaborate meeting for every incident. A recurring, well-understood failure may warrant a focused correction; an unfamiliar failure spanning several services may need interviews and a fuller analysis.

The timeline needs the information available at the time

Start with a draft that participants can correct: customer impact, measurements, logs, change records, incident notes and actions taken. Identify time zones and estimated timestamps so the sequence is useful without implying a precision the evidence cannot support. Include people whose work happened outside the main incident channel, such as support staff or dependency owners. Their account may explain a delay that otherwise appears to be empty time.

In the rollback example, knowing when someone opened the runbook is less revealing than knowing what the runbook allowed them to determine. Was compatibility information available? Did someone have permission to perform the operation? Which observations made waiting seem safer than proceeding? These questions connect the recorded sequence to the reasoning behind it.

Give responders time to recover and a way to contribute asynchronously. Separate observations from interpretations, and leave conflicting accounts visible until evidence can resolve them. A tidy narrative is not worth losing an uncertainty that mattered to the decision.

The defect and the delayed response have different causes

The failure explanation traces the trigger, the conditions that let it spread and its effect on users. The response explanation follows a different path: detection, routing, investigation, mitigation and recovery. Keeping both in view prevents a fix to the original defect from standing in for a fix to the response. An application change may remove the trigger while leaving the same permission gap for the next incident.

Look for the things that made recovery possible, too. A manual cross-check may have caught an unsafe assumption, or an effective escalation may have supplied the missing context. Understanding those adaptations helps the team decide what to preserve rather than treating every departure from the procedure as a mistake.

The blameless-review guide develops this examination of decisions without making a person the explanation for the failure. The 5 Whys method can help explore a mechanism, provided a branching incident is allowed to retain its branches. In the rollback case, unclear instructions and uncertain data compatibility could both matter without being interchangeable causes.

What would the proposed fix let someone do?

In the rollback review, the follow-up has to answer whether the older application can use the migrated data. Updating the procedure is one piece of that work; checking compatibility is another. Naming who will perform the remaining check lets the review end with an accurate account of what is ready and what the next responder still cannot assume.

A compatibility exercise needs to identify the application versions and migrated data it covers. Its result can then be incorporated into the runbook: which recovery path worked, which did not, and which conditions were outside the exercise. The owner can judge the implementation and continuing maintenance cost against that specific improvement in the next responder’s choices.

Some reviews reveal useful understanding without identifying a worthwhile new control. In that case, preserve the explanation rather than manufacture action items to make the review look productive. Where material exposure does remain, distinguish unanswered questions and accepted risk from delivery work. A changed backup configuration, for example, cannot establish a successful restore that was never tested. The risk registry provides a home for consequential uncertainty that needs a later planning decision.

Carry the explanation into follow-up work
Carry the explanation into follow-up work
Proposed review flow. A corrected timeline supports analysis, and analysis supports a specific change. Completion of the task still needs a check of its intended effect.
Read diagram description

Proposed review flow. A corrected timeline supports analysis, and analysis supports a specific change. Completion of the task still needs a check of its intended effect. Diagram labels: Corrected record: Timeline and participants’ evidence; Explanation: Contributing conditions and difficult decisions; Selected follow-up: Owner, capacity and expected result; Later check: Did the intended behavior change?.

What another team needs from the review

Share the mechanism and the consequential choices with the people who can learn from them, using appropriate access controls and removing unnecessary sensitive detail. Google's postmortem analysis chapter gives further context for learning across incidents. The local record should make it possible to recognize a related condition without requiring another team to read the entire incident channel.

Put verification into the normal planning process so it has a realistic owner and capacity. When the rollback work returns for review, the useful question is now specific: which migrated data and application versions were checked, what happened, and what remains uncertain? The next responder inherits both better instructions and an intelligible recovery decision. That is an outcome a closed documentation ticket alone could not provide.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

Related reading

Explore a related question