Guide

Who can authorize the next incident mitigation?

In brief

A rollback debate can be blocked by compatibility evidence or unreachable authority. A decision worksheet separates those problems and records the recovery checks.

6 min read

Sources
A rollback boat waits ashore between two separate unfinished connections: an open fitting and an empty receiver.
Conceptual disputed-rollback preparation: migration compatibility and reachable authorization are different questions. Both connections remain unfinished; no rollback, accepted risk or restored service is shown.
On this page5 sections

Incident command establishes shared priorities and coordinates concurrent work. It should not be assumed to confer every production permission or every right to accept risk. Separating those responsibilities lets a commander keep the response moving while the appropriate service or data owner authorizes a consequential action.

A proposed rollback can stall for two reasons that look similar in the incident channel. The team may not know whether it is compatible with a migration already in progress. Or it may understand the compatibility risk but be unable to reach someone authorized to accept it. More telemetry can help the first problem; it cannot supply the missing decision maker in the second.

Following the connections among signal, interpretation, decision, authority, action and verification helps locate that difference. This is a way to inspect the response, not a rigid sequence. Authority usually needs preparation before the incident, and verification may send the investigation back to an earlier explanation.

Coordination and authorization have different jobs

Identify who proposes the action, who may approve its scope, who executes it and who accepts the resulting evidence. One person can fill several roles in a small incident, but the decisions still need to be explicit. Otherwise a title in the channel can be mistaken for authority that the person does not hold.

Google's incident-response guidance supplies a foundation for coordination. To apply it to a rollback, check whether the assigned roles can actually reach the person authorized to accept the data-affecting consequence. A clear role chart that ends at an unreachable owner leaves the response with the same practical gap.

Prepare for actions with unusual consequences before they are urgent: destructive recovery, cross-region traffic moves, schema rollback or temporary reductions in protection. Establish scope, authorization, a reachable fallback and the conditions that stop each action. Preparation makes urgency a reason to use a known route rather than invent permission.

The compatibility question inside a rollback debate

Consider a hypothetical incident in which requests fail after a deployment. One responder proposes rollback while another suspects that the migration makes it unsafe. The immediate technical question is whether the target version remains compatible with the affected data; the timing of the failure alone does not answer it.

Name the target version, affected data and the person able to verify compatibility, with a review point appropriate to service risk. In parallel, another owner can assess bounded traffic reduction. Keeping both action owners visible allows investigation to proceed concurrently without turning competing proposals into conflicting production changes.

If compatibility remains unknown, carry that uncertainty into the risk decision instead of polishing it out of the recommendation. The commander can organize the evidence and maintain the review point, while the established authorization path identifies who may choose among the remaining options.

The evidence available when a decision was made

A concise decision entry should capture the available evidence, rejected options, expected effect and the observation that would invalidate the choice. Include authorization scope and the next verification owner. It needs enough detail for a joining responder to understand the decision without becoming a separate documentation project during recovery.

This has an immediate benefit: a new responder can see why rollback was deferred before reopening the same debate. Later, the review can ask whether the decision was reasonable with the available information instead of judging it solely against the final explanation.

If the response crosses a shift boundary, the record also supports a live handoff. Google’s incident-management guidance calls for briefing the incoming commander, obtaining explicit acceptance and telling the other responders who now leads. In the hypothetical rollback case, that briefing includes the unresolved compatibility check, the person handling it and the next review point. The new commander inherits coordination of those threads; the established service and data-risk authorization path still determines who may approve the rollback.

Record the outcome separately and describe service behavior. A successful command may show that an interface accepted an operation while customer requests remain impaired. In asynchronous work, new requests may recover before old backlog. Keeping those results separate explains both when incident coordination can reduce and which recovery work still needs an owner.

Service recovery and backlog recovery can diverge
Service recovery and backlog recovery can diverge. For asynchronous work, new requests may succeed while older jobs remain delayed. Report both states and check replay for duplicate side effects before claiming that all affected work has completed.
For asynchronous work, new requests may succeed while older jobs remain delayed. Report both states and check replay for duplicate side effects before claiming that all affected work has completed.
Read diagram description

For asynchronous work, new requests may succeed while older jobs remain delayed. Report both states and check replay for duplicate side effects before claiming that all affected work has completed. Diagram labels: Mitigation applied: Begin independent recovery checks; New work: Is the live user path healthy?; Earlier work: Which jobs remain delayed or failed?; Recovery handoff: State both results, replay needs and residual owners.

Response intervals reveal preparation gaps

Separate time spent interpreting evidence from time finding an approver, obtaining access, applying the action and checking the effect. Mark uncertain timestamps as estimates. These intervals explain the system of response rather than rank individual responders.

Repeated authorization delays may call for delegated scope or better coverage. Conflicting changes may point to unclear action ownership. Slow recovery checks may require better verification. Stating which interval changed gives a faster response a useful explanation that can guide preparation for the next incident.

For a small, familiar incident, a single concise entry may be enough. Expand the structure only when uncertainty, scope, or concurrent work warrants it. Maintaining the record should support the recovery decision rather than become more urgent than it.

Carry the record into the post-incident review to examine the options and constraints that shaped the response. Retain uncertainty where it existed; hindsight should help identify a better preparation, not rewrite the earlier decision as obvious.

A worksheet for the disputed rollback

FieldHypothetical rollback caseFill for your incident
Decision nowIs rollback safe after this migration?State one decision.
Available evidenceErrors followed deployment; migration compatibility is unknown.Separate observations from interpretation.
Proposed action and scopeRoll back one service version if compatibility is demonstrated.Specify target, scope, and preconditions.
Rejected alternativeImmediate rollback is deferred because compatibility is unresolved.Record the reason, not just the rejection.
Authority and fallbackUse the established service and data-risk authorization path.Name reachable people or roles; do not invent permission.
Executor and verifierAssign both before execution; one person may fill both where appropriate.Record acceptance of responsibility.
Stop or revisit conditionStop if compatibility fails; revisit when evidence or impact changes.State an observable trigger and next review point.
OutcomeUnknown until service and backlog checks finish.Retain uncertainty and residual owners.

The rollback case is ready for rehearsal when the worksheet identifies the compatibility evidence, reachable authorization path and service checks that establish recovery. Each empty field points to a specific preparation problem: an unknown precondition, unavailable authority or undefined result.

A compatibility check and a reachable authorization path solve different parts of the rollback delay. Preparing both lets one responder investigate the migration while another assesses traffic reduction, with the commander coordinating their work. Recovery checks then establish what happened to new requests and older backlog, so the next shift inherits the remaining work as well as the decision.

Put this into practice

Reconstruct one incident

Bring: Known timestamps, response notes and evidence of recovery. Leave unknown times blank.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

Related reading

Explore a related question