3 steps · Incident response

Review an incident

Reconstruct the response, review corrective actions and update the runbook.

BringIncident timestamps, response notes, action owners and recovery evidence.

Finish withA timing report, a follow-up action register and a revised runbook draft.

  1. Step 1

    Reconstruct the response

    Record the timestamps you know and keep missing or uncertain evidence visible.

    Before moving on: Review the timing report against your incident evidence.

    Use Prepare next draft to carry notes and the report into the action tracker. Open and review the transfer in the same browser tab.

  2. Step 2

    Review the follow-up actions

    Record the failure scenario, responsible owner, proposed change and verification evidence.

    Before moving on: Distinguish work delivered from a change whose effect has been verified.

    Select an action and use Prepare next draft to carry it into the runbook builder. Review the transferred fields before continuing.

  3. Step 3

    Update and exercise the runbook

    Turn the relevant findings into a procedure with recovery checks and escalation.

    Before moving on: Save the draft and record an exercise before relying on the procedure.

    Export the runbook for team review. The draft does not validate the procedure for your service.

Original combined worksheet and v1 drafts

Your previous toolkit files still work here. Focused tools also accept the relevant fields from older drafts.

Keep your working draft

Entries stay on this page. Download a draft to reopen later; nothing is saved automatically or sent to us.

This draft contains the timeline, incident notes, Markdown runbook and error-budget entries. Other tools have their own drafts.

01 / Reconstruct the response

Incident timeline calculator

Enter dates and times in UTC, including the date when an incident crosses midnight. Leave uncertain timestamps blank.

Detection
Unknown
Acknowledgment
Unknown
Detection to investigation
Unknown
Investigation to recovery
Unknown
Mitigation to recovery
Unknown
Total impact
Unknown

Intervals use the named endpoints. This tool assumes one response sequence; for repeated mitigations, use the worksheet to record each attempt. A single incident supplies durations, not MTTD or MTTR averages. Unknown onset means unknown detection duration, not zero.

02 / Preserve the evidence

Editable incident worksheet

Add impact, evidence, recovery checks, and follow-up. The download includes your current timeline and calculated intervals.

Read the postmortem follow-up guide
03 / Prepare the next response

Editable SRE runbook

Fill in this existing AIOpsSRE template for one service and failure mode. Exercise the procedure before using it during an incident.

Download blank templateRead the runbook guide
04 / Check the allowance

Request-based error-budget calculator

Use eligible and bad events from the same measurement window. A bad event fails your defined success or latency condition. This is an event allowance, not a downtime allowance.

Illustrative starting values: 99.9% objective, 1,000,000 eligible events, 250 bad events.

Allowed bad events
1,000
Remaining allowance
750
Budget consumed
25%
Good-event rate
99.975%
Allowance remaining

750 bad events remain in this measurement window.

25%

Use your agreed release policy to decide the next action. This indicator describes consumption, not a release recommendation.

Allowed = eligible × (1 − target / 100). Remaining = allowed − bad. Negative remaining means the allowance is exceeded. Fractional allowances are shown without rounding to whole events; display values are rounded to three decimals. Missing telemetry can make these inputs incomplete. This calculation does not authorize or block a release.

Calculate burn rateRead the error-budget guide
# Incident worksheet

All timestamps are UTC. Unknown times are left unknown.

Impact began: Unknown
Incident detected: Unknown
Page acknowledged: Unknown
Investigation began: Unknown
Mitigation applied: Unknown
Service recovered: Unknown

## Timing
Detection: Unknown
Acknowledgment: Unknown
Detection to investigation: Unknown
Investigation to recovery: Unknown
Mitigation to recovery: Unknown
Total impact: Unknown

Single-incident durations are not mean metrics.

## Incident context
Service / environment:
Incident owner and record:
User impact and affected population:
Discovery method (monitoring / engineer / customer):
Timestamp evidence, proxies and uncertainty:

## Response and recovery
Actions taken and observed results:
Customer-outcome check and observation window:
Recovery evidence / unresolved impact:

## Follow-up
Failure scenario:
Proposed treatment:
Test that would demonstrate the intended improvement:
Remaining exposure:
Owner / due date / evidence link:


# Runbook: [specific failure mode]

Status: Draft, complete and exercise before operational use.
Service / environment:
Owner and escalation route:
Related alert:
Required access:
Last edited:
Last exercised / evidence:

## Scope and stop conditions
Applies when:
Does not apply when:
Stop and escalate if:

## Confirm impact
Customer symptom:
Exact dashboard / query and time window:
Expected healthy result:
Observed failure result:
Data freshness check:

## Initial coordination
Acknowledge via:
Incident record / channel:
Response lead and communication owner:
Next update time:

## Diagnosis
1. Check [specific condition] using [read-only query or procedure].
   - If [result], continue to [step].
   - If [other result], escalate to [reachable owner].
2. Check recent changes and relevant dependencies.
   - Evidence supporting current hypothesis:
   - Evidence contradicting it:

## Bounded mitigation
Action and approved workflow:
Required preconditions:
Explicit target / environment:
Maximum scope and attempts:
Authorization required:
Expected effect and time limit:
Partial-failure handling:
Rollback or recovery procedure, if supported:
Stop condition:

## Verify recovery
Independent customer-outcome check:
Dependency / resource check:
Observation window:
Result: verified / unsuccessful / inconclusive
If unsuccessful or inconclusive, escalate via:
Evidence to hand over: timeline, impact, checks, actions, results, uncertainties.

## Follow-through
Record recovery time and remaining risks.
Assign an owner and due date to unresolved work.
Correct procedure gaps found during use.
Schedule an appropriate review and repeat exercise.


# Error-budget worksheet

Illustrative example
Service and window: [Not specified]
SLO: 99.9%
Eligible: 1000000
Bad: 250
Allowed: 1000
Remaining: 750
Consumed: 25%
Event allowance is not downtime allowance.

Understand the measures behind the tools

Runbooks

A runbook connects a specific symptom with checks, bounded actions, recovery criteria and escalation. Complete and exercise it for the service before use.

Reviewed 2026-09-29. Guide-to-tool alignment reviewed; v2 arithmetic, imports and explicit assumptions checked with illustrative fixtures.

Postmortem follow-through

A completed task and a demonstrated reduction in exposure are different results. Record the failure scenario, treatment, test, remaining exposure and responsible owner.

Reviewed 2026-09-29. Guide-to-tool alignment reviewed; v2 arithmetic, imports and explicit assumptions checked with illustrative fixtures.

Sources and calculation scope

The timeline follows the explicit endpoints in our linked metric guides. The worked example is illustrative. Calculator checks verify arithmetic and input handling; they do not establish production outcomes.

For future guides and resources, get AIOpsSRE by email or follow RSS.