Export the runbook for team review. The draft does not validate the procedure for your service.
Original combined worksheet and v1 drafts
Your previous toolkit files still work here. Focused tools also accept the relevant fields from older drafts.
Keep your working draft
Entries stay on this page. Download a draft to reopen later; nothing is saved automatically or sent to us.
This draft contains the timeline, incident notes, Markdown runbook and error-budget entries. Other tools have their own drafts.
01 / Reconstruct the response
Incident timeline calculator
Enter dates and times in UTC, including the date when an incident crosses midnight. Leave uncertain timestamps blank.
Detection
Unknown
Acknowledgment
Unknown
Detection to investigation
Unknown
Investigation to recovery
Unknown
Mitigation to recovery
Unknown
Total impact
Unknown
Intervals use the named endpoints. This tool assumes one response sequence; for repeated mitigations, use the worksheet to record each attempt. A single incident supplies durations, not MTTD or MTTR averages. Unknown onset means unknown detection duration, not zero.
02 / Preserve the evidence
Editable incident worksheet
Add impact, evidence, recovery checks, and follow-up. The download includes your current timeline and calculated intervals.
Use eligible and bad events from the same measurement window. A bad event fails your defined success or latency condition. This is an event allowance, not a downtime allowance.
Use your agreed release policy to decide the next action. This indicator describes consumption, not a release recommendation.
Allowed = eligible × (1 − target / 100). Remaining = allowed − bad. Negative remaining means the allowance is exceeded. Fractional allowances are shown without rounding to whole events; display values are rounded to three decimals. Missing telemetry can make these inputs incomplete. This calculation does not authorize or block a release.
# Incident worksheet
All timestamps are UTC. Unknown times are left unknown.
Impact began: Unknown
Incident detected: Unknown
Page acknowledged: Unknown
Investigation began: Unknown
Mitigation applied: Unknown
Service recovered: Unknown
## Timing
Detection: Unknown
Acknowledgment: Unknown
Detection to investigation: Unknown
Investigation to recovery: Unknown
Mitigation to recovery: Unknown
Total impact: Unknown
Single-incident durations are not mean metrics.
## Incident context
Service / environment:
Incident owner and record:
User impact and affected population:
Discovery method (monitoring / engineer / customer):
Timestamp evidence, proxies and uncertainty:
## Response and recovery
Actions taken and observed results:
Customer-outcome check and observation window:
Recovery evidence / unresolved impact:
## Follow-up
Failure scenario:
Proposed treatment:
Test that would demonstrate the intended improvement:
Remaining exposure:
Owner / due date / evidence link:
# Runbook: [specific failure mode]
Status: Draft, complete and exercise before operational use.
Service / environment:
Owner and escalation route:
Related alert:
Required access:
Last edited:
Last exercised / evidence:
## Scope and stop conditions
Applies when:
Does not apply when:
Stop and escalate if:
## Confirm impact
Customer symptom:
Exact dashboard / query and time window:
Expected healthy result:
Observed failure result:
Data freshness check:
## Initial coordination
Acknowledge via:
Incident record / channel:
Response lead and communication owner:
Next update time:
## Diagnosis
1. Check [specific condition] using [read-only query or procedure].
- If [result], continue to [step].
- If [other result], escalate to [reachable owner].
2. Check recent changes and relevant dependencies.
- Evidence supporting current hypothesis:
- Evidence contradicting it:
## Bounded mitigation
Action and approved workflow:
Required preconditions:
Explicit target / environment:
Maximum scope and attempts:
Authorization required:
Expected effect and time limit:
Partial-failure handling:
Rollback or recovery procedure, if supported:
Stop condition:
## Verify recovery
Independent customer-outcome check:
Dependency / resource check:
Observation window:
Result: verified / unsuccessful / inconclusive
If unsuccessful or inconclusive, escalate via:
Evidence to hand over: timeline, impact, checks, actions, results, uncertainties.
## Follow-through
Record recovery time and remaining risks.
Assign an owner and due date to unresolved work.
Correct procedure gaps found during use.
Schedule an appropriate review and repeat exercise.
# Error-budget worksheet
Illustrative example
Service and window: [Not specified]
SLO: 99.9%
Eligible: 1000000
Bad: 250
Allowed: 1000
Remaining: 750
Consumed: 25%
Event allowance is not downtime allowance.
Understand the measures behind the tools
Detection and MTTD
Detection duration runs from impact onset to discovery. MTTD averages those durations across a defined incident population; one incident is not a mean.
Reviewed 2026-09-29. Guide-to-tool alignment reviewed; v2 arithmetic, imports and explicit assumptions checked with illustrative fixtures.
Incident response
Record detection, acknowledgment, investigation, mitigation and recovery separately. Use the worksheet to retain impact, actions and recovery evidence.
Reviewed 2026-09-29. Guide-to-tool alignment reviewed; v2 arithmetic, imports and explicit assumptions checked with illustrative fixtures.
Runbooks
A runbook connects a specific symptom with checks, bounded actions, recovery criteria and escalation. Complete and exercise it for the service before use.
Reviewed 2026-09-29. Guide-to-tool alignment reviewed; v2 arithmetic, imports and explicit assumptions checked with illustrative fixtures.
Postmortem follow-through
A completed task and a demonstrated reduction in exposure are different results. Record the failure scenario, treatment, test, remaining exposure and responsible owner.
Reviewed 2026-09-29. Guide-to-tool alignment reviewed; v2 arithmetic, imports and explicit assumptions checked with illustrative fixtures.
Error budgets
For a request-based objective, allowed bad events equal eligible events times the permitted bad-event fraction. The result stays in event units; it is not an allowance for downtime.
Reviewed 2026-09-29. Guide-to-tool alignment reviewed; v2 arithmetic, imports and explicit assumptions checked with illustrative fixtures.
SRE KPIs
Keep metric endpoints, population, window and missing data explicit. Use individual incident intervals to inspect delays before combining incidents into a headline average.
Reviewed 2026-09-29. Guide-to-tool alignment reviewed; v2 arithmetic, imports and explicit assumptions checked with illustrative fixtures.
Sources and calculation scope
The timeline follows the explicit endpoints in our linked metric guides. The worked example is illustrative. Calculator checks verify arithmetic and input handling; they do not establish production outcomes.