# Runbook: [specific failure mode]

Status: Draft, complete and exercise before operational use.
Service / environment:
Owner and escalation route:
Related alert:
Required access:
Last edited:
Last exercised / evidence:

## Scope and stop conditions
Applies when:
Does not apply when:
Stop and escalate if:

## Confirm impact
Customer symptom:
Exact dashboard / query and time window:
Expected healthy result:
Observed failure result:
Data freshness check:

## Initial coordination
Acknowledge via:
Incident record / channel:
Response lead and communication owner:
Next update time:

## Diagnosis
1. Check [specific condition] using [read-only query or procedure].
   - If [result], continue to [step].
   - If [other result], escalate to [reachable owner].
2. Check recent changes and relevant dependencies.
   - Evidence supporting current hypothesis:
   - Evidence contradicting it:

## Bounded mitigation
Action and approved workflow:
Required preconditions:
Explicit target / environment:
Maximum scope and attempts:
Authorization required:
Expected effect and time limit:
Partial-failure handling:
Rollback or recovery procedure, if supported:
Stop condition:

## Verify recovery
Independent customer-outcome check:
Dependency / resource check:
Observation window:
Result: verified / unsuccessful / inconclusive
If unsuccessful or inconclusive, escalate via:
Evidence to hand over: timeline, impact, checks, actions, results, uncertainties.

## Follow-through
Record recovery time and remaining risks.
Assign an owner and due date to unresolved work.
Correct procedure gaps found during use.
Schedule an appropriate review and repeat exercise.
