On this page5 sections
An SRE runbook template organizes a service-specific procedure: the symptom to confirm, diagnostic checks, permitted actions, recovery criteria and escalation. Its purpose is to let another qualified responder follow the procedure and recognize whether it worked.
A useful runbook helps the next responder decide what to do when the familiar fix is unavailable. A database restart might be a known procedure, but an active migration can change whether it is safe. The document needs to explain how to check that condition and how to continue when the answer cannot be obtained.
Download the Markdown runbook template to structure a service-specific procedure. It needs local details, access checks and rehearsal before use; the template alone does not authorize commands in an unknown environment. The worked example below shows how to connect the fields into an investigation rather than a list of plausible actions.
The parts of a usable procedure
| Section | What the responder needs |
|---|---|
| Scope and ownership | Service, environment, related alert, owner, required access, and last tested date. |
| Confirm the symptom | User impact, exact dashboard or query, time window, expected result, and signs this is a different failure. |
| Diagnose | Ordered checks with branches that explain what each result means. |
| Act | Preconditions, approved command or workflow, target limits, expected effect, and stop conditions. |
| Verify and escalate | Independent recovery checks, observation window, reachable owner, and evidence to hand over. |
The sections progress from identifying the right problem to determining whether the chosen action worked. A symptom check should help the responder recognize when this is a different failure; an action step should explain what permits it and what stops it. Those branches matter as much as the command because they prevent a familiar procedure from being applied to an unfamiliar state.
Local terminology can help readers find the right document. Some teams use “playbook” for broader incident coordination and “runbook” for a specific operational procedure. Agree on that convention locally rather than treating the words as a universal distinction.
Is the application pool full, or is the database saturated?
In a hypothetical database-connection incident, the runbook recommends a restart, but the qualified responder cannot determine whether a migration forbids it. Before choosing another command, the procedure needs to establish both the source of the pressure and the compatibility information required for a permitted intervention.
Begin by comparing application pool occupancy and wait time with database connection counts, active queries and resource pressure. The application pool can exhaust its available connections while the database still has capacity. Conversely, several individually healthy-looking pools can together demand more connections than the database can support. Both situations can produce waiting requests while calling for different changes.
Recent application or configuration changes, transaction duration and application-instance count help distinguish those two kinds of pressure. A pool with no idle connections does not prove a leak: legitimate load and slow transactions can keep connections occupied. Investigate what is holding them and whether the number of application instances has increased total demand on the database before selecting a mitigation.
Adding capacity on the wrong side can intensify the pressure. Extra application replicas can open more connections against an already saturated database. Raising the connection limit can consume additional resources without improving useful throughput. The pool and database observations therefore matter to the choice of capacity change, not just to describing the symptom.
A sizing check makes that interaction concrete. In an illustrative configuration with one independent, direct-to-database pool per application instance, a twenty-connection maximum permits up to eighty connections across four instances and 160 across eight. These are configured upper bounds, not observed usage; other clients and reserved database slots still matter. HikariCP’s pool-size reference explains the local limit, while PostgreSQL’s connection settings describe the server-wide constraint. If a proxy uses transaction pooling, client connections need not map one-to-one to backend sessions. The runbook then needs the proxy’s backend counts as well as the application’s pool counts before a replica increase can be assessed.
Read diagram description
Illustrative diagnosis: compare application pool waits with database state before choosing a mitigation. Adding application replicas or connections can worsen a saturated database. Missing evidence needs an authorized fallback. Diagram labels: Connection-acquisition waits rise: Confirm customer impact and the affected application; Database has capacity: Investigate held connections and transaction behavior; Database is saturated: Investigate demand, queries and resource pressure; Bounded, evidence-based mitigation: Check compatibility; verify waits, transactions and DB health.
How the observations change the mitigation
If the behavior began with a rollout, inspect its approved rollback path and data compatibility. If excessive demand is the immediate pressure, an already tested traffic or concurrency limit may be appropriate. If a particular query needs cancellation, the database owner should help establish its transaction and retry consequences. Killing a query and closing a session are different operations.
Whichever path is selected, record the target, before-state, expected effect and condition that stops further changes. Observe the first result before overlapping interventions make its effect hard to interpret. If the diagnosis remains uncertain, hand over pool waits, database pressure, rollout timing and actions already attempted so the next responder can continue the reasoning.
A missing-evidence branch is especially important. When a dashboard fails or the database owner is unreachable, the runbook should name the authorized fallback and the evidence to preserve. The loss of diagnostic access should not quietly turn a conditional restart into “restart and see.” The same dependency may affect both the customer operation and the tools used to inspect it.
An unfamiliar failure may fall outside these branches. The runbook can still help by identifying the point for escalation and the observations collected so far, leaving the responder free to investigate rather than forcing the incident into the restart procedure.
Recovery includes successful customer transactions
In this case, recovery requires falling connection waits, resumed successful customer transactions and acceptable database pressure through the agreed observation window. Fewer connections alone might mean the application stopped serving traffic. Observing the user operation alongside the resource distinguishes relief from disappearance of demand.
Retain unresolved risks with their owners, and revise any link, permission, branch or expected result that proved wrong. Keep the last-edit date separate from the last-exercised date. A fresh wording change says something about the document's maintenance, not whether someone can use its recovery path.
A rehearsal with another responder
Have someone other than the author use the procedure to identify the next permitted action, its stop condition and the fallback when required evidence is missing. A safe exercise or representative staging setup should give them the access available on a real shift. Any point that requires unwritten knowledge exposes work in the procedure or its supporting access.
AI can help arrange a draft, but environment-specific review still has to establish whether its commands and assumptions are suitable. The procedure becomes a stronger basis for automated remediation only once the preconditions and recovery checks can be executed and tested.
In the database example, the useful rehearsal follows the reasoning from pool waits and database pressure to a suitable mitigation. If migration compatibility remains unknown, the responder has a route to someone who can resolve it and a record that person can use. That is the practical difference between a reusable runbook and a command the original author happens to know.
Put this into practice
Draft a procedure another responder can use
Bring: A specific failure scenario, service scope, access requirements, escalation contacts and recovery checks.
Source context
Source links appear within this article. Read them alongside the author’s analysis and evaluate the guidance against your environment.
Report an error or outdated detail