On this page4 sections
An error-budget worksheet needs to answer two questions: how much unreliability is still allowed, and what that answer changes. The first is arithmetic. The second is a release decision involving a service, a proposed change and people who can act. This template puts both on the same page so another person can reproduce the number and follow its consequence.
Begin with one user operation rather than filling in a percentage for the whole system. An error budget is the bad-outcome allowance implied by a service-level objective, or SLO. Defining that outcome establishes the units for the calculation and keeps the later release policy attached to the experience it is meant to protect.
The SLO measurement record
| Field | Fill in |
|---|---|
| Service and user operation | Named service, environment, and operation covered |
| Accountable owner | Team, escalation route, and business stakeholder |
| Eligible events | Exact population and justified exclusions |
| Good event | Success condition, including any latency threshold |
| Objective and window | Target percentage; rolling or calendar period; time zone |
| Measurement | Source, query, freshness limit, and missing-data behavior |
| Policy | Decision thresholds, approvers, exceptions, and review date |
The service-level indicator, or SLI, measures the outcome. For a latency objective, a good event might be an eligible request completed within 300 milliseconds. A target of 95% leaves an allowance of 5% bad requests. That is a request allowance, not permission for complete unavailability during 5% of a month.
The population and exclusions determine which requests enter the calculation. The query and freshness fields establish where the evidence comes from and whether it describes the relevant window. Filling them in allows two readers to reproduce the same result rather than merely agreeing that the displayed percentage looks precise.
A template cannot settle a disputed product objective. If owners disagree about which user experience matters, record that decision gap rather than disguising it with a precise-looking target.
Worked calculations in requests and minutes
For a request-based SLI, multiply the number of eligible events by one minus the target expressed as a fraction. That yields the allowed bad-event count. Subtract observed bad events to find the remaining budget. A negative remainder means the allowance has been exceeded and the SLO was not met over that window.
| Illustrative objective | Population | Allowed bad events | Observed | Remaining |
|---|---|---|---|---|
| 99.9% request success | 1,000,000 requests | 1,000 requests | 650 bad requests | 350 requests |
| 95% within 300 ms | 100,000 requests | 5,000 requests | 3,200 slow requests | 1,800 requests |
In the first worked example, one million eligible requests at 99.9% success allow 1,000 bad requests. Observing 650 leaves 350. The latency example uses the same method with a different definition of bad: requests slower than the threshold. Keeping the population, allowance, observation and remainder visible makes both results auditable.
A time-based availability SLI uses time instead. Over exactly 30 days, 99.9% permits 43.2 minutes of bad time, 99.95% permits 21.6 minutes and 99.99% permits 4.32 minutes. Calendar months of different lengths change those allowances. Partial request failures cannot be converted directly into outage minutes without changing what the SLI measures.
Incomplete collection adds a separate uncertainty that arithmetic cannot remove. Record the missing interval and the numerator or denominator values it leaves unknown. A correct sum over received records may still fail to represent the intended request population, which is why the operating policy needs a specific outcome for missing or stale telemetry.
Read diagram description
Worked example: one million eligible requests at a 99.9% success target allow 1,000 bad requests. With 650 observed bad requests, 350 remain. Missing collection coverage can invalidate the decision even when the arithmetic is correct.
Release-policy fields
Policy owner: [team and accountable decision maker]
Applies to: [services and release paths]
Normal state: [conditions and permitted changes]
Review state: [burn/remaining-budget conditions and required review]
Hold state: [conditions and changes paused]
Missing or stale telemetry: [hold, review, or other explicit behavior]
Permitted exceptions: [bounded categories and approver]
Resumption: [health evidence, observation window, and approval]
Policy review date: [date and triggers for earlier review]
The normal, review and hold states describe what people may do with a given result. The missing-telemetry field prevents an unavailable measurement from being silently treated as health. The exception and resumption fields explain how the team can admit a necessary repair and later restore ordinary delivery.
Choose thresholds using workload and incident evidence. Google’s SLO alerting guidance explains how several observation windows distinguish different rates of budget consumption, with particular care needed at low request volume. The template deliberately leaves these choices open because a number copied from another service would still need justification for this one.
Consider a hypothetical worksheet with a valid exhausted-budget calculation but no exception owner. A security repair reaches the gate, and the team must negotiate authority while the release waits. The calculation has worked; the decision path has failed at the empty field.
The disputed repair and missing-telemetry case test fields the clean arithmetic example cannot exercise. Running them through the worksheet before it controls releases reveals whether the named owners can reach and carry out the decision. A missing route limits which releases the gate can govern until the agreement is complete.
The exception record and its outcome
An exception record should retain the change identifier, reason, scope, budget evidence, expected benefit, approver, expiry and post-change outcome. These fields connect the choice made during the hold to what happened afterward. If follow-up remains, name the owner and the evidence that would demonstrate reduced risk instead of assigning an unsupported percentage improvement.
Before enabling enforcement, reproduce the worked calculation, test a no-traffic interval and walk through a mitigation requiring an exception. Check that the intended owners can be reached, rather than treating an entered team name as proof of a working escalation path. The release-gate implementation guide covers the enforcement and failure-handling details once that agreement is ready.
A completed worksheet connects the exhausted allowance to a hold, an authorized exception route and the required recovery observations. In the security-repair case, a blank approver field identifies unfinished agreement, not an arithmetic problem. Resolving that field with the relevant owners makes the template usable by the release pipeline and the responder who needs the repair.
Put this into practice
Review the allowance behind your SLO
Bring: Your SLO, measurement window, eligible population and bad outcomes, using one consistent event-based or time-based definition.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- Google’s SLO alerting guidancesre.google
