Guide

Error-budget template: calculations and release-policy fields

In brief

Worked request and time-based budgets accompany an SLO record, release-policy template and exception log, keeping each calculation in its original units.

5 min read

Sources
A filled calculation panel leads to a large empty owner field, with a wrench repair tile waiting beyond it.
Conceptual unfinished error-budget worksheet: the hypothetical calculation is valid, but the exception-owner field is blank and the security repair waits. Counting shapes are not numerical data; naming an owner would still require a usable escalation route.
On this page4 sections

An error-budget worksheet needs to answer two questions: how much unreliability is still allowed, and what that answer changes. The first is arithmetic. The second is a release decision involving a service, a proposed change and people who can act. This template puts both on the same page so another person can reproduce the number and follow its consequence.

Begin with one user operation rather than filling in a percentage for the whole system. An error budget is the bad-outcome allowance implied by a service-level objective, or SLO. Defining that outcome establishes the units for the calculation and keeps the later release policy attached to the experience it is meant to protect.

The SLO measurement record

FieldFill in
Service and user operationNamed service, environment, and operation covered
Accountable ownerTeam, escalation route, and business stakeholder
Eligible eventsExact population and justified exclusions
Good eventSuccess condition, including any latency threshold
Objective and windowTarget percentage; rolling or calendar period; time zone
MeasurementSource, query, freshness limit, and missing-data behavior
PolicyDecision thresholds, approvers, exceptions, and review date

The service-level indicator, or SLI, measures the outcome. For a latency objective, a good event might be an eligible request completed within 300 milliseconds. A target of 95% leaves an allowance of 5% bad requests. That is a request allowance, not permission for complete unavailability during 5% of a month.

The population and exclusions determine which requests enter the calculation. The query and freshness fields establish where the evidence comes from and whether it describes the relevant window. Filling them in allows two readers to reproduce the same result rather than merely agreeing that the displayed percentage looks precise.

A template cannot settle a disputed product objective. If owners disagree about which user experience matters, record that decision gap rather than disguising it with a precise-looking target.

Worked calculations in requests and minutes

For a request-based SLI, multiply the number of eligible events by one minus the target expressed as a fraction. That yields the allowed bad-event count. Subtract observed bad events to find the remaining budget. A negative remainder means the allowance has been exceeded and the SLO was not met over that window.

Illustrative objectivePopulationAllowed bad eventsObservedRemaining
99.9% request success1,000,000 requests1,000 requests650 bad requests350 requests
95% within 300 ms100,000 requests5,000 requests3,200 slow requests1,800 requests

In the first worked example, one million eligible requests at 99.9% success allow 1,000 bad requests. Observing 650 leaves 350. The latency example uses the same method with a different definition of bad: requests slower than the threshold. Keeping the population, allowance, observation and remainder visible makes both results auditable.

A time-based availability SLI uses time instead. Over exactly 30 days, 99.9% permits 43.2 minutes of bad time, 99.95% permits 21.6 minutes and 99.99% permits 4.32 minutes. Calendar months of different lengths change those allowances. Partial request failures cannot be converted directly into outage minutes without changing what the SLI measures.

Incomplete collection adds a separate uncertainty that arithmetic cannot remove. Record the missing interval and the numerator or denominator values it leaves unknown. A correct sum over received records may still fail to represent the intended request population, which is why the operating policy needs a specific outcome for missing or stale telemetry.

Keep the allowance in request units
Keep the allowance in request units. Worked example: one million eligible requests at a 99.9% success target allow 1,000 bad requests. With 650 observed bad requests, 350 remain. Missing collection coverage can invalidate the decision even when the arithmetic is correct.
Worked example: one million eligible requests at a 99.9% success target allow 1,000 bad requests. With 650 observed bad requests, 350 remain. Missing collection coverage can invalidate the decision even when the arithmetic is correct.
Read diagram description

Worked example: one million eligible requests at a 99.9% success target allow 1,000 bad requests. With 650 observed bad requests, 350 remain. Missing collection coverage can invalidate the decision even when the arithmetic is correct.

Release-policy fields

Policy owner: [team and accountable decision maker]
Applies to: [services and release paths]
Normal state: [conditions and permitted changes]
Review state: [burn/remaining-budget conditions and required review]
Hold state: [conditions and changes paused]
Missing or stale telemetry: [hold, review, or other explicit behavior]
Permitted exceptions: [bounded categories and approver]
Resumption: [health evidence, observation window, and approval]
Policy review date: [date and triggers for earlier review]

The normal, review and hold states describe what people may do with a given result. The missing-telemetry field prevents an unavailable measurement from being silently treated as health. The exception and resumption fields explain how the team can admit a necessary repair and later restore ordinary delivery.

Choose thresholds using workload and incident evidence. Google’s SLO alerting guidance explains how several observation windows distinguish different rates of budget consumption, with particular care needed at low request volume. The template deliberately leaves these choices open because a number copied from another service would still need justification for this one.

Consider a hypothetical worksheet with a valid exhausted-budget calculation but no exception owner. A security repair reaches the gate, and the team must negotiate authority while the release waits. The calculation has worked; the decision path has failed at the empty field.

The disputed repair and missing-telemetry case test fields the clean arithmetic example cannot exercise. Running them through the worksheet before it controls releases reveals whether the named owners can reach and carry out the decision. A missing route limits which releases the gate can govern until the agreement is complete.

The exception record and its outcome

An exception record should retain the change identifier, reason, scope, budget evidence, expected benefit, approver, expiry and post-change outcome. These fields connect the choice made during the hold to what happened afterward. If follow-up remains, name the owner and the evidence that would demonstrate reduced risk instead of assigning an unsupported percentage improvement.

Before enabling enforcement, reproduce the worked calculation, test a no-traffic interval and walk through a mitigation requiring an exception. Check that the intended owners can be reached, rather than treating an entered team name as proof of a working escalation path. The release-gate implementation guide covers the enforcement and failure-handling details once that agreement is ready.

A completed worksheet connects the exhausted allowance to a hold, an authorized exception route and the required recovery observations. In the security-repair case, a blank approver field identifies unfinished agreement, not an arithmetic problem. Resolving that field with the relevant owners makes the template usable by the release pipeline and the responder who needs the repair.

Put this into practice

Review the allowance behind your SLO

Bring: Your SLO, measurement window, eligible population and bad outcomes, using one consistent event-based or time-based definition.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

Related reading

Explore a related question