Commentary

What an exhausted error budget changes

In brief

An error-budget policy connects measured reliability to release and planning choices, including a repair bundled with an unrelated feature.

5 min read

Sources
Open copper shears hover over an uncut bridge joining a teal repair patch and an amber feature patch.
Conceptual mixed-release decision for the hypothetical exhausted budget. Separating the repair can make its assessment possible; the paused shears do not imply approval or execution. An inseparable bundle still needs the authorized exception described in the policy.
On this page4 sections

A 99.9% objective supplies a number; it does not decide what to do with the next release. That becomes apparent when the allowance for failed requests is exhausted and the waiting artifact contains both a repair and an unrelated feature. Holding everything prolongs the problem. Approving everything accepts exposure the repair did not require.

An error-budget policy is the agreement that turns the calculation into a usable choice. It describes which changes continue, which pause and who may authorize an exception as the allowance for unreliability is consumed. Its hardest decisions deserve discussion before an incident makes every minute of that discussion expensive.

The request population determines the allowance

A service-level indicator, or SLI, measures the user outcome, such as the fraction of eligible requests that succeed. The service-level objective, or SLO, sets a target over a defined window. At 99.9% request success, the allowance for failed eligible requests is 0.1%. An internal SLO does not automatically create a contractual service-level agreement, or SLA.

The word “eligible” carries real consequences. Document the population, justified exclusions, data sources and behavior when data is missing. If teams disagree about whether a timeout belongs in the denominator, attaching a release freeze turns that unresolved measurement question into an urgent delivery dispute.

Keep SLI changes reviewable for the same reason. Measurement defects should be corrected, but silently redefining a bad event during an incident would make the chart greener without explaining what happened to users. Retain the old and new interpretations so the release decision can be understood afterward.

How consumption affects planned work

The policy may affect discretionary releases, reliability investment, expansion into new workloads or the review needed for risky changes. Choose consequences that the named owners can carry out. “Prioritize reliability” leaves the delivery team to infer which work should stop and leaves the service team to negotiate for the work it needs.

Remaining budget and burn rate answer related questions: how much allowance is left, and how quickly it is being consumed. Rapid consumption can warrant action before exhaustion. Google's burn-rate alerting chapter explains that relationship and why low-traffic signals need careful interpretation.

The proposed change matters as much as the current balance. A repair may reduce ongoing harm while an ordinary feature introduces additional exposure. A dependency-caused breach also needs a response, even if the service team did not cause it. Attribution belongs in the review, but the immediate decision concerns the experience users are currently getting.

Connect budget consumption to a usable policy
Connect budget consumption to a usable policy
Illustrative policy flow. The SLO defines the allowance; the organization decides what consumption changes. An urgent exception still needs an authorized decision and an expiry.
Read diagram description

Illustrative policy flow. The SLO defines the allowance; the organization decides what consumption changes. An urgent exception still needs an authorized decision and an expiry. Diagram labels: SLO measurement: Eligible events, target and window; Budget consumption: Remaining allowance and burn rate; Agreed policy: Continue, pause or consider a scoped exception; Release decision: Named decision maker, checks and review point.

A repair bundled with an unrelated feature

Consider a hypothetical exhausted budget and a release that bundles a reliability fix with an unrelated feature. Separating the artifact would let the team consider the repair on its own. If that is not feasible, someone with the necessary authority must explicitly decide whether the expected recovery benefit justifies the additional feature's exposure.

The exception should identify exactly what will ship, current customer impact, expected benefit, approver and the checks that will establish the result. Bind it to this change and current situation, with an expiry. A standing exemption for anything described as urgent would avoid the actual question of why the feature should be included.

Leadership accepts risk within its authority; technical owners establish whether execution is feasible and recoverable. An incident lead coordinates the response but does not automatically acquire every approval power. Giving these roles a defined escalation route spares the on-call responder from inventing the division while also trying to restore service.

The same policy should explain who acts when a dependency consumed the budget and the team cannot repair it directly. Depending on the service, the work may involve fallback behavior, demand limits or a changed commitment with the provider. Leaving that case unassigned turns cross-service failure into a debate about fault while customer impact continues.

Different services can need different policies. Consistent measurement and explicit exceptions matter more than imposing one release freeze on every workload regardless of user impact.

What held releases and exceptions reveal

Examine held releases, granted exceptions, repeated exhaustion and customer failures the SLI missed. A gate may enforce the written rule perfectly while blocking harmless work and admitting the recurring failure mechanism. Those outcomes are evidence for revising the agreement rather than celebrating consistent enforcement.

The error-budget policy template records the agreement, the release-gate guide covers implementation and the error-budget calculation reference keeps the arithmetic explicit. Using them together should make it possible to reconstruct why a release proceeded or stopped without relying on someone's memory of the incident.

The mixed release is a useful case for the delivery and recovery owners to discuss while there is time to separate the artifact. Their policy can identify a repair-only path or the person who may accept the bundled feature’s additional exposure, together with verification and expiry. During an incident, the pipeline and responders then have an agreement they can use rather than a percentage they must negotiate around.

Put this into practice

Review the allowance behind your SLO

Bring: Your SLO, measurement window, eligible population and bad outcomes, using one consistent event-based or time-based definition.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

Related reading

Explore a related question