Explainer

Error-budget calculations, balances and burn rates

In brief

Worked request and time examples explain error-budget units, remaining allowance and burn rate, plus the release policy that gives them a practical role.

4 min read

Sources
An allowance panel has six filled and four empty recesses beside an uninserted key and separate permission receiver.
Conceptual proportions for the illustrative request budget: 600 of a 1,000-bad-request allowance used leaves 400 in the same window. The separate unjoined permission fitting shows why that balance alone cannot authorize a release or resolve migration recovery risk. No measured budget or exhaustion forecast is depicted.
On this page3 sections

At a 99.9% request-success target, one million eligible requests allow 1,000 bad requests. After 600 bad requests, 400 remain in that window's calculation. The arithmetic is straightforward. Interpreting the 400 as permission for the next release takes an additional step: deciding which risks the number describes and what the agreed policy allows.

The measured behavior is the service-level indicator, or SLI; its target and window form the service-level objective, or SLO. For an objective expressed as a fraction of good events, the allowed bad fraction is one minus the target. An error budget applies that fraction to the eligible population, so its meaning begins with what counts as an event and what makes one bad.

Requests and minutes are different budget units

For the request example, 99.9% success leaves 0.1%, or 0.001, as the permitted bad fraction. Multiplying one million requests by 0.001 yields 1,000 bad requests, and subtracting 600 gives 400. Every number refers to the same eligibility rule and window. The remainder is a request count, not a duration or a forecast.

A time-based objective starts with a different population. Exactly thirty days contain 43,200 minutes; 0.1% of that is 43.2 minutes classified as bad under the indicator's definition. This does not make 43.2 minutes interchangeable with 1,000 failed requests. A minute of peak demand can contain far more requests than a quiet minute.

A budget belongs to its own indicator and window. Unused allowance for one experience cannot be transferred to an unrelated experience or used to dismiss another objective's failure. A positive balance is meaningful evidence about the defined outcome, not a general measure of how much harm the service can afford.

Rolling windows add another reason to retain the calculation time. Older events leave as new ones enter, so the balance can change without a new failure. Put the SLI, window and measurement time beside the report to let another person reproduce it and understand why it moved.

Remaining budget is a count in a defined window
Remaining budget is a count in a defined window. Illustrative request-based objective: 99.9% of one million eligible requests allows 1,000 bad requests. After 600 bad requests, 400 remain. This is not a conversion to downtime minutes or a prediction of time to exhaustion.
Illustrative request-based objective: 99.9% of one million eligible requests allows 1,000 bad requests. After 600 bad requests, 400 remain. This is not a conversion to downtime minutes or a prediction of time to exhaustion.
Read diagram description

Illustrative request-based objective: 99.9% of one million eligible requests allows 1,000 bad requests. After 600 bad requests, 400 remain. This is not a conversion to downtime minutes or a prediction of time to exhaustion.

A 1% error rate means 10× burn against a 0.1% allowance

Burn rate divides the observed bad-event fraction by the allowed bad fraction. If the allowance is 0.1% and the observed error fraction is 1%, the ratio is ten. During that observation period, the service is consuming allowance at ten times the pace compatible with the target.

The ratio does not establish how long that behavior will persist. A brief spike and a sustained regression can produce a similar observed rate while needing different responses. Read the observation window, request volume and persistence with the number. Google’s SLO alerting chapter develops the use of multiple windows to make these alerts useful.

Balance and pace can now inform a reliability decision, but they still do not describe every consequence of a proposed change. Consider a hypothetical migration with uncertain rollback on a service that has budget remaining. The request-success balance cannot establish that a data change is recoverable. Its own preconditions and authorization still need to be satisfied.

The release policy supplies the consequence

The release policy specifies how high consumption or exhaustion affects feature releases and reliability work, including who can grant an exception and why. For the migration, it answers the release question raised by the balance; the separate compatibility assessment answers whether the proposed recovery remains possible. Both results belong in the decision because they describe different consequences of proceeding.

Exhaustion does not require stopping every change indiscriminately. A repair intended to restore reliability has a different purpose from an unrelated feature release. The policy guide and calculation template help make that distinction and its supporting calculation explicit.

For the migration, the team can now ask two properly separated questions: what does the current budget imply under the release policy, and has the migration's recovery risk been addressed by the right authority? A healthy answer to one does not supply the answer to the other.

A budget report is therefore most useful with its indicator, window and calculation time attached. Those details explain the balance and pace; the release policy explains what follows. If the product changes the outcome being measured, the SLO also needs reconsideration so the next precise calculation still describes the experience that matters.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

Related reading

Explore a related question