On this page4 sections
Pausing feature releases after an error budget is exhausted sounds straightforward until the paused pipeline contains the repair. Now the team has two risks to compare: leave the service in its current condition, or introduce a change that might restore it and might create another problem. A release gate is useful only if it can carry that distinction into the actual deployment decision.
The gate enforces part of an error-budget policy. It can read reliability evidence and stop a pipeline; it cannot supply the missing agreement about who may approve an exception, which change the exception covers or what allows ordinary releases to resume. Those choices determine whether a pause protects the service or merely sends the deployment around a different route.
Remaining budget and burn rate
The remaining error budget is the permitted unreliability left in the service-level objective's window. Burn rate expresses how quickly that allowance is being consumed relative to the objective. A service may still have budget available while spending it fast enough to warrant intervention, just as a positive balance does not make every rate of spending sustainable.
The measurement windows change the decision. A short interval can react to a burst that soon passes, while a long one can hide deterioration until substantial budget has gone. Traffic volume and the consequence of delayed action help determine useful windows and thresholds. Historical incidents then let the team ask whether those choices would have held a risky release, interrupted harmless work or reacted too late.
Google’s example error-budget policy shows how reliability performance can govern release behavior. Its conditions are useful as an example of a working agreement, rather than thresholds to copy without considering the service.
Missing telemetry cannot establish service health
The gate needs the relevant SLO, current incident state, telemetry freshness and the class of proposed change, together with the policy version used to decide. If telemetry is unavailable, a blank error count cannot establish healthy service. The policy must say whether missing or stale inputs produce a hold or require review, so an observability failure does not quietly become permission to deploy.
Consider a hypothetical paused release stream containing a repair that also changes a database schema. The incident lead wants the recovery benefit immediately, but the data owner cannot yet establish compatibility. The urgent question is whether this exact artifact can be deployed safely enough under these conditions, while the team also investigates reversible containment.
The distinction between that artifact and an ordinary feature matters. A blanket “reliability improvement” exception could admit a worthwhile refactor that does little for the current incident. The schema repair needs its own compatibility evidence, expected benefit and verification plan. Its urgency explains why the decision deserves attention; it does not answer the compatibility question.
An override needs to bind the decision to this artifact, its consequences, the authorized owner and a condition for revisiting it. If the team authorizes the repair under the emergency policy after reviewing the compatibility uncertainty, only that change proceeds; unrelated releases remain paused.
Check the evidence near execution, too. A release may wait long enough for its approval or assumptions to become stale. Direct deployment paths and emergency tools need the same treatment as the familiar pipeline, otherwise the most consequential exception can bypass the policy that was meant to govern it.
Read diagram description
Illustrative policy structure, not prescribed thresholds. Use fresh evidence and a current change classification; recheck near execution. Any exception needs an exact scope, authorized approver and verification plan. Diagram labels: Current release evidence: SLO, burn, incident state, freshness and change class; Known + within policy: Proceed through normal checks; Stale, missing or outside limits: Hold or request review per policy; Bounded exception, if permitted: Record exact change, approver, expiry and verifier.
An exception for one repair
A useful exception records the service, exact change, reason for proceeding, expected benefit, exposure, verifier and expiry. These details let another responder distinguish an approved repair from an unrelated release and understand which observations would invalidate the approval. They also prevent an emergency decision from becoming a permanent opening in the gate.
Some exceptions involve a commercial deadline as well as service health. Name the person authorized to accept that remaining tradeoff and the person who can refuse execution when technical preconditions fail. In the schema example, this gives the incident lead a route to a decision while preserving the data owner's ability to say what is not yet known.
A rigid gate can obstruct a necessary repair. Retain an emergency route with independent verification, and review its use without pretending that all risk can be settled before a live incident.
The conditions for resuming ordinary releases
A pause also needs an exit. Specify the service-health interval, required remediation and any approval needed to resume. Without that agreement, the team can repair the original problem and still disagree indefinitely about whether it has done enough.
One healthy sample may be too weak to reverse a hold designed to catch sustained deterioration. Define how much recovery evidence is needed and when the decision is reevaluated. Different entry and exit conditions can prevent repeated stopping and starting near a threshold, provided the gap reflects the service's behavior and is visible to the release owner.
The release-policy template provides fields for turning these choices into a policy. Rehearse the schema-changing repair alongside a held feature release, unavailable telemetry and a stale approval. The repair should proceed only through its scoped exception, with unrelated work still held. The recovery observations then determine when ordinary releases resume, so the emergency route does not outlast the incident it was created for.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- Google’s example error-budget policysre.google
