On this page3 sections
A reliability priority needs a place in the delivery plan. If incidents and recovery consume hours already promised to product work, approving an improvement without changing anything else leaves the same people committed twice. The SRE leader's useful contribution is to resolve that collision: what receives capacity, what moves and when a temporary compromise returns for review.
This responsibility appears at different stages of service work. Before an incident, it shapes the agreements people will use under pressure. During one, it removes obstacles to coordinated response. Afterward, it determines whether the lessons have enough time and authority behind them to change the service.
An error-budget policy is easier to use before a disputed launch
Service objectives help when product and engineering leaders understand what follows from missing them. An error-budget policy is easier to use when its consequences for releases have been accepted in advance. Introducing it during a contested launch adds a policy negotiation to an already urgent decision.
Exceptions will sometimes be reasonable. Record who accepted the risk, why, for how long and which protections remain. This preserves the context for responders and the next decision-maker. A private agreement to bypass the policy may get one launch moving while leaving everyone else unsure which rules still apply.
The agreement should also explain when the exception ends. A temporary accommodation becomes a different commitment if its exit condition repeatedly fails. Bringing that result back to the responsible leader keeps an expired compromise from becoming the team's unspoken operating model.
Read diagram description
Illustrative planning choices. Incident work and recovery consume capacity; leadership must decide which commitments to fund, defer or narrow rather than leaving the gap with responders. Diagram labels: Available team capacity: Account for response and recovery work; Fund support: Add usable capacity; Adjust commitments: Defer or narrow planned work; Revised plan: Named owners and a review point.
During the incident, access and coordination matter
Size the response around the incident: an incident lead, reachable service expertise and a communication owner appropriate to the work. Leadership can supply missing capacity, resolve an access problem or clear an organizational barrier. Those interventions help the people already coordinating the response continue their work.
Ask for the current impact, working hypothesis, action underway and next update. That information explains the present decision without requiring a premature root-cause declaration. As evidence changes, the shared account needs room to change with it. A confident explanation produced too early can make uncertainty harder to discuss just when it matters most.
Additional senior participants should have a clear purpose within that coordination path. Removing a blocked permission or finding the right specialist is concrete help. Creating another competing channel for decisions makes it harder to maintain a common picture of the incident.
After the review, the delivery plan has to change
A post-incident review needs to describe confusing interfaces, conflicting incentives and missing permissions without becoming a search for a culprit. Leaders establish that expectation through what they ask and the changes they authorize. The discussion becomes credible when a finding can lead to work on the condition that made the decision difficult.
Consider a hypothetical team that repeatedly postpones preventive engineering to respond to incidents while every product deadline stays intact. Endorsing the reliability goal leaves the collision unresolved. The next planning discussion needs the specific improvement, the hours it requires and the delivery work competing for them. Leadership can then fund the improvement, move a commitment or explicitly accept a delay. If the accepted delay outlives its exit condition, it returns for review rather than becoming an indefinite promise to fix things later.
Customer outcomes, repeated failure modes and responder workload help identify which choice is needed. Continuous exceptional effort keeping a service afloat describes a different operating problem from an isolated defect. Looking at both the service and the people carrying it prevents apparent uptime from concealing a plan that depends on repeated overload.
Finite resources mean some reliability requests should be declined or deferred. Make that tradeoff explicit rather than leaving an unfunded commitment for an engineer to fail personally. For the hypothetical team, a useful meeting ends with capacity for a named improvement or an owned deferral with a review point. The engineers can then plan around the decision the organization actually made.
Source context
This article does not include external reference links. Read it as the author’s perspective and evaluate the guidance against your environment.
Report an error or outdated detail