On this page5 sections
SRE key performance indicators, or KPIs, are most useful when service outcomes, response delays and operating workload remain distinguishable. They answer related questions, but a single combined score would conceal which part of the operating model needs attention. The goal is a small set of measurements that can support an investment decision.
Consider a hypothetical quarter in which incident volume stays the same and recovery becomes faster as every difficult case goes to the same specialist. The service result improves while that person’s engineering work stops. A reliability scorecard needs to show both effects before the organization decides whether to change staffing, spread the capability or fund work on the recurring failure.
Four questions for the scorecard
| Question | Candidate indicator | Context to retain |
|---|---|---|
| Are users receiving acceptable service? | SLI performance and SLO attainment | Eligible population, target, window, exclusions, and missing data. |
| How quickly is reliability allowance being used? | Error-budget burn rate | Observation period, request volume, and persistent versus brief changes. |
| Where does response lose time? | Detection, engagement, mitigation, and recovery durations | Timestamp definitions, incident population, outliers, and unknown start times. |
| What does operating the service cost people? | Actionable pages and active on-call effort | Night interruptions, repeat pages, specialist support, and recovery time. |
The boundaries in the table are part of each indicator's meaning. A success rate without its eligible population, or a recovery duration without its endpoint, makes comparison harder even when the underlying calculation is correct. Keep that context available in the scorecard rather than making each reviewer reconstruct it.
Deployment frequency and change failure rate can add delivery context. Read them alongside customer-facing indicators: more delivery activity does not by itself establish that users received an acceptable service. The useful comparison connects what the team changed with what customers experienced and what it cost to operate.
Read diagram description
A scorecard map, not a composite score. Service outcomes, response performance and workload describe different aspects of reliability and should remain separately interpretable. Diagram labels: Customer outcome: SLIs and service objectives; Response performance: Defined detection and recovery intervals; Operating workload: Interruptions, follow-up and specialist demand.
A faster average can conceal a changed incident mix
A mean can improve because the mix of incidents changed rather than because the response became better. Preserve incident counts and distributions, separate severities where useful and inspect long cases. Define “MTTR” locally as well: repair, recovery and resolution may end at different moments.
For an illustrative comparison using the same recovery endpoints, suppose eight minor incidents each take ten minutes and two major incidents each take sixty. The ten-incident mean is twenty minutes. In the next period, nine minor incidents and one major incident take the same respective times, producing a fifteen-minute mean. Neither class recovered faster; only their proportions changed. Reporting each class’s count and duration lets the scorecard show that difference before the lower overall mean becomes evidence for a response-process or staffing claim.
Mean time to detect, or MTTD, requires the same care with its starting point. When impact start is unknown, retain the uncertainty instead of substituting a convenient timestamp silently. A limited measurement remains interpretable when the limit is visible; an apparently precise replacement can change what the metric claims.
If specialist escalation shortens recovery, the scorecard should also show the engineering work it displaces. Together, those observations support a staffing or planning choice: change the specialist’s commitments, spread the diagnostic capability or fund work on the recurring failure. The response improvement is valuable, but its continuing cost belongs in the allocation decision.
One memorable incident does not overturn representative data on its own. Use it to test the scope of the claim and look for missing populations, then examine the wider evidence. Context should improve disciplined measurement rather than replace it with the most compelling story.
When classification or suppression moves the target
Reclassifying incidents can lower their count. Suppressing alerts can lower page volume. For each target, retain evidence about the harm those changes might conceal, such as customer-discovered failures or missed coverage. This helps distinguish a genuine improvement from movement outside the measurement.
Query, denominator and coverage changes also belong beside the trend. They may make adjacent periods incomparable even when both values are calculated correctly. Explaining the change first prevents an apparent regression or improvement from prompting the wrong intervention.
Give each row a decision and someone responsible for making it. An increase in customer-reported failures might lead the service owner to investigate a monitoring gap and bring back findings at the next review. The underlying counts and definitions remain available, while the scorecard records what the trend caused the team to do. The feedback-loop guide develops that link between a claimed improvement and its check.
Request success, budget use and detection time
The formulas below describe separate service and response indicators. In the faster-recovery quarter, those results would sit beside the record of interrupted engineering work before anyone used them to reduce staffing or claim greater efficiency. The calculations remain reproducible without turning their favorable values into a finding about workload.
| Indicator | Calculation | Example and decision |
|---|---|---|
| Request success rate | Good eligible requests ÷ all eligible requests × 100 | 999,500 successful requests out of 1,000,000 gives 99.95%. Compare it with the agreed service target. |
| Error-budget consumption | Bad eligible requests ÷ allowed bad requests × 100 | A 99.9% SLO allows 1,000 bad requests in that population. 500 bad requests consume 50% of the budget. |
| Error-budget burn rate | Observed bad-request fraction ÷ allowed bad-request fraction | A 0.5% error rate against a 0.1% allowance is 5× burn for that observation window. |
| Mean time to detect | Sum of measured detection durations ÷ incidents with valid measurements | 4, 32, and 11 minutes give 15.7 minutes. Investigate the 32-minute case and report excluded incidents. |
Keep the same population and observation window within each calculation. Request errors cannot be converted directly to downtime minutes. A latency SLI also needs its threshold, since a technically successful request can still be too slow to qualify as good service.
The guide to defining service level objectives helps select the outcome being measured, and the error-budget worksheet records the allowance and policy. Google's SLO implementation chapter explains the relationship among indicator, objective and budget. Together these definitions let another reviewer reproduce the number before debating what it means for the plan.
What faster recovery means for the specialist’s work
The faster-recovery quarter presents two findings the team should be able to examine together: shorter incident durations and more specialist work. The on-call workload record supplies the second view, including off-shift help and follow-up that page counts miss. With comparable incident populations and timing definitions, the review can distinguish a better response process from one that has become more dependent on a single engineer.
In the faster-recovery quarter, shared diagnostic knowledge or broader appropriate access may be a better investment than a more aggressive response-time target. The scorecard has become useful when it lets the team make that distinction, assign the change and check whether the improvement can be sustained without quietly abandoning other commitments.
Put this into practice
Inspect your detection delays
Bring: Incident onset and discovery timestamps, a service and reporting window, and explicit exclusions.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- Google SRE: Implementing SLOssre.google
