Explainer

SRE KPIs for service quality and the work behind it

In brief

Worked formulas and a scorecard connect customer outcomes, response delays and specialist workload to reliability investment.

6 min read

Sources
Several incident tabs converge on one bobbin while an output ribbon emerges beside an unfinished engineering cloth.
Conceptual illustration of the hypothetical recovery quarter: concentrated specialist effort can improve response while displacing engineering work. The scene depicts that tradeoff without measuring recovery time, staffing efficiency or anyone’s health.
On this page5 sections

SRE key performance indicators, or KPIs, are most useful when service outcomes, response delays and operating workload remain distinguishable. They answer related questions, but a single combined score would conceal which part of the operating model needs attention. The goal is a small set of measurements that can support an investment decision.

Consider a hypothetical quarter in which incident volume stays the same and recovery becomes faster as every difficult case goes to the same specialist. The service result improves while that person’s engineering work stops. A reliability scorecard needs to show both effects before the organization decides whether to change staffing, spread the capability or fund work on the recurring failure.

Four questions for the scorecard

QuestionCandidate indicatorContext to retain
Are users receiving acceptable service?SLI performance and SLO attainmentEligible population, target, window, exclusions, and missing data.
How quickly is reliability allowance being used?Error-budget burn rateObservation period, request volume, and persistent versus brief changes.
Where does response lose time?Detection, engagement, mitigation, and recovery durationsTimestamp definitions, incident population, outliers, and unknown start times.
What does operating the service cost people?Actionable pages and active on-call effortNight interruptions, repeat pages, specialist support, and recovery time.

The boundaries in the table are part of each indicator's meaning. A success rate without its eligible population, or a recovery duration without its endpoint, makes comparison harder even when the underlying calculation is correct. Keep that context available in the scorecard rather than making each reviewer reconstruct it.

Deployment frequency and change failure rate can add delivery context. Read them alongside customer-facing indicators: more delivery activity does not by itself establish that users received an acceptable service. The useful comparison connects what the team changed with what customers experienced and what it cost to operate.

Keep three kinds of reliability evidence visible
Keep three kinds of reliability evidence visible
A scorecard map, not a composite score. Service outcomes, response performance and workload describe different aspects of reliability and should remain separately interpretable.
Read diagram description

A scorecard map, not a composite score. Service outcomes, response performance and workload describe different aspects of reliability and should remain separately interpretable. Diagram labels: Customer outcome: SLIs and service objectives; Response performance: Defined detection and recovery intervals; Operating workload: Interruptions, follow-up and specialist demand.

A faster average can conceal a changed incident mix

A mean can improve because the mix of incidents changed rather than because the response became better. Preserve incident counts and distributions, separate severities where useful and inspect long cases. Define “MTTR” locally as well: repair, recovery and resolution may end at different moments.

For an illustrative comparison using the same recovery endpoints, suppose eight minor incidents each take ten minutes and two major incidents each take sixty. The ten-incident mean is twenty minutes. In the next period, nine minor incidents and one major incident take the same respective times, producing a fifteen-minute mean. Neither class recovered faster; only their proportions changed. Reporting each class’s count and duration lets the scorecard show that difference before the lower overall mean becomes evidence for a response-process or staffing claim.

Mean time to detect, or MTTD, requires the same care with its starting point. When impact start is unknown, retain the uncertainty instead of substituting a convenient timestamp silently. A limited measurement remains interpretable when the limit is visible; an apparently precise replacement can change what the metric claims.

If specialist escalation shortens recovery, the scorecard should also show the engineering work it displaces. Together, those observations support a staffing or planning choice: change the specialist’s commitments, spread the diagnostic capability or fund work on the recurring failure. The response improvement is valuable, but its continuing cost belongs in the allocation decision.

One memorable incident does not overturn representative data on its own. Use it to test the scope of the claim and look for missing populations, then examine the wider evidence. Context should improve disciplined measurement rather than replace it with the most compelling story.

When classification or suppression moves the target

Reclassifying incidents can lower their count. Suppressing alerts can lower page volume. For each target, retain evidence about the harm those changes might conceal, such as customer-discovered failures or missed coverage. This helps distinguish a genuine improvement from movement outside the measurement.

Query, denominator and coverage changes also belong beside the trend. They may make adjacent periods incomparable even when both values are calculated correctly. Explaining the change first prevents an apparent regression or improvement from prompting the wrong intervention.

Give each row a decision and someone responsible for making it. An increase in customer-reported failures might lead the service owner to investigate a monitoring gap and bring back findings at the next review. The underlying counts and definitions remain available, while the scorecard records what the trend caused the team to do. The feedback-loop guide develops that link between a claimed improvement and its check.

Request success, budget use and detection time

The formulas below describe separate service and response indicators. In the faster-recovery quarter, those results would sit beside the record of interrupted engineering work before anyone used them to reduce staffing or claim greater efficiency. The calculations remain reproducible without turning their favorable values into a finding about workload.

IndicatorCalculationExample and decision
Request success rateGood eligible requests ÷ all eligible requests × 100999,500 successful requests out of 1,000,000 gives 99.95%. Compare it with the agreed service target.
Error-budget consumptionBad eligible requests ÷ allowed bad requests × 100A 99.9% SLO allows 1,000 bad requests in that population. 500 bad requests consume 50% of the budget.
Error-budget burn rateObserved bad-request fraction ÷ allowed bad-request fractionA 0.5% error rate against a 0.1% allowance is 5× burn for that observation window.
Mean time to detectSum of measured detection durations ÷ incidents with valid measurements4, 32, and 11 minutes give 15.7 minutes. Investigate the 32-minute case and report excluded incidents.

Keep the same population and observation window within each calculation. Request errors cannot be converted directly to downtime minutes. A latency SLI also needs its threshold, since a technically successful request can still be too slow to qualify as good service.

The guide to defining service level objectives helps select the outcome being measured, and the error-budget worksheet records the allowance and policy. Google's SLO implementation chapter explains the relationship among indicator, objective and budget. Together these definitions let another reviewer reproduce the number before debating what it means for the plan.

What faster recovery means for the specialist’s work

The faster-recovery quarter presents two findings the team should be able to examine together: shorter incident durations and more specialist work. The on-call workload record supplies the second view, including off-shift help and follow-up that page counts miss. With comparable incident populations and timing definitions, the review can distinguish a better response process from one that has become more dependent on a single engineer.

In the faster-recovery quarter, shared diagnostic knowledge or broader appropriate access may be a better investment than a more aggressive response-time target. The scorecard has become useful when it lets the team make that distinction, assign the change and check whether the improvement can be sustained without quietly abandoning other commitments.

Put this into practice

Inspect your detection delays

Bring: Incident onset and discovery timestamps, a service and reporting window, and explicit exclusions.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

Related reading

Explore a related question