Explainer

Calculating mean time to detect from incident records

In brief

Worked incident records explain the MTTD formula, uncertain onset times and the effect of including failures first reported by customers.

5 min read

Sources
Two differently sized ribbons run from the bounds of a possible-onset interval to one discovery bead.
Conceptual illustration of the article’s bounded-onset example: the same discovery time can yield a duration range when onset is uncertain. No measured MTTD, clock reading or sample average is shown.
On this page5 sections

Mean time to detect is the sum of measured incident detection durations divided by the number of incidents with usable measurements. The calculation is short. The consequential choices are which failures enter the sum, when each began and what counts as discovery. A lower result can reflect faster detection, or it can reflect the disappearance of difficult incidents from the report.

Keep those choices visible and MTTD becomes a way to investigate monitoring. Separate acknowledgment, investigation, mitigation and recovery intervals so the team can also distinguish a late discovery from slow work after discovery. Each interval needs stated endpoints before it can support a comparison.

Three incidents produce a 15.7-minute mean

The following illustrative record measures onset to discovery for three incidents. It retains the discovery method because a monitoring event, an engineer's observation and a customer report describe different paths by which a failure became known.

IncidentOnsetDiscoveryMethodDuration
A14:0014:04Monitoring4 minutes
B09:1509:47Engineer observation32 minutes
C22:3022:41Customer report11 minutes
MTTD = (4 + 32 + 11) / 3
     = 15.666… minutes
     ≈ 15.7 minutes

The three durations total 47 minutes, producing a mean of approximately 15.7 minutes. The median is 11 minutes and the range is 4–32 minutes. With just three incidents, examining those cases is more informative than presenting a finely calculated percentile. The Atlassian metric reference supplies the conventional detection definition.

The 32-minute case now poses a useful question: which evidence existed before discovery, and could an earlier signal have revealed it? The average helps locate the issue, but the row and timeline are what make improvement possible.

An uncertain start produces a range

Onset should represent the failure condition being measured, using the best available evidence. The first anomalous metric can be a proxy, but an anomaly may precede user impact or be unrelated to it. Name the proxy and its uncertainty rather than calling it an exact start.

Discovery needs a definition too. If an automated event counts when the detector fires, retaining a separate delivery timestamp reveals any delay before the responder receives it. Engineer and customer discovery routes need recorded times as well. Normalize time zones and examine clock skew or delayed ingestion before comparing intervals.

Some incidents have no defensible onset timestamp. Mark that uncertainty or analyze a separate population; a precise guessed value makes comparisons cleaner while making the conclusion less trustworthy.

An onset interval may be more defensible than a point. In an illustrative case, the last confirmed healthy observation is at 10:00, the first confirmed failure at 10:06 and discovery at 10:10. Under those observations, detection took between four and ten minutes. Reporting exactly four selects the latest possible onset, so it should be identified as an estimation choice rather than the only result supported by the record.

Uncertain onset creates a duration range
Uncertain onset creates a duration range. Illustrative observations: last confirmed healthy at 10:00, first confirmed failure at 10:06, discovery at 10:10. Under these observations, detection took between four and ten minutes; reporting only four chooses the latest possible onset.
Illustrative observations: last confirmed healthy at 10:00, first confirmed failure at 10:06, discovery at 10:10. Under these observations, detection took between four and ten minutes; reporting only four chooses the latest possible onset.
Read diagram description

Illustrative observations: last confirmed healthy at 10:00, first confirmed failure at 10:06, discovery at 10:10. Under these observations, detection took between four and ten minutes; reporting only four chooses the latest possible onset.

Keep the interval or disclose the estimation rule. Otherwise a change in observation frequency could alter reported MTTD without an equivalent change in detection. Unknown onset should not become zero; show how many incidents lack usable timestamps and why they are excluded.

The missing durations also raise a population question: which incidents remain visible in the report? A failure first reported by a customer still tells the team something about monitoring, even when its onset cannot be timed precisely.

Customer reports belong in the detection population

Consider a hypothetical month in which monitoring misses two long incidents that customers eventually report. Removing them from the calculation makes automated discovery appear faster while omitting consequential coverage gaps. Keep the discovery routes visible and explain missing timestamps instead of letting the denominator hide those failures.

Adding eligible customer-reported cases may make measured MTTD rise. That increase can reflect more complete reporting of previously invisible failures rather than deteriorating detection. Preserve enough context to distinguish those explanations before interpreting the trend.

Where sample size permits, compare similar severity, service and failure populations. A universal five-minute target says little without evidence that it fits the consequences and available intervention time. A slow capacity warning and a failed critical transaction can leave different windows for useful action.

What the 32-minute case can reveal

Trace one slow-detected incident to the observable signal, when it became available and what earlier discovery would have enabled. A synthetic check might reveal a failed user journey sooner. Better ownership might instead improve action after discovery. The timeline establishes which change addresses the observed constraint.

Anomaly detection may help identify deviations, but it also needs evaluation for false-alert cost and coverage on known failures before replacing a threshold. There is no requirement to optimize detection before recovery. Choose the interval with the strongest evidence of avoidable harm.

The companion article on MTTD and response delay explores the post-discovery side. Used together, the intervals should explain a decision the team can improve rather than compete for the most flattering number.

Where MTTD ends and response timing begins

MTTD measures discovery. Mean time to acknowledge, or MTTA, usually starts at an alert or notification and ends at acknowledgment. MTTR is less consistent: teams use it for repair, recovery, resolution or response. State the endpoints rather than assuming the acronym makes two reports comparable.

In an illustrative incident, impact starts at 10:00, detection occurs at 10:04, acknowledgment at 10:06 and recovery at 10:25. Detection took four minutes; acknowledgment took two after detection. Recovery took 25 minutes from impact or 21 from detection. Both recovery durations are understandable once their start is stated. A single incident contributes a duration, while a mean requires the population.

“Mean time to identify” may refer to discovery or identifying a cause later in the response. Record which event it measures in the SRE KPI scorecard, then use the guide to reducing incident response delays with AIOps to select an intervention.

The three-row example shows how much the incident record contributes to a small sample. The 15.7-minute mean summarizes the delay, while the 32-minute row gives the team somewhere to investigate. Keeping excluded incidents beside that result also prevents a cleaner calculation from concealing a failure the monitoring never found.

Put this into practice

Inspect your detection delays

Bring: Incident onset and discovery timestamps, a service and reporting window, and explicit exclusions.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

Related reading

Explore a related question