On this page4 sections
In an illustrative incident, the alert arrives one minute after latency starts rising. Twenty minutes later, users can finally complete their work normally. Both statements can be true, and the space between them is where a useful incident review begins. A fast detection result tells the team something important, but it has not yet explained the response.
That matters when choosing the next improvement. A failure that remains unseen needs better detection. A page that reaches the wrong team needs a different repair. Separating those conditions prevents a better mean time to detect from becoming credit for work that still happens slowly after the clock stops.
Where MTTD stops on the timeline
Mean time to detect, or MTTD, averages the interval from incident onset to discovery. It ordinarily ends before mitigation reduces the impact and before the user-facing service recovers. Atlassian’s incident metric definitions distinguish those events, though the team still needs local definitions of onset and discovery to make its own records comparable.
Consider this hypothetical latency incident. Monitoring finds the condition promptly, but the page does not identify the team equipped to investigate it:
10:02 Latency first crosses the incident threshold
10:03 Monitoring detects the condition and sends an alert
10:05 Responder acknowledges the page
10:11 Investigation begins with the responsible team
10:18 A mitigation is applied
10:23 User-facing performance recovers
The detection duration is one minute. Another eight minutes elapse before the responsible team begins investigating, so onset to investigation is nine minutes. Twelve more minutes pass before recovery. Calling the nine-minute interval MTTD would erase the distinction between seeing the problem and getting useful investigation underway. And because this is one incident, its one-minute duration is an input to a mean, not a mean by itself.
Read diagram description
Illustrative incident: detection takes one minute; investigation starts eight minutes after detection. The full onset-to-recovery interval is 21 minutes. One incident is not a mean.
What happened between acknowledgment and investigation?
The six minutes between acknowledgment and investigation are tempting to label as delay. The timestamps do not say what happened inside them. A responder might be arranging coverage for another incident, checking whether rollback could damage data, or finding the only team authorized to change a dependency. Each explanation suggests a different response to the same apparently quiet interval.
Reconstruction therefore needs the information and options available at the time. A page naming the affected operation and reaching a staffed rotation gives the responder a better starting point than an unexplained latency threshold. A runbook helps only if its first check separates plausible causes. The responder and supporting records can establish which conditions held; the chart cannot infer their motive for waiting.
Separating detection from response shows which condition kept a credible page from producing a useful action. Wrong ownership calls for routing work. Expired access calls for a permission repair and rehearsal. Recurring non-actionable pages call for alert review. Generating a more confident summary would not remove any of those particular obstacles.
Some investigation time protects the service. A high-impact action that is difficult to reverse may deserve additional checking; the improvement target is avoidable uncertainty rather than the shortest possible interval. A review should distinguish necessary evidence gathering from work the system needlessly forced a responder to repeat.
For the hypothetical incident, evidence of activity outside the main channel may explain part of the six-minute gap, such as checking the affected account or arranging coverage. Where the records do not explain it, retain that uncertainty. Filling a blank with an assumption about responder performance can make an accurate timeline support a false staffing conclusion.
Which incidents made it into the average?
Combining timelines introduces a second way to lose the story. A report limited to automatically detected incidents omits failures discovered by customers or later investigation. Adding many easy-to-detect minor cases can also lower the mean while severe failures remain no easier to find. Population, severity mix and sample size need to travel with the comparison.
Even onset may be an estimate. A timestamp can record arrival at a collector rather than the underlying event. An estimated interval with its supporting evidence is more honest and useful than a precise-looking value the record cannot justify. The MTTD calculation guide develops these measurement choices without folding them into the separate question of response delay.
Different delays call for different repairs
Across recent incidents, reconstruct onset, detection, acknowledgment, investigation, mitigation and recovery wherever the evidence supports them. Recurring obstacles then suggest concrete experiments. Repeated misrouting points to an escalation repair. Confusion about a dependency suggests adding its relevant view to the page. Missing rollback access calls for a rehearsal with the actual on-call permissions, rather than an assumption that the documented command is usable.
If AI assembles that context, its links to underlying signals and contradictory evidence let responders judge the proposed explanation. After the change, revisit the interval it was intended to improve and examine the resulting decisions as well as elapsed time. Choosing a wrong hypothesis sooner would shorten a chart without resolving the original problem.
If better routing gets the latency page to the responsible team sooner, the improvement belongs in detection-to-investigation time. If a rehearsed rollback restores service faster, it belongs later in the timeline. Keeping those changes separate makes a good MTTD result useful without asking it to describe recovery work outside its clock.
Put this into practice
Inspect your detection delays
Bring: Incident onset and discovery timestamps, a service and reporting window, and explicit exclusions.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- Atlassian’s incident metric definitionswac-cdn.atlassian.com
