Guide

What makes an alert worth paging for

In brief

A capacity warning needs enough lead time and a reachable responder who can act. The alert review separates useful urgency from a missing response arrangement.

6 min read

Sources
A near-full reservoir connects to a warning bell, while a separate response linkage stops short of a capacity lever.
Conceptual depiction of the hypothetical capacity warning: credible urgency can coexist with an unavailable intervention. The gap remains an exposure to resolve, not a reason to silence the signal. No exhaustion time or measured incident is pictured.
On this page5 sections

An accurate capacity page can still be a poor interruption. It may correctly predict that storage will fill before morning, yet reach someone who can only confirm the graph while the only person allowed to add capacity is unavailable. The warning is valuable; the response arrangement is unfinished.

That distinction makes alert review more useful than a campaign to lower page counts. Examine what waiting would cost, what a reachable responder can accomplish and whether the signal provides credible grounds for that work. The result may be a page, a staffed queue or a response gap that needs repair before either route is adequate.

The recovery option that waiting would remove

The first comparison is immediate response versus the next staffed review. What becomes harder, more expensive or impossible if the team waits? Deteriorating service, capacity exhaustion and a shrinking recovery window can each create a deadline before customers encounter an outage.

For the storage warning, reaching someone before the disk fills is only the first part of the response. The responder needs time to check the growth, reach whoever can authorize more capacity and complete the intervention. If those steps extend beyond the projected exhaustion point, moving the alert earlier may help; if the authorized person cannot be reached at all, earlier detection leaves the same response gap. Following that path shows which part of the arrangement needs to change.

Treat the interruption as a claim that needs evidence: waiting has a meaningful cost, a reachable responder can do something useful, and the signal is credible enough to justify that work. Keeping those claims separate prevents an accurate measurement from standing in for an actionable response.

Google's SLO alerting guidance evaluates precision, recall, detection time and reset time. Those signal properties matter, and the operational question adds another check: can the intended recipient use this warning before the relevant option disappears?

Check urgency and the available response
Ask whether waiting can lose the chance to prevent harm and whether a useful intervention is reachable in that window. Urgent with a usable response: page the reachable responder. Can wait: route to a staffed work queue. Urgent without a usable response: retain an unresolved coverage gap and repair escalation, delegated scope or the explicit risk decision; do not assume the signal can simply be suppressed.
Illustrative routing decision. Urgency alone is insufficient for a usable page: the intervention must be reachable before the response window closes. Work that can wait still needs a staffed owner; an urgent signal without a usable response remains an unresolved coverage gap.
Read diagram description

Ask whether waiting can lose the chance to prevent harm and whether a useful intervention is reachable in that window. Urgent with a usable response: page the reachable responder. Can wait: route to a staffed work queue. Urgent without a usable response: retain an unresolved coverage gap and repair escalation, delegated scope or the explicit risk decision; do not assume the signal can simply be suppressed.

An accurate warning with nobody able to act

Consider a hypothetical capacity alert whose current projection puts exhaustion before the next staffed period. The primary responder can verify the growth rate, but only an unavailable specialist can add capacity. Suppressing the page would hide a real exposure; repeating it unchanged would keep interrupting someone who cannot complete the needed response.

A reachable escalation route would let the responder bring in the specialist; a previously delegated, bounded intervention might let the responder add capacity without them. Where the service supports safe load shedding, reducing demand could preserve time to act. If none of those options is available, the accountable owner has an exposure to accept or repair. The review should make that choice visible while the warning remains meaningful, rather than treating fewer interruptions as evidence that the storage problem has gone away.

Now change one condition in the example. If stable growth puts exhaustion several staffed periods away, a tracked item for the capacity owner may leave enough time to act. The observation has not become unimportant; its urgency has changed because the available intervention can wait. Revisit the route if the projection changes.

Why a total score can hide a disqualifying gap

A rubric can help reviewers compare alerts, but its total should not average away a missing condition. Excellent documentation does not make an unavailable responder reachable. High precision cannot compensate for a destructive failure the rule misses. Record page, staffed queue, duplicate or unresolved coverage gap with the evidence behind that disposition.

Also distinguish repeated notifications from different symptoms that appear related. Deduplication can remove copies of one condition; correlation can combine two simultaneous failures. Responders need the original signals and a way to split the latter. The correlation evaluation guide examines how to test that separate behavior.

Coverage after a quieter routing change

Test a bounded routing change against known failures, normal conditions and missing telemetry. Compare interruption burden with independent coverage evidence, including customer reports, incident reviews and signals outside the new suppression rule. Fewer pages are useful when important failure detection survives the change.

Rare catastrophic conditions may have little replay evidence. Do not remove them merely because they seldom fire. Record the uncertainty, the consequence of missing them and the independent means of checking the failure before deciding how their warning should be handled.

A trial can establish useful behavior without answering every rare case. Keeping the unknowns explicit lets the team choose a limited change and a condition for reversing it, rather than treating the absence of replay examples as proof that the alert has no value.

An alert routing record

Complete a record for one alert or a family of genuinely duplicate alerts. Its fields are separate because each addresses a different condition the routing decision depends on.

Decision inputRecord the evidenceConsequence if absent
Protected operationName the affected user operation or recovery option.Investigate meaning before assigning urgency.
Cost of waitingDescribe the deadline, uncertainty, and next staffed review.Do not assume that an unusual metric needs a page.
Usable interventionName the first check, permitted action, and required execution time.Record an unresolved response gap.
Reachable authorityName the rotation, escalation route, and permission owner.Repair coverage rather than silently suppressing the risk.
Signal evidenceRecord known failures, normal cases, missing-data behavior, and independent coverage checks.Preserve uncertainty and choose a bounded trial.
DispositionChoose page, staffed queue, duplicate, or unresolved coverage gap; name the decision owner.The review is incomplete.
Reversal conditionSpecify the missed failure or changed assumption that reopens the decision.The decision has no usable feedback path.

The two capacity warnings therefore produce different records. Exhaustion before the next staffed review leaves a response gap to resolve through a reachable intervention or explicit acceptance of the exposure. A slower projection can leave enough time for a staffed work item, with a change in that projection reopening the routing decision. Completing the record makes the reason for each route visible: what must happen before storage fills, who can make it happen and what would change that assessment. That is a more useful outcome of alert review than a lower page count on its own.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

Related reading

Explore a related question