Sustainable on-call: make workload relief change the plan
Connect on-call interruptions, specialist demand, and recovery work to a concrete capacity decision using an on-call workload worksheet.
Practical AI operations. Reliable systems.
Nate’s take on AI in production, reliability metrics, and how engineering teams work.
Connect on-call interruptions, specialist demand, and recovery work to a concrete capacity decision using an on-call workload worksheet.
Read validated architectures as a starting point for workload, recovery, and support testing, not a promise of automatic reliability.
Compare detection, correlation, investigation, and automation using your incidents, review costs, and failure paths.
Locate the delay, choose a bounded intervention, and measure recovery quality alongside elapsed time.
Make a useful AI request without assuming that tool approval, anonymization, or a training opt-out settles every data-handling question.
Separate incident coordination from mitigation authority, compare competing actions, and use a worksheet that preserves decisions and verification.
Define who does the work, who decides, and who accepts risk so ownership remains useful during incidents and follow-up.
Make AI-assisted routing and remediation accountable through visible evidence, meaningful overrides, and review of who bears the errors.
Protect time away through explicit coverage, recovery policies, and realistic planning instead of relying on individual boundary-setting alone.
Use interruption patterns and responder feedback to change paging, coverage, and recovery expectations when on-call work becomes unsustainable.
Give teams clear service ownership, decision authority, and capacity to act on reliability evidence before the next incident.