On-call overload: respond to the workload before it becomes normal
Use interruption patterns and responder feedback to change paging, coverage, and recovery expectations when on-call work becomes unsustainable.
Practical AI operations. Reliable systems.
Nate’s take on AI in production, reliability metrics, and how engineering teams work.
Use interruption patterns and responder feedback to change paging, coverage, and recovery expectations when on-call work becomes unsustainable.
Give teams clear service ownership, decision authority, and capacity to act on reliability evidence before the next incident.
Choose an SLI, target, and evaluation window that reflect user experience, then agree on the decisions an SLO result will change.
Connect telemetry, customer reports, decisions, and verification so improvements in a dashboard reflect improvements in the service.
An error budget translates an SLO into an allowed amount of bad service. Its value comes from the decisions attached to consumption and burn rate.
Choose a small set of reliability indicators with clear definitions, useful distributions, and an explicit decision attached to each.