Protecting time off means changing the SRE workload
Protect time away through explicit coverage, recovery policies, and realistic planning instead of relying on individual boundary-setting alone.
Practical AI operations. Reliable systems.
SRE and platform engineering practitioner. Writing about what actually works in production.
Protect time away through explicit coverage, recovery policies, and realistic planning instead of relying on individual boundary-setting alone.
Use interruption patterns and responder feedback to change paging, coverage, and recovery expectations when on-call work becomes unsustainable.
Give teams clear service ownership, decision authority, and capacity to act on reliability evidence before the next incident.
Use AIOps to connect operational evidence, then test whether it improves investigation without hiding missing signals or unsupported conclusions.
Choose a baseline, evaluate false positives and missed incidents, and connect anomalies to an operational decision before paging on them.
Start tracing with a customer-critical path, then test propagation, sampling, and whether the trace supports a real investigation.
Build a ticket-routing evaluation around real ownership, uncertain cases, and correction cost. Start with an honest baseline before adding a model.
Design continuous monitoring around freshness, coverage, delivery, and response ownership before adding analysis or automation.
Choose an SLI, target, and evaluation window that reflect user experience, then agree on the decisions an SLO result will change.
Understand spans, parent-child relationships, retries, and missing evidence so traces support a careful diagnosis.
Connect telemetry, customer reports, decisions, and verification so improvements in a dashboard reflect improvements in the service.
An error budget translates an SLO into an allowed amount of bad service. Its value comes from the decisions attached to consumption and burn rate.