Canary deployments: limit exposure and test the comparison
Choose representative traffic, define promotion criteria, and verify rollback compatibility before expanding a release.
Practical AI operations. Reliable systems.
Practical methods you can take back to your team.
Choose representative traffic, define promotion criteria, and verify rollback compatibility before expanding a release.
Use consistent onset and discovery timestamps, include customer-reported incidents, and interpret the average with its sample and uncertainty.
Counters, gauges, and distributions answer different questions. Build metrics that preserve scope, denominators, and the evidence behind an incident.
Design logs around useful events, searchable context, and a collection path whose failures are visible.
Define coverage, escalation, training, and recovery time before assigning the calendar. Review workload as well as shift counts.
Containers package processes; orchestration reconciles desired state. Reliability still depends on probes, capacity, dependencies, and application behavior.
Start tracing with a customer-critical path, then test propagation, sampling, and whether the trace supports a real investigation.
Build a ticket-routing evaluation around real ownership, uncertain cases, and correction cost. Start with an honest baseline before adding a model.
Understand spans, parent-child relationships, retries, and missing evidence so traces support a careful diagnosis.
Use the free SRE books as references for a concrete service problem, then adapt and test the practices against your own constraints.