How-To
Read a distributed trace without mistaking it for the whole system
Understand spans, parent-child relationships, retries, and missing evidence so traces support a careful diagnosis.
Practical AI operations. Reliable systems.
Walkthroughs for observability, incident response, and AI-assisted operations.
Start with what AI can do for operations, and where judgment still matters.
Read AIOps fundamentals ↗Make metrics, logs, and traces work together.
Read Observability for SRE ↗Build an incident response practice that learns.
Read Incident management with AI ↗Instructions, examples, and implementation notes.
Understand spans, parent-child relationships, retries, and missing evidence so traces support a careful diagnosis.
Use the free SRE books as references for a concrete service problem, then adapt and test the practices against your own constraints.