OpenTelemetry for SRE: connect the signals that explain a request
Understand instrumentation, context propagation, and the Collector, and why consistent telemetry still needs careful signal design.
Practical AI operations. Reliable systems.
Walkthroughs for observability, incident response, and AI-assisted operations.
Start with what AI can do for operations, and where judgment still matters.
Read AIOps fundamentals ↗Make metrics, logs, and traces work together.
Read Observability for SRE ↗Build an incident response practice that learns.
Read Incident management with AI ↗Instructions, examples, and implementation notes.
Understand instrumentation, context propagation, and the Collector, and why consistent telemetry still needs careful signal design.
A practical ownership model for shared platforms, service reliability, incident command, and the work that falls between teams.
Separate delivered postmortem actions from demonstrated risk reduction, close only tested scope, and record residual exposure with a treatment worksheet.
Follow evidence through technical failure and incident response, then test whether the proposed correction changes the mechanism.
Use a curated notebook to compare runbooks and incident evidence, while keeping citations, freshness, and operational authority in view.
A design guide for an incident assistant that handles duplicate events, partial failures, and reviewed AI summaries across tools.
A reusable prompt pattern for incident analysis, with checks for missing facts, unsafe recommendations, and unsupported certainty.
Adapt a complete SLO record, worked budget examples, and an exception log without confusing bad requests with minutes of downtime.
Use directory purpose, mount boundaries, and read-only checks to investigate missing files, full disks, and unexpected runtime state.
Measure a representative workload, identify the limiting resource, and test one reversible change instead of applying a universal sysctl recipe.
Translate customer symptoms into an owned investigation while preserving impact, uncertainty, and a useful communication loop.
Decide which alerts earn an interruption using urgency, reachable authority, coverage evidence, and a practical alert acceptance record.