AIOps anomaly detection: unusual does not always mean actionable
Choose a baseline, evaluate false positives and missed incidents, and connect anomalies to an operational decision before paging on them.
Practical AI operations. Reliable systems.
Nate’s take on AI in production, reliability metrics, and how engineering teams work.
Choose a baseline, evaluate false positives and missed incidents, and connect anomalies to an operational decision before paging on them.
Design continuous monitoring around freshness, coverage, delivery, and response ownership before adding analysis or automation.
Choose an SLI, target, and evaluation window that reflect user experience, then agree on the decisions an SLO result will change.
Connect telemetry, customer reports, decisions, and verification so improvements in a dashboard reflect improvements in the service.
An error budget translates an SLO into an allowed amount of bad service. Its value comes from the decisions attached to consumption and burn rate.
Choose a small set of reliability indicators with clear definitions, useful distributions, and an explicit decision attached to each.