Start here
AIOpsSRE covers the work of running reliable systems, including systems that use AI and teams using AI to help with operations. The paths below lead to explanations, worked examples and tools for the problem in front of you.
New to AIOps and SRE?
AIOps fundamentals distinguishes anomaly detection, event correlation, investigation assistance, and automated action. Next, use the observability guide for SRE to connect user outcomes with operational evidence. The AIOps and SRE glossary explains unfamiliar terms as you encounter them.
Investigate an incident with better evidence
A slow request can cross several services before it reaches a customer. These guides explain how metrics establish the affected population, logs preserve relevant events and traces connect the execution path. Together, they help a responder move from a dashboard symptom to a specific investigation.
- Choose useful service metrics and denominators.
- Preserve investigation context with structured Python logging and distributed tracing.
- Evaluate observability with AIOps against the questions responders actually need answered.
For AI applications, include token cost and latency alongside task quality. Those measures explain the resources used; task-quality checks establish whether the user received a useful result.
Make incident response more repeatable
Begin with AI incident management for response roles and workflow. Then adapt the SRE runbook template to a common failure. The example connects the symptom with diagnostic choices, mitigation and the evidence a responder needs to establish recovery.
Use MTTD and its related incident intervals to find delays without confusing detection with recovery. Compare those timings with SRE KPIs and the impact users experienced. When a repeated action is ready for a pilot, follow the automated-remediation guide.
Operate AI agents with clear boundaries
Agent systems introduce state, tool calls, and actions that can outlive a single request. Read AI agents as production systems before choosing an execution architecture. Define the permitted actions and stop conditions with the agent skills execution contract.
For teams already supervising coding assistants, the Herdr guide explains how a session workspace helps people find and resume their conversations, and what survives a restart.
Build a sustainable reliability practice
Define service level objectives, then use the error budget worksheet to connect reliability consumption to decisions. Agree on who can authorize an exception before a release is under pressure.
For the human side, examine on-call workload, SRE burnout, and reliability leadership. After incidents, turn review findings into changes with owners and verification criteria.
Keep a reference close at hand
Browse tools, templates and references when you need a working document, or explore the SRE and AIOps tool stack when you are evaluating a capability. The news section covers recent releases and their operational implications. For all topics and publication dates, use the complete article archive.