From AIOps signals to agent actions: design the handoff
Decide which workflow deserves an agent, where a fixed automation is sufficient, and how to measure the work left for responders.
Practical AI operations. Reliable systems.
Original AIOpsSRE articles, newest first. Find current releases and industry coverage in News.
Decide which workflow deserves an agent, where a fixed automation is sufficient, and how to measure the work left for responders.
Connect on-call interruptions, specialist demand, and recovery work to a concrete capacity decision using an on-call workload worksheet.
Evaluate AIOps grouping by detection coverage, investigation effort, and recoverable mistakes, not alert reduction alone.
Turn reliability policy into a release decision with fresh evidence, limited exceptions, and a clear path to resuming normal changes.
A design guide for an incident assistant that handles duplicate events, partial failures, and reviewed AI summaries across tools.
Build a view that answers who is affected, what changed, and where to investigate, then test its queries and missing-data behavior.
Read validated architectures as a starting point for workload, recovery, and support testing, not a promise of automatic reliability.
Check market definitions, forecast periods, and growth arithmetic before turning a market estimate into an operational investment case.
A reusable prompt pattern for incident analysis, with checks for missing facts, unsafe recommendations, and unsupported certainty.
Connect an SLO’s allowance to planning and release behavior, with explicit ownership, exceptions, and treatment of measurement gaps.
Adapt a complete SLO record, worked budget examples, and an exception log without confusing bad requests with minutes of downtime.
Use directory purpose, mount boundaries, and read-only checks to investigate missing files, full disks, and unexpected runtime state.