Agent skills for SRE: write the execution contract
Define inputs, permissions, idempotency, verification, and stop conditions for one operational capability before expanding its authority.
Search titles and article text.
Define inputs, permissions, idempotency, verification, and stop conditions for one operational capability before expanding its authority.
Connect AI-assisted diagnosis to bounded runbooks, explicit approval gates, and independent recovery checks. Start with a practical SRE rollout checklist.
Define the action, contain its scope, and verify what changed at the destination.
Use a curated notebook to compare runbooks and incident evidence, while keeping citations, freshness, and operational authority in view.
Build repeatable cases and check recovery independently, with a runnable Python grader.
Check agent status, terminal persistence, and recovery before adopting Herdr.
A design guide for an incident assistant that handles duplicate events, partial failures, and reviewed AI summaries across tools.
Design incident channels, bot actions, evidence links, and a fallback path that still works when the chat system is unavailable.
Use language-model assistance for evidence review, draft communication, and practice, with a clear boundary around production decisions.
Connect an AIOps investment to a measurable workflow, its full operating cost, and the evidence needed to expand or stop the trial.
NVIDIA reports low confidential-computing overhead on an eight-GPU test. Separate that measured performance result from your own security and recovery claims.
Lower model prices and new caching controls make agent pilots cheaper to run. Measure accepted work before expanding the workload.