AI token usage: plan for cost, latency, and context limits
Measure tokens per completed task, account for retries and tool output, and preserve the evidence an operational assistant needs.
Practical AI operations. Reliable systems.
SRE and platform engineering practitioner. Writing about what actually works in production.
Measure tokens per completed task, account for retries and tool output, and preserve the evidence an operational assistant needs.
Follow evidence through technical failure and incident response, then test whether the proposed correction changes the mechanism.
Use a curated notebook to compare runbooks and incident evidence, while keeping citations, freshness, and operational authority in view.
When inference slows without obvious application errors, test physical constraints alongside queueing, workload changes, and software regressions.
Test evidence retrieval, long-session consistency, tool boundaries, and cost before letting a model influence production changes.
Define inputs, permissions, idempotency, verification, and stop conditions for one operational capability before expanding its authority.
Decide which workflow deserves an agent, where a fixed automation is sufficient, and how to measure the work left for responders.
Connect on-call interruptions, specialist demand, and recovery work to a concrete capacity decision using an on-call workload worksheet.
Evaluate AIOps grouping by detection coverage, investigation effort, and recoverable mistakes, not alert reduction alone.
Turn reliability policy into a release decision with fresh evidence, limited exceptions, and a clear path to resuming normal changes.
A design guide for an incident assistant that handles duplicate events, partial failures, and reviewed AI summaries across tools.
Build a view that answers who is affected, what changed, and where to investigate, then test its queries and missing-data behavior.