Herdr: how to supervise coding agents
Herdr keeps coding agents visible and their terminals persistent. Here is how to evaluate its status signals, recovery behavior, and place in SRE work.
Practical AI operations. Reliable systems.
Nate’s take on AI in production, reliability metrics, and how engineering teams work.
Herdr keeps coding agents visible and their terminals persistent. Here is how to evaluate its status signals, recovery behavior, and place in SRE work.
Separate detection from response, reconstruct the missing minutes, and find the delays that an improving average can conceal.
Tool-using agents need bounded authority, observable actions, and a recovery path when a technically successful request produces the wrong result.
Measure tokens per completed task, account for retries and tool output, and preserve the evidence an operational assistant needs.
When inference slows without obvious application errors, test physical constraints alongside queueing, workload changes, and software regressions.
Test evidence retrieval, long-session consistency, tool boundaries, and cost before letting a model influence production changes.
Define inputs, permissions, idempotency, verification, and stop conditions for one operational capability before expanding its authority.
Decide which workflow deserves an agent, where a fixed automation is sufficient, and how to measure the work left for responders.
Count interruptions, overnight disruption, escalation work, and the unfinished follow-up that page totals leave out.
Evaluate AIOps grouping by detection coverage, investigation effort, and recoverable mistakes, not alert reduction alone.
Turn reliability policy into a release decision with fresh evidence, limited exceptions, and a clear path to resuming normal changes.
Read validated architectures as a starting point for workload, recovery, and support testing, not a promise of automatic reliability.