Long agent sessions need a prompt-cache SLI
Measure cached input against total input, locate the prefix miss boundary, and judge cache changes beside latency, task cost and correction work.
AI in production. Reliability in practice.
Latest edition ·The latest analysis
Read coding-agent benchmarks as system comparisons when harnesses differ, then use controlled contrasts and independent tests before attributing a result to the model.
Read the storyFrom the publication
Measure cached input against total input, locate the prefix miss boundary, and judge cache changes beside latency, task cost and correction work.
Review the claim each test or document protects, challenge the evidence, and record a keep, rewrite or retire decision before deleting repository history.
A host that launches codex exec must own the process, thread lease, JSONL trace, cancellation and artifact checks. Resume is a lifecycle decision, not a shell shortcut.
Basalt narrows local inference to Qwen3.8-Flash-Next on Blackwell. Its memory and concurrency controls matter more than copying the largest throughput number.
Field guides
Follow a request. Review the code. Find out what actually happened.
Explore the guidesKeep the traces that explain a failure
Tail sampling selects recorded traces by error, latency and other evidence. A checkout timeout shows how routing, waiting time and memory affect the trace a responder can retrieve.
Give AI code review a useful job
CodeRabbit brings AI feedback into pull requests, editors and the CLI. A retry loop that exceeds its caller’s deadline shows the kind of finding that can repay review time.
Recover when a tool call times out
A missing response can leave a worker already created. Durable operation IDs and the destination’s idempotency contract let a replacement executor recover the original request.
Put it into practice
Free reliability tools. No account needed.
Open the workbenchThe AI tools guide
Coding assistants, code review and agent supervision.
Browse the guideCompanies & products