Prompt caching: measure the cost of useful work
Build a prompt-caching pilot that counts retries, failed work and unknown charges, with a local calculator and sample ledger.
Practical guidance for SRE and platform teams operating reliable services and production AI. Independent analysis, field guides, and working resources.
Selected reading
Plan an inference-service node drain with a scoped disruption budget, replacement-capacity checks and a practical maintenance worksheet.
Test whether an inference service can carry its workload during node maintenance.
Read the article
Original writing / Newest first
Build a prompt-caching pilot that counts retries, failed work and unknown charges, with a local calculator and sample ledger.
Build repeatable cases and check recovery independently, with a runnable Python grader.
Check agent status, terminal persistence, and recovery before adopting Herdr.
Separate detection from response and find delays that an improving average can conceal.
Define the action, contain its scope, and verify what changed at the destination.
Use instrumentation and context propagation to connect the evidence in an investigation.
Releases, research & operational impact
AI infrastructureNews analysis
NVIDIA introduces workload-driven GPU cluster validation. Choose the acceptance claim and thresholds before interpreting a completed run as readiness.
Read the analysisKubernetes operationsNews analysis
NodeWright brings Kubernetes-aware host changes into a declarative workflow. Check node selection, interruption limits and recovery before a fleet rollout.
AI modelsNews analysis
Anthropic releases Opus 5.5 with lower token prices. Evaluate operational exceptions and integration behavior before changing an agent workflow.
From reading to doing
Incident response
Work through a database connection example, from the first symptom to a verified recovery.
Reliability planning
Define the service objective, calculate the allowance, and agree when a release needs to wait.
Observability
Start with one important request path. Check propagation, sampling, and the gaps in your evidence.