Prometheus missing data: test what your error-rate alert cannot see
A zero fallback can hide missing telemetry. Use a small PromQL test matrix to check what your error-rate alert can and cannot establish.
Practical guidance for SRE and platform teams operating reliable services and production AI. Independent analysis, field guides, and working resources.
Selected reading
Keep unknown write results visible, preserve operation identity and rehearse a lost response with a local Python fixture.
Keep an uncertain write owned until the destination can resolve it.
Read the article
Original writing / Newest first
A zero fallback can hide missing telemetry. Use a small PromQL test matrix to check what your error-rate alert can and cannot establish.
Plan an inference-service node drain with a scoped disruption budget, replacement-capacity checks and a practical maintenance worksheet.
Build a prompt-caching pilot that counts retries, failed work and unknown charges, with a local calculator and sample ledger.
Build repeatable cases and check recovery independently, with a runnable Python grader.
Check agent status, terminal persistence, and recovery before adopting Herdr.
Separate detection from response and find delays that an improving average can conceal.
Releases, research & operational impact
AI data infrastructureNews analysis
Google’s advisor and batch updates connect storage findings to bulk actions. Review the workload, selected objects and recovery boundary before execution.
Read the analysisAgent evaluationNews analysis
Google’s internal security project validates agent hypotheses before routing reports. Its evidence path is useful; incomplete coverage still needs an owner.
Agent infrastructureNews analysis
Google’s read-only agent compute preview separates query execution from production compute. Evaluate bursts, return from idle and data permissions.
From reading to doing
Incident response
Work through a database connection example, from the first symptom to a verified recovery.
Reliability planning
Define the service objective, calculate the allowance, and agree when a release needs to wait.
Observability
Start with one important request path. Check propagation, sampling, and the gaps in your evidence.