Prometheus missing data: test what your error-rate alert cannot see
A zero fallback can hide missing telemetry. Use a small PromQL test matrix to check what your error-rate alert can and cannot establish.
In-depth explanations, practical guides and independent analysis. Choose a topic and format to find your next read.
Guides show how. Explainers unpack a concept. Commentary makes an argument. Explore Production notes.
A zero fallback can hide missing telemetry. Use a small PromQL test matrix to check what your error-rate alert can and cannot establish.
Plan an inference-service node drain with a scoped disruption budget, replacement-capacity checks and a practical maintenance worksheet.
Turn reliability policy into a release decision with fresh evidence, limited exceptions, and a clear path to resuming normal changes.
Connect an SLO’s allowance to planning and release behavior, with explicit ownership, exceptions, and treatment of measurement gaps.
Adapt a complete SLO record, worked budget examples, and an exception log without confusing bad requests with minutes of downtime.
Choose representative traffic, define promotion criteria, and verify rollback compatibility before expanding a release.
Choose an SLI, target, and evaluation window that reflect user experience, then agree on the decisions an SLO result will change.
Connect telemetry, customer reports, decisions, and verification so improvements in a dashboard reflect improvements in the service.
An error budget translates an SLO into an allowed amount of bad service. Its value comes from the decisions attached to consumption and burn rate.
Choose a small set of reliability indicators with clear definitions, useful distributions, and an explicit decision attached to each.