On this page6 sections
A latency objective such as “99% of search requests finish within 300 milliseconds” contains a useful question: how many requests crossed that threshold? A p99 chart approaches it from the other direction, estimating the duration at a particular rank. Both views can help an SRE, but the fraction meeting the threshold is the more direct input to a request-based error budget.
Prometheus histograms retain enough information to calculate that fraction across service replicas. The practical work is choosing which requests belong in the calculation, putting the threshold into the measurement, and combining the observations without changing their meaning. This guide develops that calculation for an illustrative search API, then shows how to introduce it alongside an existing latency dashboard.
Give the search objective a measurable meaning
Consider a hypothetical internal search service used by an incident assistant to retrieve runbooks. For this example, the team wants 99% of eligible search requests to return a successful result within 300 ms over a rolling 30 days. A successful empty result is valid; a server error is not. These are illustrative choices, not recommended targets for every search service.
The application measures from entry into the search handler until the response is ready. That includes the handler’s dependency calls, but excludes time before the request reaches it and time delivering the response to the caller. A client-side objective would need observations that include those intervals. If the assistant spends several seconds deciding to call search, improving this handler’s histogram will not explain that delay.
Choose the eligible requests and completion condition before writing the query. Here, a fast error remains a bad event because the assistant did not receive the result it needed. This applies the distinction in our guide to defining service level objectives: the measurement must see the outcome the objective protects. Google’s SLO implementation chapter develops the good-events-to-total-events approach and the tradeoffs between measurement locations.
What a histogram keeps about each request
A classic Prometheus histogram exposes cumulative bucket counters. With an upper boundary of 0.3 seconds, that bucket counts every observed duration at or below 300 ms. A 120 ms request increments that bucket and every higher bucket, while a 700 ms request does not. The histogram also exposes a total observation count and sum. The bucket counts overlap, so adding all buckets would count a request more than once.
For the example, imagine 10,000 completed requests in a test dataset. Of those, 9,850 are successful and at most 300 ms, 100 are successful but slower, and 50 fail. The combined success-and-latency indicator is 9,850 ÷ 10,000, or 98.5%. There are 150 bad events, not merely the 100 slow successes. A 99% objective would allow 100 bad events in that population, so this dataset consumes 150% of that allowance.
The exact threshold count is one reason to retain a classic bucket at the objective’s boundary when working with classic instrumentation. A percentile between boundaries requires estimation. Prometheus’s histogram documentation explains the distinction and now recommends native histograms where the collection path supports them. The classic example here is useful for services that already expose bucket series; it does not require a format migration before the SLI can be understood.
Instrument one completion path
You need a scrapeable application metric, permission to change its instrumentation, and a Prometheus-compatible query backend. Confirm the metric name, duration units and labels in the actual exposition before copying a query. The names below are deliberately specific to the example and are not built-in Prometheus metrics.
The Python client allows explicit histogram boundaries. This illustrative wrapper records one duration per completed attempt, with only two outcome values. It assumes exceptions represent unsuccessful searches and a normal return represents a valid response. An application that returns error objects normally must classify those responses as errors too.
from time import perf_counter
from prometheus_client import Histogram
SEARCH_SECONDS = Histogram(
"runbook_search_duration_seconds",
"Search handler duration, including unsuccessful attempts",
["outcome"],
buckets=(0.05, 0.1, 0.2, 0.3, 0.5, 1.0, 2.0),
)
def measured_search(query):
started = perf_counter()
outcome = "error"
try:
result = search_backend(query)
outcome = "success"
return result
finally:
SEARCH_SECONDS.labels(outcome=outcome).observe(
perf_counter() - started
)
This is synchronous application scaffolding, not a runnable search implementation. The application supplies search_backend and its existing metrics endpoint. The Python histogram reference documents bucket configuration and observation. For asynchronous handlers, keep the timer around the awaited work rather than timing only coroutine creation.
The finally block covers ordinary returns and exceptions, but cannot record a request after its process is forcibly killed. Likewise, a request that never reaches the handler is absent. Compare the observation count with an independent ingress or client count during rollout. If those missing requests are part of the objective, measure at a location that can account for them instead of presenting the handler ratio as end-to-end reliability.
Calculate the fraction before estimating the percentile
Assume the scrape configuration supplies job="runbook-search" and all replicas use the same bucket layout. This query returns the fraction of observed requests that succeeded within the threshold over the last five minutes:
sum(rate(runbook_search_duration_seconds_bucket{
job="runbook-search", outcome="success", le="0.3"
}[5m]))
/
sum(rate(runbook_search_duration_seconds_count{
job="runbook-search"
}[5m]))
The numerator selects fast successes. The denominator includes both outcomes. Each counter becomes a per-second rate before the replicas are summed, allowing counter resets to be handled at the individual series. Dividing rates cancels the time unit. Prometheus’s rate documentation explains reset handling and why aggregation follows the rate calculation.
A five-minute indicator shows recent behavior; it is not the 30-day objective. For the longer window, replace both rate(...[5m]) expressions with increase(...[30d]), provided the backend retains that history and can evaluate the query. The result uses estimated counter increases across the window. Do not average five-minute percentages to obtain a monthly percentage: a quiet interval would receive the same influence as a busy one. Add good and total events over the intended window, then divide.
Keep a request-volume panel beside the ratio. With no observations, the ratio is undefined; replacing it with 100% would turn missing activity into apparent success. Missing scrapes need their own collection-health check. Our Prometheus missing-data guide covers that distinction without treating an absent series as a healthy service.
A companion p99 query can help diagnose the successful requests’ timing:
histogram_quantile(0.99,
sum by (le) (
rate(runbook_search_duration_seconds_bucket{
job="runbook-search", outcome="success"
}[5m])
)
)
This combines bucket observations before estimating a percentile. It does not average replica percentiles. Notice the population: this panel describes successes, while the combined SLI includes errors in its denominator. A faster p99 alongside a worse SLI can therefore be consistent, for example when slow requests increasingly fail. Label the panels so the next responder can see which question each answers.
Where native histograms change the query
A native histogram carries its distribution in histogram samples rather than a separate floating-point series for each bucket. If the service and its entire telemetry path support that representation, the corresponding successful-request fraction can use histogram_fraction. The following example assumes the base metric contains native histogram samples:
histogram_fraction(0, 0.3,
sum(rate(runbook_search_duration_seconds{
job="runbook-search", outcome="success"
}[5m]))
)
*
sum(histogram_count(rate(runbook_search_duration_seconds{
job="runbook-search", outcome="success"
}[5m])))
/
sum(histogram_count(rate(runbook_search_duration_seconds{
job="runbook-search"
}[5m])))
The first term estimates the fast fraction among successes. Multiplying by the successful request rate and dividing by the all-outcome rate puts failures back into the denominator. The threshold may fall between native bucket boundaries, so the fraction can be an estimate. Check the resolution and resulting uncertainty against the decision you intend to make. The function reference describes that interpolation.
A team needing exact classification at a contractual threshold may prefer an explicitly counted good-event counter alongside total events. That trades away the histogram’s flexibility for a direct answer to one fixed question. Keep the histogram for diagnosis if its distribution is useful, but do not make a format migration a prerequisite for counting the outcome the service actually promises.
Introduce the measurement without changing the paging policy
Start with one service and one measurement location. Run the new calculation beside the current dashboard through a representative busy period and an instrumentation restart. Use controlled successful, slow and unsuccessful requests to confirm which counters move. The first acceptance condition is agreement with the intended event population, not a more favorable percentage.
There is a small arithmetic experiment you can run without a service. It verifies the illustrative definition and guards against the easy mistake of dividing only by successful requests:
fast_success = 9850
slow_success = 100
failed = 50
all_requests = fast_success + slow_success + failed
indicator = fast_success / all_requests
bad = all_requests - fast_success
assert indicator == 0.985
assert bad == 150
assert bad / (all_requests * 0.01) == 1.5
print(f"Good: {indicator:.1%}; bad events: {bad}")
This tests arithmetic, not PromQL execution or production coverage. For a real rollout, inspect the scraped le="0.3" series on every replica. Changing a bucket layout partway through the fleet can leave a threshold present on only some instances while the denominator still includes all of them. Introduce a versioned metric when the layout must change, compare it independently, and retire the old query only after the new population is complete.
Costs follow the measurement dimensions. Seven finite buckets plus the infinite bucket, count and sum produce ten core series per outcome per target, before any additional client metrics. Two outcomes across twenty targets therefore produce 400 core series. Extra labels multiply that count. Raw search text, user IDs and request IDs have no place in these labels: they both expand the series population and can expose sensitive information. Prometheus’s instrumentation guidance discusses label cardinality; keep detailed searches in appropriately controlled logs instead.
Record the added series, scrape payload size and query duration during the trial. If a repeated dashboard query is expensive, a recording rule can precompute it, with an explicit name and evaluation interval. Preserve the counts needed for longer-window calculations rather than keeping only a convenient percentage. If instrumentation causes unacceptable overhead or disagrees with the independent request count, revert the instrumentation change and leave existing alerts active while investigating.
Once the measurements agree, the search team can connect its 30-day result to its existing error-budget policy and design shorter-window alerts separately. The useful outcome is a calculation the team can explain: which searches were eligible, which completed successfully within 300 ms, and how failures affected the result. That makes a latency regression something responders can investigate without first reconstructing what the chart meant.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- SLO implementation chaptersre.google
- histogram documentationprometheus.io
- Python histogram referenceprometheus.github.io
- rate documentationprometheus.io
- function referenceprometheus.io
- instrumentation guidanceprometheus.io
- recording ruleprometheus.io
