Guide

OpenTelemetry tail sampling: keep the traces that explain failures

In brief

Configure and size a tail-sampling pilot, then test whether known errors retain the span relationships needed for diagnosis.

10 min read

Sources
Two saffron thread paths stand out among lavender paths on plum fabric, each preserved from beginning to end.
Original AI-generated conceptual illustration; not a telemetry result or product interface.
On this page7 sections

A trace-storage bill can fall for two very different reasons: you stopped keeping routine requests, or you lost the requests that would have explained the next failure. Tail sampling gives you a way to choose between them. Instead of deciding at the beginning of a request, a collector waits for records of its individual operations, called spans, to arrive. It can then use those records to decide whether to retain the trace that connects them.

That makes it possible to keep errors, slow operations and a smaller sample of ordinary traffic. The useful question for an SRE is whether the resulting dataset still answers a particular investigation. Start with one request path and a known failure, then verify its retrieval under the proposed policy before expanding collection. This guide develops that test into a small, measurable OpenTelemetry pilot.

What the sampler sees, and what it cannot recover

Consider a hypothetical checkout request that calls a payment service and then receives a timeout. The checkout span, payment-client span and payment-server span share a trace ID. If the timeout status arrives at the collector before its decision, an error policy can retain the associated trace, including the earlier successful work needed to understand the sequence.

A head sampler makes its choice earlier, commonly in the application SDK. It can reduce the work of recording and exporting traces, but it cannot base that initial choice on a failure that has not happened yet. A tail sampler receives recorded spans and makes a later choice. It cannot restore spans discarded upstream. A 10% head sample followed by an error-retention policy therefore means keeping qualifying errors from that incoming sample, not all errors in the application. The OpenTelemetry sampling overview explains this distinction.

For the checkout pilot, first establish which spans the applications actually export. Check the SDK configuration and the parent sampling behavior across the payment call. Record a trace ID from a controlled request and find all expected spans with sampling disabled on the pilot path. If the payment-server span is already missing, changing a collector policy will not fix its instrumentation or propagation.

A staging checkout deployment with synthetic transactions and no real payment data gives you a manageable place to collect that baseline. Keeping full collection on this one path makes it easier to see whether the checkout and payment records connect, without increasing collection across the fleet. The existing tracing implementation guide covers choosing and instrumenting that first path.

Sampling a trace requires a view of its related spans. In a multi-instance deployment, sending each batch to an arbitrary collector can split the checkout and payment observations. One instance may see an error and keep its part, while another sees only successful work and discards the context.

A common architecture has an intake layer that routes by trace ID into a separate sampling layer. The OpenTelemetry load-balancing exporter supports this routing. Every intake instance must resolve a consistent set of sampling destinations. Its current documentation also notes that changing that destination set reroutes some trace IDs. Adding a sampler is therefore a stateful topology change, even when the new collector is healthy.

Checkout and payment spans sharing trace A are routed to sampler one; trace B goes to sampler two. Each sampler buffers, decides and exports retained traces.
Route by trace ID before applying a trace-level policy. The diagram shows logical paths, not a guarantee that a restart or membership change preserves buffered spans.

For a first experiment, one sampling instance removes the routing variable and makes the outcome easier to explain. It is a pilot simplification, not a high-availability design. Before expanding, repeat the same request cases through multiple intake instances and deliberately change sampling membership in staging. Compare the retrieved span sets before, during and after that change.

Keep resource enrichment and sensitive-data removal upstream of the policy when the policy depends on their output. For example, if you remove a customer attribute for privacy, do not expect a downstream policy to match its original value. The gateway deployment documentation describes separating collector responsibilities; your diagram should identify which hop changes attributes and which one makes the sampling decision.

Choose a policy you can explain with three requests

Begin with an error policy, a latency policy and a modest probabilistic sample of otherwise ordinary requests. This combination gives the responder detailed failures and a comparison population. The percentage is a starting assumption to test against cost and usefulness, not a universal recommendation.

The following is a configuration fragment for a Collector distribution that includes the tail-sampling processor. Merge it into your existing traces pipeline, preserving the receiver, memory protection, authentication and exporter configuration. Validate it with the exact Collector build you deploy; it is not a complete production configuration or a reported hands-on benchmark.

processors:
  tail_sampling:
    decision_wait: 30s
    num_traces: 20000
    expected_new_traces_per_sec: 300
    policies:
      - name: checkout-errors
        type: status_code
        status_code:
          status_codes: [ERROR]
      - name: slow-traces
        type: latency
        latency:
          threshold_ms: 1500
      - name: ordinary-comparison
        type: probabilistic
        probabilistic:
          sampling_percentage: 5

The processor reference documents these policy types and the full decision rules. With the simple policies shown here, an error or qualifying slow trace can be retained independently of the 5% policy. Adding explicit drop policies changes that reasoning, so review combined behavior whenever the configuration grows. The capacity and arrival-rate values are illustrative inputs, not measured sizing advice.

Now send a fast successful checkout, a slow successful checkout and a checkout with an explicitly recorded error status. Give each a unique trace ID and keep those IDs in the test output. The slow and error cases should be retrievable after the decision and export delay. A single fast request may or may not be kept, which is the expected behavior of the probabilistic branch. Use a larger set of distinct trace IDs to inspect its approximate selection rate.

Check the recorded status rather than assuming an HTTP response automatically sets every span to error. A timeout caught by application code may be represented differently from an uncaught exception. Likewise, a trace latency policy should not be described as a promise to retain every request exceeding your service-level objective: the spans received, their timestamps and the configured threshold determine what the processor can evaluate. Looking at the resulting trace lets you check what the policy actually retained and whether it explains the request.

Size the waiting window from incoming work

The sampler needs space for work awaiting a decision. A useful first estimate is the rate of new trace IDs multiplied by the waiting interval. At an illustrative 300 new traces per second and a 30-second wait, about 9,000 traces may be awaiting decisions under steady arrivals. A burst to 900 per second over that interval raises the estimate to 27,000, already beyond the example's 20,000-trace setting.

This arithmetic tells you how many traces may be waiting, but estimating RAM takes another step. A checkout trace with six small spans and an agent execution with hundreds of tool spans have very different footprints. Measure resident memory and CPU with representative span counts and attribute sizes, including exporter queues and other processors in the container budget. That extra capacity matters because an out-of-memory restart can discard the failure evidence the sampler was waiting to collect.

The waiting interval also creates a tradeoff in the investigation itself. A longer interval gives delayed spans more opportunity to influence the decision, but holds more data and delays availability. For the payment timeout, measure the interval between the first received span and the last span needed to explain that timeout. Request duration alone does not include every SDK batching or transport delay.

Do not stretch one short-request sampler to cover jobs that run for hours. If a workflow's decisive evidence normally arrives after the sampling window, use a separately designed policy or tracing path and correlate its stages through appropriate context. Otherwise the pilot can appear successful on synchronous checkouts while failing precisely where background work matters.

Test the failures of the telemetry pipeline

Once the three basic cases work, introduce conditions that challenge the collection path. Delay the payment span until after the decision window, increase new-trace arrivals in a safe test environment, and restart the sampler during outstanding requests. These are separate experiments: changing all three at once makes a missing span difficult to explain.

The processor documentation describes late spans, decision caches and premature trace removal. A remembered decision can help later spans follow the original choice, but it does not make the original decision aware of evidence that arrived afterward. Monitor the component's early-drop and late-span telemetry using the names exposed by your pinned version. Match that telemetry with the known test trace IDs so a quiet dashboard cannot hide a failed retrieval test.

ExperimentInspectWhat the result changes
Payment error arrives on timeCheckout and payment spans under the saved trace IDConfirms the intended investigation survives the normal path
Error span arrives after the windowInitial decision, late-span behavior and final retrieved traceShows whether the wait or workload split needs changing
Arrival burstKnown trace retrieval, memory, refusals and early removalsTests capacity rather than average-volume assumptions
Sampler restart or membership changeMissing span sets and recovery timeEstablishes the evidence loss to expect during maintenance

Export failures need their own observation. A trace selected successfully can still fail to reach the backend. The Collector resiliency guidance explains queues and retries; those mechanisms operate at different stages from a pending sampling decision. Track both selection and export, then perform the backend lookup that the responder will actually use.

For the checkout example, the acceptance test is deliberately concrete: retrieve a known failed request and distinguish a payment timeout from time spent before the payment call. If the trace only proves that checkout was slow, the pilot has not yet preserved the evidence it was built to keep. This is why the test begins with an investigation and works backward to collection settings.

Measure savings without misreading the sample

Compare retained bytes and backend ingestion cost with the baseline, then add the collector resources and maintenance work needed to obtain those savings. Tail sampling may reduce downstream storage while leaving the application-to-collector traffic substantial. Use actual billing units, compression behavior and deployment prices from your environment; a retention percentage by itself is not a cost estimate.

Alongside cost, report how many controlled failure cases yielded the span relationships needed for diagnosis, how long those traces took to become searchable, and what happened during the burst and restart experiments. Those results make the tradeoff understandable to an operational lead deciding whether another collector tier is worth running.

Keep request-rate and error-rate measurements independent of this selected trace dataset. If you deliberately retain errors more often than successes, the fraction of stored traces showing errors will overstate their share of all requests. Use population-level metrics for that denominator and sampled traces to investigate particular executions. The guide to reading distributed traces explains how missing observations and overlapping spans also limit interpretation.

Privacy applies before storage reduction. Attributes can pass through the application, network and sampler even when the trace is eventually discarded. Remove secrets and unnecessary personal data at instrumentation or the earliest appropriate processing point, restrict collector access, and use authenticated encrypted transport across trust boundaries. Review the OpenTelemetry sensitive-data guidance with the person responsible for the attributes your applications emit.

Roll out one path and keep a way back

Adopt the pilot only when its savings justify the measured collection cost and the known failure trace remains useful. For a low-volume service where full retention is affordable, continuing to retain everything may be simpler. For a path whose decisive spans cannot reliably reach the same sampler in time, fix that limitation or choose a different collection design before relying on tail sampling.

Keep the previous collector configuration and a documented route back to the baseline. Rehearse that change in staging, including how much extra backend traffic it creates. Restoring full retention without checking ingestion limits can replace one telemetry problem with another. Prefer a scoped rollback of the pilot path to a fleet-wide emergency change.

Start the production trial with a limited service slice and a defined review period that includes representative load. Ask the team that investigates checkout failures to retrieve the controlled trace using its normal tooling, rather than handing it a prepared screenshot. Expand when that lookup explains the timeout and the operating measurements support the cost. The outcome worth preserving is the ability to follow a failed request through the work that explains it.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

A useful next step

Continue the work