On this page3 sections
For an illustrative service that accepts background jobs, completed results arrive late. Did each job wait in the queue, or did the worker spend the time executing it? That is a useful first tracing question because instrumenting only the HTTP request shows acceptance while leaving completion out of view. The first project can follow this small path before expanding across the service.
Distributed tracing connects records of work across service boundaries. The implementation challenge is to preserve the relationships at the boundaries that matter to the investigation. A small, deliberately chosen path can teach the team more than a broad deployment of spans whose gaps nobody has checked.
The queue handoff connects acceptance to completion
Start at the entry request and identify the database operations, remote calls and queue handoffs needed to explain the chosen behavior. Their operation records, called spans, need stable service identity and deployment context. Those details let a responder connect the request's execution with the version and configuration serving it.
An asynchronous boundary requires particular care. The later work needs an explicit relationship to the request, rather than an assumed synchronous parent-child chain. OpenTelemetry’s context propagation guidance describes how identifying context crosses processes. Both producer and consumer can emit spans while still appearing disconnected if that context is not passed or accepted.
Check whether the related records identify when the job was handed off and when the worker began and finished. Those observations let this trace separate waiting from execution. Additional spans belong where they explain a remaining interval, rather than simply increasing the number of operations displayed.
Read diagram description
Illustrative checkout path: publish and consumer work need an explicit context relationship. The HTTP response does not establish fulfillment. Represent retries, batches and links according to the instrumentation model. Diagram labels: Checkout entry: Record the user operation and service identity; Queue publication: Propagate the appropriate context with the message; Consumer execution: Continue or link related work, including retries; Fulfillment result: Check actual downstream completion.
A narrow path can leave important cross-service interactions unobserved. Expand when the evidence reveals that limitation, rather than treating one successfully instrumented journey as service-wide coverage. This preserves the value of a focused beginning without turning its scope into a permanent blind spot.
Following a known delayed job through the trace
Send a known request through a safe environment and compare its expected path with the retained trace. Missing spans and unexpected service names show where the connection or identity is wrong. Examine errors and retries as well: a repeated attempt should remain recognizable as related work rather than appearing to be another independent customer operation.
A controlled slow dependency provides a second test. Does the trace show enough of the path to locate the wait? Overlapping child durations cannot simply be added to compute the total request time, and uninstrumented work can leave unexplained gaps. Those limitations should help identify the next check instead of disappearing behind a precise-looking waterfall.
If the retained trace ends at successful HTTP acceptance, the exercise has not answered why the result arrived late. The downstream interval remains unexplained. Check the instrumentation and collection path alongside whether the job reached a worker; absent spans alone do not establish fast execution or completion. Once the handoff and worker can be followed, the next question is which of those records the sampling policy will retain.
Sampling trades retained evidence for collection cost
Sampling lowers retained trace volume by discarding some evidence. A decision made early in a request may exclude it before a rare downstream error occurs. Waiting until more of the trace is available enables a different selection, at the cost of buffering and capacity. Choose the policy in relation to the failures and delays the team needs to investigate.
The selection also changes what a collection of traces can establish. A sample enriched for slow or failed requests is useful diagnostic material, but cannot directly estimate a fleet error rate without accounting for that enrichment. Keep metrics for the population-level view and use traces to inspect execution within it.
Attributes need a similar cost and evidence judgment. Sensitive payloads should stay out, and unbounded data should not accumulate simply because it can be attached to a span. Review storage, access and retention along with the instrumentation so the intended responder can use the evidence appropriately.
The companion guide to reading a distributed trace develops interpretation of the retained record. For this first implementation, success is more concrete: repeat the delayed job and find its handoff and downstream work. Explain where the delay accumulated and which gaps remain. That answer gives the team both a working investigation path and a reasoned basis for choosing where further instrumentation will be worth its cost.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- OpenTelemetry’s context propagation guidanceopentelemetry.io
