On this page
A Collector can recover its connection before a responder recovers a trustworthy view of the service. The exporter starts sending again, the queue gets smaller, and old observations arrive in a convincing burst. That burst establishes progress in delivery. It does not establish what is happening now.
For a hypothetical inference service, imagine that the telemetry backend is unavailable for four minutes while requests continue. A persistent queue preserves many of those observations. When the backend returns, an incident agent reads the newly ingested records and describes an error spike as current. The records are real; the time interpretation is wrong.
Define when telemetry becomes too old for its intended decision, and keep that loss of visibility separate from service health. This guide applies our monitoring-freshness principle to OpenTelemetry Collector queues: size the waiting room, understand its failure boundaries, and rehearse the return to useful evidence. The same pipeline can support historical diagnosis while temporarily being unsuitable for automated recovery decisions.
Find the boundary that actually holds the data
The Collector receives telemetry, passes it through configured processors and exports it to a destination. An exporter sending queue buffers work that cannot yet be delivered. Retry policy handles retryable export failures. Persistence can preserve queued work across a process restart. These are separate protections, described in the OpenTelemetry resilience documentation; none creates unlimited capacity or protects every earlier stage.
Draw your actual path before tuning it. Include the application SDK, local Collector if present, gateway, backend ingestion and the query or analysis step. A queue at the gateway cannot recover an event that the SDK discarded before sending. A disk-backed exporter queue does not automatically persist a processor’s private memory.
Use an independent observation path for the decision that cannot wait. For example, a service owner might use a narrowly scoped synthetic request whose result does not depend on the unavailable telemetry backend. That probe establishes only the behavior it exercises. It cannot establish complete service health or replace traces needed to explain a complex failure.
Calculate capacity and age separately
The exporter-helper configuration distinguishes queue sizing in requests, items or serialized bytes. A request count is not a span count: batch sizes can vary. Record the unit beside every capacity calculation and confirm that the exporter and installed Collector version support the selected settings.
For a first approximation, let arrival rate be λ items per second, delivery rate be μ items per second and the backlog be B items. During a complete delivery outage lasting T seconds, additional demand for queue space is λ × T. After recovery, estimated drain time is B ÷ (μ − λ), provided μ is greater than λ. These are planning equations for steady rates, not predictions of Collector scheduling or memory use.
In an illustrative calculation, 1,500 spans per second over a 240-second outage creates 360,000 spans of backlog if none is lost or exported. A backend accepting 2,000 spans per second afterward has only 500 spans per second left for catch-up, so draining takes about 720 seconds. If acceptance falls to 1,500 per second, the backlog does not shrink. More buffer postpones overflow; it does not create delivery capacity.
Do not infer a fixed freshness delay from that drain-time estimate. Different requests may take different retry paths, and a fresh record can arrive before older work. Measure event age where the downstream decision reads the evidence. For an AI incident assistant, retain the source event window in its evidence package instead of treating retrieval time as event time.
| Question | Evidence to collect | What it does not prove |
|---|---|---|
| Can the queue absorb the planned outage? | Arrival rate, sizing unit, capacity and overflow behavior | That data remains timely enough to act on |
| Can the pipeline catch up? | Sustained delivery capacity above new arrivals | That the backend has indexed every accepted event |
| Is the decision using current evidence? | Source event age and expected coverage at query time | That an unrelated service dependency is healthy |
Build one bounded pilot before changing the fleet
Use one nonproduction Collector and a disposable backend or controlled test endpoint. You need the exact Collector distribution and version, its component list, a writable test volume, access to its internal metrics, and permission to interrupt the test destination. Establish a maximum test rate and disk allowance so the experiment cannot consume shared storage unexpectedly. Choose a pipeline owner and the person who can stop the test.
The following exporter fragment illustrates the relationship between settings. It is not a complete Collector configuration or a version-pinned deployment manifest. Merge it into a known working test configuration, retaining the receiver, processors, backend authentication and TLS settings appropriate to that environment. Validate it with the installed binary before starting the pilot.
extensions:
file_storage/pilot:
directory: /var/lib/otelcol/pilot-queue
exporters:
otlp/pilot:
endpoint: telemetry-test.example.com:4317
sending_queue:
enabled: true
storage: file_storage/pilot
sizer: items
queue_size: 500000
num_consumers: 4
retry_on_failure:
enabled: true
initial_interval: 5s
max_interval: 30s
max_elapsed_time: 10m
service:
extensions: [file_storage/pilot]
# Retain your test pipeline and reference otlp/pilot there.
The capacity and consumer count are experiment inputs, not recommended production defaults. The ten-minute retry setting is a limit on retrying a failed request, not a ten-minute retention guarantee measured from the original event. Enqueue rejection occurs before export retry logic. A retry policy therefore cannot rescue work that never entered the queue.
The file-storage extension requires a writable directory. In Kubernetes, make the mount’s lifecycle match the restart scenario you intend to survive. Give each Collector instance its own appropriate storage; do not improvise multiple writers against one directory. Verify volume attachment and permissions after replacement, rather than assuming that an unchanged path string identifies the same retained data.
Persisting telemetry also persists its contents. Use synthetic payloads for the drill, avoid credentials and customer prompts, and apply your access, encryption and retention requirements to the volume. Account for storage, I/O, network egress and backend ingestion charges during replay. A larger disk queue can retain more sensitive data and generate a larger recovery burst.
Observe failures without counting every retry as loss
Use the Collector’s internal telemetry to distinguish pressure from loss. Inspect queue size and capacity, enqueue failures, receiver refusals, export failures and successfully sent items. Export failures can be retried, so their count alone is not a count of permanently lost records. Scrape and inspect the metrics exposed by your exact build; names, labels and units need to agree with your queries.
Record a small sequence of synthetic event identifiers and source timestamps at the producer and query the destination for them. This provides a concrete reconciliation sample. Keep per-event identifiers in the test dataset, not in unbounded metric labels. Compare unique identifiers, duplicates, arrival delay and omissions after a defined wait. The exercise does not establish an exactly-once delivery guarantee.
Monitor memory as well as disk. The memory limiter can refuse work when memory pressure is high; recovery then depends on upstream retry behavior. Persistent export storage does not remove memory needed by receivers, processors and active requests. Increasing queue size without observing the rest of the process can move the failure to another boundary.
Run a failure drill with two acceptance decisions
Write the acceptance conditions before injecting failure. For this illustrative pilot, historical evidence must remain available after the planned short interruption, while an immediate recovery decision must visibly abstain whenever its evidence exceeds the service owner’s freshness limit. Choose that limit from the consequence of waiting, not from the queue’s available space.
- Establish a baseline at the agreed test rate. Confirm that the synthetic identifiers appear in the destination and that event age is comfortably inside the decision limit.
- Interrupt only the test destination. Record queue growth, export failures and the moment the downstream evidence becomes stale. Verify that the reader or agent sees unavailable evidence rather than a healthy-service claim.
- Restart the test Collector while the destination is unavailable. Verify that the retained queue resumes from the intended storage. Record any pre-queue loss separately.
- Restore the destination while keeping new arrivals active. Measure catch-up rate and destination query freshness. Check whether replay delays or throttles fresh work.
- In a separate bounded case, exceed the queue limit or deny writes to the disposable volume. Confirm that the loss signal reaches an owner through a functioning path. Never fill a shared production disk for this test.
- Reconcile the synthetic sequence and review the decision history. Account for missing and duplicate events; distinguish an expected loss boundary from an unexplained gap.
Success requires both a delivery result and an honest decision state. A shrinking queue with stale evidence is still a visibility incident for time-sensitive automation. Conversely, a fresh synthetic check does not close unexplained historical gaps. Record the remaining gap and its owner if the missing interval matters to diagnosis or audit.
Choose the protection the decision needs
Roll out to one owned workload, compare with the baseline, and expand only when storage use, refusal rate, catch-up time and event age meet the agreed conditions. Preserve the prior configuration and a stop procedure. During rollback, stop or redirect new test input deliberately and account for the retained queue before changing its identity or deleting its volume. Removing the storage would destroy the evidence needed to evaluate the pilot.
If the evidence deadline is shorter than the outage you plan to buffer, persistence cannot make that evidence suitable for immediate action. Keep historical collection, but give the urgent decision an independent observation path or suspend automation until visibility returns. For a low-consequence hourly trend, accepting a longer delay may be simpler and cheaper than adding another durable system.
Start with one alert or agent decision that consumes this pipeline. Put its permitted event age, fallback observation, owner and recovery check beside the Collector runbook. Those four entries tell the responder when a recovered exporter has actually restored the evidence they need.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- OpenTelemetry resilience documentationopentelemetry.io
- exporter-helper configurationgithub.com
- file-storage extensiongithub.com
- Collector’s internal telemetryopentelemetry.io
- memory limitergithub.com
