On this page3 sections
“No matching records” is an answer from a search system. It leaves open whether nothing happened, the query looked in the wrong place or the event never reached the store. A logging strategy has to help responders distinguish those possibilities, as well as make the events themselves useful.
Take an illustrative worker that restarts with an unfinished operation. Its last message says “retrying request,” but contains no operation identifier or previous result. Even perfect collection would leave the responder unable to tell whether another attempt completes the work or repeats a change already made. More messages are not automatically more evidence; the fields must support the decision someone will later need to make.
Fields that explain an unfinished operation
Operations worth investigating deserve stable event names and structured context. A failed dependency call, exhausted retry limit, configuration change or completed job describes a meaningful transition. An event timestamp and enough environment context let the reader place it in time and distinguish a production operation from a test.
Fields need equally stable meanings. A duration is useful only with a known unit. Transport status and application outcome should have different names if both are recorded, since a successful exchange can report an unsuccessful operation. A query may continue to run after a field's meaning changes, leaving its reader with an undetected misunderstanding.
For the restarted worker, the operation identifier and previous outcome enable a comparison with the destination. A trace identifier can connect the event to request execution without including the entire request body. Redact or omit secrets and personal data before collection where possible, and inspect exception messages as well as the fields intentionally supplied by the application. Context should make the operation understandable without copying everything it touched.
How collection loss appears in an investigation
The collector is part of this investigation path. Its ingestion freshness, rejected records, queue growth and drops reveal whether a quiet result might be a delivery problem. Preserve event time and observation time where supported: when an event happened and when it was observed answer different questions if delivery was delayed.
Slow destinations require an explicit tradeoff. Waiting indefinitely to deliver diagnostic logs can block application work and turn a logging fault into a service outage. A bounded buffer limits that exposure, but it can discard events once full. Counting and exposing those drops tells the responder that an apparent gap may reflect lost collection rather than absent activity.
This is why a logging pipeline needs to report its own condition alongside the records it carries. It cannot promise a complete history in every failure, but it can preserve evidence about how incomplete that history may be.
Read diagram description
Before inferring that no event occurred, check the collection path and the query. Compare a known event with timestamps, index, access and retention; expose drops and stale delivery. Diagram labels: No records returned: An observation about the search, not the service; Collection check: Known event, queue, rejects, drops and freshness; Query check: Time zone, interval, index, access and retention; Interpret the evidence: Distinguish no events from missing or unqueried events.
Searching for a known event
Begin an investigation with the affected operation or user outcome. Metrics can establish population and scale before individual events explain the mechanism. If the logs were sampled, a count of matching messages describes those records and does not automatically count every failure in the service.
When the search is empty, look for an event that should be present. A known request identifier or controlled test event provides a useful reference. Compare its timestamp with the query's time zone and interval, then check the index, access and retention window. That can distinguish missing collection from a search that never had a chance to find the evidence.
Retention and permissions turn this into an operational test rather than a configuration exercise. A record available to the logging administrator today may not be available to an on-call responder investigating an older incident. Choose the fields around the later response decision, then verify retrieval using that responder's actual permissions and the expected incident window. For the worker, both the identity of the first attempt and its availability after the restart matter.
Unfamiliar failures also need some diagnostic context, even when nobody anticipated the exact question. A proportionate baseline retains information that has proved useful to investigations while limiting sensitive detail and storage cost. The restarted worker illustrates the balance: an operation identifier and prior outcome are more useful here than a larger copy of the request.
For the restarted worker, the decisive log improvement is the ability to reconstruct the first attempt and identify the destination that holds its outcome. If collection was incomplete, that fact belongs with the search result. The next responder can then investigate the remaining uncertainty rather than mistake an empty search for permission to repeat the operation.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- event time and observation timeopentelemetry.io
