Observability for SRE
A dashboard can show that checkout is slow without explaining where the time went. Observability is the ability to infer internal system behavior from available outputs. In software operations, metrics, logs and traces help responders move from that visible symptom toward an explanation. Monitoring tracks conditions and can raise the alert that starts the investigation; observability also supports questions the team did not anticipate.
Consider an illustrative checkout service whose latency rises after a release. The release is a useful lead, but several explanations are still possible: a changed application path, a slow dependency or an overloaded resource. The signals become useful when they help distinguish those possibilities.
Which checkout requests became slow?
The first question concerns the affected population. Is latency elevated across all checkout requests, in one region or only for a particular operation? Request volume, error rate and latency distributions establish the extent of the symptom. Successful and failed requests need separate attention because a fast error response can make an aggregate latency figure look better while customers still cannot complete checkout.
The guide to service metrics explains these measurement choices and denominators. Once the affected route and interval are clear, an individual slow request becomes worth examining.
From a slow request to its execution path
A trace connects the instrumented operations involved in that request. It can show whether time accumulated in the application or around a dependency call. The tracing implementation guide follows the instrumentation needed to make that path visible. Sampling and missing spans affect which requests you can inspect and how much of their execution you can explain.
Structured logs add selected events that a duration alone cannot explain. A request or trace identifier lets a responder find the corresponding event and retain its service and release identity. The logging guide discusses what context is useful without collecting unnecessary payloads. Metrics summarize the affected population; traces and logs provide detail about particular executions. Profiles and other evidence may be needed when those signals leave the question unresolved.
Testing the release hypothesis
In the checkout example, the next step follows the evidence already found. If representative traces show extra time in one dependency, compare that dependency's behavior and the affected request cohorts. If the traces omit the slow operation, the instrumentation gap itself explains why the dashboard cannot answer the question. An empty query may reflect lost or delayed collection, so data freshness belongs in the investigation too.
After an intervention, the customer-facing measurements provide the recovery check: successful checkout completion, latency and errors for the affected population. The investigation also needs to account for adverse effects on other paths. Google's monitoring chapter explains latency, traffic, errors and saturation; the OpenTelemetry primer develops the relationship between telemetry and investigation.
How the measurements inform reliability decisions
A service level objective describes the experience the team intends to protect. Its error budget uses the indicator's units. For example, a time-based 99.9% objective over exactly thirty days allows 43.2 minutes of bad time. A request-based objective instead counts eligible requests and bad outcomes; the two allowances answer different questions.
The error-budget policy guide connects that consumption with release decisions, exception authority and recovery work. In a review, the measurements explain what happened to the service; the agreed policy determines what that means for the next release.
When checkout depends on an AI step
An AI-dependent request adds more behavior to explain. Token use, time to first token where relevant, total latency and tool execution can help locate cost or delay. Task outcome still needs its own check: an HTTP success says the request was handled, not that an answer was correct or the requested work finished.
The AI token guide covers cost, latency and context limits, whose accounting and truncation behavior depend on the model and interface. The production-agent guide follows the state and external actions that can extend beyond a model call. Keeping these signals attached to the same user journey makes the AI step investigable with the rest of the service.
Closing the gap the investigation exposed
The missing evidence determines the implementation work. Use Python collection and aggregation when inputs cannot be assembled, logging configuration when events lack context, or the OpenTelemetry guide when responders cannot connect a request across services. For container platforms, resource and lifecycle behavior supplies another part of the explanation. With those inputs available, AIOps investigation assistance can be assessed on how well it helps answer the question that remained difficult.