Choosing tools for an SRE workflow

An incident moves through more than one tool: a measurement becomes an alert, an alert reaches a responder, and an investigation leads to a change. Choosing a stack means understanding those connections as well as the individual products. If a dashboard cannot identify the service named in a page, responders have to reconstruct the relationship themselves.

This guide follows that work from instrumentation through delivery. Its examples and selection criteria draw on the linked documentation and AIOpsSRE guides; they are not a ranked product comparison or a report of hands-on testing.

Instrumentation and investigation

OpenTelemetry provides instrumentation and collection components. It still needs a destination where the team can store and query the telemetry. That distinction matters when comparing a collector with a complete monitoring service: installing the former does not supply all the retention, query and investigation capabilities of the latter. The OpenTelemetry architecture guide explains how those parts connect.

For an illustrative slow-checkout investigation, the metrics view needs to identify the affected operation and time window, then let a responder find relevant request evidence. Query behavior, retention and cardinality affect whether that work is feasible at the service's scale. The Grafana dashboard guide follows the investigation view, while the metrics guide explains the populations and measurements behind it. Collection loss and stale data need to remain visible so an empty panel does not silently end the investigation.

Paging and incident coordination

The next connection is between a detected condition and someone able to respond. Scheduling, escalation and acknowledgment decide who receives the page and what happens if they cannot take it. The incident record then holds the impact, current actions and next decision as people move between monitoring, tickets and chat. The PagerDuty, Jira and Slack integration guide examines where those records can diverge and how delayed notifications or unavailable chat affect the response.

For existing Opsgenie users, replacement planning has a concrete deadline. Atlassian states that Opsgenie will shut down on April 5, 2027. Migration work therefore includes schedules, routing and permissions, together with evidence that the replacement can reach the intended responder.

AI assistance and execution

An assistant could assemble the checkout incident's source material or help compare possible explanations. The useful comparison is whether the resulting account helps the responder investigate, including omissions and time spent correcting it. A larger context window may allow more input, but it does not explain whether the assistant uses that input well.

The Claude Opus evaluation article considers model configurations on operational work, while the NotebookLM guide follows a source-backed incident dossier. Data handling and failure behavior belong in the comparison because the assistant becomes another dependency in the response. Generating a patch or executing a recovery procedure adds separate consequences; the automated-remediation guide covers action limits and recovery checks.

Delivery and failure testing

Deployment tools turn declared intent into changes in the running service. A partial rollout, drift or an incompatible data change can complicate recovery even when the tool reports that its own step completed. The canary deployment guide explains traffic selection, promotion criteria and rollback compatibility. Feature controls and rollback solve different problems, so the recovery path depends on what the change actually did.

A failure-injection tool can help examine an assumption about that path under controlled conditions. The team still supplies the hypothesis, scope, observation and stop condition. The AIOps evaluation guide brings these questions into a comparison with the existing workflow, including the integration and maintenance effort. For the checkout response, a useful stack is one whose outputs remain interpretable as the work passes from detection to investigation and recovery.