Explainer

Using ChatGPT for SRE analysis and writing

In brief

Runbook comparisons, timelines and draft updates provide checkable inputs for assistance. Incident closure also needs evidence about service recovery and unfinished work.

4 min read

Sources
A lifted drafting sheet repeats an intact teal loop and an open coral loop from the incident source folio beneath it.
Conceptual illustration of the hypothetical incident draft: recovered new traffic and unresolved failed backlog both belong in the account. The pictured draft is not a verified recovery or authorization to close an incident.
On this page4 sections

In a hypothetical incident, new requests are succeeding again while failed backlog still needs attention. A summary can be accurate about the new traffic and incomplete about recovery. ChatGPT can help organize that distinction in a draft; giving the same assistant permission to close the incident creates a different task with different evidence requirements.

That is a useful starting point for SRE adoption. Choose work whose output can be checked against known inputs, then measure preparation and review together. A permitted timeline and selected observations can become a useful draft without assuming the assistant can continuously monitor the service or decide when its recovery is complete.

Runbooks, timelines and draft updates

Comparing runbook revisions, extracting unanswered questions from a timeline and drafting status updates are suitable candidates because their sources provide a basis for review. The assistant can organize material that already exists and make its relationships easier to see. The responder still checks whether the draft represents that material faithfully.

In the hypothetical backlog case, the draft needs to account for two outcomes: new requests are succeeding again, and earlier failed work remains unresolved. A responder checking only the first can miss the omission. A useful draft links both outcomes to the supplied timeline, so the next responder can see the recovery work still open before a general service-restored statement erases it.

Closing an incident requires a defined recovered state, permission to change the record and a way to verify that state beyond the generated draft. A successful summary trial can justify continued writing assistance while ticket closure remains contingent on explicit service and backlog checks outside the summary.

For low-stakes drafting, a lightweight review may be sufficient. A helpful writing result cannot justify access to production data or permission to execute a recovery action.

Apply the same distinction to code and query generation. Ask for assumptions and test against a safe fixture with inspectable inputs and expected results. Valid syntax does not establish that a query selects the intended requests, and a command can succeed in the wrong environment. Checking output and target is part of the task, rather than overhead to exclude from the claimed saving.

Generated training scenarios can also help when clearly labeled as simulations. An experienced reviewer should check their expected behavior so the exercise does not turn a plausible-sounding recovery shortcut into something trainees are taught to repeat.

Keep source review in the drafting workflow
Keep source review in the drafting workflow
Example of language assistance: supplied incident records become a draft update. A responder checks the draft against those records before it is used; drafting does not require production write access.
Read diagram description

Example of language assistance: supplied incident records become a draft update. A responder checks the draft against those records before it is used; drafting does not require production write access. Diagram labels: Approved source material: Timeline and relevant observations; ChatGPT draft: Organize and explain the supplied record; Responder review: Check omissions, claims and uncertainty; Reviewed update: Use the existing communication process.

A supplied snapshot does not establish ongoing coverage

Explaining a supplied metrics snapshot can help a responder understand one point in an incident. Continuous anomaly detection requires an ongoing data pipeline, a way to identify unusual behavior, evaluation and an alert-delivery path. Those requirements do not appear merely because the assistant gave a useful explanation of the snapshot.

A suggested cause likewise remains a hypothesis until supported by the incident evidence. Reading a dashboard may provide another observation; changing a deployment adds execution consequences. The production-agent guide develops the permissions, action records and containment needed for write-capable integrations.

Preparation time and review time belong together

Compare the assisted workflow with the existing process using redacted historical cases and constructed examples whose requirements are known. Include the unresolved-backlog case as well as conflicting and missing evidence. A summary can be easy to read precisely because it left out the complication the responder most needed to notice.

Record supported claims, omissions, corrections, review time and whether the result was useful. These observations reveal where preparation became easier and where checking became harder. Measuring only generation speed would leave the team unable to tell whether the full drafting task improved.

OpenAI's evaluation documentation covers defining tasks and checking outputs. A manual rubric or evaluation service can support the comparison, provided the cases and criteria remain stable enough to interpret changes. Repeat important cases and record the configuration: model, prompt and retrieval changes can alter behavior after a successful initial trial.

The account receiving the operational records

Confirm the approved product, account, connected tools and data class before adding workplace material. ChatGPT and an application built on an API are different deployments with different controls. The workplace AI guide provides the questions to resolve for the environment the team actually uses.

Start with a timeline whose uncertain claims and outstanding work are known. If the draft makes those easier to understand with less total preparation and checking, the language assistance has demonstrated value. If automated closure is proposed later, return to the recovered-requests and failed-backlog distinction: the new workflow must establish both the state it is allowed to declare and the observations that justify declaring it.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

More on OpenAI

All OpenAI coverage

Related reading

Explore a related question