On this page5 sections
A successful Slack API call can leave an incident integration in an uncertain state. The channel exists and responders are using it, but the process crashes before saving its identifier. On restart, “try again” might create another channel for the same incident. The integration has recovered its ability to run while losing track of the work it already did.
This reference design connects PagerDuty, Jira and Slack around that problem. It is not an installable product or a tested deployment package. Its purpose is to show how to preserve incident identity and finish incomplete work before adding generated summaries or more integrations. The coordination should remain understandable when only some of its API calls succeed.
Which system owns each part of the incident?
PagerDuty might own the incident lifecycle, Jira the follow-up work and Slack the discussion. Whatever division fits the organization, decide which system owns identity, acknowledgment, severity and resolution. If each integration derives these independently, a resolved notification in one system can contradict the active response in another.
The integration also needs its own durable record of progress: a stable internal incident identifier, the linked destination IDs, incoming event IDs, completed steps and pending work. Store it outside the running assistant so a replacement process can determine which actions need reconciliation. A saved channel identifier can be checked directly. If the process crashed before saving it, recovery must locate the existing channel using the recorded incident identity through a supported lookup, or leave the result unresolved for an operator. The missing identifier does not establish that channel creation failed.
Verified event → durable intake → deduplication
→ incident-state update
→ queued destination actions
→ reconcile results and record destination IDs
→ reviewed summary with links to evidence
The sequence separates receiving an event from completing every destination action. Once intake and intended work are durably recorded, a slow destination can remain pending while the rest of the incident response continues. The record also gives a responder an answer more useful than a single success flag: which system is current, and which still needs attention?
Read diagram description
Reference design: a durable incident record links provider IDs and completed steps. Slack and Jira may succeed independently; reconcile each destination rather than replaying the whole workflow. Diagram labels: Verified incoming event: Validate, persist and deduplicate; Authoritative incident state: Incident ID, lifecycle, event IDs and queued work; Slack destination: Create or locate channel; store ID; Jira destination: Create required follow-up; store ID; Reconcile + report: Expose pending sync; review summaries against evidence.
Event retries and destination writes
Start in test workspaces and services with approved data handling and narrowly scoped identities. Credentials belong in an appropriate secret store, not checked-in configuration or a Kubernetes ConfigMap. Verify incoming signatures using the provider's supported method and reject invalid or stale requests before they become queued work.
Slack Events API documents retries and best-effort delivery. Acknowledge a valid event promptly under that contract, then do longer work asynchronously. Save the event identifiers already handled, so receiving the same delivery again does not automatically repeat its writes. This is the difference between recognizing an old request and discovering whether its earlier action completed; the latter still requires destination-state reconciliation.
Rate limits create another kind of delay. Jira Cloud’s documentation explains HTTP 429 and retry guidance. Follow the relevant response, bound retries and keep failed work available for review. During a storm, an unbounded queue can make an apparently responsive integration fall further behind every minute.
Recovery after Slack succeeds and Jira times out
- Persist a test incident's identity before creating its dependent artifacts.
- Create or locate the discussion channel and record its destination identifier.
- Create a follow-up record only when policy requires one, linked to the same incident.
- Write links back to the authoritative record and check that responders can open them.
- If a write times out, establish destination state before deciding whether to repeat it.
In a hypothetical partial failure, the Slack channel is created but the Jira update times out. Responders begin investigating in that channel while the integration waits. Recovery should preserve the first channel and establish the ticket's state; replaying the entire workflow could create a second coordination space, while deleting the first would discard the investigation already underway.
The recovery logic needs to track and reconcile each consequential side effect independently. Its state now includes the responders’ work in the first channel, so an apparently administrative cleanup can be just as consequential as repeating the ticket write.
This makes “incident exists; Jira synchronization pending” a legitimate, visible outcome. A workflow should not report completion merely because Slack returned success, but it need not prevent responders from using the valid channel while Jira is unavailable. Deliberately disable each destination in testing to verify that the implementation can represent that distinction.
For a low-consequence notification, occasional duplication may be acceptable. Make that choice explicit per operation; do not extend it to assignments, incident closure, or other changes that redirect response.
Summaries draw on the durable incident record
A generated summary can make a long discussion easier to enter by separating observed impact, current hypothesis, actions taken and the next update time. Links to the underlying events and explicit unknowns let a responder check its account. Consequential status changes, including a new recovery claim or incident resolution, need review by an authorized responder.
Maintained runbooks provide a more dependable reference than asking the model to invent a recovery procedure for each event. Keep the discussion under the organization's retention policy as well. A dedicated incident channel is not automatically disposable, and archiving it is different from deleting the evidence it contains.
Restarts, rate limits and synchronization lag
Exercise duplicate delivery, out-of-order resolution, expired credentials, rate limiting and a restart between a destination write and local persistence. Monitor queue age and failed synchronization alongside process uptime. A healthy process can still be carrying an old incident state into every destination.
The Slack operations guide covers communication fallback when Slack is unavailable; the agent skill contract addresses later remediation capabilities. Responders should be able to continue if Slack or the assistant disappears. The actual implementation needs to demonstrate these behaviors in its environment before production use.
Restarting just after channel creation exercises the hardest gap in the design: Slack has changed, but the local record may not yet contain its identifier. Recovery must locate that existing channel, reconcile the ticket and finish only the missing work. If lookup cannot establish the result, the pending state gives a responder a specific problem to investigate without creating a second incident channel.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- Slack Events APIdocs.slack.dev
- Jira Cloud’s documentationdeveloper.atlassian.com
