Explainer

Where SRE and platform engineering need an explicit handoff

In brief

A stalled rollout shows how platform repair and application recovery depend on different teams, with a handoff that keeps both investigations moving.

5 min read

Sources
Two hands firmly support the same copper baton as one owner accepts it while the sending owner maintains a grip.
Conceptual accepted handoff in the hypothetical stalled rollout. The work remains supported through transfer; platform investigation, application rollback compatibility and incident coordination still have distinct responsibilities.
On this page4 sections

A rollout stalls inside a shared deployment platform. The platform team can investigate the controller, but the application team knows whether a rollback is compatible with the data already written. In this hypothetical case, both teams have something essential to do. The customer’s recovery depends on those investigations meeting, not on finding a job title broad enough to own every detail.

That is the practical boundary between SRE and platform engineering. Their emphases differ, while their work often overlaps. A useful agreement says which service is owned, who can make the next decision and how work moves when another team’s knowledge or access is needed. Dividing engineers into people who build and people who run does not answer those questions.

Different purposes still create shared incidents

Platform engineering commonly treats internal capabilities as products for developers: deployment workflows, templates, identity integration, infrastructure access and supported defaults. Reliability engineering asks whether services meet user expectations and how to reduce the work required to keep doing so. Both can involve writing software, operating infrastructure and taking on-call shifts.

Those emphases do not remove ownership. A platform team remains responsible for its platform’s reliability, while a product team remains responsible for application behavior. SRE involvement can range from advice to substantial operational ownership depending on staffing and maturity. A small organization may combine the functions without losing their purposes, provided its actual commitments remain clear.

Google’s guidance on implementing SLOs is relevant because reliability objectives need stakeholder agreement. SRE cannot independently choose acceptable customer impact, and a common platform template cannot infer every application’s tolerance. The agreement has to reach the people carrying those consequences.

What, exactly, is the receiving team accepting?

For the stalled rollout, controller repair and application rollback compatibility are separate questions. Coordination can arrange both investigations without giving the coordinator every production permission. The application owner still has to establish whether the older workload can use data left by the newer version. The examples below show how the accountable owner follows the failed component and required decision.

SituationAccountable ownerRequired collaboration
Deployment service unavailablePlatform service ownerApplication owners identify blocked or partial releases
Application regression after a successful rolloutApplication service ownerPlatform supplies rollback capability; SRE assists when engaged
Customer journey crosses several unhealthy servicesNamed incident commander coordinates responseEach service owner investigates and implements changes
Reliability policy exceptionDesignated business and technical decision makersSRE or service owner explains exposure and recovery options

An incident commander needs a way to summon the owners who can act and keep their work aligned with customer impact. That does not transfer ownership of every affected component to command. It gives each investigation a shared context without obscuring who can authorize and implement its change.

Define a handoff by the decision the receiving team accepts, the evidence accompanying it and the fallback when the needed owner is unreachable. In the rollout case, the application owner might show that execution never reached application startup and attach the platform error. The platform owner can explicitly accept investigation of that stage while the incident lead continues coordinating customer impact.

Until that acceptance is established, the sending team keeps the investigation moving. A ticket reassignment records a routing action, but it does not prove that a person has taken up the next check. The distinction prevents work from falling into the space between two otherwise reasonable team boundaries.

Recurring incidents can change platform defaults

If several services repeatedly misconfigure timeouts, local fixes may leave the default causing the pattern untouched. Incident evidence can make a platform-backlog proposal concrete: identify the affected workflow, explain the mechanism, propose the default change and account for migration work among existing users. Reliability work then influences the shared product instead of recurring as separate repairs.

The transfer can run the other way. A deployment guard may require service-specific health checks, while a telemetry convention can increase storage cost. Application owners need to participate in those choices and know who will maintain the integration after launch. A platform improvement that creates hidden application work has not made that work disappear.

Adoption and reliability outcomes together help expose the tradeoff. Teams may willingly use a platform that still fails too often, or avoid controls that make routine delivery difficult. In the latter case, a nominal control may push work into unsupported paths. Neither adoption alone nor a well-designed policy on paper establishes the result.

When the expected owner is unavailable

A recent case that bounced between teams is a useful rehearsal. Walk through who receives the page, who can pause deployments, who restores the platform and who verifies the user journey. Check actual permissions along with names. The exercise should expose a missing capability or unavailable owner before the same gap appears during the next rollout.

Explicit acceptance must not delay urgent mitigation; appropriate bounded actions can be authorized in advance, with emergency transfers recorded so continuity does not depend on one person. An unresolved cross-team exposure can enter the risk registry with an owner and review date, where it can influence planning rather than be rediscovered in another incident.

A reassigned ticket is not an accepted handoff
A reassigned ticket is not an accepted handoff. Illustrative failed rollout: the sending team keeps investigation moving until the receiving team accepts it. The incident lead continues coordinating customer impact across that transfer.
Illustrative failed rollout: the sending team keeps investigation moving until the receiving team accepts it. The incident lead continues coordinating customer impact across that transfer.
Read diagram description

Illustrative failed rollout: the sending team keeps investigation moving until the receiving team accepts it. The incident lead continues coordinating customer impact across that transfer. Diagram labels: Application owner: Identify failed stage; attach platform evidence; Platform owner: Explicitly accepts investigation of that stage; Accepted workstream: Named owner, next check and escalation route; Incident lead throughout: Coordinates customer impact and shared priorities.

In the stalled rollout, the platform owner investigates the failed controller while the application owner checks whether rollback can use the current data. The incident lead keeps both workstreams connected to customer recovery. An explicit handoff lets that parallel work continue; a ticket moving between unstaffed queues leaves the original problem waiting.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

Related reading

Explore a related question