Commentary

SRE accountability needs authority, capacity, and a clear owner

In brief

A service owner may depend on another team’s access, risk decision or engineering time. Following a shared-database change reveals where that arrangement needs help.

4 min read

Sources
A tagged repair roll lacks a completed connection to its dependency tool, beside an empty planning sleeve.
Conceptual ownership rehearsal: naming a service owner does not supply a reachable dependency decision or time for the agreed repair. The missing connection and empty allocation are hypothetical organizational gaps, not personal blame or measured workload.
On this page4 sections

Putting a team name beside a service answers who owns the outcome. It does not necessarily answer who can make the change needed to restore it. The service owner may depend on a shared database team, lack permission to accept a data risk or have no capacity left for the agreed repair. An ownership chart becomes useful when those dependencies lead to reachable decisions.

Accountability therefore needs to be tested against actual work. Follow one consequential action from evidence through authorization, execution and verification. The missing connection will usually be more actionable than another reminder that someone should take ownership.

The duties behind responsibility and accountability

Responsibility describes work a person or team performs. Accountability identifies who answers for an outcome or decision. Organizations use these terms differently, so specify the actual duties instead of relying on a matrix letter to carry their meaning.

For a rollback, a service owner may assess compatibility, an authorized responder execute the change and an incident commander coordinate timing. Leadership may accept the commercial consequences of a release pause. Assign those duties according to the organization's real authority, rather than assuming one job title includes them all.

The SRE and platform engineering guide applies this distinction to shared systems. A platform team operates its platform, while the application team retains application responsibilities even when deployment is available through a platform button. Convenience at the interface does not remove the need to decide who evaluates the application's behavior.

Doing, deciding and coordinating can belong to different people
Doing, deciding and coordinating can belong to different people
Role comparison, not a required organizational chart. Name who fills each role for the service and check that their access and availability match the work.
Read diagram description

Role comparison, not a required organizational chart. Name who fills each role for the service and check that their access and availability match the work. Diagram labels: Performs the work: Investigates or applies the change; Authorizes the change: Can accept its specific consequences; Coordinates the response: Maintains priorities and handoffs; Verifies recovery: Checks the affected user operation.

An owner need not personally control every dependency. The operating model needs a usable escalation and acceptance path, rather than unlimited authority concentrated in one team.

What the service owner needs from other teams

A reliability commitment needs access to evidence, a route to permission, time to do the work and help from dependency owners. Naming which of these is missing turns a vague conversation about commitment into a concrete request for a decision or resource.

Consider a hypothetical service owner responsible for recovery whose mitigation requires a shared database change. The owner can coordinate investigation but cannot accept the data risk. The response model needs the dependency owner and a reachable fallback; asking for more initiative would leave the missing authority exactly where it was.

Test ownership against the consequential action and its escalation path, then repair the missing right or dependency. In the database case, the service owner's accountability becomes workable when it connects to the people who can establish compatibility, authorize the scope and execute the change.

Capacity can be the missing condition too. If the organization defers the work, record the remaining risk, the reason, who reviewed that choice and what will reopen it. An overdue ticket with unsettled priority gives the responder less useful information than an explicit decision to accept a particular exposure for now.

Disagreement across a boundary provides another useful test. If the service owner wants a rollout pause and the delivery owner disagrees, both need to know who resolves the conflict and when. Bring service evidence and the cost of the pause to that person. The accountable owner then has a path to protect the objective instead of carrying responsibility without recourse.

Decisions made under operational constraints

Accountability includes explaining decisions honestly and completing agreed follow-up. It need not turn the last person to touch a system into the explanation for every failure. Examine the information, constraints and alternatives available when the action was taken.

A postmortem can identify a mistaken action while also investigating why it was easy to take, difficult to detect or hard to reverse. Personnel or misconduct concerns belong in the organization's appropriate process. Letting them silently govern an engineering learning review makes it harder to understand the system conditions that shaped the response.

A follow-up commitment includes time to do the work

“Improve monitoring” leaves completion open to interpretation. Detecting a particular failed user operation and demonstrating that the page reaches a staffed rotation describes a result the team can check. Add the accountable team and delivery owner so the commitment includes both the work and acceptance of its outcome.

The forum allocating work should also resolve capacity conflicts among these commitments. The risk registry retains the relationship between tasks and remaining exposure, while the follow-through guide explains verification. Together they keep task completion from becoming a substitute for the intended service improvement.

For the shared-database change, the service owner needs a reachable database decision maker, a fallback and time allocated for the agreed work. A gap in any of those arrangements is a concrete organizational problem to resolve. Recording the missing permission or capacity gives the planning and service owners something they can change, beyond the name already printed beside the service.

Source context

This article does not include external reference links. Read it as the author’s perspective and evaluate the guidance against your environment.

Report an error or outdated detail

Related reading

Explore a related question