On this page3 sections
Which event means an upload succeeded: receiving the file, accepting it into a queue or saving it so the customer can retrieve it? The answer determines what a service level objective should measure. A healthy receiving endpoint can coexist with failed downstream work, leaving an apparently reliable service short of the result its users need.
An SLO sets a target for a measured behavior over a defined window. The measurement is its service level indicator, or SLI. An SLA is a separate agreement that may attach consequences to performance. Keeping the engineering target and contractual agreement distinct allows a team to define a useful objective without accidentally implying a different kind of promise.
For queued uploads, success occurs after acceptance
Consider a hypothetical upload service that queues files before saving them durably. Its endpoint continues returning success while the downstream processing fails. Counting accepted requests describes the front door, but cannot establish whether the files will be available afterward.
If durable completion is the intended outcome, the SLI needs evidence at that state. If speed is part of the promise, it also needs completion timing. A synchronous service might measure the fraction of requests finishing within a threshold; asynchronous work may require a completion-time or freshness measure that continues beyond the initial response.
Observation points have different blind spots. Server-side counts cannot see requests that never reached the server, where client observations or synthetic transactions may help. A downstream completion record answers another question: did accepted work finish? Explain what the selected source establishes and which population it includes before deciding what its percentage means.
An indicator at durable completion also changes the recovery question. The team needs the queued files to be saved successfully, even if the receiving endpoint has been healthy throughout. Restoring that endpoint alone cannot resolve the incident. The target and response policy should protect the completed upload, which is the task the customer needs.
One objective may not adequately represent every customer journey. Use a small set when necessary and inspect consequential populations that overall success could conceal. Simplifying the dashboard is not a sufficient reason to hide a different user outcome inside a single average.
What a 99.9% target allows
Once the outcome is clear, choose a target using customer needs, historical performance, dependency constraints and engineering cost. A tighter percentage can demand substantial work while leaving the user's main difficulty elsewhere in the journey. The target should express the experience the team has decided is worth improving, rather than reward a more impressive number.
Suppose the target is 99.9% request success in a window with one million eligible requests. The permitted bad-event fraction is 0.1%, yielding an allowance of 1,000 bad events. This is a request-based error budget. It does not automatically become a fixed number of downtime minutes, because a minute at high traffic and a minute at low traffic contain different amounts of work.
Google’s SLO implementation chapter offers examples for choosing indicators and objectives. Applying them locally still requires deciding what completion means and which target fits the customers. Those choices give the arithmetic its operational significance.
When failed uploads consume the budget
If queued files are failing and consuming the budget, the objective should inform repair priority, release risk and incident response. An error-budget policy records the agreed consequences before a release becomes disputed. It gives product and engineering a shared basis for choosing what changes next.
The policy depends on trustworthy measurement. Validate the pipeline and revisit the definition when the product or user population changes. Retain earlier definitions so a trend can be interpreted as a change in observed service behavior or a change in what was counted. Those are different reasons for the same graph to move.
Read diagram description
Define the measurement boundary and eligible population first. An SLI measures behavior; an SLO sets its target and window; the resulting error budget informs an agreed policy. An SLA is a separate agreement. Diagram labels: User operation: What experience must the service protect?; SLI: Good eligible events ÷ all eligible events; SLO: Target for that indicator over a defined window; Error budget + policy: Calculate the allowance and agree what changes.
The first upload SLO is now more than a percentage: it identifies durable completion, the observations and eligible population that establish it, the acceptable target and the response to threatened performance. When the front door remains healthy but files stop being saved, the team knows why the objective is failing and which user outcome the repair must restore.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- Google’s SLO implementation chaptersre.google
