On this page4 sections
A canary can finish its observation window without testing the change it was deployed to examine. If the new version changes writes but receives only read requests, a quiet error chart establishes very little about that release. Limited exposure has reduced the number of people at risk, but the traffic still needs to exercise the relevant behavior.
A canary deployment sends part of the production workload to a candidate while retaining an existing version as a baseline. Its value comes from learning under bounded exposure. Designing what to learn, who supplies the test traffic and what permits the next increase is as important as choosing the starting percentage.
The changed operation determines the measurements
For a write-path change, identify the expected outcome of an eligible write: correct completion, acceptable latency and any required follow-on work. Those expectations suggest successful-transaction measurements, latency distributions, correctness checks and asynchronous-completion evidence. Request count and workload mix show how much of that behavior actually reached the candidate.
Google's canarying chapter discusses comparing candidate and control behavior. Choose an observation interval that can expose the failure in question. A memory leak or interaction with a daily batch may not appear during a brief rollout, even when the traffic mix is otherwise representative.
Assignment by request, user, tenant or instance produces different groups. Session affinity may need to keep related requests together; shared state and customer-specific behavior may affect both candidate and baseline. A traffic percentage therefore does not by itself establish either isolation or a fair comparison.
Those measurements explain the hypothetical write-path case. A candidate that receives only reads has supplied no write-completion evidence, however quiet its error chart. Representative eligible writes through a safe test path can answer that missing question; until then, the promotion decision must retain the untested write risk.
What withdrawing the candidate leaves behind
Keep a known baseline artifact and sufficient capacity to take traffic back. Check schema compatibility, messages already published and external side effects. Sending requests to the old code does not reverse a payment, restore deleted data or undo an incompatible migration.
Kubernetes Deployments support rolling updates, but precise weighted routing and automated canary analysis may require additional ingress, service-mesh or rollout tooling. Verify what the actual stack does rather than treating replica count as an exact division of requests.
Before increasing exposure, show that the candidate handled the behavior being changed and identify the conditions still untested. For the write-path release, representative write observations justify expansion while recovery checks explain how the team can withdraw the candidate; completing a rollout stage supplies neither answer by itself.
Some defects require scale, time, or interactions a small canary cannot reproduce. Keep those limits in the release decision and use additional controls where the untested consequence matters.
Promotion, insufficient evidence and failure
- State the change, expected effect, candidate population, baseline and stopping criteria.
- Check instrumentation and recovery through a safe test before exposure.
- Admit a limited relevant workload and compare outcomes over the agreed interval.
- Hold when evidence is insufficient; stop or mitigate when failure criteria are met.
- Increase exposure in defined steps and repeat the checks at each step.
The distinction between insufficient evidence and failure matters. A candidate that has not yet handled enough relevant work has not necessarily failed, but the comparison cannot support promotion. Waiting for a timer to expire does not fill that gap.
An unrelated incident can also change the workload or baseline enough to undermine the original comparison. Preserve the observations and pause when they no longer answer the planned question. This gives the team a chance to restore a meaningful comparison before broadening exposure.
Read diagram description
Candidate and baseline must receive comparable, meaningful work. A shared database can degrade both, so check the service objective as well as the difference. Insufficient evidence means hold; failure means stop or mitigate. Diagram labels: Bounded representative traffic: Choose assignment and preserve recovery capacity; Candidate: New artifact; Baseline: Known comparison artifact; Compare outcomes: User signals, counts, workload mix and absolute SLO; Decision: Promote in steps · hold · stop / mitigate.
Compare absolute service health as well as the difference between versions. A shared database can degrade both groups, making the candidate look similar to an unhealthy baseline. Relative agreement is useful only with enough context to know whether either group is meeting the service objective.
What remains untested after promotion
A successful canary establishes that monitored behavior passed under the tested conditions. It does not establish the absence of rare errors, compatibility with every customer configuration or safety at full scale. Continued monitoring and an appropriate recovery path remain necessary after promotion.
Review missed failure modes and improve future canaries where additional evidence is worth the delay. The release-gate guide connects these choices to policy, while the AI release-engineering article explains how model assistance can contribute without replacing independent checks.
For a write-path change, the promotion record should show which eligible writes reached the candidate, whether they completed correctly and what happened to their follow-on work. Read-only traffic leaves that question open however long the canary runs. Representative writes can support expansion, with the remaining scale, timing and recovery limits carried into the next stage.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- canarying chaptersre.google
