News

Claude Sonnet 5.5 keeps token prices but changes the migration work

Sonnet 5.5 promises faster, cheaper tasks at unchanged token prices. Check the API changes and measure retries before switching operational workflows.

Nate Reuck3 min read

Sources
Layered coral and ivory paper shapes fit closely together with small offcuts, illustrating efficient use of material.
Original AI-generated conceptual illustration; not documentary evidence or a product interface.
On this page

Claude Sonnet 5.5 arrived on September 28 with a claim that matters to teams running many small coding jobs: lower cost per task without lower token prices. Anthropic says it generates output more than 30% faster than Sonnet 5 and costs up to 30% less per task in its testing because it uses fewer tokens. Those are vendor results, not savings measured by AIOpsSRE. Anthropic’s announcement

The published rates remain $2 per million input tokens, $10 per million output tokens and $0.20 per million cache-read tokens. Anthropic reports availability on its platform and through AWS, Google Cloud and Microsoft Azure; the Claude Platform model ID is claude-sonnet-5-5. Its positioning is everyday, well-scoped work such as bug fixes, with Opus remaining its choice for more complex, open-ended work.

Check the request before comparing the answer

Existing integrations need more attention than a model-name swap. The Sonnet 5.5 migration guide says thinking.type: disabled now returns an error. The replacement for keeping up-front thinking off is between_tools, at low, medium or high effort. Responses can still contain thinking blocks, so a client must read blocks by type instead of assuming the first block contains answer text.

Forced tool choices of type any or tool are also rejected. If the workflow depends on a forced tool call, changing only the model name is not a valid migration. Moving to auto means the model may answer without calling a tool. The application therefore needs to recognize that outcome instead of treating every response as a completed lookup or action.

A task may include an initial attempt, retries and corrections before acceptance. Cost comparison includes all attempts, not only the successful final response.
Unchanged token rates can still produce a different bill. Count the work required to reach an accepted result.

Compare a finished job, including its corrections

Consider a hypothetical job that drafts fixes for recurring configuration errors. Give each model the same failing configuration and ask for a patch. Run the normal validation against the patch, record any additional attempts, and have the usual reviewer decide whether it is usable. A faster first answer helps only to the extent that it shortens this complete job.

Count the attempts that produced rejected or incomplete work when comparing the cost of an accepted result. Keep the prompt, tool permissions and acceptance test fixed for the initial comparison. Report how many jobs passed alongside total model spend and review time; otherwise a lower bill can simply describe a model that finished less work. Our task-level token-cost guide explains why retries belong in that accounting.

Start with a small set of routine jobs that your team already knows how to assess. Confirm the request format works, then compare effort settings deliberately. Sonnet 5.5 may make those jobs cheaper and quicker; this experiment shows whether that advantage survives your tools, validation and review process.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

A useful next step

Continue the work