On this page4 sections
The final answer is only a paragraph. Behind it, an agent may have read ten documents, resent conversation history, called tools and retried a failed step. The visible response makes a poor meter for that work. A useful token budget follows the task from its first request to a checked result, including the human effort still needed to use it.

What gets counted before the short answer appears?
A tokenizer converts text into tokens: units that may be words, word fragments, punctuation or other sequences. The mapping depends on the tokenizer and language. Word count can therefore help estimate size, but it cannot establish exact billed usage or whether the request fits a particular model’s context limit.
The user’s question is only part of the input. Instructions, tool schemas, retrieved passages, prior conversation and tool results may also go to the model. Generated content adds output usage; reasoning, images and audio can have provider-specific accounting. The endpoint’s actual usage fields are the place to determine which categories apply.
For an agent task, those records need to be added across every model call, including retries, with cached input retained separately when reported. Otherwise the cost of an extended investigation can disappear behind the size of its final paragraph. A task identifier provides the connection between that paragraph and the attempts that produced it.
From token counts to a per-attempt charge
With per-million-token rates, input cost is input tokens × input rate ÷ 1,000,000; output cost uses the corresponding output count and rate. Add the two, then include separately priced cache, reasoning, tool or storage usage as applicable. The worked diagram uses hypothetical prices so its arithmetic can be inspected without treating it as a current vendor quote.
Which part of the request grows also matters. A longer prompt may leave output unchanged, reuse cached computation or cross a pricing tier; tools may contribute charges of their own. Following all those categories to the accepted task keeps the comparison about the result the team needed, rather than the cheapest individual call.
A ledger should connect the task to all its attempts, including failures that never produce a visible answer. Billed usage and estimates remain distinct. Where a failed call lacks a usage record, the charge stays unresolved until reconciliation rather than becoming zero. Applying that treatment consistently keeps missing accounting from looking like a candidate integration’s advantage.
Read diagram description
Hypothetical rates from the worked example: 12,000 input tokens at $2 per million plus 1,000 output tokens at $8 per million cost $0.032 per attempt. Four identical attempts cost $0.128 before other charges. These are not vendor prices. Diagram labels: Input: 12,000 tokens: 12,000 × $2 / 1,000,000 = $0.024; Output: 1,000 tokens: 1,000 × $8 / 1,000,000 = $0.008; One attempt: $0.032: Input cost + output cost; Four attempts: $0.128: Include failed attempts and retries in the task ledger; Task outcome: Useful accepted result, or unresolved work.
Latency and context limits affect the same task
Time to first token, complete-response time and tool latency describe different waits. A response can start quickly and finish after the incident decision it was supposed to support. Queueing, throttling, retrieval and sequential tools can dominate elapsed time even when the prompt is short. A token budget alone does not explain those delays.
The application also needs known behavior at its context limit, the amount of material it can send in a request. Depending on the API and integration, oversized input may fail, truncate or trigger summarization. A non-sensitive fixture near that boundary lets the team inspect what is actually sent and which behavior follows.
Summarization can let work continue while losing an exception or uncertainty that changes the recommendation. Stable constraints and source references should remain accessible outside the rolling summary, with a test that the assistant can recover the evidence required for its decision. The ten-record rollback case gives that test something more concrete than a target reduction in tokens.
Model charges and responder effort need separate budgets, with the rollback exception preserved whichever cost is being reduced. The ledger and quality review must therefore refer to the same task: they show whether a smaller model bill came with more source reconstruction for the responder.
Exploratory analysis may justify more budget than a routine status summary; a universal cap can discard the evidence that makes a difficult decision defensible. The useful limit follows the job’s purpose rather than the desire to make every request equally small.
Limits on calls, elapsed time and repeated work
Before execution, set purpose-appropriate limits on model calls, elapsed time, output length and tool-result size. A loop repeating an unsuccessful investigation without new evidence should eventually stop with a clear account of attempts and unresolved questions. That partial result gives the responder a starting point for continuing, instead of consuming the rest of the budget to repeat itself.
For the ten-record rollback case, retrieval changes need both the usage ledger and a check that the exception survives. The operator prompting guide explains how explicit observations and constraints make that check possible. Use the current model and endpoint documentation to interpret the ledger: Anthropic’s Opus 4.6 release description describes capability and context options, while actual endpoint pricing needs rechecking before a purchase.
Running the original and shorter contexts against the same exception-bearing case reveals whether the cheaper request still retrieves the rollback condition. The comparison needs usage records from every attempt and the effort spent rereading sources. A reduction in both is a clear saving. When one rises as the other falls, the accounting shows the tradeoff instead of hiding it behind the length of the final answer.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- Opus 4.6 release descriptionwww.anthropic.com
