On this page5 sections
A long agent session can resend a large rendered context on every turn, but that does not mean every input token is billed or processed the same way. Prompt caching may reuse an exact prefix. The useful operational question is not whether a session is long. It is how much of each request matches a reusable prefix, where that match breaks, and whether the complete task is getting cheaper or merely harder to observe.
OpenAI's prompt-caching guide describes caching as automatic reuse of repeated prompt prefixes. The rendered input includes developer messages, tool definitions, prior messages, and the current turn. Keeping a conversation identifier does not by itself guarantee a hit. Exact prefix agreement is the mechanism that matters.
Model charges and responder effort need separate budgets, with the rollback exception preserved whichever cost is being reduced. For a coding or incident agent, cached input can lower model cost while the team still pays for retries, tool time, review, and a context that has become difficult to reason about.
Measure the cacheable prefix
Define a prompt-cache service-level indicator as cached input tokens divided by total input tokens for eligible requests. OpenAI reports cached-token details in the usage object, so the ratio can be computed per call and aggregated by workflow, model, prompt version, and session age. Track cache-write tokens where the selected model reports them.
Use distributions rather than one fleet average. A median can look healthy while the expensive longest sessions miss repeatedly. Report the ratio at several percentiles and for cohorts such as fresh sessions, sessions after tool-schema changes, sessions after compaction, and sessions restored from storage. Pair it with input tokens, latency, model cost, and task completion.
The SLI is a diagnostic, not an objective to maximize. A workflow can improve its hit rate by carrying an enormous stable prefix that contains irrelevant history. It can lower the ratio after compaction while reducing total input enough to save time and money. Read the ratio beside the total work.
Find where the prefix started changing
When the ratio drops, compare the rendered request structure, not only the visible user messages. A reordered tool list, revised developer instruction, dynamic timestamp, changing repository summary, or different serialization can alter an early token and make everything after it ineligible for the same prefix match.
Instrument stable segments with hashes. Record separate hashes for policy text, tool schema, repository context, durable conversation history, and the volatile tail. Do not log secrets or raw private content to make this work. A salted or access-controlled digest can show that a segment changed without copying the segment into a broadly readable telemetry system.
Place stable material first and volatile material later when the API and application semantics allow it. Keep tool definitions deterministically ordered. Version system instructions intentionally. Avoid inserting “current time” or a random request identifier near the front of every prompt when that value is only needed for a later tool call. These are request-construction changes, so verify that they preserve behavior before treating a higher hit rate as an improvement.
A prompt-cache key can help route related requests and separate accounting, but it should not be treated as a command that forces a hit. The rendered prefix still has to match. Alerting solely on key presence would certify the label rather than the cache outcome.
Connect cache misses to agent lifecycle events
Long-running agents cross boundaries that ordinary chat dashboards hide. They restart after a worker crash, compact history, add a tool, switch a model, refresh repository state, or hand work to another process. Emit a lifecycle event for each boundary and align it with the first request after the change.
Consider a hypothetical repair agent whose cached-token ratio falls from a normal band after the sixth turn. The user did not change the task. A trace shows that the host injected a newly generated list of repository files before the stable policy on every turn. The list's ordering varied with filesystem traversal. Sorting the list and placing the volatile worktree delta after stable instructions restores reusable prefixes in a controlled replay. This example is illustrative, not a measured OpenAI result.
The same symptom can require the opposite fix. If a repository changed substantially, reusing an old prefix may preserve stale assumptions. The right action is to invalidate or move the boundary, then measure whether fewer retries and corrections offset the new input. Cache efficiency never outranks correctness.
The site's supervised Codex worker guide treats restarts, cancellation, and unknown outcomes as host responsibilities. Add prompt construction and cache telemetry to that lifecycle record so a restart that changes request shape is visible rather than mistaken for random model latency.
Budget the complete task
Calculate model input cost from the provider's actual usage fields and current pricing rules. Then keep a second task budget for wall time, retries, tool calls, reviewer minutes, and failed attempts. The token-usage guide explains why the request that produced the final answer is only one part of agent cost.
For each completed task, retain at least the uncached input, cached input, cache writes if reported, output, number of model calls, tool duration, retries, and final disposition. A cache optimization passes when it reduces the intended resource without increasing correction work or losing decision-relevant context.
Do not set a universal cache-hit target. Workloads differ. A repetitive code review may have a large stable prefix. A research agent that continuously incorporates new documents should not be forced into the same band. Establish a baseline for one workflow and alert on unexplained changes within that workflow.
Use a small operational gate
A practical rollout can use three checks. First, the rendered-prefix hashes explain expected invalidations. Second, cached and total input move within the workflow's normal band under a fixed replay. Third, complete-task cost and correction effort do not regress. If only the cache ratio improves, hold the change.
Stop optimizing the prefix when preserving a cache hit would keep stale tools, policy, or repository evidence. That boundary changes the recommendation: accept the miss, record its cause, and judge the request by the resulting task outcome. A prompt-cache SLI is valuable because it makes reuse inspectable. It should never turn reuse into the goal of the agent.
Source context
Source links appear within this article. Read them alongside the author’s analysis and evaluate the guidance against your environment.
Report an error or outdated detail