Explainer

Long agent sessions need a prompt-cache SLI

In brief

Measure cached input against total input, locate the prefix miss boundary, and judge cache changes beside latency, task cost and correction work.

5 min read

Sources
A long paper ribbon passes through a printmaking frame, repeating a blue pattern until one splice changes the downstream pattern to yellow.
Original AI-generated conceptual illustration; not documentary photography or an OpenAI product image.
On this page5 sections

A long agent session can resend a large rendered context on every turn, but that does not mean every input token is billed or processed the same way. Prompt caching may reuse an exact prefix. The useful operational question is not whether a session is long. It is how much of each request matches a reusable prefix, where that match breaks, and whether the complete task is getting cheaper or merely harder to observe.

OpenAI's prompt-caching guide describes caching as automatic reuse of repeated prompt prefixes. The rendered input includes developer messages, tool definitions, prior messages, and the current turn. Keeping a conversation identifier does not by itself guarantee a hit. Exact prefix agreement is the mechanism that matters.

Model charges and responder effort need separate budgets, with the rollback exception preserved whichever cost is being reduced. For a coding or incident agent, cached input can lower model cost while the team still pays for retries, tool time, review, and a context that has become difficult to reason about.

Measure the cacheable prefix

Define a prompt-cache service-level indicator as cached input tokens divided by total input tokens for eligible requests. OpenAI reports cached-token details in the usage object, so the ratio can be computed per call and aggregated by workflow, model, prompt version, and session age. Track cache-write tokens where the selected model reports them.

Use distributions rather than one fleet average. A median can look healthy while the expensive longest sessions miss repeatedly. Report the ratio at several percentiles and for cohorts such as fresh sessions, sessions after tool-schema changes, sessions after compaction, and sessions restored from storage. Pair it with input tokens, latency, model cost, and task completion.

The SLI is a diagnostic, not an objective to maximize. A workflow can improve its hit rate by carrying an enormous stable prefix that contains irrelevant history. It can lower the ratio after compaction while reducing total input enough to save time and money. Read the ratio beside the total work.

Two requests reuse a stable prompt prefix, while an early change moves the cache boundary and makes later input new; the prompt-cache SLI is cached tokens divided by total input tokens.
Prompt caching follows the exact rendered prefix. An early change can move the miss boundary even when the conversation still looks like one continuous session.

Find where the prefix started changing

When the ratio drops, compare the rendered request structure, not only the visible user messages. A reordered tool list, revised developer instruction, dynamic timestamp, changing repository summary, or different serialization can alter an early token and make everything after it ineligible for the same prefix match.

Instrument stable segments with hashes. Record separate hashes for policy text, tool schema, repository context, durable conversation history, and the volatile tail. Do not log secrets or raw private content to make this work. A salted or access-controlled digest can show that a segment changed without copying the segment into a broadly readable telemetry system.

Place stable material first and volatile material later when the API and application semantics allow it. Keep tool definitions deterministically ordered. Version system instructions intentionally. Avoid inserting “current time” or a random request identifier near the front of every prompt when that value is only needed for a later tool call. These are request-construction changes, so verify that they preserve behavior before treating a higher hit rate as an improvement.

A prompt-cache key can help route related requests and separate accounting, but it should not be treated as a command that forces a hit. The rendered prefix still has to match. Alerting solely on key presence would certify the label rather than the cache outcome.

Connect cache misses to agent lifecycle events

Long-running agents cross boundaries that ordinary chat dashboards hide. They restart after a worker crash, compact history, add a tool, switch a model, refresh repository state, or hand work to another process. Emit a lifecycle event for each boundary and align it with the first request after the change.

Consider a hypothetical repair agent whose cached-token ratio falls from a normal band after the sixth turn. The user did not change the task. A trace shows that the host injected a newly generated list of repository files before the stable policy on every turn. The list's ordering varied with filesystem traversal. Sorting the list and placing the volatile worktree delta after stable instructions restores reusable prefixes in a controlled replay. This example is illustrative, not a measured OpenAI result.

The same symptom can require the opposite fix. If a repository changed substantially, reusing an old prefix may preserve stale assumptions. The right action is to invalidate or move the boundary, then measure whether fewer retries and corrections offset the new input. Cache efficiency never outranks correctness.

The site's supervised Codex worker guide treats restarts, cancellation, and unknown outcomes as host responsibilities. Add prompt construction and cache telemetry to that lifecycle record so a restart that changes request shape is visible rather than mistaken for random model latency.

Budget the complete task

Calculate model input cost from the provider's actual usage fields and current pricing rules. Then keep a second task budget for wall time, retries, tool calls, reviewer minutes, and failed attempts. The token-usage guide explains why the request that produced the final answer is only one part of agent cost.

For each completed task, retain at least the uncached input, cached input, cache writes if reported, output, number of model calls, tool duration, retries, and final disposition. A cache optimization passes when it reduces the intended resource without increasing correction work or losing decision-relevant context.

Do not set a universal cache-hit target. Workloads differ. A repetitive code review may have a large stable prefix. A research agent that continuously incorporates new documents should not be forced into the same band. Establish a baseline for one workflow and alert on unexplained changes within that workflow.

Use a small operational gate

A practical rollout can use three checks. First, the rendered-prefix hashes explain expected invalidations. Second, cached and total input move within the workflow's normal band under a fixed replay. Third, complete-task cost and correction effort do not regress. If only the cache ratio improves, hold the change.

Stop optimizing the prefix when preserving a cache hit would keep stale tools, policy, or repository evidence. That boundary changes the recommendation: accept the miss, record its cause, and judge the request by the resulting task outcome. A prompt-cache SLI is valuable because it makes reuse inspectable. It should never turn reuse into the goal of the agent.

Source context

Source links appear within this article. Read them alongside the author’s analysis and evaluate the guidance against your environment.

Report an error or outdated detail

More on OpenAI

All OpenAI coverage

Related reading

Explore a related question