Prompt Caching Effects on Agent Loop Cost and Behavior
Prompt caching can slash agent costs, but only with the right prompt structure.

Agent loop costs grow faster than their token counts imply because every additional step compounds the cost of every step that came before it. A ten-step agent task does not just read its system prompt and tool definitions once; it reads them ten times, and it pays full price each time unless something intervenes.
Why agent loop costs grow faster than token counts
The arithmetic of an agent loop is unforgiving. Each turn resends the conversation history built up so far, so the model pays again for the same system prompt, the same tool schemas, and the same accumulated context on every call. A task that takes ten steps does not cost ten times what a one-step task costs; it costs more, because the prompt itself grows with each turn while the stable portions get recomputed from scratch every time. This is a quadratic dynamic, not a linear one, and it sits beneath the sticker price most teams quote when they estimate what an agent will cost to run at scale.
The imbalance becomes sharper when input and output tokens are weighed against each other. Manus, during the period before and through its brief acquisition by Meta (an acquisition that Chinese regulators subsequently forced Meta to unwind), reported an average input-to-output token ratio of around 100 to 1. A cost model that treats input and output symmetrically will miscalculate where the actual leverage is, and it will send engineering effort toward output optimization when the input side is where the money is being spent.
The quadratic growth in cost tracks a parallel growth in latency. The cost and latency problems share the same underlying redundancy, so solving one without addressing the other leaves half the stakes on the table.
What prompt caching covers
Prompt caching exists to eliminate that redundancy. During the prefill stage of inference, the model has to compute, for every token in the prompt and across every attention layer, a key and value matrix projection. Once a prefix has been cached, an identical prefix in a subsequent request can skip straight to decoding, bypassing the expensive prefill step for the tokens that match, and billing at a fraction of the standard input rate.
The match has to be exact. Semantic caching returns a stored response for a prompt judged similar in meaning, but because conversation history and tool outputs accumulate with each turn, the full prompt in an agent loop is never exactly the same twice, so semantic caching cannot work for agent loops. Prefix caching matches token sequences rather than meanings, and that is what fixes this problem.
The three major providers (OpenAI, Anthropic, and Google) have arrived at structurally different implementations of this mechanism, a divergence the "Don't Break the Cache" study examined directly by evaluating full-context caching, system-prompt-only caching, and a strategy that excludes dynamic tool results across all three, using the DeepResearch Bench multi-turn agentic benchmark. One structural detail carries forward into later engineering decisions: Anthropic's cache has a two-tier architecture with a sharp size threshold, and prefixes that fall below that threshold see their hit rate climb only toward a sub-maximum plateau below the theoretical ceiling. That threshold matters later on, because it decides whether a compression strategy applied to a cached prefix helps the cache or quietly undermines it.
Savings from correct caching
The upside of getting this right is large, and the gap between doing it well and doing it badly is larger still. Deriv's production AI assistant, Amy, reached a cache-hit rate well above four in five, and that cut input-token costs by a large majority. The feature itself is not where the value lives. The structure of the prompt around it is.
ProjectDiscovery's experience makes the same point from a different angle. ProjectDiscovery saw a very large volume of tokens served from cache instead of recomputed, driving a substantial drop in overall LLM cost. Thomson Reuters Labs reported a 60% cost reduction on its document analysis pipeline, a separate data point that reinforces the same pattern: when the prefix is engineered correctly, the savings are not marginal.
Across the "Don't Break the Cache" study's more than 500 agent sessions, cost savings varied substantially from one provider to another, and time to first token also improved meaningfully whenever a cache hit occurred, tying the cost benefit to a user-facing latency benefit. But the benefit is conditional. On Anthropic's model, a cache-write premium applies, and if the hit rate on an otherwise-stable prompt stays too low, the cost of writing to the cache exceeds what the reads save. Turning caching on without restructuring the prompt around it can therefore produce a net cost increase rather than a reduction, which is the opposite of what the feature promises on paper.
The savings also apply only to the input side of a request. Workloads that are output-heavy and built around short prompts will see little benefit from this, however carefully the prefix is structured, because there is little prefix there to cache to begin with.
The four structural rules that determine whether a prefix stays cacheable
Cache hit rates are set almost entirely by decisions made before a request is ever sent. The ordering principle holds across providers: arrange content from least variable to most variable, with the live user message placed last.
The first rule follows directly from that principle. Stable tool definitions come first, then system instructions, then reference documents shared across conversations, then conversation-specific history, and finally the current user message, in that order. Deriv applied this rule concretely by separating Amy's stable instructions into a static file passed through the pipeline as static_instruction, with dynamic context rendered separately through a DYNAMIC_CONTEXT_TEMPLATE. The stable foundation stayed byte-identical across requests, letting the cache treat it as a repeatable prefix each time.
The second rule concerns rendering, not content. Repeated content has to produce identical token sequences every time you call it, because the cache cannot tell a meaningful change from an accidental one. The OpenAI Codex CLI shows the pattern done correctly: system instructions and tool definitions stay identical and consistently ordered between requests, and when sandbox configuration or environment context changes mid-conversation, Codex appends a new message, leaving the original one untouched. The prefix is never mutated, so it never stops matching.
The third rule is the one behind both the Deriv and ProjectDiscovery results. Working memory is any state that changes from turn to turn, such as intermediate reasoning, retrieved documents, or tool outputs, and it belongs after the cache boundary, not inside it. ProjectDiscovery's roughly 60-fold cost difference came from exactly this change: the same token volume and the same workload, with working memory moved from inside the system prompt to a trailing user message.
The fourth rule concerns tools specifically. Tool definitions are serialized near the front of the prompt, so any change to them invalidates everything that follows. Dynamic tool discovery through protocols such as MCP, where the set of available tools can shift depending on which servers are connected or what the runtime context looks like, can reset an entire cached prefix mid-session without anyone intending it to. If you are building on dynamic tool sets, load new tools only at session boundaries, since doing it mid-session trades a small flexibility gain for a cache reset that can erase far more value than it adds.
Anthropic adds one more constraint: its cache breakpoints search backward through at most 20 content blocks. A single agentic turn that issues many parallel tool calls can exceed that limit on its own, forcing a full cache miss even when nothing about the underlying prompt has changed in a way a human would call meaningful.
Three silent failure modes that break caching without raising an error
Caching fails quietly. There is no error thrown when a cache is missed, no warning raised when a write premium is being paid without a matching read benefit, and no flag when a provider-side anomaly is distorting the numbers. A team can believe its caching setup is working when it is doing nothing, or when it is actively making costs worse.
The first failure mode involves TTL expiry during slow tool execution. If an agent's steps include long-running tool calls, pauses for human review, or polling against queued jobs, the gap between one request and the next can exceed the provider's TTL window even though the underlying prefix would otherwise still match perfectly. The fix is to match TTL tiers to request frequency: short-TTL tiers for high-QPS workloads, longer-TTL tiers for medium-frequency agents despite the higher write premium that comes with them, and extended-retention options for slow-burn batch agents where requests are spaced far apart.
The second failure mode comes from caching everything by default. Caching the full context, including dynamic tool results that will never recur, triggers cache writes for content with essentially no chance of being reused, and that overhead can offset or exceed whatever the cache reads save. Dynamic tool calls and results change on nearly every turn, so caching them charges you for a write with close to zero chance of a matching hit, and that write becomes a pure cost.
Cache hits are probabilistic at the provider level, so a provider-side anomaly can cause an entire model-route pair to get zero cached tokens across every phase of a workload for an extended stretch. The "Don't Break the Cache" study documented exactly this: one model-route pair received zero cached tokens across every phase for three days, a provider-side anomaly that would silently distort any cost comparison or A/B test run during that window without anyone realizing the baseline itself had shifted. Hit-rate metrics alone cannot confirm that a caching setup is functioning; they can also reflect a temporary provider condition that has nothing to do with how the prompt is built.
A related and more consequential problem is the cache-invalidation bug, where the cache keeps serving a previous version's behavior after a prompt prefix has changed, continuing to do so until the TTL expires on its own. The only way to catch it is to compare the agent's actual trace against what the current prompt says it should be doing. Detecting this class of failure requires active comparison, not passive monitoring of hit rates or latency.
How caching interacts with prompt compression
People market prompt caching and prompt compression as cost-reduction techniques, and teams often assume they stack. In practice, the dominant family of compression methods actively undermines caching. Query-aware compression, the standard approach in the compression literature, produces a different compressed prefix for every query by design, since the compression is built around what that specific query needs. So the cache requires an exact token-for-token match, but a query-aware method guarantees the prefix will never be the same twice, and that mechanically invalidates the prefix cache on every single call.
The resolution is to decouple compression from the query. Query-agnostic compression produces the same token sequence no matter what the current query asks, so you can cache the compressed prefix once and reuse it across calls. Cache-Aware Prompt Compression, or CAPC, pairs this query-agnostic approach with explicit cache control markers and a tier-preserving ratio bound, so the compression cannot get so aggressive that it pushes the cached prefix below the size threshold where hit rates start to degrade. Tested on a standard long-context benchmark, CAPC produced the cheapest outcome in every document-size and compression-ratio configuration examined.
That ratio bound exists because of the two-tier architecture described earlier. If you compress a prefix down below roughly 3,500 tokens on Anthropic's cache, it falls into the tier where hit rates plateau below the theoretical maximum, so a compression strategy can shrink token count even as it reduces the value the cache would otherwise deliver. Anyone compressing a long system prompt or tool schema needs to check two things before applying it: whether the compressed form will stay above the provider's caching-efficient size threshold, and whether the compression method is query-agnostic enough to keep the prefix stable across calls. Skipping either check tends to produce exactly the outcome the Sonnet 4.6 study found, a technique that looks like savings on paper and costs more in practice.
Why caching changes agent behavior
Caching can keep an agent appearing to function normally while it is quietly reasoning from the wrong instructions, which is the deeper risk in all of this. A cache-invalidation bug produces no error message and no latency spike that would draw attention to it; the agent simply continues to act on a previous version of its prompt until the TTL expires and the stale entry finally falls out of rotation. From the outside, and even from most standard monitoring dashboards, that agent looks correct. Its outputs are coherent, its tool calls succeed, and its cost profile looks healthy, because the token counts billed under a cache hit are smaller than they would be otherwise. What has actually happened is that the system has been executing against instructions nobody currently intends for it to follow, and the only way to catch that gap is to compare the agent's behavior against what the live prompt says it should be doing, turn by turn, rather than trusting that a stable hit rate and a lower bill mean the system is behaving as designed.