Quadratic Token Cost Growth in Multi-Step Agent Runs
Agent loops bill you quadratically, not linearly, because each step replays accumulated history.

Multi-step agent runs cost more than engineering teams expect because the standard way of estimating that cost, multiplying a single call's price by the number of steps, is structurally wrong. An engineer pricing out a ten-step agent task reaches for the same math used to price ten separate API calls: take the cost of one call, multiply by ten, and treat the result as a budget. That arithmetic holds for independent requests. It fails for agents, because every call in an agent loop carries the full accumulated history of everything that came before it. A ten-step run isn't ten units of cost. It follows a triangular progression, where each new step rebills every token of context the run has already generated.
The gap this produces isn't marginal. A twenty-step loop in which each step appends even a modest, non-trivial number of tokens can produce a total input-token bill that exceeds the naive per-step estimate by an order of magnitude. The surprise compounds because agents, unlike chatbots, don't make one call per user turn. They make dozens of calls per task: planning calls, tool calls, verification calls, retries. An unconstrained software-engineering agent working a single task can run up several dollars in model fees alone, a figure that looks alarming only once the accumulation mechanism that produces it is made explicit. That mechanism is the subject of the rest of this piece.
The precise mechanism: how accumulated history turns linear steps into quadratic spend
The quadratic growth isn't a loose analogy borrowed from computer science to sound rigorous. It follows directly from how stateless LLM APIs work: every call is independent, carrying no memory of its own, so the entire conversation so far has to be resent as input on every single turn. Most agent frameworks implement this the simplest way possible, appending each new model response and tool result onto the growing prompt and sending the whole thing back in on the next step. That design choice, reasonable as a starting point, produces prompt growth with no ceiling.
A worked example from a context engineering benchmark makes the effect concrete. In a naive ten-step file-reading agent loop, where each step reads a file and appends the result to history before the next call goes out, the total input tokens processed across the run run many times higher than what a single-pass version of the same task would require. The gap between those two numbers, a full history versus a constrained window, is the entire argument of this piece, expressed as tokens.
For a naive N-step loop, total input-token cost works out to Total_naive = N×S + u×N(N+1)/2 + r×N(N-1)/2, where S is the fixed system prompt resent every call, u is the new input tokens introduced at each step, and r is the output tokens generated at each step. The N×S term grows linearly and is the term most engineers intuit correctly. The N(N+1)/2 and N(N-1)/2 terms are triangular numbers, and they are where the real cost lives. Expressed as a continuous approximation, cumulative prompt-token cost over n steps becomes T_prompt(n) = n·s₀ + p·n(n−1)/2, which is Θ(n²) with leading coefficient p/2, where p is the average tokens appended per hop and s₀ is the fixed base prompt. Practitioners have taken to calling this the State Snowball, a name that captures the mechanism better than the math does: each step doesn't just add its own cost, it adds its cost to every step that follows.
Two hidden multipliers that make the quadratic problem worse than the formula implies
Even engineers who have internalized the N² formula and budgeted for it tend to underspend their estimate, because two further cost drivers sit on top of the base accumulation and are easy to miss.
The first is tool-definition overhead carried in by the protocol an agent uses to connect to external tools. That overhead is invisible to anyone reasoning only about message-history growth, because it doesn't accumulate as new content added over turns, but sits in the system prompt from the first call. It inflates the fixed S term in the formula directly: the baseline every subsequent step is multiplied against is larger than the system prompt alone would suggest, and it stays inflated for the life of the session regardless of how disciplined the history management is.
The second is repository context files, the AGENTS.md-style documents increasingly common in coding agent setups. These files are meant to improve agent behavior by supplying project conventions and guidance up front. In practice, they can substantially increase inference cost per session while delivering minimal improvement in task outcomes, and on complex tasks they sometimes reduce success rates. The mechanism is that richer upfront context encourages agents to explore more broadly, branching into more tool calls and more reasoning steps than a leaner prompt would prompt them to take. That exploration costs tokens whether or not it leads anywhere useful, and on harder tasks it frequently doesn't.
Both multipliers operate on the fixed and semi-fixed parts of the cost equation, compounding against the quadratic term that is already dominating the bill. Neither appears on a cost dashboard as its own line item. They raise the total above what the formula alone predicts, and they're the reason teams who've done the quadratic math correctly still get surprised by the invoice.
How unconstrained context growth silently degrades agent quality, not just cost
The accumulation driving up cost is the same accumulation degrading the quality of the agent's decisions, and it does both at once. Chroma Research's technical report on what researchers term "context rot," cited in a separate work on harness token economics called "The Harness Effect," documents that model performance degrades as input length increases, and that this degradation is often surprising and non-uniform across the eighteen models the report tested. A longer context doesn't just cost more to process. It measurably changes what the model attends to and how reliably it reasons over the material it's been given, so every additional token appended to a snowballing history is simultaneously a cost increment and a reliability risk.
What makes this dangerous in production rather than merely wasteful is that the failures it produces are silent. As entropy accumulates in a long-running agent's context, the agent keeps returning success signals even as its internal reasoning quietly deteriorates. No exception fires. No error gets logged. The system reports the same confidence it always has, even as the substance underneath that confidence erodes. Three specific patterns follow from this. Instruction burial happens when the original task goal is still technically present in the context window but gets buried under accumulated tool output, so the agent deprioritizes it in later steps while it remains in context. Silent truncation happens when the context grows past the model's window limit and the framework quietly drops the oldest messages, with the agent carrying on as though nothing was lost, because nothing in its interface tells it otherwise. Looping happens when an agent gets stuck retrying the same failed action, each attempt appending more tokens to the history until a timeout or a budget cap ends the run, again with no exception raised along the way.
None of this is abstract risk. A benchmark analysis of coding agent runs found that the majority of total tokens consumed came from tool results alone, and that roughly half of those tokens were removable with no loss in task performance. That figure says something precise: much of what gets billed, and much of what gets loaded into the model's attention matrix on every subsequent call, is waste. The State Snowball accumulates noise at a cost, and that noise actively works against the reasoning it's supposed to support.
Why prompt caching, the standard first response, only partially solves the problem
Prompt caching is the fix most engineering teams reach for first, and it's a legitimate one. The problem is that it addresses only one term in the cost formula and leaves the dominant one untouched. On Claude Sonnet 4.6, a cache hit on the input prefix costs $0.30 per million tokens against a standard input price of $3.00 per million tokens, a 90 percent discount that makes repeatedly resending a stable system prompt dramatically cheaper than it would be uncached. That's a real saving, and any team running multi-step agents should be claiming it.
What caching can't touch is the conversation history itself. Each new tool output, each new reasoning trace, each new turn of accumulated context is unique to that point in the run, generated fresh and never seen before, so there's nothing for a cache to hit against. In the formula, caching reduces the N×S term, the fixed system prompt resent on every call, but it leaves the N(N+1)/2 triangular term completely intact. That triangular term is the actual cost trap this piece has been describing, and it's the part of the bill that caching was never designed to solve. A framework built by TokenPilot illustrates the point by treating the problem as having two genuinely separate parts: a static prefix, addressed through what the paper calls Ingestion-Aware Compaction to stabilize cache hits, and a dynamic history, addressed separately through Lifecycle-Aware Eviction. Both pieces are treated as significant, independent cost drivers in the framework's own ablations, which found that stabilizing the prefix alone cut cost from $8.31 to $4.35 in one test, with lifecycle eviction contributing further savings on top of that. Caching and history management solve different halves of the same problem, and a team that implements only the first half will watch the uncached, ever-growing second half eventually dominate the bill at any run length worth worrying about.
Five structural patterns that constrain quadratic growth at the source
Caching changes how existing context gets priced. The durable fix changes what gets put into the context window in the first place, and five patterns address that at different points in the architecture, from the least invasive to the most structurally complete.
The first and simplest is sliding window context. Instead of carrying forward the full history of a run, the agent keeps only a fixed number of recent steps, for instance the two most recent iterations, and discards the rest. The file-reading benchmark described earlier showed this approach producing a small fraction of the token volume a naive full-history loop generates, for comparable task output. It's the easiest pattern to bolt onto an existing agent, and the tradeoff is equally easy to state: a run that depends on something established many steps earlier will lose access to it once the window slides past that point, which makes this pattern a poor fit for tasks with long dependency chains.
The second is selective tool-output retention. Rather than appending raw tool output to history wholesale, the agent summarizes each result immediately after receiving it, keeping only the essential content and discarding the rest before the next call goes out. A benchmark analysis found a substantial share of tool-result tokens removable with no effect on task performance, which makes this plausibly the single highest-yield change available to most agents, because it targets exactly that category of token spend shown to be waste.
The third is dynamic turn limits paired with loop detection. Rather than letting a run continue until it exhausts a fixed budget or hits a timeout, the system estimates the probability that further iterations will succeed and exits once that probability drops low enough to make continuing not worth the spend. Where the first two patterns work by shrinking p, the average tokens appended per step, this pattern works by capping N itself, the number of steps the triangular term gets to compound across, making it a complement to windowing.
The fourth is server-side context compaction. Anthropic's reference implementation for this pattern automatically summarizes older context once a run approaches its window limit, and in a 100-turn web search evaluation this reduced token consumption by 84 percent while letting the agent complete workflows that would otherwise have hit the context ceiling. In one reported case, a context that had grown past a hundred thousand tokens was collapsed down to a few thousand. This pattern directly addresses the accumulation of dynamic history, but compaction itself requires an LLM call to generate the summary, trading some of the cost it saves for additional latency and a modeling overhead that has to be accounted for.
The fifth and most structurally transformative is coordinator-specialist design, sometimes called isolated subagent architecture. Instead of one agent accumulating history across an entire run, a coordinator delegates narrow pieces of the task to specialist subagents, each of which starts with a clean, narrow context of its own. Because history doesn't accumulate across the full run, only within each specialist's bounded slice of it, this pattern doesn't shrink the coefficient on the quadratic term the way the other four do. It eliminates the term structurally, at the cost of being the most invasive pattern to implement, since it requires redesigning the agent's topology and adjusting its memory management. A routing layer that classifies incoming task complexity and sends simple queries to a fast, cheap path while reserving the coordinator-specialist path for complex ones prevents the quadratic regime from activating at all on the tasks that don't need it. Applied together rather than selected from individually, these five patterns have been shown in production deployments to cut agent costs substantially, and they're best understood as a stack to layer.
The compression trap: when aggressive context reduction itself becomes a cost and quality risk
Every pattern in the previous section works by removing something from the context window, and that's precisely where the risk reenters. A sliding window that's too narrow discards the one piece of information a later step actually needed, producing the same instruction burial and silent degradation this piece described as a consequence of context rot, this time self-inflicted by the mitigation. Selective tool-output retention carries an equivalent risk in the other direction: a summarization step aggressive enough to meaningfully cut tokens is also aggressive enough to drop a detail that later reasoning depends on, and because the agent has no way of knowing what it no longer has, that loss produces a quiet wrong answer rather than a flagged failure, the same way context overflow does.
Server-side compaction carries its own version of the same tension, compounded by the fact that the summarization call itself costs tokens and adds latency: a compaction policy tuned to trigger too often can erode the very savings it's meant to produce while adding round-trip delay to every run it touches. Dynamic turn limits, if tuned to exit too early, will terminate runs that would have succeeded with a few more steps, trading a real cost saving for a real drop in task completion. None of these risks argue against the patterns described in the previous section. They argue for treating every token removed from context as a design decision with a measurable quality cost attached to it, not a free efficiency gain. The teams that get this right are the ones who track completion rate alongside token spend when they tune any of these five patterns, because a cheaper run that fails is a failed run that happened to cost less to fail.
Sources
- The Hidden Economics of AI Agents: Managing Token Costs and Latency Trade-offs
- TokenPilot: Cache-Efficient Context Management for LLM Agents
- The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
- Token Economics for LLM Agents: A Dual-View Study from Computing and Economics
- Silent Failure in LLM Agent Systems: The Entropy Principle and the Inevitable Disorder of Autonomous Agents
- How to optimize token efficiency in agentic systems
- ICML Poster ACON: Optimizing Context Compression for Long-horizon LLM Agents
- Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents

