I have been measuring what agents actually send to the model. The number that surprised me, even having built an agent layer, is the ratio. Roughly 70 percent of the tokens in a production agent call are not the user's question and not the answer. They are operational overhead - system prompts, tool definitions, retrieved context, accumulated history. The model reads all of it, you pay for all of it, and most of it is the same on turn 50 as it was on turn 1.

This is what I have started calling the stateless tax, and it is the single biggest reason agent bills are bigger than people expect. It also explains why adding an MCP server to an agent can double its token cost even when the server is doing useful work. This note is the math and the fixes.

Why the tax exists: APIs have no memory

LLM APIs are stateless. There is no session on the provider's side that remembers what you sent last turn. Every single call has to carry the full context the model needs to behave correctly: the system prompt with your guardrails, the definitions of every tool it is allowed to call, the retrieved documents for this query, and the entire conversation history so far. On turn 1 of a task that is a few thousand tokens. On turn 50 of a coding agent that has been working through a repo, it is the entire prior conversation re-sent, every turn, for 50 turns.

The compounding is the part people miss. A 50-turn run is not 50 times the cost of turn 1. It is turn 1 once, plus turn 1 plus turn 2 once, plus turn 1 plus turn 2 plus turn 3 once, and so on. That is roughly n-squared-over-two in history re-transmission cost. A long agentic loop does not scale linearly with its length. It scales quadratically.

Where the 70 percent goes

Typical token composition of an agentic turn
System prompts - 2K to 8K tokens (behavioral guardrails, safety)
Chain-of-thought - 1K to 10K tokens (hidden reasoning, planning)
Retrieved context / RAG - 10K to 100K tokens (vector DB chunks)
History accumulation - 100K+ tokens (all prior turns)
Value - the actual response delivered to the user

Add it up and the user-visible response is the small slice. Everything else is the cost of making the model behave the same way it behaved last turn. That is the stateless tax.

Why MCP makes the tool-definition tax bigger

MCP is a win for portability and a tax for tokens. An MCP server that exposes 30 tools to an agent contributes a JSON definition for each of those tools to the context on every turn - 5,000 to 15,000 tokens of tool schema that the model reads whether it calls the tool or not. A well-organized MCP server with 50 tools can be the single largest line item in your prompt before the user has even asked their question.

This is not an argument against MCP. We build and ship MCP servers - the protocol is the right abstraction for connecting agents to tools. It is an argument for how you expose them. The default - dump every tool the server knows about into the context on every turn - is the expensive default.

If your agent can see every tool on every turn, you are paying for the tools it did not call. The fix is not fewer tools. It is loading the ones the agent is likely to call.

The five fixes that actually compound

I am listing these in order of payoff, from biggest to smallest. The first two are the ones I see teams skip most often, and they are the ones that matter most.

1. Prefix caching - the 90 percent discount nobody takes

Most of your prompt is static. The system prompt does not change between turns. The tool definitions do not change. The guardrails do not change. The providers will cache that static prefix on their side and discount it up to 90 percent - if you give them a stable prefix to cache. If your framework reorders the prompt, injects dynamic content at the top, or rebuilds the tool list per call, you break the cache and pay full price every turn. Lock the static content at the absolute beginning of the prompt. This is free money and it is the largest single lever.

2. Just-in-time tool loading

Instead of sending all 50 tool definitions every turn, send the 5 the agent is likely to need. A local vector search over tool descriptions against the current task is enough to pick them, and the routing itself is cheap. This typically cuts the static tool overhead by about 80 percent per turn. For an agent with a large MCP surface area, this is the difference between a profitable agent and an unprofitable one.

3. Summarized memory rollups

The quadratic history cost is the one that eats long runs. The fix is compaction: an async process that summarizes earlier turns into a compact representation and replaces them in the context window. You keep the recent turns at full fidelity for the working memory the model needs, and you compress the older turns into a summary. This turns the quadratic history tax back into something close to linear.

4. Prompt compression on retrieved context

RAG chunks come back verbose. Tools like LLMLingua strip the redundant boilerplate from retrieved passages before they hit the context window, and they typically cut doc-chunk overhead 50 to 75 percent. The retrieved content loses almost no signal. The context window loses a third of its weight.

5. Speculative execution and model shielding

Run the agentic loop on a cheap, fast model. Escalate only the synthesis step - the judgment call where the frontier model actually earns its price - to the expensive one. Most of an agent's turns are routing, formatting, and tool-call plumbing. None of that needs a frontier model. Shielding the expensive model for the step that needs it is the lever that compounds with the caching and routing work above.

How this changes what I build

I think about this every time I add a tool to an agent now. The question used to be "does this tool make the agent more capable." The question now is "does this tool earn its per-turn token cost across every turn the agent runs." A tool the agent calls once in a hundred runs still costs you on the other ninety-nine, because the model reads its definition every turn whether it calls it or not.

The agent layer we built - Sugar - runs coding agents in the background, and background agents are exactly the workload where the stateless tax shows up on the invoice. Long runs, deep history, many tools. The fixes above are not theoretical. They are the difference between a queue that is cheap to run and one that quietly costs more than the work it produces.

Where the numbers come from

The 70 percent overhead figure and the token composition breakdown come from a Google Cloud executive briefing on AI tokenomics by Eric Lam. The Goldman Sachs forecast on agentic token consumption (roughly 24-fold) is from May 2026. The optimizations are the standard set in the literature - prefix caching, just-in-time tool loading, memory compaction, LLMLingua-style compression, and speculative routing - but the order and the payoff ranking are my own, from running coding agents and seeing where the money actually goes.