Skip to main content
●Measured savings across 11 LLMs, from Claude Opus 4.7 to Gemini Flash.→ See per-model data
Connect your client
ResearchToken costsMCP tooling

Agentic coding can use 1,000× more tokens than code chat — and where the tokens actually go

A Microsoft Research / Stanford paper from April 2026 quantified what most engineering teams feel but never measure: agentic coding tasks consume roughly 1,000× more tokens than chat on the same model. Input drives the cost. Token counts alone do not determine task accuracy.

James Hollingsworth(Contributor)Published 8 min read

The baseline nobody measures

Most teams running AI coding agents watch their Anthropic invoice climb and make intuitive guesses about why. They swap models. They shorten system prompts. They wonder if the agent is being “verbose.” What they rarely do is measure the actual token distribution across a real agentic run.

In April 2026, researchers at Microsoft Research and Stanford published a systematic study of token consumption across agentic coding benchmarks (arXiv:2604.22750). The headline number: agentic tasks consume approximately 1,000× more tokens than code-reasoning and code-chat baselines on the same model. This is a benchmark comparison, not a 1,000× dollar-cost estimate. Billing also depends on model rates, input/output mix, caching and retries.

As an illustration, at Claude Opus 4.7 standard rates of $5 per million input tokens and $25 per million output tokens, a 200-token chat exchange costs fractions of a cent. A single agentic coding task at 200,000 tokens costs roughly $1.00 to $5.00 depending on the input/output split. Run a hundred tasks a day and that becomes $100 to $500 daily, before any caching. At team scale, across a sprint, the numbers get uncomfortable fast.

The paper also found that individual runs for the same task differ by up to 30×. That variance matters if you are trying to budget, or if you are building a product that charges downstream users on token consumption.

Where the tokens go: input, not output

The study's second finding is the one that should change how you think about optimization: input tokens, not output tokens, dominate the cost in agentic coding workloads. This runs against the intuition that the expensive part is generating code.

Three input categories account for the bulk of consumption:

Input categoryRepresentative sizeSource
Tool-schema injection (MCP)47,300 tokens/turn in a 6-server / 120-tool sim; 17,600 for GitHub alone (94 tools); 10,000+ for Jira + ConfluencearXiv:2604.21816; Atlassian Labs
Stateless history replayFull conversation replayed each turn in most agent frameworks; compounds as tasks extendarXiv:2604.22750
Intermediate tool resultsFull output of every prior tool call fed back verbatim as context for subsequent callsarXiv:2604.22750

The tool-schema number deserves a concrete anchor. Atlassian Labs measured their MCP server in March 2026: the GitHub MCP server alone sent 17,600 tokens of tool descriptions per turn, across 94 tools. Adding Jira and Confluence pushed the overhead past 27,000 tokens before a single character of actual task context had been included. Their fix was lazy schema loading, which reported 70–97% lower schema overhead and almost no quality impact in their evaluation setup. Those results depend on the workload.

A more extreme case: an April 2026 MCPProxy case study measured a 14-server / 187-tool setup. At session start, the full tool manifest consumed 54,707 tokens, verified via Anthropic's count_tokens API. With a lazy meta-tool replacing eager injection, that dropped to 818 tokens: approximately 98.5% by arithmetic, although the source labels it 97%. The underlying tools were still accessible; they just were not described until the model explicitly asked for them.

Eager tool schemas are a predictable input overhead. Their cost depends on whether the client exposes the full catalog, discovers schemas on demand, and caches repeated prefixes. Measure the actual request payload and billing.

The accuracy-cost paradox

Here is the finding most teams have not absorbed: higher token usage does not reliably improve task accuracy.

The Microsoft Research / Stanford study found that accuracy peaks at an intermediate level of token consumption, then plateaus or declines. Past a certain threshold, you are not getting proportional return on your token spend. You are paying for context the model cannot effectively use.

This is consistent with what we know about context degradation. The “Tool Attention” paper (arXiv:2604.21816) simulates schema overhead for 120 tools across six servers. It reports a 95% reduction in schema tokens, from 47,300 to 2,400 per turn. Its end-to-end quality, latency and cost improvements are projections, not measured agent runs; its roughly 70% context-fill discussion is not a new universal degradation threshold established by this simulation.

The benchmark suggests diminishing returns; it does not prove that more spending universally causes worse results. The practical question for a team building with agents is where that level sits for their workload, and what fraction of their current token budget falls above it.

Model variability and the estimation problem

The study's third finding compounds the first two: different models consume very different token counts for the same task, and the models themselves are bad at predicting their own consumption.

The correlation between a model's self-estimated token usage and its actual usage was 0.39 or lower across the tasks studied. A model that tells you “this task will take about 50,000 tokens” is producing a number with limited predictive value. Building a routing layer on that self-report will not give you what you expect.

The 30× run-to-run variance compounds this further. Even with the same model on the same task, token consumption is not stable. Some of that variance comes from stochastic generation; some comes from branching in the task graph that sends the agent down longer or shorter paths. The practical implication: empirical measurement on your actual task distribution is the only reliable basis for cost modeling.

What it means for teams spending real money

Given that input drives the cost, and that spending more does not buy accuracy past a threshold, the obvious place to look is at what is going into the context. Three surfaces account for most of the compressible input in a typical agentic coding setup:

Tool-schema lazy-loading. The Atlassian and MCPProxy results both point at the same intervention: do not inject the full manifest of every connected MCP server on every turn. Inject a minimal descriptor and load full schemas on demand. Lower schema counts can reduce uncached input charges, but do not imply the same percentage reduction in total invoices. Include cache rates and discovery calls.

History compression. Most agent frameworks replay the full conversation on every turn. For a long-running task, that replay grows linearly as turns accumulate. A compressed history that preserves the causal structure of prior decisions — what was tried, what worked, what state changed — can reduce replay tokens. Evaluate task correctness and total billed usage; compression can also cause retries or longer outputs.

Retrieved-context semantic compression. Agents routinely pull whole files, documentation pages, or tool-call results into context. Most of that content is relevant in aggregate but dense in low-value detail. Semantic compression extracts the salient sections and discards the rest, with savings that scale with document size.

All three surfaces are candidates for reducing the input budget. Compare them with model selection and caching using the same tasks, quality gates and total billed usage; no study here establishes one universal best intervention.

gotcontext addresses all three. The gc_compress_manifest MCP tool compresses tools/list output directly; the compression and session-filtering tools handle the other two. If your team is spending material money on agentic workloads and has not measured the input composition, that is the place to start.

Method and sources

All quantitative claims in this post come from the cited primary sources below. No numbers were extrapolated or estimated by gotcontext.