Claude Code vs OpenCode: 33k vs 7k Token Overhead
13 Jul 2026
The hidden token tax of coding agents
A benchmark comparing two popular coding-agent harnesses — Claude Code and OpenCode — found a substantial gap in how many tokens each burns before a user's prompt even reaches the model. Testing was conducted with Claude Code 2.1.207 and OpenCode 1.17.18, both pinned to claude-sonnet-4-5, in July 2026.
The headline figure: Claude Code loaded roughly 33,000 tokens of system prompt, tool schemas, and injected scaffolding before any user input arrived. OpenCode, running the same underlying model, used about 7,000 tokens for the equivalent setup.
Where the overhead comes from
The report breaks the baseline down into two main components:
- System prompt size: Claude Code's system prompt runs 26,891 characters (~6,500 tokens); OpenCode's is 8,811 characters (~2,000 tokens).
- Tool schemas: Claude Code's 27 tools account for roughly 24,000 of its ~33,000 baseline tokens. OpenCode's 10 classic coding tools account for roughly 4,800 of its ~6,900 baseline.
On actual tasks, the gap showed up in metered usage. For one task (T2), Claude Code made 6 HTTP requests totaling ~199,000 cumulative input tokens, while OpenCode made 4 requests totaling ~41,000 tokens. On another task (T3), OpenCode made 9 tool calls versus Claude Code's 3 — a reminder that tool-call count and token volume don't always move in the same direction. Separately, Claude Code was reported to write up to 54x more cache tokens than OpenCode on the same task.
Overhead compounds fast
The baseline numbers are just a starting point. The report notes several additive factors in real-world setups:
- A 72KB instruction file in a production repo adds ~20,000 tokens to every request.
- Five modest MCP servers add 5,000–7,000 tokens.
- Put together, a realistic working setup can already be 75,000–85,000 tokens deep before a user types anything.
- A local LLM gateway can add its own ~6,200 token envelope overhead.
The most dramatic multiplier came from subagent fan-out: a task that cost 121,000 tokens when run directly cost 513,000 tokens when the same work was split across two subagents.
Why founders should care
For teams building or heavily using agentic coding tools, these figures suggest — though don't prove in every context — that harness choice may meaningfully affect metered costs. A few probabilistic takeaways:
- Teams running high-volume agentic workflows may find that a leaner harness (like OpenCode's reported lower baseline) reduces per-request costs, though actual savings likely depend on task type and usage patterns.
- The 54x cache-token gap could indicate that similar-looking tasks produce very different bills depending on harness, making it worth benchmarking before committing to one tool at scale.
- Multi-agent architectures may carry a real cost penalty for simple tasks — the 121k-to-513k jump suggests fan-out patterns should probably be tested before being adopted broadly, rather than assumed to be cost-neutral.
- Instruction files and MCP server count appear to be tunable levers: trimming either could plausibly lower baseline overhead, based on the additive figures reported.
- Founders operating in the EU should note that Article 12 of the EU AI Act is expected to require logging and understanding of agentic system behavior — a requirement that high, opaque token overhead could make harder to satisfy without added observability tooling.
What's still unclear
The report leaves some gaps worth flagging. The exact content and difficulty of the benchmarked tasks (T1, T2, T3) aren't described beyond character and token counts, and no dollar-cost or latency figures accompany the token numbers. It's also not established whether these findings generalize beyond claude-sonnet-4-5, how many test runs were performed, or whether the different overhead figures (33k baseline, 75k–85k "real setup," 20k instruction file, 5–7k MCP) fully reconcile into one consistent total. Founders should treat these numbers as directional signals rather than definitive cost models — and consider running their own benchmarks against their actual workloads before switching harnesses or architectures.