All news
aiproduct

Study: Cleaner Code May Cut Coding Agent Token Costs

11 Jul 2026

Researchers ran a controlled study to test a question many founders building on coding agents have likely wondered about informally: does the cleanliness of a codebase actually affect how efficiently an AI coding agent performs? The answer, based on early results, appears to be yes—at least for efficiency metrics.

What the study found

Using a minimal-pair design across six repository pairs, researchers evaluated Claude Code on 33 tasks, running 660 total trials. Each pair presumably compared a "cleaner" version of a codebase against a less clean counterpart, though the report does not specify the exact criteria used to define cleanliness.

The results showed that when Claude Code operated on cleaner code, it:

  • Used 7–8% fewer tokens
  • Revisited files 34% less often

The researchers frame code cleanliness as a factor that materially affects agent behavior—placing it alongside more commonly discussed variables like model choice, harness, and prompting strategy.

What's still unclear

Several important questions remain open. The report does not specify:

  • What specific metrics defined "clean" versus "unclean" code
  • Which repositories or programming languages were used
  • Whether task success rate or correctness was measured, or only efficiency (token usage and revisitation)
  • Whether the results are statistically significant, or how much variance existed across trials
  • Whether these findings would generalize beyond Claude Code to other agents
  • When the study was published

Notably, the absence of success-rate data means it's unclear whether cleaner code actually helps agents complete tasks correctly, or merely helps them do so with fewer tokens and less backtracking. Efficiency and correctness are not the same thing, and founders should be cautious about conflating the two based on this data alone.

Risks and limitations

The study has a narrow scope: it relies on a single agent (Claude Code) and a limited set of 33 tasks, which may limit how broadly the findings apply to other agents or to larger, more complex codebases. There's also a reasonable possibility that coding agents working on messy or legacy codebases—common in many real-world startups—could incur even higher token costs and inefficiency than what this study captured, given that the cleaner-code condition already showed measurable gains.

Why founders should care

For teams building products on top of coding agents, or using them internally for development, this study suggests—though does not conclusively prove—that codebase quality could be a meaningful lever for controlling operational costs. If token usage and file revisitation scale with cleanliness in the way this study indicates, founders running agents at scale may want to consider:

  • Investing in refactoring or code-quality tooling as a potential way to reduce token consumption, since the study's results point in that direction, though the magnitude of savings at scale is not yet established.
  • Exploring linting or code-quality integrations alongside coding agent workflows, as a plausible efficiency lever worth testing rather than a guaranteed fix.
  • Treating code cleanliness as a benchmarking variable, similar to model choice or prompt design, when evaluating or building coding agent tools—this could inform how startups building agent infrastructure design their own benchmarks.

That said, because the study is limited to one agent, 33 tasks, and does not report success-rate outcomes, these findings should be treated as preliminary directional evidence rather than an operational rule. Founders would likely benefit from monitoring whether follow-up research confirms these patterns across other agents, larger codebases, and—critically—whether cleaner code also improves task correctness, not just efficiency.

Sources