Quesma's Deep-Research Pipeline Hits Token Limits Fast
20 Jul 2026
Quesma is digging into a question that matters to every founder building with AI: what does agentic coding actually cost? Their answer came fast—and painfully. The company's first attempt at a multi-agent deep-research pipeline burned through the entire Claude Max 5x plan limit in just 30 minutes.
What happened
The initial run launched 111 agents and queued 123 claims for verification. Before the token limit hit, only 25 of those claims had actually been verified—leaving the vast majority of the research unvalidated.
Rather than scale back ambition, Quesma optimized the pipeline. They extended the claude-mem plugin to support Codex and Antigravity with shared memory, and began spreading workload across multiple paid subscriptions. That shift alone let research run roughly 10x longer than relying on Fable subscription alone.
The combined effect: research runtime stretched from 30 minutes to a few hours before hitting a limit again.
The final setup
After optimization, a subsequent /deep-research run used just 61 agents—nearly half the original count—and completed in 22 minutes. One week later, the resulting knowledge base held hundreds of validated notes.
For context, the report also flags Headroom, a project with 56k stars, though its exact relationship to Quesma's pipeline isn't detailed in available sources.
Why founders should care
This experiment is a live case study in the hidden costs of scaling agentic AI systems, and it carries a few probable lessons:
- Agent count doesn't always correlate with output quality. Dropping from 111 to 61 agents while also cutting runtime suggests leaner pipelines may likely perform as well or better than brute-force scaling—an important cost lever for teams burning through token budgets.
- Verification is likely a structural bottleneck. With only 25 of 123 claims verified in the initial run, teams building research or fact-checking agents should probably plan verification capacity as carefully as generation capacity, or risk shipping unvalidated output at scale.
- Multi-subscription resource pooling may be a viable cost-management strategy. The 10x runtime extension from spreading load across subscriptions suggests founders operating token-hungry agent systems could reduce single-point token exhaustion by diversifying providers—though this likely increases operational complexity and billing overhead.
- Architecture changes can meaningfully shift token economics. The jump from 30 minutes to several hours of runtime, achieved through plugin extension and shared memory, indicates that engineering investment in memory/context-sharing may pay off faster than simply buying more compute.
What's still unclear
Several open questions remain in the reporting: what "Fable" is exactly and how it relates to the other subscriptions used, the dollar or token cost of the initial versus final runs, how the Headroom project connects to this work, the specific criteria used for claim verification, and what changes beyond the claude-mem plugin extension drove the runtime improvement.
Bottom line
Quesma's experience is a useful data point for any founder building multi-agent AI products: token limits arrive faster than expected, verification pipelines can lag badly behind generation, and thoughtful architecture—not just more agents or more subscriptions—is probably the more sustainable path to scaling agentic research.