Wattage: A Cost-Regression Gate for AI Agent Tokens
28 Jul 2026
What launched
Wattage has launched as a token-spend profiler and cost-regression gate for AI agents. The tool analyzes agent traces to identify where tokens are being burned or wasted, prices each waste pattern in real dollars, and prescribes fixes. Notably, it can fail a CI build if code changes make an agent more expensive to run — treating cost the way teams already treat test failures or security regressions.
How it works
Wattage runs fully offline, requiring no config file or API key, and operates directly on OTLP JSON trace exports. At its core is a convergence engine designed to catch "non-convergence" patterns that quietly inflate costs: agent loops that make no progress, retries that reappear with fresh timestamps, oscillation between competing strategies, and stalls that look productive but aren't.
The platform ships with eight detectors for trace analysis, discovered through a Python entry-point group — meaning new detectors can be added without touching the core pipeline. This gives teams an extensible framework to build custom checks for their own agent stacks.
For reporting, Wattage generates a Token Efficiency grade on a 0–100 scale, suitable for README badges or CI gates. Its cost-regression gate, invoked via wattage ci, fails builds when agents regress past configured thresholds and posts per-detector delta tables as pull-request comments. Output is available in SARIF and JUnit XML formats, easing integration into existing CI dashboards and code review workflows.
The numbers behind the claims
Wattage's benchmarking so far rests on a small dataset: 10 hand-reviewed, labeled synthetic loops. Within that set, a simulated fix for a "prefix_churn" pattern — using prompt caching — produced a 44.7% cost reduction, cutting cost from $0.000199 to $0.000110 in one captured trace example.
These are compelling numbers, but they come from a limited sample. The benchmark size and the single-trace cost-savings example mean the headline figures should be read as illustrative rather than representative of typical production savings.
Risks and open questions
Several gaps are worth flagging for anyone evaluating the tool:
- The 10-loop benchmark may not generalize to the messier, more varied behavior of production agents.
- The 44.7% cost-reduction figure is drawn from one trace example, not an aggregate or repeated test.
- Offline-only operation, while privacy-friendly, could complicate integration with cloud-based trace pipelines that require authentication.
- Reliance on OTLP JSON exports may limit compatibility with agents that aren't instrumented in this format.
The report also notes several missing details: there's no confirmed release date or version history, no disclosed methodology for how the 10-loop benchmark was constructed, no pricing or licensing information, no data on real-world adoption or independent validation, and no comparison against existing agent-observability or cost-monitoring tools.
Why founders should care
For teams running AI agents in production, token waste is likely an underappreciated cost center — one that could scale unnoticed as usage grows. Wattage's approach of pricing waste patterns in dollars and gating CI on cost regressions may offer a low-friction way to catch these issues early, before they compound.
The offline, no-API-key design could plausibly lower the barrier to adoption for teams wary of sending trace data to third-party services, though this remains a design choice rather than a proven security guarantee. Similarly, the entry-point-based detector architecture suggests founders could extend the tool for domain-specific agent behaviors, which may be valuable for teams with non-standard agent architectures.
That said, the small benchmark size means founders should treat the cost-savings claims as directional rather than definitive. Teams considering Wattage would likely benefit from validating its detectors and savings estimates against their own production traces rather than relying solely on the reported figures.
Bottom line
Wattage introduces a novel framing — treating token cost as a CI-gated metric alongside tests and security checks — which could resonate with founders trying to control AI infrastructure spend. But with benchmarking still limited to a small synthetic dataset and key details like pricing, adoption, and methodology unspecified, this is a tool to watch and test rather than adopt on faith.