Opus 5 Tops SlopCodeBench, Still Fails Full Challenges
28 Jul 2026
Opus 5 posted the best score yet on SlopCodeBench, a long-horizon coding benchmark from @GOrlanski's lab at UW Madison — but the results underline how far agentic coding still has to go before founders can trust it to run unsupervised.
What happened
SlopCodeBench, created in March 2026, tests models on multi-checkpoint coding challenges designed to catch both correctness failures and "slop" — sloppy, low-quality code that technically passes but isn't clean. In the original paper, GPT-5.4 scored an 11% strict pass rate and Opus 4.6 scored 17%.
In a newer, smaller test, three models — Opus 4.8, Sonnet 5, and Opus 5 — were run against a subset of the benchmark: 3 problems and 17 checkpoints total (8 easy, from circuit_eval; 5 medium, from database_migration; and 4 hard, from dynamic_config_service_api).
Opus 5 came out on top with a 24% pass rate, passing 4 of the 17 checkpoints strictly. Opus 4.8 and Sonnet 5 each managed just 1 of 17. Opus 5 also wrote roughly five times as many functions/callables as Opus 4.8 across the same challenges.
Despite Opus 5's lead, none of the three models completed any challenge with all checkpoints passing — not even the easy-difficulty problem.
The code-quality picture
SlopCodeBench tracks 41 code quality metrics alongside pass/fail results, using a library of over 200 Python slop detectors (76 for a TypeScript subset). Across all three models, the share of code lines flagged by at least one slop rule was high: 98% for Opus 4.8, 93% for Opus 5, and 89% for Sonnet 5.
Verbose-flagged lines also crept upward as challenges progressed — from about 65% at checkpoint 1 to 80% by checkpoint 8 — across all models, though the report doesn't explain why.
As the report's author summarized: "every dollar bought correctness. nobody bought enough of it."
What's still unclear
A few gaps in the data are worth flagging. The original paper reports results for Opus 4.6, while this new test reports results for Opus 4.8 — it isn't clear whether these are the same model line, different versions, or why the comparison shifted. It's also unclear whether the original 11%/17% figures were measured on the same 17-checkpoint subset or the full benchmark, and the total size of the full SlopCodeBench benchmark beyond this subset isn't stated. No cost or token-usage data accompanies the author's dollar-cost comment, and there's no explanation for the rising verbose-flag trend.
Why founders should care
These results are based on a small sample — 17 checkpoints across just 3 problems — so they likely offer directional rather than definitive signal. Still, a few patterns are probably worth founders' attention:
- The jump from Opus 4.6's 17% to Opus 5's 24% strict pass rate may suggest incremental, not breakthrough, progress on long-horizon coding correctness.
- Since no model fully passed any challenge — including the easy one — founders building autonomous coding agents may still need human review loops rather than lights-off automation, at least for now.
- Persistently high slop-flag rates (89–98%) across all models suggest that code quality and correctness may be separate axes worth tracking independently when evaluating coding agents.
- Because SlopCodeBench appears unsaturated, it could serve as a useful ongoing benchmark for founders deciding which models to integrate into coding-agent tooling as new versions ship.
Given the small sample size, it's reasonable to treat these specific numbers as an early signal rather than a settled ranking — but the underlying pattern of high effort, low completion is one founders evaluating AI coding tools may want to watch closely.