Fable 5 Beats GPT-5.6 Sol on NP-Hard Routing Test
20 Jul 2026
A newly published benchmark compares Fable 5 against GPT-5.6 Sol on a notoriously hard combinatorial optimization task — and Fable 5 comes out ahead, with its native "/goal" mode adding a modest but inconsistent boost.
The setup: an NP-hard routing puzzle from Paris
The test problem, referred to as the KIRO problem, involves routing across 532 terminals spread over 11 distribution hubs in Paris. It's the kind of task where brute-force search is a non-starter: one restricted solution family alone — built from exactly 19 loops of 28 terminals each — has a search space of roughly 10^1223 possible configurations.
The author has personal history with this exact problem, having first encountered it as an engineering student in 2018 and later spending about a week writing C++ to solve it. That background frames the benchmark as a deliberate stress test rather than a generic coding exercise.
The results
Both Fable 5 and GPT-5.6 Sol were run three matched times each, with and without /goal mode enabled, and the full results — code, prompts, tables, exclusions, and trajectory notes — were published on CLIArena.
Key numbers from the report:
- Fable 5's plain (non-goal) mean score beat Sol's plain mean by 1,875 points.
- With /goal mode enabled, Fable 5's mean advantage grew to 1,984 points.
- Fable 5 was notably more consistent: its plain-mode scores stayed within a 319-point range, while Sol's plain-mode scores spanned 1,958 points.
- Fable 5's best overall result — 31,934 — was achieved in goal mode.
- Across the head-to-head comparison, goal mode won 4 of 6 trials.
Fable 5 produced the best overall solution in the benchmark, and it did so more reliably than Sol across repeated runs.
/goal mode: helpful, but not a guaranteed win
The /goal feature is described by the author as changing the model's control loop and search path — not a simple "try harder" switch. That framing matters, because the results back it up only partially. While goal mode won a majority of trials, it wasn't uniformly beneficial: in some runs, it gave a weak solution path more time to develop, which caused large score regressions. The author's own characterization is that /goal is "not a game changer" — a claim that sits somewhat uneasily next to the 4-of-6 win rate, and the report does not fully reconcile the two.
Caveats that matter
Several factors complicate a clean read of these results:
- Infrastructure mismatch: the test containers exposed 8 CPUs despite task metadata declaring only 1, which may have unfairly advantaged Fable 5's parallel search strategies.
- Sequential testing drift: runs went through subscription services one after another, raising the possibility of performance drift between trials that could affect comparability.
- Small sample size: with only three runs per model, statistical confidence in the results is limited.
Sources do not quantify how much the CPU discrepancy affected final scores, nor do they explain the scoring units behind the reported point gaps — so the headline numbers should be read as directional rather than precise.
Why founders should care
For founders evaluating AI coding or reasoning tools, this benchmark is a useful — if narrow — data point, not a verdict:
- The performance gap between Fable 5 and Sol on this task may indicate real differences in how each model handles large combinatorial search spaces, but with only three trials per model, that signal should be treated as suggestive rather than conclusive.
- The uneven payoff from /goal mode suggests that specialized "reasoning" or "try harder" features could plausibly underperform standard operation in some cases, even if they win more often than not — a pattern worth testing on your own workloads before assuming a feature toggle is a strict upgrade.
- The CPU allocation discrepancy is a reminder that benchmark environment configuration can meaningfully skew comparative results, especially for tasks that reward parallel search. Founders comparing vendor claims should ask how test infrastructure was configured, not just what scores were reported.
- Because the full benchmark artifacts — code, prompts, tables, and trajectory notes — are published on CLIArena, teams with their own NP-hard or optimization-heavy workloads have a realistic opportunity to independently verify these findings or adapt the test harness to their own use case, rather than relying solely on the summary scores.
Bottom line
Fable 5 outperformed GPT-5.6 Sol on this single, hard optimization benchmark, both with and without its /goal feature, and did so with tighter score consistency. But the result rests on a small number of trials, an acknowledged infrastructure imbalance, and an unresolved tension between the model's win rate and the author's own characterization of /goal as incremental rather than transformative. Founders should treat this as one useful signal among many — worth digging into via the published artifacts — rather than a definitive ranking of model capability.