CrucibleBench: Evaluating LLMs in a Text-Based MUD
24 Jul 2026
A telnet-era world for testing frontier AI
CrucibleBench is a proof-of-concept benchmark that evaluates large language models not through photorealistic simulation or browser automation, but through a text-based MUD (multi-user dungeon) — the kind of command-line, RPG-style virtual world that predates the modern web. The project describes this choice as "lateral thinking with withered technology," leaning on decades-old interactive fiction infrastructure to probe how models behave under social and rule-based pressure.
The environment itself is deliberately compact: 7 command types, 12 rooms, and 14 items, populated by 4 NPCs whose trust/suspicion state ranges from 0 to 100. Models are dropped into this persistent world with hidden social objectives, and their behavior — including how they navigate rooms, interact with NPCs, and handle dialogue — is logged and scored.
For context, MUDs are text-only virtual worlds, historically accessed via telnet clients, where players are represented as avatars inside "rooms" and environments can simulate time, weather, economies, and NPC schedules. Interactive fiction is essentially the single-player cousin of the format. Scholars like Rheingold and Turkle have described MUDs as imaginary collaborative worlds and early forms of community and literature — a lineage CrucibleBench is explicitly drawing on rather than building a new 3D simulation from scratch.
What the testing found
CrucibleBench ran 50 runs per model (5 seeds × 2 objectives × 5 repetitions) and released 650 transcripts as a full artifact for outside review. The results raise some flags about both model behavior and evaluation methodology:
- Judge inconsistency: Per-model agreement with an independent judge varied widely, from 21.7% to 84.8% — a spread the report says suggests judge-dependent scoring may not be consistent across models.
- Judge sensitivity: Swapping in a single LLM-judge component reordered the leaderboard by up to 6 positions, and removing classifier-dependent dimensions shifted 6 rankings beyond scenario-sampling noise (at 90% paired block bootstrap confidence).
- Model behavior quirks: Dialogue looping showed up in 14–66% of frontier runs, and Grok 4 exhibited wrong-room interaction in 12% of cases — both flagged as potential signals of reliability or navigation/state-tracking limitations in current frontier models.
The project's own framing is notably self-critical: it argues that benchmarks using LLM judges should report per-subject agreement and ranking stability under judge ablation — not just aggregate reliability scores. In other words, a single leaderboard number may be hiding a lot of judge-dependent noise.
What's missing
Several details remain unclear from the current release. It's not specified which LLMs beyond Grok 4 were tested, or how many models were included in the 50-runs-per-model methodology. The report also doesn't explain how hidden social objectives are defined or scored, or precisely what the NPC "trust and suspicion state" measures in practice. Funding sources, team size, and a concrete timeline for a next phase aren't detailed either — only that a provisional budget of $3,500 has been floated. Sources also don't clarify how a related background piece on MUD history and telepresence connects authorship-wise to the CrucibleBench project itself.
What's next
The project is framed explicitly as a proof-of-concept, with an invitation for outside parties to "fund it, build it, or run it" as part of a Phase 2 that would turn the current setup into a fuller benchmark. No specific timeline for Phase 2's start has been given.
Why founders should care
If your product roadmap includes evaluating or selecting LLMs — for agents, customer-facing tools, or internal automation — this report suggests a few things worth weighing:
- Benchmarks that rely on a single LLM judge may be more fragile than they appear: a 21.7%–84.8% agreement spread and leaderboard shifts of up to 6 positions from one judge swap indicate that rankings could plausibly change depending on which evaluation method you trust.
- Frontier models showing dialogue looping in up to two-thirds of runs, and specific navigation errors like Grok 4's 12% wrong-room rate, may signal that reliability issues persist even in top-tier models under sustained, multi-turn interaction — something founders building long-running agents should probably test for directly rather than assume away.
- The low-cost, text-based MUD approach hints that founders with tight budgets might not need expensive simulation pipelines to stress-test model behavior; a simplified, rule-based environment could offer a reasonably informative — if narrower — signal.
- Because the project is openly inviting funders, builders, and operators for Phase 2, founders working on AI evaluation tooling may have an early window to shape how judge-based benchmarking standards evolve, rather than reacting to them later.
As with any early-stage benchmark, these are directional signals rather than settled conclusions — the report itself flags several methodological gaps that make definitive claims premature.