LeMario: A JEPA World Model Learns Super Mario Physics
16 Jul 2026
LeMario: Testing What a JEPA World Model Actually Learns
A new experiment, dubbed LeMario, applies a Joint-Embedding Predictive Architecture (JEPA) — a self-supervised "world model" approach — to Super Mario Bros. The goal: see whether a model trained purely on pixels and button inputs can learn the game's underlying dynamics well enough to predict outcomes and plan actions, without any hand-coded rules about jumping, gravity, or level layout.
What was built
LeMario reproduces an architecture called LeWorldModel, designed to learn dynamics directly from raw pixels and action inputs. The team trained it on:
- 737,134 frames across 280 episodes and 32 Mario levels
- A vision encoder that compresses each frame into a 192-dimensional latent representation
- Observation pairs separated by 5 emulator frames
- 6 possible button states (Left, Right, Up, Down, A, B)
- A causal predictor built from 6 transformer blocks
The pipeline mirrors a now-familiar recipe in world-model research: encode the visual world into a compact latent space, then predict how that latent space evolves under different actions.
Does it actually use the actions?
One core risk in this kind of architecture is representation collapse — a failure mode where the model's latent vectors all become similar or identical, making predictions look deceptively accurate while carrying no real information. To rule this out, the team ran an action-shuffling ablation: they scrambled which button inputs were paired with which frames and measured how much worse the predictions got.
The results suggest the model is genuinely using action information:
- Shuffling actions increased one-step prediction error by 20.2%
- Over five recursive steps, shuffled-action error was 47.5% worse than with correct actions
- Across those same five steps, LeMario beat a simple persistence baseline by 45.5%
This is a meaningful sanity check: if scrambling the actions hadn't hurt performance, it would have signaled the model was ignoring inputs and just pattern-matching on visuals alone.
Where the latent space is strong — and where it isn't
To test whether the 192-dimensional latent space captured Mario's actual position in the world, the team trained simple probes to recover his coordinates from the latent representation:
- Horizontal position: MAE of 9.30 pixels, R² of 0.997 — a strong, near-linear recovery
- Vertical position: MAE of 21.62 pixels, R² of only 0.188 — a much weaker signal
The report does not explain why vertical position is so much harder to recover than horizontal position, flagging it as an open question. It's a notable asymmetry: the model appears to track where Mario is along the level far better than how high he is — which matters a lot in a game built around jumping.
Planning: probes help, but limits remain
The most interesting results come from planning experiments, where the model was asked to steer Mario from a starting position toward a goal frame using the Cross-Entropy Method (CEM):
- Raw JEPA + CEM: starting at x=40 with a goal at x=72, the model only reached x=44 — far short of the target
- Probe-scored CEM: using the position probe to score candidate action sequences, the model reached x=71 against a goal of x=72 — a dramatic improvement
- Local replanning: iteratively replanning in shorter segments, the model reached x=176 against a goal of x=177
In short, raw latent-space planning alone performed poorly, but adding an auxiliary probe signal — and replanning frequently — closed the gap almost entirely on these test cases.
Despite these gains, the model failed to reliably jump over the first major obstacle in test levels and struggled to navigate toward distant goal images, pointing to limits in handling complex game mechanics and long-horizon planning.
What's still unclear
The report leaves several open questions: there's no information on training duration, compute resources, or hardware used; no comparison against other world models beyond the persistence baseline; no explanation for the vertical-position probing gap; and no detail on which levels were used for training versus evaluation. There's also no timeline indicating when this project was conducted.
Why founders should care
For founders building on top of self-supervised world models — whether for robotics, simulation, gaming, or agentic AI — this report offers a few probabilistic signals worth weighing:
- It's plausible that JEPA-style world models can be adapted to game or simulation environments for short-horizon prediction tasks, given the strong one-step and five-step results here.
- The gap between raw JEPA+CEM planning (x=44) and probe-scored planning (x=71) suggests it's likely that auxiliary supervision signals meaningfully improve planning outcomes in latent-space world models — a pattern that could generalize beyond Mario.
- The obstacle-jumping and long-horizon navigation failures suggest current world models may still face real limitations for complex, sequential decision-making applications, and founders should be cautious about assuming these architectures generalize to harder control problems without further validation.
- The action-shuffling ablation is a useful template: it's a reasonably cheap and important validation step for any team building predictive or world models, to confirm the model is actually using its inputs rather than exploiting shortcuts that only look accurate.
No funding, product launch, or commercial application is mentioned in connection with LeMario — this is a research-stage architecture test, not a product announcement. But the methodology, particularly around probing latent representations and validating action-sensitivity, offers a useful blueprint for teams evaluating their own world-model-based systems.