All news
aiproduct

Rich Sutton Warns of 'One-Step Trap' in AI Prediction

13 Jul 2026

Reinforcement learning pioneer Rich Sutton published an essay on X on July 18, 2024, warning AI researchers about what he calls the 'one-step trap'—a structural flaw he argues undermines many predictive AI systems that rely on chaining together single-step predictions to forecast the future.

What Sutton Argues

Sutton's core claim is deceptively simple: if a one-step prediction model were perfectly accurate, you could theoretically chain it forward to generate perfect longer-term predictions. But in practice, one-step predictions are never perfectly accurate—and once that assumption fails, Sutton says "all bets are off" for the reliability of any longer-term forecast built on top of them.

The problem compounds in stochastic environments. Sutton notes that in a stochastic world—or under a stochastic policy—the future isn't a single trajectory but a branching tree of possibilities. Computing long-term predictions from one-step models across that tree carries exponential computational complexity relative to the length of the prediction, making the approach generally infeasible at scale, according to Sutton.

Despite this, Sutton calls one-step models "hopeless, yet extremely appealing," pointing out they remain widely used in POMDPs (partially observable Markov decision processes), Bayesian analyses, control theory, and compression-based theories of AI.

Sutton's Proposed Fix

Sutton proposes that the way out of the trap is building temporally abstract models of the world using options and General Value Functions (GVFs)—rather than relying purely on iterated one-step predictions. This isn't a new idea for Sutton: it builds on a research line stretching back over two decades.

Timeline of Related Work

  • 1999: Sutton, Precup, and Singh publish foundational work on temporal abstraction in reinforcement learning.
  • 2011: Sutton and colleagues publish the Horde architecture paper.
  • 2023: Sutton and colleagues publish a paper on reward-respecting subtasks for model-based RL.
  • July 18, 2024: Sutton publishes the one-step trap essay on X.

What's Missing

The essay itself doesn't include empirical evidence or experiments quantifying the exponential complexity claim, nor does it specify concretely how options and GVFs mitigate the trap in practice. There's also no indication yet of reactions or critiques from the broader AI research community, and it remains unclear how directly this argument applies to large-scale systems like LLMs versus classical RL and control settings.

Why Founders Should Care

For founders building AI products—particularly in reinforcement learning, planning, forecasting, or control—Sutton's essay raises questions worth taking seriously, even without hard data attached:

  • If your system depends on chaining one-step predictions to make long-horizon decisions, accuracy may plausibly degrade as the prediction horizon grows, per Sutton's reasoning.
  • Strong short-term model performance may not reliably imply trustworthy long-term forecasting—a distinction that could matter for products marketed on long-range planning capability.
  • As prediction horizons extend, computational cost—not just accuracy—could become a practical bottleneck, according to Sutton's complexity argument.
  • Startups working on planning-heavy AI may find it worthwhile to evaluate temporal abstraction techniques (options, GVFs) as a potential alternative architecture, building on Sutton's 1999, 2011, and 2023 research lines.

None of this is confirmed by independent benchmarks in the report—Sutton's essay is a conceptual argument, not a study—but for teams whose products hinge on long-horizon AI predictions, it's a signal worth factoring into architecture decisions and due diligence conversations with technical advisors.

Sources