All news
aiproduct

Encord, Zander Labs Trial Brain Waves for Robot Training

28 Jul 2026

The data bottleneck behind humanoid robots

Encord, a company that builds data tooling for training AI models, is running a trial with German neuroscience startup Zander Labs to test whether brain-wave headsets can generate better training data for robotics models. The trial is happening at Encord's San Leandro facility, where roughly a dozen human "pilots" — Encord's term for its robotic trainers — maneuver robotic arms while wearing headsets that measure brain activity.

Encord was originally founded to help machine-vision companies annotate data and evaluate models. It has since expanded into collecting egocentric data from factories around the globe, using San Leandro as a testbed for new data modalities — brain waves being the latest.

Why brain waves, and why now

The logic behind the trial comes from Lucas Gehrke, the Zander neuroscientist supervising the work: the amount of brain activity a pilot exhibits during a task may offer clues to model builders about when a robot needs to switch into a higher-effort mode. In other words, human cognitive load during a task could become a training signal for teaching robots when a task is hard versus routine.

That's a meaningful proposition if it holds up, but Encord's own plan treats it as unproven. According to the report, the company intends to build an initial brain-wave-tagged data set, run it through customer robotics models, and evaluate whether it actually improves performance — before deciding whether to scale the approach at all.

"The data simply does not exist"

Underpinning the experiment is a blunt assessment from Vineeth Velmurugan, Encord's head of robot learning and a veteran of OpenAI's robot lab and Berkshire Grey: robotics training data, at the volume and quality needed, largely doesn't exist yet. Velmurugan estimates the field needs a data set roughly five times the size of YouTube's entire video corpus to make a real breakthrough in robotics — a staggering figure that underscores how early this market still is.

He also pointed to specific demand: "Every humanoid company has asked us for these pieces," he said, referring to leader-follower robotic arm data, suggesting a concentrated and urgent need among humanoid robotics builders for this particular kind of training data.

The economics of data quality

The trial also surfaces a cost-versus-value tradeoff that founders building in this space will recognize. Velmurugan estimates that densely annotated data — the more rigorously labeled kind — is worth around 100 times as much as "junky" ego data for training specific tasks. But that quality comes at a price: dense annotation costs roughly 20 times more to produce than the lower-quality alternative.

That gap — 100x value against 20x cost — is the kind of math that could shape how robotics and AI founders prioritize their own data pipelines: investing more per data point may still pay off if it meaningfully accelerates model performance, but it's a bet on quality over volume that not every startup can afford to make.

What's still unknown

The report leaves several open questions. There's no detail yet on how large the initial brain-wave-tagged data set will be, how long the trial will run, or which customer robotics models will actually be tested against it. It's also unclear how "improved performance" will be measured, or what the brain-wave headsets themselves are technically capable of. And even if the trial succeeds, there's no stated timeline or cost estimate for scaling brain-wave data collection beyond San Leandro.

There are also open technical risks worth flagging: it's not yet established whether brain activity reliably correlates with task difficulty or the effort a model actually needs to expend, and the high cost of dense annotation could limit how much brain-wave-tagged data gets produced even if the signal proves valuable.

Why founders should care

For founders building in robotics, physical AI, or adjacent data infrastructure, this trial is a signal worth watching rather than a proven playbook. It's likely that data scarcity — not model architecture — remains one of the biggest constraints on humanoid and physical AI progress right now, given Velmurugan's estimate that a YouTube-scale (5x) data set is needed to move the field forward. It's plausible that specialized data pipelines, including novel modalities like brain-wave signals or leader-follower arm capture, become a differentiated niche for startups, especially given explicit demand from humanoid robotics companies. At the same time, founders should treat the brain-wave angle itself with caution: it remains unconfirmed whether the signal will meaningfully improve model performance, and reliance on human pilots for data generation could become a scalability bottleneck if demand for training data continues to outpace what pilot-based collection can supply. Finally, the 100x-value-versus-20x-cost tradeoff on dense annotation suggests founders will need to make deliberate, math-backed decisions about where to invest in data quality rather than assuming more data is automatically better.

Sources: No conflicting reports were identified in the underlying material for this story.

Sources