28.9M-Parameter LLM Runs on an $8 ESP32-S3 Chip
28 Jul 2026
A language model with 28.9 million parameters has been successfully run entirely on an ESP32-S3 microcontroller—a piece of hardware that costs roughly $8. No server round-trip, no cloud API call. The inference happens fully on-device.
What happened
The project pushes a language model onto one of the cheapest, most widely available microcontrollers on the market. The ESP32-S3 has just 512KB of SRAM, so the model leans on flash memory to store a large embedding table—25 million rows—pulling roughly 450 bytes per token and reading 6 rows per token from that table during generation.
To make this work, the model uses Per-Layer Embeddings, a technique also used in Google's Gemma 3n and Gemma 4 models. Its presence here suggests the technique isn't just useful for scaling large models down for phones—it may also help stretch tiny models onto severely memory-constrained chips.
The model was trained on the TinyStories dataset, and the result matches that scope: it writes short, simple stories and mostly keeps them coherent. That's the extent of it. It cannot answer questions, follow instructions, write code, or recall facts.
This isn't the first attempt at squeezing a language model onto this class of hardware. A prior model with just 260,000 parameters was run on a similar chip. The new model is about 100x larger—a jump that raises the question of how much further this scaling trend can go on similarly priced hardware.
The numbers at a glance
- 28.9 million parameters
- ~$8 hardware cost (ESP32-S3)
- 9 tokens per second generation speed
- 512KB SRAM on-chip
- 25 million rows in the embedding table (stored in flash)
- 450 bytes read per token from flash
- 100x parameter increase over the prior on-chip model (260K parameters)
Where it falls short
The report flags several real limitations worth weighing before getting too excited:
- Narrow capability: the model can't answer questions, follow instructions, write code, or reason about facts—it only generates simple stories.
- Speed: at 9 tokens per second, it may be too slow for real-time or interactive use cases.
- Flash dependency: reading 450 bytes per token from flash memory could introduce latency or wear over time, though the report doesn't quantify this.
- Narrow training data: training exclusively on TinyStories suggests the model likely won't generalize well beyond simple narrative generation.
Important context is also missing. There's no date attached to either this project or the earlier 260K-parameter comparison, no named developer, no power consumption or battery-life data, and no clear methodology for how "coherence" of the generated stories was judged. Training time, hardware, and dataset size in tokens are also unspecified.
Why founders should care
For founders building in edge AI, embedded systems, or privacy-first products, this likely signals that on-device inference on extremely cheap hardware is becoming more technically plausible than it was even a model generation ago. The 100x jump in parameter count on the same class of chip suggests further scaling on similar microcontrollers is plausible, though not guaranteed, as memory and speed constraints will likely bite harder as models grow.
The use of Per-Layer Embeddings—shared with Google's Gemma line—may indicate that memory-efficient architectural techniques developed for consumer-scale models are becoming transferable down to the smallest edge hardware. Startups exploring offline, privacy-preserving AI (think: sensors, toys, appliances, or field devices with no connectivity) may want to track this technique closely.
That said, the current capability ceiling is real: this is not a model that can power a chatbot, a coding assistant, or any product requiring factual accuracy or instruction-following. Founders should treat this as an early proof-of-concept for narrow, low-stakes generative tasks on constrained hardware—not as evidence that general-purpose on-device LLMs are imminent at this price point. Whether the underlying techniques scale to more useful capability levels, and at what hardware cost, remains an open question the report does not answer.