Inflect-Micro-v2: A Sub-10M-Param Local TTS Model
28 Jul 2026
A tiny TTS model with outsized ambitions
An independent developer known as Owen has published Inflect-Micro-v2, a local text-to-speech (TTS) model that fits under 10 million parameters, on Hugging Face. The release is part of a broader project called Inflect v2, which offers one public API across two model sizes: Micro, which prioritizes output quality while staying below 10M parameters, and Nano, which prioritizes a minimal footprint below 4M parameters.
Owen built and funded Inflect v2 independently, without institutional backing.
The numbers
Inflect-Micro-v2 ships with:
- 9,356,513 deployable parameters
- A 37.53 MB footprint in FP32 format
- 24 kHz mono audio output
- Fixed-voice, English-only synthesis with deterministic seeds, long-text handling, and both CPU and CUDA inference support
An official verified FP32 ONNX export, Inflect-Micro-v2-ONNX, is published separately.
How it stacks up in testing
The model was evaluated against three comparison systems — KittenTTS Nano, Piper Low, and Supertonic 3 — using two test sets: a 400-prompt "Modern400" evaluation corpus and a 500-prompt UTMOS22 test set, each using identical unseen English prompts per system.
Reported results include:
- A 66.2% preference rate in a final anonymous community study (21 wins, 10 losses, 3 ties)
- A UTMOS22 score of 4.395 (95% bootstrap confidence interval: 4.381–4.408)
- Word error rates (WER) of 2.52% on Qwen3-ASR, 5.45% on Nemotron 3.5, and 2.73% on Whisper large-v3
These figures suggest the model produces intelligible, well-rated speech relative to the listed comparison systems — though the report notes several gaps in methodology detail (see below).
What's missing from the picture
The report flags a handful of open questions that founders evaluating this model should keep in mind:
- No release date or publication timeline is specified.
- Licensing terms and permitted commercial use are not described.
- The hardware and setup used to measure WER and UTMOS22 scores are not detailed.
- The methodology behind the "final anonymous community study" — participant count, selection process, blinding — is not explained.
- It's unclear whether non-English languages are supported now or only planned for a future v3.
- There's no quantitative comparison of Nano's footprint/parameter tradeoffs against Micro.
Risks worth weighing
A few structural risks stand out. The model's fixed-voice, English-only design may limit its usefulness for multilingual or multi-voice applications. The preference metrics come from a community-based study whose results may not generalize to broader user populations. Perhaps most notably, Owen has stated that continuation into a broader v3 — which might include more languages, voices, and stability improvements — is contingent on this release "finding a real audience." As an independently funded project, long-term maintenance and support resources may also be limited.
Why founders should care
For early-stage teams building voice features on tight budgets, a sub-10M-parameter TTS model that runs on CPU or CUDA could plausibly reduce infrastructure costs and simplify on-device or embedded deployment — a meaningful advantage if resource constraints are a real bottleneck. The reported WER figures across multiple ASR benchmarks may indicate usable transcription accuracy, though independent verification is likely warranted given the limited methodology disclosure. The preference-study results could suggest the model is competitive against the listed alternatives, but given the small, single-study sample size, founders should treat these results as directional rather than conclusive. Most importantly, because future development (a broader v3 with more languages and stability improvements) is explicitly conditional on adoption, teams considering Inflect-Micro-v2 for production should probably build in contingency plans rather than assume ongoing support is guaranteed.
The bottom line
Inflect-Micro-v2 is a notable entrant in the lightweight, local TTS space — small enough to potentially run on constrained hardware, with benchmark numbers that look competitive against a few named alternatives. But as an independent, single-developer project with unresolved licensing terms and an uncertain roadmap, it's better suited for experimentation and evaluation right now than for founders seeking a fully de-risked, long-term voice infrastructure bet.