All news
aiproduct

Fermion Research Launches Neutrino-1 8B Ternary LLM

28 Jul 2026

Fermion Research has released Neutrino-1, a three-model family built in a proprietary ternary weight format, with the flagship 8B model and a companion draft model available for immediate download under Apache License 2.0—no waitlist, no gated preview.

What was announced

Neutrino-1 8B is an 8.19B-parameter decoder-only transformer that uses grouped-query attention across 252 transformer linears, yet ships as a compact 3.88 GB file. It's a derivative of Qwen3-8B (itself Apache-2.0 licensed). The model card reports that 62.63% of its 6.95B coded weights sit at zero, a hallmark of the ternary quantization approach Fermion Research is using.

The family also includes a 328 MB, 0.6B-parameter draft model designed to pair with the 8B model for speculative decoding, plus a third model that the report does not name or describe. All three are dated 2026 in their citation and released under Apache 2.0, which permits commercial use, modification, fine-tuning, and redistribution.

Performance and memory footprint

The model card details a range of throughput and memory benchmarks, run through a llama.cpp fork that offloads all 36 transformer layers:

  • 24.9 tok/s on an Apple M5 (CPU-only, 9 threads)
  • 33.7 tok/s on a base M5 MacBook using the Apple-silicon runtime
  • 30.7 tok/s on an NVIDIA L4 at 4k context (4.68 GiB memory)
  • KV cache costs of 144 KiB per token at fp16, totaling roughly 0.60 GB for a 4k-token session

With speculative decoding enabled, the combined draft+8B setup uses about 1 GB of cache at 4k shared context, with peak combined memory of 4.3 GiB on a 16 GB Apple M5 (0.53 GiB of which is attributed to the draft model alone).

Draft-acceptance rates—and resulting speedups—vary notably by prompt type. On the Apple M5, the drafted rate reached 25.71 tok/s versus a 22.00 tok/s plain rate, with a 0.744 acceptance rate. Factual prompts pushed acceptance to 96.5%. Counting prompts told a different story: plain-rate throughput jumped to roughly 396 tok/s, with only about 7 tokens emitted per 8B forward pass—a pattern that signals speculative decoding's benefit is highly workload-dependent rather than uniform.

What's missing

The report is notably light on quality signals. There are no accuracy, reasoning, or standard benchmark comparisons (such as MMLU or HumanEval) against other 8B-class models. Training data, fine-tuning methodology, and safety evaluation aren't disclosed. The third model in the family remains unidentified. And while the ternary format enables the reported efficiency gains, its compatibility with ML tooling outside the specific llama.cpp fork isn't detailed—nor is it clear how the format interacts with Apache 2.0's redistribution terms.

Why founders should care

For teams evaluating local or on-device inference, Neutrino-1's small file size and low KV-cache footprint likely make it worth a look—especially if you're constrained to consumer or edge hardware. A model running near 30+ tok/s on an Apple M5 or an NVIDIA L4, with combined memory under 6 GiB in speculative-decoding mode, could plausibly reduce infrastructure costs for certain deployment scenarios.

The Apache 2.0 license is likely to appeal to founders who want to avoid licensing negotiations before shipping a commercial product, though it's worth independently confirming how that license applies given the proprietary ternary weight format.

At the same time, the complete absence of accuracy benchmarks means founders should not assume quality parity with other 8B models—efficiency gains reported here say nothing about output quality, and that gap will probably need to be closed with in-house evaluation before any production commitment. The sharp variation in speculative-decoding speedups between counting and factual prompts also suggests real-world performance could be inconsistent across use cases, so benchmarking against your actual workload is probably a safer bet than trusting the headline throughput numbers alone.

Finally, because Neutrino-1 relies on a proprietary format and a specialized llama.cpp fork, teams standardized on typical ML inference pipelines should expect some integration overhead—likely a bigger factor for larger engineering teams than for solo founders experimenting locally.

Sources