All news
aiproduct

Qwen3.8-2.4T-A95B: A Max-Class Model Goes Open

31 Aug 2026

Qwen has released Qwen3.8-2.4T-A95B, described as the most capable generation in the Qwen open-model family to date. The headline feature: this is the first time a Qwen-Max-class model has been made available as open weights, alongside a quantized variant and an official hosted version with expanded features.

What's in the release

The base model, Qwen3.8-2.4T-A95B, carries 2.4 trillion total parameters with 95 billion activated per inference pass — a Mixture-of-Experts (MoE) design spanning 92 layers, 512 total experts, and 11 activated experts (10 routed plus 1 shared) at any given time.

Context handling is a standout spec: a native context length of 262,144 tokens, extensible up to 1,010,000 tokens.

Three versions accompany the launch:

  • Qwen3.8-2.4T-A95B — the base open-weight model, text-only, requiring thinking mode for all interactions.
  • Qwen3.8-2.4T-A95B-FP8 — a quantized variant using fine-grained FP8 quantization with a block size of 128, reported to perform nearly identically to the full-precision model.
  • Qwen3.8-Max — the official, feature-expanded version, adding vision input, non-thinking support, a 1M-token context length, and built-in tools.

The model also supports adjustable reasoning_effort at three levels — xhigh (default), medium, and low — and enables preserve_thinking by default across workloads. It's compatible with vLLM, SGLang, and TokenSpeed inference frameworks; the FP8 model page also references unspecified 'other inference frameworks.'

What's missing

No benchmark scores or head-to-head comparisons against other open or closed models have been published. There's also no release date, licensing terms, pricing, availability, or hardware requirements disclosed, and no information on training data, training compute, or safety evaluation methodology. It remains unclear which additional frameworks beyond the three named ones are supported.

Why founders should care

This release plausibly signals rising competitive pressure among AI providers to open-source higher-capability models — a dynamic that, if it continues, could give founders more cost-effective alternatives to closed APIs over time. The FP8-quantized variant may make comparable performance accessible at lower compute cost, which could matter for teams without budgets for full-precision infrastructure. The tiered reasoning_effort settings could also let founders tune cost-versus-quality trade-offs across different product tiers or features.

That said, the gap between the 2.4T total parameter count and 95B activated parameters means founders will likely need to assess whether their infrastructure supports MoE-style deployment before committing engineering time. And because no independent benchmarks have been published, claims of this being the 'most capable' Qwen model to date should be treated as unverified until third-party testing emerges.

Risks to weigh

  • The base model's text-only, thinking-mode-only design may add latency or complexity for use cases needing multimodal input or non-thinking responses — those needs are only addressed in the separate Qwen3.8-Max version.
  • The sheer scale (2.4T total parameters) could create real infrastructure and cost hurdles for teams lacking specialized inference frameworks.
  • FP8 quantization is described as 'nearly identical' in performance to the full model, but no quantified trade-off data has been shared.
  • Without published benchmarks, there's currently no independent way to verify comparative performance claims.

Bottom line

Qwen3.8-2.4T-A95B marks a notable moment: a Max-class model, previously reserved for hosted/paid tiers, is now available in open form. For founders evaluating open-weight options, the extended context length and FP8 variant are worth watching — but given the absence of benchmarks, licensing details, and hardware specs, independent testing before production adoption is advisable.

Sources