All news
aiproduct

Dev Fixes 3 Bugs to Run Qwen 3.5 122B on Mac Studio

12 Jul 2026

The problem: a 3-5 minute wait for the first token

An independent developer running Qwen 3.5's 122B-parameter model (roughly 10B active parameters, mixture-of-experts architecture) on an M3 Mac Studio Ultra with 96GB of unified memory found that long conversations became nearly unusable. Once a conversation hit around 50,000 tokens, asking a follow-up question triggered a 3-5 minute wait before the first token appeared.

The developer, who bought the Mac Studio two months ago and worked on the issue during parental leave, described the experience bluntly: "That is not a chatbot, it is a batch job, and you go and make a cup of coffee while it thinks." For workflows like pair programming, that delay is a dealbreaker: "You cannot pair program with a model that makes you wait for a cup of tea."

What was actually broken

The setup used a system prompt of 130,000 tokens, and the agent framework wrote a background checkpoint every 256 generated tokens. Digging into the serving stack, the developer found:

  • 27GB of unmatchable checkpoint bodies sitting in the cache store, unable to be reused.
  • Zero in-memory cache hits versus 109 disk hits in a normal measurement window — meaning the system was constantly falling back to slow disk reads instead of fast memory lookups.

After three weeks of debugging a cache leak, the developer fixed three bugs in total (though only the cache-leak issue is described in detail) and released the patched code as qMLX, a fork of the rapid-mlx serving stack specialized for hybrid attention, on GitHub.

The result

Post-fix, one session (uid=58) showed 53,267 cached tokens against just 670 prefill tokens — a sharp improvement in cache efficiency compared to the pre-fix state. The developer also references DS4 Flash as an alternative option for running large models on consumer hardware, though no direct comparison is given.

It's worth noting these are internal cache and token metrics, not end-to-end latency or throughput benchmarks, so the practical speedup in real conversation response time isn't independently quantified beyond the anecdotal before/after wait-time description.

Why founders should care

This is a single-developer anecdote on one specific hardware configuration, so it likely won't generalize cleanly to every setup — but it's a useful signal for a few reasons:

  • Local inference for large MoE models may be more viable than assumed. If similar caching fixes can be replicated, founders building dev tools or AI products may increasingly be able to consider consumer Apple Silicon hardware for long-context workflows, potentially reducing cloud inference spend for certain use cases.
  • Serving-stack inefficiencies may be common, not unique. The cache leak and checkpoint bugs described here suggest that other open serving stacks could harbor similar issues, which may represent a real opportunity for founders building inference-optimization tooling.
  • Caution is warranted before adopting similar patches. Without published throughput or tokens-per-second benchmarks, and without clarity on qMLX's maintenance status or license, teams evaluating this approach should treat it as an early, unverified signal rather than a production-ready solution.

Founders in AI infrastructure or devtools may want to watch this space, particularly around Apple Silicon–optimized serving stacks and MoE model deployment, but should apply their own benchmarking before betting workflows on forked, single-maintainer projects like this one.

Sources