Kimi K3 (2.8T Params) Runs on a MacBook and AMD MI355X
02 Aug 2026
Two paths to running a 2.8-trillion-parameter model
Kimi K3, a mixture-of-experts model reported at 2.78 trillion parameters (per WASTE) or 2.8 trillion parameters (per wafer.ai), has been demonstrated running in two very different environments: a consumer-grade 64GB MacBook Pro and AMD's MI355X GPUs. Sources differ slightly on the exact parameter count — likely a rounding difference rather than two distinct model versions, though neither source confirms this explicitly.
Running on a MacBook Pro via WASTE
WASTE, an embeddable inference engine written in C, packages Kimi K3 for on-device inference. On a 64GB MacBook Pro, it achieves 0.49–0.54 tokens per second — a rate driven by the model's mixture-of-experts design, where roughly 4% of the model activates per token, touching 16 experts across 92 layers.
Key technical details from the report:
- The converted model weighs in at 982 GiB, down from an originally published 1.42 TB.
- A 27.28 GB resident trunk occupies RAM, with a 46 GB memory budget and 17.56 GB expert cache on the 64GB machine.
- Each token requires reading 17 GB of expert data.
- Experts are stored using residual vector quantization at 3.00 bits per weight.
- WASTE validated its outputs against a PyTorch reference: final logits agreed to 3.6e-06, and the vision tower matched its oracle to 2.3e-06.
- Overlapping expert reads with arithmetic operations delivered roughly a 1.6x speedup, since reads consume about 55% of a decode step versus 27% for arithmetic.
- CRC32 checking adds about 5% overhead on the related Kimi-Linear model and 1% on K3.
A smaller sibling, Kimi-Linear-48B-A3B-Instruct, runs far more practically: a 19 GB container achieving 10.7 tok/s with a 1.87 GB RAM floor.
Streaming model weights introduces its own bottleneck: internal SSDs stream at 12.78 GB/s, while USB enclosures manage only 0.94 GB/s — a gap that could matter for anyone trying to replicate this setup with external storage.
Running on AMD MI355X GPUs
Separately, wafer.ai reports running Kimi K3 on AMD's MI355X GPUs, benchmarking throughput and cost against NVIDIA's B200 and B300 chips.
Cost comparison:
- MI355X: 288GB VRAM per GPU at $2.50/GPU-hr
- NVIDIA B300: $6.00/GPU-hr
- NVIDIA B200: $4.25/GPU-hr
- MI355X works out to roughly 2.4x cheaper per GPU than B300, and about 1.7x cheaper than B200.
Throughput comparison:
- MI355X: 952 tok/s/node aggregate, 118 tok/s single-stream
- Against a TP16 B200 configuration (498 tok/s aggregate across two nodes), MI355X delivered 3.8x aggregate throughput and 1.3x single-stream decode
- B300, however, outperformed MI355X by 1.65x on aggregate throughput
Software-side optimizations also mattered: speculative decoding produced a 2.2x single-stream improvement, 1.7x at moderate load, and an 18% peak aggregate throughput gain. Prefill optimization yielded a 2–3x speedup on cold prefill, with AMD's AITER MLA prefill kernel reaching roughly 13,000 tokens/s in steady state versus 4,000–7,000 tok/s for a Triton fallback. Even so, a 172k-token cold prefill took about 51 seconds on MI355X versus roughly 23 seconds on B300.
The risks and caveats
The report flags several unresolved issues:
- AMD's software maturity gap: MI355X reportedly has slower kernels and less day-0 support on inference frameworks compared to NVIDIA, which could add engineering overhead for teams adopting it.
- Throughput limits on consumer hardware: Sub-1 tok/s on a MacBook Pro is far from viable for real-time production use, positioning WASTE more as an experimentation tool than a deployment path.
- Storage bottlenecks: Streaming model weights from external USB storage (0.94 GB/s) versus internal SSDs (12.78 GB/s) could meaningfully slow down any similar setup.
- No accuracy benchmarks under quantization: The report includes no data on output quality or task performance degradation from the 3-bit quantization scheme, despite close logit-matching to a PyTorch reference.
- No unified benchmark: It's unclear whether the MI355X benchmarks and the WASTE CPU deployment used identical quantization or parameter counts, making direct cost/performance comparisons across the two setups speculative.
For context, other large models mentioned in the report include GLM5.2 (753 billion parameters) and DeepSeek V4-Pro (1.6 trillion parameters) — both smaller than Kimi K3.
Why founders should care
For early-stage teams evaluating large language model infrastructure, these developments plausibly matter in a few ways:
- Running trillion-parameter MoE models on consumer hardware may lower the barrier to experimentation — founders could plausibly prototype with very large models without immediate cloud GPU spend, though throughput constraints (sub-1 tok/s) suggest this is unlikely to support production inference anytime soon.
- MI355X's reported cost advantage over B200/B300 could make it a viable alternative for startups optimizing inference spend, though AMD's noted software support gaps may likely mean added engineering time before deployments are production-ready.
- The quantization and validation techniques shown here (3-bit RVQ, tight logit-matching) suggest that teams considering similar compression strategies for their own models should likely evaluate accuracy trade-offs carefully — the report offers no benchmark data on how quantization affects real-world output quality.
What's still unclear
The report leaves several open questions: no release dates are given for either the WASTE engine or the MI355X benchmarks, so the timeline is approximate. There's also no total-cost-of-ownership comparison between the CPU-based Mac setup and the GPU-based MI355X setup, and no confirmation that both benchmarking efforts used the same underlying quantization or parameter count. Founders evaluating either path should treat these figures as early signals rather than settled production benchmarks.