All news
aiproductsaas

Flash-MSA: Open-Source Sparse Attention Kernels Debut

13 Jul 2026

What happened

A new project called Flash-MSA has released what its author calls the world's first performant open-source training kernels for Minimax Sparse Attention (MSA), built in CuTeDSL for Hopper and Blackwell GPUs. The kernels target long-context training, using blockwise sparsity to select key-value (KV) pairs rather than processing full quadratic attention.

No release date, versioning, or licensing details were provided, and the identity or organizational affiliation of the author behind the announcement is not specified.

How MSA works

MSA selects KVs in blocks of 128 using max-pooling over proxy scores — a blockwise sparsity approach designed to reduce computation at long context lengths. The kernel design caches block indices so that only the proxy forward pass remains quadratic with respect to context length; everything downstream relies on cached sparse blocks instead of recomputing full attention.

A notable architectural choice: MSA replaces MLA (Multi-head Latent Attention) with GQA (Grouped-Query Attention) for the main attention mechanism. This differs from Deepseek Sparse Attention, which uses MLA, even though the author describes MSA as similar to Deepseek's approach with some core changes. MSA also implements group-wise specialization of proxy heads, letting different subsets of KVs be selected by each proxy head.

The report notes that frontier models GLM-5.2 and DSv4 reportedly use sparse attention formulations fit to MLA — meaning MSA's GQA-based design diverges from what appears to be the current frontier-model approach. The author states that, to their knowledge, no western labs have adopted MLA into their training.

Verification and limitations

Correctness was verified by comparing the kernel implementation against an eager PyTorch implementation using cosine similarity sweeps in bf16 precision, with a typical tolerance of 0.01 cited for these checks. This is a relatively narrow validation method — it may not capture all edge cases or numerical instabilities that could surface in production training runs.

Several other caveats stand out. The kernels currently only run on Hopper and Blackwell GPUs, limiting accessibility for teams without that hardware. The claim of being the "first performant" open-source kernel for MSA is self-reported and not independently benchmarked in the available material. And no performance benchmarks — throughput, latency, or memory — comparing Flash-MSA to existing MLA-based or dense attention implementations have been published.

Why founders should care

For teams working on long-context or sparse-attention training, Flash-MSA's open-source availability likely lowers the barrier to experimentation, since the kernels are freely accessible rather than proprietary. The blockwise caching design suggests it could reduce computational cost relative to fully quadratic attention, though this remains unverified by independent benchmarks.

Founders evaluating this technology should treat performance claims as unconfirmed until third-party benchmarks emerge. The architectural divergence from MLA — the approach reportedly used by frontier models like GLM-5.2 and DSv4 — means MSA's compatibility and performance at scale are largely untested outside the author's own validation. Teams already invested in MLA-based infrastructure may find limited immediate value, while those exploring alternative architectures could see this as a viable path worth monitoring.

The hardware restriction to Hopper and Blackwell GPUs is also a practical consideration: smaller teams without access to this infrastructure may not be able to adopt these kernels regardless of their technical merits.

What's missing

The report leaves several open questions: there's no release date, versioning, or licensing information; no data on community adoption or downstream usage since release; and no quantitative comparison of how GLM-5.2 and DSv4's MLA-based sparse attention formulations perform relative to MSA. Founders considering these kernels should treat the announcement as an early, unverified signal rather than a production-ready benchmark-backed release.

Sources

Flash-MSA: Open-Source Sparse Attention Kernels Debut — Xcelit