All news
aiproduct

Lossless Compression Trims GLM-5.2 Size by ~25-30%

20 Jul 2026

Researchers have reported a bit-for-bit lossless compression scheme applied to a 1.4TB checkpoint shard of GLM-5.2, a large language model, achieving size reductions in the 25-30% range without altering the underlying weight values.

What was tested

Standard BF16 weights use 16 bits per parameter: 1 sign bit, 8 exponent bits, and 7 mantissa bits. The core technique, called K15 accounting, replaces the 9-bit sign-and-exponent symbol with a 4-bit code that points into a 15-entry table.

Two separate measurements were reported:

  • A full model scan using K15 accounting found a 30.168% size reduction, pricing weights at 11.173 bits each.
  • A separate byte-split representation, decoded across 59,509 BF16 tensors, achieved a 24.967% reduction at 12.005 bits per weight.

All 59,509 tensors were verified to match bit-for-bit between the reconstructed and source versions — the basis for calling the compression lossless in this test.

On the performance side, a dense 12-bit prototype was benchmarked for GEMV (matrix-vector multiplication) speed on an A40 GPU, measuring 0.733x the time of BF16 — suggesting the compressed format could run faster, not just smaller, though this remains unconfirmed at full model scale.

The work builds on prior lossless BF16 exponent-compression research, citing ZipNN and DFloat11 as established precedents in this size range, and ZipServ as the closest existing runtime design demonstrating fixed-length coding with direct register reconstruction.

What's still open

The report is explicit about its limitations:

  • An exact lossless speedup for the compressed format has not yet been established.
  • A physical, GLM-scale K15 container has not yet been built, so real-world feasibility at full model size remains unproven.
  • End-to-end serving integration is an open engineering task — this is not yet a production-ready deployment.

Sources also leave some questions unresolved. It's unclear whether the 30.168% (K15) and 24.967% (byte-split) figures apply to the same weight set or represent different partial compressions of the model. Similarly, it's not confirmed whether the 0.733x GEMV benchmark reflects the lossless K15 format itself or the separate dense 12-bit prototype, and whether that speed advantage would hold at GLM-5.2's full scale. No data on inference accuracy or quality impact was provided beyond the bit-for-bit reconstruction check, and details on the 15-entry table's construction, hardware requirements, or decode overhead in production were not disclosed.

Why founders should care

For founders building or evaluating model-serving infrastructure, this result is worth watching but not yet worth building on. A 25-30% reduction in memory footprint could plausibly translate into meaningful storage and hosting cost savings for teams running large models like GLM-5.2 — if the technique holds up at full scale and integrates cleanly into serving pipelines. The GEMV benchmark hints at a possible inference speed upside alongside the memory savings, which would be a notable double win, but this is likely to require independent verification before it should inform any product roadmap.

Because the approach builds on established prior work (ZipNN, DFloat11, ZipServ), the technical risk of extending it toward production may be lower than for a fully novel method — a reason for infra-focused teams to track this space rather than dismiss it. That said, with no physical GLM-scale container built and end-to-end serving integration still undone, this should currently be treated as an early-stage research signal rather than a deployable solution.

Sources