Neural Weight Compression: lossless LLM weight compression decoded inside the CUDA matvec. BF16 31 % smaller and faster than cuBLAS; fp8 13 % smaller at native fp8 speed. Bit-identical weights.
-
Updated
Sep 19, 2026 - Cuda
Neural Weight Compression: lossless LLM weight compression decoded inside the CUDA matvec. BF16 31 % smaller and faster than cuBLAS; fp8 13 % smaller at native fp8 speed. Bit-identical weights.
Post-training weight compression for low-RAM machines: Q4/Q8 quantization, green-format repair, AVX2 CPU inference, optional CUDA. ~45% less RAM at ~99.9% quality.
A deterministic, random-access archive format and toolkit for compressing, storing, diffing and patching large language model weights.
To associate your repository with the weight-compression topic, visit your repo's landing page and select "manage topics."