CFP last date
20 August 2026
Reseach Article

Practical Limits of Lossless Compression for bf16 Transformer LLM Weights, with Companion Measurements on Q4 K Tensors in GGUF Q4 K M Files

by Nimrod Rothstein, Ariel Rothstein
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Number 120
Year of Publication: 2026
Authors: Nimrod Rothstein, Ariel Rothstein
10.5120/ijcaa71cdd2650e6

Nimrod Rothstein, Ariel Rothstein . Practical Limits of Lossless Compression for bf16 Transformer LLM Weights, with Companion Measurements on Q4 K Tensors in GGUF Q4 K M Files. International Journal of Computer Applications. 187, 120 ( Jun 2026), 1-12. DOI=10.5120/ijcaa71cdd2650e6

@article{ 10.5120/ijcaa71cdd2650e6,
author = { Nimrod Rothstein, Ariel Rothstein },
title = { Practical Limits of Lossless Compression for bf16 Transformer LLM Weights, with Companion Measurements on Q4 K Tensors in GGUF Q4 K M Files },
journal = { International Journal of Computer Applications },
issue_date = { Jun 2026 },
volume = { 187 },
number = { 120 },
month = { Jun },
year = { 2026 },
issn = { 0975-8887 },
pages = { 1-12 },
numpages = {9},
url = { https://ijcaonline.org/archives/volume187/number120/practical-limits-of-lossless-compression-for-bf16-transformer-llm-weights-with-companion-measurements-on-q4_k-tensors-in-gguf-q4_k_m-files/ },
doi = { 10.5120/ijcaa71cdd2650e6 },
publisher = {Foundation of Computer Science (FCS), NY, USA},
address = {New York, USA}
}
%0 Journal Article
%1 2026-07-01T03:10:10+05:30
%A Nimrod Rothstein
%A Ariel Rothstein
%T Practical Limits of Lossless Compression for bf16 Transformer LLM Weights, with Companion Measurements on Q4 K Tensors in GGUF Q4 K M Files
%J International Journal of Computer Applications
%@ 0975-8887
%V 187
%N 120
%P 1-12
%D 2026
%I Foundation of Computer Science (FCS), NY, USA
Abstract

This paper quantifies how far bf16 transformer large language model (LLM) weights—and Q4 K-typed tensors inside GGUF Q4 K M files—can be compressed with no loss of information. Every method is scored under one accounting discipline: a trained profile is billed to its method on each file, each run is checked for a byte-exact roundtrip, and the train/test partition is taken over source models rather than individual tensors. For bf16 weights the marginal-byte ceiling sits at ≈1.495× (model-level 95% confidence interval [1.487, 1.502]); the strongest coder that can be run end-to-end under the present accounting—the bf16 split coder introduced here—reaches 1.488×, just below that ceiling. For Q4 K-typed tensors the marginal-byte ceiling is ≈1.076× at the tensor-stream level, dropping to 1.041–1.045× once whole GGUF artifacts are scored, because Q4 K M files also carry Q6 K tensors that barely compress (near 1.01×); the mixture-CDF coder presented here attains 1.052×, whereas a dictionary-trained zstd ends up enlarging the stream once its dictionary is billed per file. Between adjacent same-role weight matrices at bf16 precision no exploitable linear redundancy is found (median Pearson +0.0004 over 250 layer pairs drawn from two Qwen2.5 models). Across 11,942 verified-roundtrip method-evaluations spanning 7,960 distinct benchmark rows, not a single roundtrip failed; the contribution is the comparison protocol itself rather than another compressor. Each ceiling is a compression-ratio bound—hardware independent and confirmed on weights up to 7B parameters (Qwen2.5-7B-Instruct); the single-core x86 (Sapphire Rapids) host constrains only the throughput figures, which are reported for comparability rather than as deployment numbers.

References
  1. Zhang, T., Sui, Y., Zhong, S., Chaudhary, V., Hu, X., and Shrivastava, A. 2025. 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float. In Advances in Neural Information Processing Systems (NeurIPS). arXiv:2504.11651.
  2. Nikulin, I. 2026. Unweight: Lossless MLP Weight Compression for LLM Inference. Tech. Rep. Cf-TR-2026.04.v1, Cloudflare Research.
  3. Hershcovitch, M., Choshen, L., Wood, A., Enriquez, I., Loaiza-Ganem, G., Peleg, T., Kim, H., Klein, T., and Harnik, D. 2024. ZipNN: Lossless Compression for AI Models. arXiv:2411.05239.
  4. Hao, Y., Cao, Y., and Mou, L. 2024. NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks. arXiv:2410.20650.
  5. Yang, Z., Zhang, T., Xie, Y., Li, B., Xu, Y., and Shrivastava, A. 2026. To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration. In International Conference on Learning Representations (ICLR). arXiv:2510.02676.
  6. Lindstrom, P. 2014. Fixed-Rate Compressed Floating-Point Arrays. IEEE Transactions on Visualization and Computer Graphics 20, 12, 2674–2683.
  7. Lindstrom, P., and Isenburg, M. 2006. Fast and Efficient Compression of Floating-Point Data. IEEE Transactions on Visualization and Computer Graphics 12, 5, 1245–1250.
  8. Liang, X., Zhao, K., Di, S., Li, S., Underwood, R., Gok, A. M., Tian, J., Deng, J., Calhoun, J. C., Tao, D., Chen, Z., and Cappello, F. 2023. SZ3: A Modular Framework for Composing Prediction-Based Error-Bounded Lossy Compressors. IEEE Transactions on Big Data 9, 2, 485–498.
  9. Galicer, M., Nikulin, I., and Branch, C. 2026. Unweight: how we compressed an LLM 22% without sacrificing quality. Cloudflare Blog, Apr. 17, 2026.
  10. Shannon, C. E. 1948. A Mathematical Theory of Communication. Bell System Technical Journal 27, 3, 379–423.
  11. Hu, E. J., Shen, Y.,Wallis, P., Allen-Zhu, Z., Li, Y.,Wang, S., Wang, L., and Chen, W. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations (ICLR). arXiv:2106.09685.
  12. Kingma, D. P., and Ba, J. 2015. Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations (ICLR). arXiv:1412.6980.
Index Terms

Computer Science
Information Sciences

Keywords

Lossless weight compression bf16 transformer LLMs entropy ceiling reproducible benchmarking protocol GGUF Q4 K quantization source coding