Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction

Jatin Chhugani, Geonhwa Jeong, Bor-Yiing Su, Yunjie Pan, Hanmei Yang, Aayush Ankit, Jiecao Yu, Summer Deng, Yunqing Chen, Nadathur Satish, Changkyu Kim
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:19376-19394, 2026.

Abstract

Large Language Models (LLMs) have intensified the need for low-precision formats for efficient inference. The Open Compute Project Microscaling (MX) standard is attractive due to its favorable hardware efficiency, but its 4-bit variant (MXFP4) lags behind NVIDIA’s NVFP4 in accuracy, limiting adoption. We introduce two software-only techniques, Overflow-Aware Scaling (OAS) and Macro Block Scaling (MBS), that improve MXFP4 quantization fidelity without requiring hardware changes. OAS reduces overall errors by increasing effective dynamic range under power-of-two block scaling, while MBS allocates higher-precision scaling at a coarser granularity to better preserve outliers. Across multiple LLMs and standard downstream benchmarks, OAS and MBS reduce the end-to-end accuracy gap between MXFP4 and NVFP4 from about 10% to below 1% on average, while incurring modest GEMM overhead (6.2% on average). These results re-establish MXFP4 as a practical alternative to NVFP4, enabling near-NVFP4 accuracy while retaining MX’s hardware-efficiency advantages (e.g., 12% relative area savings in tensor cores).

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chhugani26a, title = {Unveiling the Potential of Quantization with {MXFP}4: Strategies for Quantization Error Reduction}, author = {Chhugani, Jatin and Jeong, Geonhwa and Su, Bor-Yiing and Pan, Yunjie and Yang, Hanmei and Ankit, Aayush and Yu, Jiecao and Deng, Summer and Chen, Yunqing and Satish, Nadathur and Kim, Changkyu}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {19376--19394}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chhugani26a/chhugani26a.pdf}, url = {https://proceedings.mlr.press/v306/chhugani26a.html}, abstract = {Large Language Models (LLMs) have intensified the need for low-precision formats for efficient inference. The Open Compute Project Microscaling (MX) standard is attractive due to its favorable hardware efficiency, but its 4-bit variant (MXFP4) lags behind NVIDIA’s NVFP4 in accuracy, limiting adoption. We introduce two software-only techniques, Overflow-Aware Scaling (OAS) and Macro Block Scaling (MBS), that improve MXFP4 quantization fidelity without requiring hardware changes. OAS reduces overall errors by increasing effective dynamic range under power-of-two block scaling, while MBS allocates higher-precision scaling at a coarser granularity to better preserve outliers. Across multiple LLMs and standard downstream benchmarks, OAS and MBS reduce the end-to-end accuracy gap between MXFP4 and NVFP4 from about 10% to below 1% on average, while incurring modest GEMM overhead (6.2% on average). These results re-establish MXFP4 as a practical alternative to NVFP4, enabling near-NVFP4 accuracy while retaining MX’s hardware-efficiency advantages (e.g., 12% relative area savings in tensor cores).} }
Endnote
%0 Conference Paper %T Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction %A Jatin Chhugani %A Geonhwa Jeong %A Bor-Yiing Su %A Yunjie Pan %A Hanmei Yang %A Aayush Ankit %A Jiecao Yu %A Summer Deng %A Yunqing Chen %A Nadathur Satish %A Changkyu Kim %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chhugani26a %I PMLR %P 19376--19394 %U https://proceedings.mlr.press/v306/chhugani26a.html %V 306 %X Large Language Models (LLMs) have intensified the need for low-precision formats for efficient inference. The Open Compute Project Microscaling (MX) standard is attractive due to its favorable hardware efficiency, but its 4-bit variant (MXFP4) lags behind NVIDIA’s NVFP4 in accuracy, limiting adoption. We introduce two software-only techniques, Overflow-Aware Scaling (OAS) and Macro Block Scaling (MBS), that improve MXFP4 quantization fidelity without requiring hardware changes. OAS reduces overall errors by increasing effective dynamic range under power-of-two block scaling, while MBS allocates higher-precision scaling at a coarser granularity to better preserve outliers. Across multiple LLMs and standard downstream benchmarks, OAS and MBS reduce the end-to-end accuracy gap between MXFP4 and NVFP4 from about 10% to below 1% on average, while incurring modest GEMM overhead (6.2% on average). These results re-establish MXFP4 as a practical alternative to NVFP4, enabling near-NVFP4 accuracy while retaining MX’s hardware-efficiency advantages (e.g., 12% relative area savings in tensor cores).
APA
Chhugani, J., Jeong, G., Su, B., Pan, Y., Yang, H., Ankit, A., Yu, J., Deng, S., Chen, Y., Satish, N. & Kim, C.. (2026). Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:19376-19394 Available from https://proceedings.mlr.press/v306/chhugani26a.html.

Related Material