CoCoQuant: Breaking the Bandwidth Wall via Co-Optimized Communication and Computation Quantization

Haojie Duanmu, Jifeng Ding, Size Zheng, Xuegui Zheng, Jiangfei Duan, Xingcheng Zhang, Li-Wen Chang, Xin Liu, Dahua Lin
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:27069-27088, 2026.

Abstract

The rapid scaling of large language models (LLMs) has made distributed inference indispensable, yet end-to-end latency is increasingly dominated by communication, forming a critical bandwidth wall that fundamentally limits the practical gains of existing quantization techniques. Existing approaches typically treat communication and computation in isolation, failing to exploit their coupled nature and introducing limited system-level acceleration and accuracy degradation. To address this, we propose CoCoQuant, a co-designed framework that jointly optimizes communication and computation as a unified end-to-end design space. CoCoQuant introduces a precision-aligned graph-rewriting that enables zero-overhead fusion between low-precision communication and computation. CoCoQuant formulates a hardware-aware mixed-precision allocation problem that integrates roofline-based cost modeling with relative sensitivity calibration, solved via global integer linear programming. Extensive experiments on LLMs of varing scales demonstrate that CoCoQuant achieves Pareto-optimal accuracy-latency trade-offs, delivering up to 2.92 end-to-end speedup with a negligible increase in perplexity (0.22).

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-duanmu26a, title = {{C}o{C}o{Q}uant: Breaking the Bandwidth Wall via Co-Optimized Communication and Computation Quantization}, author = {Duanmu, Haojie and Ding, Jifeng and Zheng, Size and Zheng, Xuegui and Duan, Jiangfei and Zhang, Xingcheng and Chang, Li-Wen and Liu, Xin and Lin, Dahua}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {27069--27088}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/duanmu26a/duanmu26a.pdf}, url = {https://proceedings.mlr.press/v306/duanmu26a.html}, abstract = {The rapid scaling of large language models (LLMs) has made distributed inference indispensable, yet end-to-end latency is increasingly dominated by communication, forming a critical bandwidth wall that fundamentally limits the practical gains of existing quantization techniques. Existing approaches typically treat communication and computation in isolation, failing to exploit their coupled nature and introducing limited system-level acceleration and accuracy degradation. To address this, we propose CoCoQuant, a co-designed framework that jointly optimizes communication and computation as a unified end-to-end design space. CoCoQuant introduces a precision-aligned graph-rewriting that enables zero-overhead fusion between low-precision communication and computation. CoCoQuant formulates a hardware-aware mixed-precision allocation problem that integrates roofline-based cost modeling with relative sensitivity calibration, solved via global integer linear programming. Extensive experiments on LLMs of varing scales demonstrate that CoCoQuant achieves Pareto-optimal accuracy-latency trade-offs, delivering up to 2.92 end-to-end speedup with a negligible increase in perplexity (0.22).} }
Endnote
%0 Conference Paper %T CoCoQuant: Breaking the Bandwidth Wall via Co-Optimized Communication and Computation Quantization %A Haojie Duanmu %A Jifeng Ding %A Size Zheng %A Xuegui Zheng %A Jiangfei Duan %A Xingcheng Zhang %A Li-Wen Chang %A Xin Liu %A Dahua Lin %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-duanmu26a %I PMLR %P 27069--27088 %U https://proceedings.mlr.press/v306/duanmu26a.html %V 306 %X The rapid scaling of large language models (LLMs) has made distributed inference indispensable, yet end-to-end latency is increasingly dominated by communication, forming a critical bandwidth wall that fundamentally limits the practical gains of existing quantization techniques. Existing approaches typically treat communication and computation in isolation, failing to exploit their coupled nature and introducing limited system-level acceleration and accuracy degradation. To address this, we propose CoCoQuant, a co-designed framework that jointly optimizes communication and computation as a unified end-to-end design space. CoCoQuant introduces a precision-aligned graph-rewriting that enables zero-overhead fusion between low-precision communication and computation. CoCoQuant formulates a hardware-aware mixed-precision allocation problem that integrates roofline-based cost modeling with relative sensitivity calibration, solved via global integer linear programming. Extensive experiments on LLMs of varing scales demonstrate that CoCoQuant achieves Pareto-optimal accuracy-latency trade-offs, delivering up to 2.92 end-to-end speedup with a negligible increase in perplexity (0.22).
APA
Duanmu, H., Ding, J., Zheng, S., Zheng, X., Duan, J., Zhang, X., Chang, L., Liu, X. & Lin, D.. (2026). CoCoQuant: Breaking the Bandwidth Wall via Co-Optimized Communication and Computation Quantization. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:27069-27088 Available from https://proceedings.mlr.press/v306/duanmu26a.html.

Related Material