AD-BTS: Adaptive Dual-Branch Token Sparsification via Spatial Information Density

Xinpei Gao, Xin Luo, Ming Liu, Chunjiang Wang, S Kevin Zhou
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:33367-33384, 2026.

Abstract

High-resolution visual encoders in multimodal large language models (MLLMs) substantially improve fine-grained perception, yet incur prohibitive computational costs.Existing token pruning methods are effective on natural images but struggle with spatially sparse structured inputs (e.g., charts), where critical high-frequency information is sparse, localized, and structurally essential. To address this challenge, we propose Adaptive Dual-Branch Token Sparsification (AD-BTS), a density-aware framework that dynamically allocates computation according to input signal characteristics. Specifically, AD-BTS introduces a Gradient-based Routing Gate (GRG) that uses lightweight pixel-level gradient statistics to estimate structural flatness and guide routing. Then, AD-BTS activates either a Redundancy Selection Branch (RSB) for aggressive token pruning with a frozen encoder, or a Structural Fusion Branch (SFB) with conditional LoRA and context fusion to preserve sparse structural information.Extensive experiments on Qwen2.5-VL demonstrate that AD-BTS establishes a new Pareto frontier between efficiency and accuracy. Under extreme compression (20% token retention), AD-BTS outperforms the strongest baseline by 12.1% on ChartQA while achieving a 1.8$\times$ prefill speedup, effectively reconciling computational efficiency with structural robustness.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-gao26r, title = {{AD}-{BTS}: Adaptive Dual-Branch Token Sparsification via Spatial Information Density}, author = {Gao, Xinpei and Luo, Xin and Liu, Ming and Wang, Chunjiang and Zhou, S Kevin}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {33367--33384}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/gao26r/gao26r.pdf}, url = {https://proceedings.mlr.press/v306/gao26r.html}, abstract = {High-resolution visual encoders in multimodal large language models (MLLMs) substantially improve fine-grained perception, yet incur prohibitive computational costs.Existing token pruning methods are effective on natural images but struggle with spatially sparse structured inputs (e.g., charts), where critical high-frequency information is sparse, localized, and structurally essential. To address this challenge, we propose Adaptive Dual-Branch Token Sparsification (AD-BTS), a density-aware framework that dynamically allocates computation according to input signal characteristics. Specifically, AD-BTS introduces a Gradient-based Routing Gate (GRG) that uses lightweight pixel-level gradient statistics to estimate structural flatness and guide routing. Then, AD-BTS activates either a Redundancy Selection Branch (RSB) for aggressive token pruning with a frozen encoder, or a Structural Fusion Branch (SFB) with conditional LoRA and context fusion to preserve sparse structural information.Extensive experiments on Qwen2.5-VL demonstrate that AD-BTS establishes a new Pareto frontier between efficiency and accuracy. Under extreme compression (20% token retention), AD-BTS outperforms the strongest baseline by 12.1% on ChartQA while achieving a 1.8$\times$ prefill speedup, effectively reconciling computational efficiency with structural robustness.} }
Endnote
%0 Conference Paper %T AD-BTS: Adaptive Dual-Branch Token Sparsification via Spatial Information Density %A Xinpei Gao %A Xin Luo %A Ming Liu %A Chunjiang Wang %A S Kevin Zhou %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-gao26r %I PMLR %P 33367--33384 %U https://proceedings.mlr.press/v306/gao26r.html %V 306 %X High-resolution visual encoders in multimodal large language models (MLLMs) substantially improve fine-grained perception, yet incur prohibitive computational costs.Existing token pruning methods are effective on natural images but struggle with spatially sparse structured inputs (e.g., charts), where critical high-frequency information is sparse, localized, and structurally essential. To address this challenge, we propose Adaptive Dual-Branch Token Sparsification (AD-BTS), a density-aware framework that dynamically allocates computation according to input signal characteristics. Specifically, AD-BTS introduces a Gradient-based Routing Gate (GRG) that uses lightweight pixel-level gradient statistics to estimate structural flatness and guide routing. Then, AD-BTS activates either a Redundancy Selection Branch (RSB) for aggressive token pruning with a frozen encoder, or a Structural Fusion Branch (SFB) with conditional LoRA and context fusion to preserve sparse structural information.Extensive experiments on Qwen2.5-VL demonstrate that AD-BTS establishes a new Pareto frontier between efficiency and accuracy. Under extreme compression (20% token retention), AD-BTS outperforms the strongest baseline by 12.1% on ChartQA while achieving a 1.8$\times$ prefill speedup, effectively reconciling computational efficiency with structural robustness.
APA
Gao, X., Luo, X., Liu, M., Wang, C. & Zhou, S.K.. (2026). AD-BTS: Adaptive Dual-Branch Token Sparsification via Spatial Information Density. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:33367-33384 Available from https://proceedings.mlr.press/v306/gao26r.html.

Related Material