Beyond Blind Noising: Disentangled Visual Rectification for Hallucination Mitigation in MLLMs

Yujia Chen, Rui Sun, Bingzhou Wang, Huayu Mai, Wangkai Li, Zhaoyang Li, Aibing Li, Wenzhang Sun
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:16768-16805, 2026.

Abstract

Visual Contrastive Decoding (VCD) mitigates hallucinations in Multimodal Large Language Models (MLLMs) by penalizing the output shift from noise-perturbed images, assuming this shift captures the hallucination direction. We prove this assumption flawed: noise-induced drift in Language-Image Pretrained (LIP) encoders is a coupled vector entangling (i) structural degradation from corrupted visual information with (ii) hallucination induction from linguistic prior activation. VCD’s indiscriminate penalty inevitably suppresses valid visual semantics. Our key insight is that Self-Supervised Learning (SSL) encoders exhibit only structural degradation under noise—geometrically orthogonal to hallucination paths—enabling principled disentanglement via LIP–SSL differential response. We propose Disentangled Visual Rectification (DVR), a training-free dual-stream framework performing visual-layer rectification and decoding-layer contrast on purified representations. DVR achieves approximately $5\times$ theoretical error reduction over VCD and establishes SOTA performance on POPE, MME, LLaVA-Bench and CHAIR benchmarks.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26dz, title = {Beyond Blind Noising: Disentangled Visual Rectification for Hallucination Mitigation in {MLLM}s}, author = {Chen, Yujia and Sun, Rui and Wang, Bingzhou and Mai, Huayu and Li, Wangkai and Li, Zhaoyang and Li, Aibing and Sun, Wenzhang}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {16768--16805}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26dz/chen26dz.pdf}, url = {https://proceedings.mlr.press/v306/chen26dz.html}, abstract = {Visual Contrastive Decoding (VCD) mitigates hallucinations in Multimodal Large Language Models (MLLMs) by penalizing the output shift from noise-perturbed images, assuming this shift captures the hallucination direction. We prove this assumption flawed: noise-induced drift in Language-Image Pretrained (LIP) encoders is a coupled vector entangling (i) structural degradation from corrupted visual information with (ii) hallucination induction from linguistic prior activation. VCD’s indiscriminate penalty inevitably suppresses valid visual semantics. Our key insight is that Self-Supervised Learning (SSL) encoders exhibit only structural degradation under noise—geometrically orthogonal to hallucination paths—enabling principled disentanglement via LIP–SSL differential response. We propose Disentangled Visual Rectification (DVR), a training-free dual-stream framework performing visual-layer rectification and decoding-layer contrast on purified representations. DVR achieves approximately $5\times$ theoretical error reduction over VCD and establishes SOTA performance on POPE, MME, LLaVA-Bench and CHAIR benchmarks.} }
Endnote
%0 Conference Paper %T Beyond Blind Noising: Disentangled Visual Rectification for Hallucination Mitigation in MLLMs %A Yujia Chen %A Rui Sun %A Bingzhou Wang %A Huayu Mai %A Wangkai Li %A Zhaoyang Li %A Aibing Li %A Wenzhang Sun %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26dz %I PMLR %P 16768--16805 %U https://proceedings.mlr.press/v306/chen26dz.html %V 306 %X Visual Contrastive Decoding (VCD) mitigates hallucinations in Multimodal Large Language Models (MLLMs) by penalizing the output shift from noise-perturbed images, assuming this shift captures the hallucination direction. We prove this assumption flawed: noise-induced drift in Language-Image Pretrained (LIP) encoders is a coupled vector entangling (i) structural degradation from corrupted visual information with (ii) hallucination induction from linguistic prior activation. VCD’s indiscriminate penalty inevitably suppresses valid visual semantics. Our key insight is that Self-Supervised Learning (SSL) encoders exhibit only structural degradation under noise—geometrically orthogonal to hallucination paths—enabling principled disentanglement via LIP–SSL differential response. We propose Disentangled Visual Rectification (DVR), a training-free dual-stream framework performing visual-layer rectification and decoding-layer contrast on purified representations. DVR achieves approximately $5\times$ theoretical error reduction over VCD and establishes SOTA performance on POPE, MME, LLaVA-Bench and CHAIR benchmarks.
APA
Chen, Y., Sun, R., Wang, B., Mai, H., Li, W., Li, Z., Li, A. & Sun, W.. (2026). Beyond Blind Noising: Disentangled Visual Rectification for Hallucination Mitigation in MLLMs. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:16768-16805 Available from https://proceedings.mlr.press/v306/chen26dz.html.

Related Material