MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

Huiyi Chen, Jiawei Peng, Dehai Min, Changchang Sun, Kaijie Chen, Yan Yan, Xu Yang, Lu Cheng
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:16324-16341, 2026.

Abstract

Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and responsible deployment. However, existing robustness benchmarks largely focus on hallucination or misleading textual inputs, overlooking the critical challenge posed by misleading visual inputs in assessing visual understanding. To fill this gap, we introduce MVI-Bench, the first comprehensive benchmark specially designed for evaluating how Misleading Visual Inputs undermine the robustness of LVLMs. Grounded in fundamental visual primitives, the design of MVI-Bench centers on three hierarchical levels of misleading visual inputs: Visual Concept, Visual Attribute, and Visual Relationship. Using this taxonomy, we curate six representative categories and compile 1,248 expertly annotated VQA instances. To facilitate fine-grained robustness evaluation, we further introduce MVI-Sensitivity, a novel metric that characterizes LVLM robustness. Empirical results across 18 state-of-the-art LVLMs uncover pronounced vulnerabilities to misleading visual inputs, and our in-depth analyses on MVI-Bench provide actionable insights that can guide the development of more reliable and robust LVLMs.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26dh, title = {{MVI}-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in {LVLM}s}, author = {Chen, Huiyi and Peng, Jiawei and Min, Dehai and Sun, Changchang and Chen, Kaijie and Yan, Yan and Yang, Xu and Cheng, Lu}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {16324--16341}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26dh/chen26dh.pdf}, url = {https://proceedings.mlr.press/v306/chen26dh.html}, abstract = {Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and responsible deployment. However, existing robustness benchmarks largely focus on hallucination or misleading textual inputs, overlooking the critical challenge posed by misleading visual inputs in assessing visual understanding. To fill this gap, we introduce MVI-Bench, the first comprehensive benchmark specially designed for evaluating how Misleading Visual Inputs undermine the robustness of LVLMs. Grounded in fundamental visual primitives, the design of MVI-Bench centers on three hierarchical levels of misleading visual inputs: Visual Concept, Visual Attribute, and Visual Relationship. Using this taxonomy, we curate six representative categories and compile 1,248 expertly annotated VQA instances. To facilitate fine-grained robustness evaluation, we further introduce MVI-Sensitivity, a novel metric that characterizes LVLM robustness. Empirical results across 18 state-of-the-art LVLMs uncover pronounced vulnerabilities to misleading visual inputs, and our in-depth analyses on MVI-Bench provide actionable insights that can guide the development of more reliable and robust LVLMs.} }
Endnote
%0 Conference Paper %T MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs %A Huiyi Chen %A Jiawei Peng %A Dehai Min %A Changchang Sun %A Kaijie Chen %A Yan Yan %A Xu Yang %A Lu Cheng %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26dh %I PMLR %P 16324--16341 %U https://proceedings.mlr.press/v306/chen26dh.html %V 306 %X Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and responsible deployment. However, existing robustness benchmarks largely focus on hallucination or misleading textual inputs, overlooking the critical challenge posed by misleading visual inputs in assessing visual understanding. To fill this gap, we introduce MVI-Bench, the first comprehensive benchmark specially designed for evaluating how Misleading Visual Inputs undermine the robustness of LVLMs. Grounded in fundamental visual primitives, the design of MVI-Bench centers on three hierarchical levels of misleading visual inputs: Visual Concept, Visual Attribute, and Visual Relationship. Using this taxonomy, we curate six representative categories and compile 1,248 expertly annotated VQA instances. To facilitate fine-grained robustness evaluation, we further introduce MVI-Sensitivity, a novel metric that characterizes LVLM robustness. Empirical results across 18 state-of-the-art LVLMs uncover pronounced vulnerabilities to misleading visual inputs, and our in-depth analyses on MVI-Bench provide actionable insights that can guide the development of more reliable and robust LVLMs.
APA
Chen, H., Peng, J., Min, D., Sun, C., Chen, K., Yan, Y., Yang, X. & Cheng, L.. (2026). MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:16324-16341 Available from https://proceedings.mlr.press/v306/chen26dh.html.

Related Material