BabyVision: Visual Reasoning Beyond Language

Liang Chen, Weichu Xie, Yiyan Liang, Hongfeng He, Haozhe Zhao, Zhibo Yang, Zhiqi Huang, Haoning Wu, Haoyu Lu, Y. Charles, Yiping Bao, Yuantao Fan, Guopeng Li, Haiyang Shen, Xuanzhong Chen, Wendong Xu, Shuzheng Si, Zefan Cai, Wenhao Chai, Ziqi Huang, Fangfu Liu, Tianyu Liu, Baobao Chang, Ming Wu, Xiaobo Hu, Kaiyuan Chen, Yixin Ren, Yang Liu, Yuan Gong, Kuan Li
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:17692-17713, 2026.

Abstract

While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit. Code and data are released at https://github.com/UniPat-AI/BabyVision.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26fj, title = {{B}aby{V}ision: Visual Reasoning Beyond Language}, author = {Chen, Liang and Xie, Weichu and Liang, Yiyan and He, Hongfeng and Zhao, Haozhe and Yang, Zhibo and Huang, Zhiqi and Wu, Haoning and Lu, Haoyu and Charles, Y. and Bao, Yiping and Fan, Yuantao and Li, Guopeng and Shen, Haiyang and Chen, Xuanzhong and Xu, Wendong and Si, Shuzheng and Cai, Zefan and Chai, Wenhao and Huang, Ziqi and Liu, Fangfu and Liu, Tianyu and Chang, Baobao and Wu, Ming and Hu, Xiaobo and Chen, Kaiyuan and Ren, Yixin and Liu, Yang and Gong, Yuan and Li, Kuan}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {17692--17713}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26fj/chen26fj.pdf}, url = {https://proceedings.mlr.press/v306/chen26fj.html}, abstract = {While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit. Code and data are released at https://github.com/UniPat-AI/BabyVision.} }
Endnote
%0 Conference Paper %T BabyVision: Visual Reasoning Beyond Language %A Liang Chen %A Weichu Xie %A Yiyan Liang %A Hongfeng He %A Haozhe Zhao %A Zhibo Yang %A Zhiqi Huang %A Haoning Wu %A Haoyu Lu %A Y. Charles %A Yiping Bao %A Yuantao Fan %A Guopeng Li %A Haiyang Shen %A Xuanzhong Chen %A Wendong Xu %A Shuzheng Si %A Zefan Cai %A Wenhao Chai %A Ziqi Huang %A Fangfu Liu %A Tianyu Liu %A Baobao Chang %A Ming Wu %A Xiaobo Hu %A Kaiyuan Chen %A Yixin Ren %A Yang Liu %A Yuan Gong %A Kuan Li %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26fj %I PMLR %P 17692--17713 %U https://proceedings.mlr.press/v306/chen26fj.html %V 306 %X While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit. Code and data are released at https://github.com/UniPat-AI/BabyVision.
APA
Chen, L., Xie, W., Liang, Y., He, H., Zhao, H., Yang, Z., Huang, Z., Wu, H., Lu, H., Charles, Y., Bao, Y., Fan, Y., Li, G., Shen, H., Chen, X., Xu, W., Si, S., Cai, Z., Chai, W., Huang, Z., Liu, F., Liu, T., Chang, B., Wu, M., Hu, X., Chen, K., Ren, Y., Liu, Y., Gong, Y. & Li, K.. (2026). BabyVision: Visual Reasoning Beyond Language. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:17692-17713 Available from https://proceedings.mlr.press/v306/chen26fj.html.

Related Material