ImpText: A Benchmark and Tool-Augmented Framework for Implicit Text Reasoning

Litao Guo, Jinsong Zhou, Shuaibo Li, Man Chen, Xinli Xu, Zixin Zhang, Harold Haodong Chen, Ying-Cong Chen
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:38701-38719, 2026.

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated exceptional proficiency in standard text extraction, but they encounter significant challenges when confronting real-world implicit text. Such content typically contains malicious information, intentionally concealed through physical deformation, visual camouflage, or cognitive suggestion. These concealment techniques circumvent content moderation systems and pose severe risks to user safety. To bridge the research gap in text recognition under real-world adversarial scenarios, we define the task of Implicit Text Reasoning and introduce ImpText-Bench, a meticulously constructed benchmark. Extensive evaluations on this benchmark reveal significant vulnerability in current systems; even advanced proprietary models achieve a maximum Text Match Score of only 35.79%. In response, we propose ImpText-Reader, a tool-augmented framework. It employs a three-stage training strategy utilizing capability-boundary data to collaboratively optimize tool selection and semantic reasoning, thereby effectively extracting hidden text. Extensive experiments demonstrate that our approach achieves SOTA performance, significantly enhancing model robustness in adversarial environments.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-guo26aj, title = {{I}mp{T}ext: A Benchmark and Tool-Augmented Framework for Implicit Text Reasoning}, author = {Guo, Litao and Zhou, Jinsong and Li, Shuaibo and Chen, Man and Xu, Xinli and Zhang, Zixin and Chen, Harold Haodong and Chen, Ying-Cong}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {38701--38719}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/guo26aj/guo26aj.pdf}, url = {https://proceedings.mlr.press/v306/guo26aj.html}, abstract = {Multimodal Large Language Models (MLLMs) have demonstrated exceptional proficiency in standard text extraction, but they encounter significant challenges when confronting real-world implicit text. Such content typically contains malicious information, intentionally concealed through physical deformation, visual camouflage, or cognitive suggestion. These concealment techniques circumvent content moderation systems and pose severe risks to user safety. To bridge the research gap in text recognition under real-world adversarial scenarios, we define the task of Implicit Text Reasoning and introduce ImpText-Bench, a meticulously constructed benchmark. Extensive evaluations on this benchmark reveal significant vulnerability in current systems; even advanced proprietary models achieve a maximum Text Match Score of only 35.79%. In response, we propose ImpText-Reader, a tool-augmented framework. It employs a three-stage training strategy utilizing capability-boundary data to collaboratively optimize tool selection and semantic reasoning, thereby effectively extracting hidden text. Extensive experiments demonstrate that our approach achieves SOTA performance, significantly enhancing model robustness in adversarial environments.} }
Endnote
%0 Conference Paper %T ImpText: A Benchmark and Tool-Augmented Framework for Implicit Text Reasoning %A Litao Guo %A Jinsong Zhou %A Shuaibo Li %A Man Chen %A Xinli Xu %A Zixin Zhang %A Harold Haodong Chen %A Ying-Cong Chen %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-guo26aj %I PMLR %P 38701--38719 %U https://proceedings.mlr.press/v306/guo26aj.html %V 306 %X Multimodal Large Language Models (MLLMs) have demonstrated exceptional proficiency in standard text extraction, but they encounter significant challenges when confronting real-world implicit text. Such content typically contains malicious information, intentionally concealed through physical deformation, visual camouflage, or cognitive suggestion. These concealment techniques circumvent content moderation systems and pose severe risks to user safety. To bridge the research gap in text recognition under real-world adversarial scenarios, we define the task of Implicit Text Reasoning and introduce ImpText-Bench, a meticulously constructed benchmark. Extensive evaluations on this benchmark reveal significant vulnerability in current systems; even advanced proprietary models achieve a maximum Text Match Score of only 35.79%. In response, we propose ImpText-Reader, a tool-augmented framework. It employs a three-stage training strategy utilizing capability-boundary data to collaboratively optimize tool selection and semantic reasoning, thereby effectively extracting hidden text. Extensive experiments demonstrate that our approach achieves SOTA performance, significantly enhancing model robustness in adversarial environments.
APA
Guo, L., Zhou, J., Li, S., Chen, M., Xu, X., Zhang, Z., Chen, H.H. & Chen, Y.. (2026). ImpText: A Benchmark and Tool-Augmented Framework for Implicit Text Reasoning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:38701-38719 Available from https://proceedings.mlr.press/v306/guo26aj.html.

Related Material