$G^2$-Reader: Dual Evolving Graphs for Multimodal Document QA

Yaxin Du, Junru Song, Yifan Zhou, Cheng Wang, Jiahao Gu, Zimeng Chen, Menglan Chen, Wen Yao, Yang Yang, Ying Wen, Siheng Chen
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:26624-26648, 2026.

Abstract

Retrieval-augmented generation is a practical paradigm for question answering over long documents, but it remains brittle for multimodal reading where text, tables, and figures are interleaved across many pages. First, flat chunking breaks document-native structure and cross-modal alignment, yielding semantic fragments that are hard to interpret in isolation. Second, even iterative retrieval can fail in long contexts by looping on partial evidence or drifting into irrelevant sections as noise accumulates, since each step is guided only by the current snippet without a persistent global search state. We introduce $G^2$-Reader, a dual-graph system, to address both issues. It evolves a Content Graph to preserve document-native structure and cross-modal semantics, and maintains a Planning Graph, an agentic directed acyclic graph of sub-questions, to track intermediate findings and guide stepwise navigation for evidence completion. On VisDoMBench across five multimodal domains, $G^2$-Reader with Qwen3-VL-32B-Instruct reaches 66.21% average accuracy, outperforming strong baselines and a standalone GPT-5 (53.08%). Code is available: https://github.com/DorothyDUUU/G2_Reader.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-du26k, title = {$G^2$-Reader: Dual Evolving Graphs for Multimodal Document {QA}}, author = {Du, Yaxin and Song, Junru and Zhou, Yifan and Wang, Cheng and Gu, Jiahao and Chen, Zimeng and Chen, Menglan and Yao, Wen and Yang, Yang and Wen, Ying and Chen, Siheng}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {26624--26648}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/du26k/du26k.pdf}, url = {https://proceedings.mlr.press/v306/du26k.html}, abstract = {Retrieval-augmented generation is a practical paradigm for question answering over long documents, but it remains brittle for multimodal reading where text, tables, and figures are interleaved across many pages. First, flat chunking breaks document-native structure and cross-modal alignment, yielding semantic fragments that are hard to interpret in isolation. Second, even iterative retrieval can fail in long contexts by looping on partial evidence or drifting into irrelevant sections as noise accumulates, since each step is guided only by the current snippet without a persistent global search state. We introduce $G^2$-Reader, a dual-graph system, to address both issues. It evolves a Content Graph to preserve document-native structure and cross-modal semantics, and maintains a Planning Graph, an agentic directed acyclic graph of sub-questions, to track intermediate findings and guide stepwise navigation for evidence completion. On VisDoMBench across five multimodal domains, $G^2$-Reader with Qwen3-VL-32B-Instruct reaches 66.21% average accuracy, outperforming strong baselines and a standalone GPT-5 (53.08%). Code is available: https://github.com/DorothyDUUU/G2_Reader.} }
Endnote
%0 Conference Paper %T $G^2$-Reader: Dual Evolving Graphs for Multimodal Document QA %A Yaxin Du %A Junru Song %A Yifan Zhou %A Cheng Wang %A Jiahao Gu %A Zimeng Chen %A Menglan Chen %A Wen Yao %A Yang Yang %A Ying Wen %A Siheng Chen %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-du26k %I PMLR %P 26624--26648 %U https://proceedings.mlr.press/v306/du26k.html %V 306 %X Retrieval-augmented generation is a practical paradigm for question answering over long documents, but it remains brittle for multimodal reading where text, tables, and figures are interleaved across many pages. First, flat chunking breaks document-native structure and cross-modal alignment, yielding semantic fragments that are hard to interpret in isolation. Second, even iterative retrieval can fail in long contexts by looping on partial evidence or drifting into irrelevant sections as noise accumulates, since each step is guided only by the current snippet without a persistent global search state. We introduce $G^2$-Reader, a dual-graph system, to address both issues. It evolves a Content Graph to preserve document-native structure and cross-modal semantics, and maintains a Planning Graph, an agentic directed acyclic graph of sub-questions, to track intermediate findings and guide stepwise navigation for evidence completion. On VisDoMBench across five multimodal domains, $G^2$-Reader with Qwen3-VL-32B-Instruct reaches 66.21% average accuracy, outperforming strong baselines and a standalone GPT-5 (53.08%). Code is available: https://github.com/DorothyDUUU/G2_Reader.
APA
Du, Y., Song, J., Zhou, Y., Wang, C., Gu, J., Chen, Z., Chen, M., Yao, W., Yang, Y., Wen, Y. & Chen, S.. (2026). $G^2$-Reader: Dual Evolving Graphs for Multimodal Document QA. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:26624-26648 Available from https://proceedings.mlr.press/v306/du26k.html.

Related Material