LINGUA: Bridging the Grounding Gap in VideoQA via Typed Memory and Belief-State Reasoning

Saman Forouzandeh, Wei Peng, Xinghuo Yu, Mahdi Jalili
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:31437-31490, 2026.

Abstract

VideoQA models can be accurate yet often fail to align answers with the correct video segments (the grounding gap). We introduce LINGUA (Language-based INference for Grounded Video Understanding Agent), a memory-based agent that performs grounded VideoQA by reasoning in an explicit linguistic belief state. LINGUA uses five mechanisms: (1) event-driven perception (retains 8–12% of frames while preserving 94% of question-relevant events); (2) typed memory for episodic narratives, semantic affordances, and procedural scripts; (3) Belief-Action-Verification loops with postcondition and temporal checks; (4) meta reflection with contrastive refinement; and (5) Bayesian reliability tracking for continual learning without gradient updates. Built with Gemma3-4B (Ollama, 4-bit), LINGUA outperforms strong baselines on five VideoQA benchmarks, reaching 82.4% on NExT-QA and 42.3% Acc@GQA on NExT-GQA (answer + IoU$\geq$0.5 temporal localization), while running 2.6$\times$ faster than dense-frame methods. In continual learning over 100 videos, accuracy rises from 45.2% (first 10) to 61.8% (last 10) without catastrophic forgetting, indicating online adaptation via memory refinement.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-forouzandeh26a, title = {{LINGUA}: Bridging the Grounding Gap in {V}ideo{QA} via Typed Memory and Belief-State Reasoning}, author = {Forouzandeh, Saman and Peng, Wei and Yu, Xinghuo and Jalili, Mahdi}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {31437--31490}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/forouzandeh26a/forouzandeh26a.pdf}, url = {https://proceedings.mlr.press/v306/forouzandeh26a.html}, abstract = {VideoQA models can be accurate yet often fail to align answers with the correct video segments (the grounding gap). We introduce LINGUA (Language-based INference for Grounded Video Understanding Agent), a memory-based agent that performs grounded VideoQA by reasoning in an explicit linguistic belief state. LINGUA uses five mechanisms: (1) event-driven perception (retains 8–12% of frames while preserving 94% of question-relevant events); (2) typed memory for episodic narratives, semantic affordances, and procedural scripts; (3) Belief-Action-Verification loops with postcondition and temporal checks; (4) meta reflection with contrastive refinement; and (5) Bayesian reliability tracking for continual learning without gradient updates. Built with Gemma3-4B (Ollama, 4-bit), LINGUA outperforms strong baselines on five VideoQA benchmarks, reaching 82.4% on NExT-QA and 42.3% Acc@GQA on NExT-GQA (answer + IoU$\geq$0.5 temporal localization), while running 2.6$\times$ faster than dense-frame methods. In continual learning over 100 videos, accuracy rises from 45.2% (first 10) to 61.8% (last 10) without catastrophic forgetting, indicating online adaptation via memory refinement.} }
Endnote
%0 Conference Paper %T LINGUA: Bridging the Grounding Gap in VideoQA via Typed Memory and Belief-State Reasoning %A Saman Forouzandeh %A Wei Peng %A Xinghuo Yu %A Mahdi Jalili %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-forouzandeh26a %I PMLR %P 31437--31490 %U https://proceedings.mlr.press/v306/forouzandeh26a.html %V 306 %X VideoQA models can be accurate yet often fail to align answers with the correct video segments (the grounding gap). We introduce LINGUA (Language-based INference for Grounded Video Understanding Agent), a memory-based agent that performs grounded VideoQA by reasoning in an explicit linguistic belief state. LINGUA uses five mechanisms: (1) event-driven perception (retains 8–12% of frames while preserving 94% of question-relevant events); (2) typed memory for episodic narratives, semantic affordances, and procedural scripts; (3) Belief-Action-Verification loops with postcondition and temporal checks; (4) meta reflection with contrastive refinement; and (5) Bayesian reliability tracking for continual learning without gradient updates. Built with Gemma3-4B (Ollama, 4-bit), LINGUA outperforms strong baselines on five VideoQA benchmarks, reaching 82.4% on NExT-QA and 42.3% Acc@GQA on NExT-GQA (answer + IoU$\geq$0.5 temporal localization), while running 2.6$\times$ faster than dense-frame methods. In continual learning over 100 videos, accuracy rises from 45.2% (first 10) to 61.8% (last 10) without catastrophic forgetting, indicating online adaptation via memory refinement.
APA
Forouzandeh, S., Peng, W., Yu, X. & Jalili, M.. (2026). LINGUA: Bridging the Grounding Gap in VideoQA via Typed Memory and Belief-State Reasoning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:31437-31490 Available from https://proceedings.mlr.press/v306/forouzandeh26a.html.

Related Material