VideoTrace-R1: Long Video-based Retrieval-Augmented Generation via Reinforcement Learning

Zongsheng Cao, Anran Liu, Jun Xie, Feng Chen, Lang Chen, Jing Li, Zhepeng Wang, Zigan Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:11365-11376, 2026.

Abstract

Long-video temporal reasoning remains a bottleneck for Large Video Language Models (LVLMs). Existing reinforcement-learning approaches reward only final-answer correctness, so they cannot distinguish answers reached through grounded reasoning from those reached through fabricated chronology; the intermediate temporal claims that constitute the reasoning are never verified. We trace this gap to a structural correspondence between two kinds of traces: a video has its own temporal trace, an ordered sequence of how events unfold, while a model’s answer is built up through a reasoning trace, an ordered sequence of intermediate temporal claims. Correct reasoning requires the latter to mirror the former, claim by claim. We act on this correspondence with two contributions. We introduce Temporal Reasoning Traces (TRT), a structured index of a video’s ordered event chains that exposes a small set of deterministic verification primitives, materializing the temporal trace as a programmatically queryable object. We then propose temporal-enhanced GRPO, a reinforcement-learning procedure whose reward decomposes into per-block components, each computed by a TRT primitive on a typed think block of the reasoning trace. Because the reward is fully symbolic, fabricated temporal claims are caught at the per-claim level rather than masked by a correct final answer. Across long-video reasoning benchmarks, our model achieves state-of-the-art performance, with the largest gains on out-of-domain reasoning tasks such as Video-Holmes, CG-Bench-Reasoning, and VRBench.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cao26m, title = {{V}ideo{T}race-R1: Long Video-based Retrieval-Augmented Generation via Reinforcement Learning}, author = {Cao, Zongsheng and Liu, Anran and Xie, Jun and Chen, Feng and Chen, Lang and Li, Jing and Wang, Zhepeng and Wang, Zigan}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {11365--11376}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cao26m/cao26m.pdf}, url = {https://proceedings.mlr.press/v306/cao26m.html}, abstract = {Long-video temporal reasoning remains a bottleneck for Large Video Language Models (LVLMs). Existing reinforcement-learning approaches reward only final-answer correctness, so they cannot distinguish answers reached through grounded reasoning from those reached through fabricated chronology; the intermediate temporal claims that constitute the reasoning are never verified. We trace this gap to a structural correspondence between two kinds of traces: a video has its own temporal trace, an ordered sequence of how events unfold, while a model’s answer is built up through a reasoning trace, an ordered sequence of intermediate temporal claims. Correct reasoning requires the latter to mirror the former, claim by claim. We act on this correspondence with two contributions. We introduce Temporal Reasoning Traces (TRT), a structured index of a video’s ordered event chains that exposes a small set of deterministic verification primitives, materializing the temporal trace as a programmatically queryable object. We then propose temporal-enhanced GRPO, a reinforcement-learning procedure whose reward decomposes into per-block components, each computed by a TRT primitive on a typed think block of the reasoning trace. Because the reward is fully symbolic, fabricated temporal claims are caught at the per-claim level rather than masked by a correct final answer. Across long-video reasoning benchmarks, our model achieves state-of-the-art performance, with the largest gains on out-of-domain reasoning tasks such as Video-Holmes, CG-Bench-Reasoning, and VRBench.} }
Endnote
%0 Conference Paper %T VideoTrace-R1: Long Video-based Retrieval-Augmented Generation via Reinforcement Learning %A Zongsheng Cao %A Anran Liu %A Jun Xie %A Feng Chen %A Lang Chen %A Jing Li %A Zhepeng Wang %A Zigan Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cao26m %I PMLR %P 11365--11376 %U https://proceedings.mlr.press/v306/cao26m.html %V 306 %X Long-video temporal reasoning remains a bottleneck for Large Video Language Models (LVLMs). Existing reinforcement-learning approaches reward only final-answer correctness, so they cannot distinguish answers reached through grounded reasoning from those reached through fabricated chronology; the intermediate temporal claims that constitute the reasoning are never verified. We trace this gap to a structural correspondence between two kinds of traces: a video has its own temporal trace, an ordered sequence of how events unfold, while a model’s answer is built up through a reasoning trace, an ordered sequence of intermediate temporal claims. Correct reasoning requires the latter to mirror the former, claim by claim. We act on this correspondence with two contributions. We introduce Temporal Reasoning Traces (TRT), a structured index of a video’s ordered event chains that exposes a small set of deterministic verification primitives, materializing the temporal trace as a programmatically queryable object. We then propose temporal-enhanced GRPO, a reinforcement-learning procedure whose reward decomposes into per-block components, each computed by a TRT primitive on a typed think block of the reasoning trace. Because the reward is fully symbolic, fabricated temporal claims are caught at the per-claim level rather than masked by a correct final answer. Across long-video reasoning benchmarks, our model achieves state-of-the-art performance, with the largest gains on out-of-domain reasoning tasks such as Video-Holmes, CG-Bench-Reasoning, and VRBench.
APA
Cao, Z., Liu, A., Xie, J., Chen, F., Chen, L., Li, J., Wang, Z. & Wang, Z.. (2026). VideoTrace-R1: Long Video-based Retrieval-Augmented Generation via Reinforcement Learning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:11365-11376 Available from https://proceedings.mlr.press/v306/cao26m.html.

Related Material