RELO: Reinforcement Learning to Localize for Visual Object Tracking

Xin Chen, Chuanyu Sun, Jiao Xu, Houwen Peng, Dong Wang, Huchuan Lu, Kede Ma
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:16806-16822, 2026.

Abstract

Conventional visual object trackers localize targets using handcrafted spatial priors, often in the form of heatmaps. Such priors provide only surrogate supervision and are poorly aligned with tracking optimization and evaluation metrics, such as intersection over union (IoU) and area under the success curve (AUC). Here, we introduce RELO, a REinforcement-learning-to-LOcalize method for visual object tracking that formulates target localization as a Markov decision process. Specifically, RELO replaces handcrafted spatial priors with a localization policy learned over spatial positions via reinforcement learning, with rewards combining frame-level IoU and sequence-level AUC. We additionally introduce layer-aligned temporal token propagation to improve semantic consistency across frames, with negligible computational overhead. Across multiple benchmarks, RELO achieves superior results, attaining $57.5$% AUC on LaSOT$_\mathrm{ext}$ without template updates. This confirms that reward-driven localization provides an effective alternative to prior-driven localization for visual object tracking.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26ea, title = {{RELO}: Reinforcement Learning to Localize for Visual Object Tracking}, author = {Chen, Xin and Sun, Chuanyu and Xu, Jiao and Peng, Houwen and Wang, Dong and Lu, Huchuan and Ma, Kede}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {16806--16822}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26ea/chen26ea.pdf}, url = {https://proceedings.mlr.press/v306/chen26ea.html}, abstract = {Conventional visual object trackers localize targets using handcrafted spatial priors, often in the form of heatmaps. Such priors provide only surrogate supervision and are poorly aligned with tracking optimization and evaluation metrics, such as intersection over union (IoU) and area under the success curve (AUC). Here, we introduce RELO, a REinforcement-learning-to-LOcalize method for visual object tracking that formulates target localization as a Markov decision process. Specifically, RELO replaces handcrafted spatial priors with a localization policy learned over spatial positions via reinforcement learning, with rewards combining frame-level IoU and sequence-level AUC. We additionally introduce layer-aligned temporal token propagation to improve semantic consistency across frames, with negligible computational overhead. Across multiple benchmarks, RELO achieves superior results, attaining $57.5$% AUC on LaSOT$_\mathrm{ext}$ without template updates. This confirms that reward-driven localization provides an effective alternative to prior-driven localization for visual object tracking.} }
Endnote
%0 Conference Paper %T RELO: Reinforcement Learning to Localize for Visual Object Tracking %A Xin Chen %A Chuanyu Sun %A Jiao Xu %A Houwen Peng %A Dong Wang %A Huchuan Lu %A Kede Ma %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26ea %I PMLR %P 16806--16822 %U https://proceedings.mlr.press/v306/chen26ea.html %V 306 %X Conventional visual object trackers localize targets using handcrafted spatial priors, often in the form of heatmaps. Such priors provide only surrogate supervision and are poorly aligned with tracking optimization and evaluation metrics, such as intersection over union (IoU) and area under the success curve (AUC). Here, we introduce RELO, a REinforcement-learning-to-LOcalize method for visual object tracking that formulates target localization as a Markov decision process. Specifically, RELO replaces handcrafted spatial priors with a localization policy learned over spatial positions via reinforcement learning, with rewards combining frame-level IoU and sequence-level AUC. We additionally introduce layer-aligned temporal token propagation to improve semantic consistency across frames, with negligible computational overhead. Across multiple benchmarks, RELO achieves superior results, attaining $57.5$% AUC on LaSOT$_\mathrm{ext}$ without template updates. This confirms that reward-driven localization provides an effective alternative to prior-driven localization for visual object tracking.
APA
Chen, X., Sun, C., Xu, J., Peng, H., Wang, D., Lu, H. & Ma, K.. (2026). RELO: Reinforcement Learning to Localize for Visual Object Tracking. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:16806-16822 Available from https://proceedings.mlr.press/v306/chen26ea.html.

Related Material