CorrectionPlanner: Self-Correction Planner with Reinforcement Learning in Autonomous Driving

Yihong Guo, Dongqiangzi Ye, Sijia Chen, Anqi Liu, Xianming Liu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:38576-38595, 2026.

Abstract

Autonomous driving requires safe planning, but most learning-based planners lack explicit self-correction ability: once an unsafe action is proposed, there is no mechanism to correct it. Thus, we propose CorrectionPlanner, an autoregressive planner with self-correction that contains a propose, evaluate, and correct loop in the motion-token generation process. At each planning step, the policy proposes an action, namely a motion token, and a learned collision critic predicts whether it will induce a collision within a short horizon. If the critic predicts a collision, we retain the sequence of historical unsafe motion tokens as a self-correction trace, generate the next motion token conditioned on it, and repeat this process until the safe motion token is proposed or the safety criterion is met. This self-correction trace, consisting of all the unsafe motion tokens, represents the planner’s correction process in motion-token space. We train the planner with imitation learning followed by model-based reinforcement learning using rollouts from a pretrained world model that realistically models agents’ reactive behaviors. Closed-loop evaluations show that CorrectionPlanner reduces the collision rate by over $20%$ on Waymax and obtains state-of-the-art planning scores on nuPlan.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-guo26ae, title = {{C}orrection{P}lanner: Self-Correction Planner with Reinforcement Learning in Autonomous Driving}, author = {Guo, Yihong and Ye, Dongqiangzi and Chen, Sijia and Liu, Anqi and Liu, Xianming}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {38576--38595}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/guo26ae/guo26ae.pdf}, url = {https://proceedings.mlr.press/v306/guo26ae.html}, abstract = {Autonomous driving requires safe planning, but most learning-based planners lack explicit self-correction ability: once an unsafe action is proposed, there is no mechanism to correct it. Thus, we propose CorrectionPlanner, an autoregressive planner with self-correction that contains a propose, evaluate, and correct loop in the motion-token generation process. At each planning step, the policy proposes an action, namely a motion token, and a learned collision critic predicts whether it will induce a collision within a short horizon. If the critic predicts a collision, we retain the sequence of historical unsafe motion tokens as a self-correction trace, generate the next motion token conditioned on it, and repeat this process until the safe motion token is proposed or the safety criterion is met. This self-correction trace, consisting of all the unsafe motion tokens, represents the planner’s correction process in motion-token space. We train the planner with imitation learning followed by model-based reinforcement learning using rollouts from a pretrained world model that realistically models agents’ reactive behaviors. Closed-loop evaluations show that CorrectionPlanner reduces the collision rate by over $20%$ on Waymax and obtains state-of-the-art planning scores on nuPlan.} }
Endnote
%0 Conference Paper %T CorrectionPlanner: Self-Correction Planner with Reinforcement Learning in Autonomous Driving %A Yihong Guo %A Dongqiangzi Ye %A Sijia Chen %A Anqi Liu %A Xianming Liu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-guo26ae %I PMLR %P 38576--38595 %U https://proceedings.mlr.press/v306/guo26ae.html %V 306 %X Autonomous driving requires safe planning, but most learning-based planners lack explicit self-correction ability: once an unsafe action is proposed, there is no mechanism to correct it. Thus, we propose CorrectionPlanner, an autoregressive planner with self-correction that contains a propose, evaluate, and correct loop in the motion-token generation process. At each planning step, the policy proposes an action, namely a motion token, and a learned collision critic predicts whether it will induce a collision within a short horizon. If the critic predicts a collision, we retain the sequence of historical unsafe motion tokens as a self-correction trace, generate the next motion token conditioned on it, and repeat this process until the safe motion token is proposed or the safety criterion is met. This self-correction trace, consisting of all the unsafe motion tokens, represents the planner’s correction process in motion-token space. We train the planner with imitation learning followed by model-based reinforcement learning using rollouts from a pretrained world model that realistically models agents’ reactive behaviors. Closed-loop evaluations show that CorrectionPlanner reduces the collision rate by over $20%$ on Waymax and obtains state-of-the-art planning scores on nuPlan.
APA
Guo, Y., Ye, D., Chen, S., Liu, A. & Liu, X.. (2026). CorrectionPlanner: Self-Correction Planner with Reinforcement Learning in Autonomous Driving. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:38576-38595 Available from https://proceedings.mlr.press/v306/guo26ae.html.

Related Material