FlowMAP: Flow Matching for Generalizable Agent Planning

Jiarun Fu, Lizhong Ding, Ye Yuan, Qiuning Wei, Zhaohuan Linghu, Yurong Cheng, Changsheng Li, Tianlong Gu, Liang Chang, Guoren Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:31756-31771, 2026.

Abstract

Agent planning faces dynamic heterogeneity—nonstationary observations, dynamics, and objectives with sparse, delayed rewards—which dominant methods largely ignore, leading to poor generalization under environment shifts. We propose Flow-Matching for Agent Planning (FlowMAP), which formulates planning as a continuous-time flow-matching problem by learning a planning-time velocity field that transports an initial meta-state distribution toward a task-conditioned target. FlowMAP introduces Value-Transport Flow Matching to provide a distribution-level planning objective that steers transport toward high-value regions in the meta-state distribution, mitigating error accumulation under environmental shifts. To enforce alignment between meta-state distribution transport and action–environment interaction, FlowMAP further proposes Flow–Policy Co-Training, which jointly optimizes the planning flow and policy so that the flow transport directly regularizes the policy-induced meta-distribution dynamics. Across diverse agent planning benchmarks, FlowMAP consistently outperforms strong baselines, yielding improvements in planning generalization.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-fu26b, title = {{F}low{MAP}: Flow Matching for Generalizable Agent Planning}, author = {Fu, Jiarun and Ding, Lizhong and Yuan, Ye and Wei, Qiuning and Linghu, Zhaohuan and Cheng, Yurong and Li, Changsheng and Gu, Tianlong and Chang, Liang and Wang, Guoren}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {31756--31771}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/fu26b/fu26b.pdf}, url = {https://proceedings.mlr.press/v306/fu26b.html}, abstract = {Agent planning faces dynamic heterogeneity—nonstationary observations, dynamics, and objectives with sparse, delayed rewards—which dominant methods largely ignore, leading to poor generalization under environment shifts. We propose Flow-Matching for Agent Planning (FlowMAP), which formulates planning as a continuous-time flow-matching problem by learning a planning-time velocity field that transports an initial meta-state distribution toward a task-conditioned target. FlowMAP introduces Value-Transport Flow Matching to provide a distribution-level planning objective that steers transport toward high-value regions in the meta-state distribution, mitigating error accumulation under environmental shifts. To enforce alignment between meta-state distribution transport and action–environment interaction, FlowMAP further proposes Flow–Policy Co-Training, which jointly optimizes the planning flow and policy so that the flow transport directly regularizes the policy-induced meta-distribution dynamics. Across diverse agent planning benchmarks, FlowMAP consistently outperforms strong baselines, yielding improvements in planning generalization.} }
Endnote
%0 Conference Paper %T FlowMAP: Flow Matching for Generalizable Agent Planning %A Jiarun Fu %A Lizhong Ding %A Ye Yuan %A Qiuning Wei %A Zhaohuan Linghu %A Yurong Cheng %A Changsheng Li %A Tianlong Gu %A Liang Chang %A Guoren Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-fu26b %I PMLR %P 31756--31771 %U https://proceedings.mlr.press/v306/fu26b.html %V 306 %X Agent planning faces dynamic heterogeneity—nonstationary observations, dynamics, and objectives with sparse, delayed rewards—which dominant methods largely ignore, leading to poor generalization under environment shifts. We propose Flow-Matching for Agent Planning (FlowMAP), which formulates planning as a continuous-time flow-matching problem by learning a planning-time velocity field that transports an initial meta-state distribution toward a task-conditioned target. FlowMAP introduces Value-Transport Flow Matching to provide a distribution-level planning objective that steers transport toward high-value regions in the meta-state distribution, mitigating error accumulation under environmental shifts. To enforce alignment between meta-state distribution transport and action–environment interaction, FlowMAP further proposes Flow–Policy Co-Training, which jointly optimizes the planning flow and policy so that the flow transport directly regularizes the policy-induced meta-distribution dynamics. Across diverse agent planning benchmarks, FlowMAP consistently outperforms strong baselines, yielding improvements in planning generalization.
APA
Fu, J., Ding, L., Yuan, Y., Wei, Q., Linghu, Z., Cheng, Y., Li, C., Gu, T., Chang, L. & Wang, G.. (2026). FlowMAP: Flow Matching for Generalizable Agent Planning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:31756-31771 Available from https://proceedings.mlr.press/v306/fu26b.html.

Related Material