Laplacian Flows for Policy Learning from Experience

Xingrui Gu, Chuyi Jiang
Proceedings of GRaM: the Second Edition of the Workshop on Geometry-grounded Representation Learning and Generative Modeling, PMLR 326:227-261, 2026.

Abstract

Many learning and decision-making systems output conditional distributions rather than point predictions, yet are trained via locally “reasonable” myopic gradient updates that implicitly assume their composition remains globally stable and feasible. In RL (policy gradient/actor–critic) and LLMs (cross-entropy), drift is typically controlled by KL/Fisher trust regions, which need not reflect the true behavioral scale of policy change, so small per-step moves can accumulate into large transport-scale shifts that break stability, long-horizon evidence integration, and robustness (like a millimeter map error causing a catastrophic fall in physical space). We propose the Policy Laplacian Trace (PLT): retrieved historical policies define an OT-induced local graph, and each update solves a variational OT+KL proximal step coupling a Wasserstein barycenter term with KL regularization, yielding experience-induced Laplacian smoothing of task-gradient drift. Geometrically, PLT connects to Laplace learning in Wasserstein space: its discrete graph energy approximates a $p$-Dirichlet/Laplace–Beltrami energy on the realizable policy subset. Empirically, PLT is plug-and-play and improves PPO/MAPPO stability, sample efficiency, and robustness under controlled shifts, and strengthens LLM-as-policy performance on counterfactual trust, long-range factual recall, and few-shot novel-category learning across GPT-family models, while maintaining or improving base performance and calibration.

Cite this Paper


BibTeX
@InProceedings{pmlr-v326-gu26a, title = {Laplacian Flows for Policy Learning from Experience}, author = {Gu, Xingrui and Jiang, Chuyi}, booktitle = {Proceedings of GRaM: the Second Edition of the Workshop on Geometry-grounded Representation Learning and Generative Modeling}, pages = {227--261}, year = {2026}, editor = {Pouplin, Alison and Vadgama, Sharvaree and Bekkers, Erik and Kaba, Sékou-Oumar and Lawrence, Hannah and Lecha, Manuel and Baker, Elizabeth and Suk, Julian and Walters, Robin and Tomczak, Jakub and Jegelka, Stefanie}, volume = {326}, series = {Proceedings of Machine Learning Research}, month = {26 Apr}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v326/main/assets/gu26a/gu26a.pdf}, url = {https://proceedings.mlr.press/v326/gu26a.html}, abstract = {Many learning and decision-making systems output conditional distributions rather than point predictions, yet are trained via locally “reasonable” myopic gradient updates that implicitly assume their composition remains globally stable and feasible. In RL (policy gradient/actor–critic) and LLMs (cross-entropy), drift is typically controlled by KL/Fisher trust regions, which need not reflect the true behavioral scale of policy change, so small per-step moves can accumulate into large transport-scale shifts that break stability, long-horizon evidence integration, and robustness (like a millimeter map error causing a catastrophic fall in physical space). We propose the Policy Laplacian Trace (PLT): retrieved historical policies define an OT-induced local graph, and each update solves a variational OT+KL proximal step coupling a Wasserstein barycenter term with KL regularization, yielding experience-induced Laplacian smoothing of task-gradient drift. Geometrically, PLT connects to Laplace learning in Wasserstein space: its discrete graph energy approximates a $p$-Dirichlet/Laplace–Beltrami energy on the realizable policy subset. Empirically, PLT is plug-and-play and improves PPO/MAPPO stability, sample efficiency, and robustness under controlled shifts, and strengthens LLM-as-policy performance on counterfactual trust, long-range factual recall, and few-shot novel-category learning across GPT-family models, while maintaining or improving base performance and calibration.} }
Endnote
%0 Conference Paper %T Laplacian Flows for Policy Learning from Experience %A Xingrui Gu %A Chuyi Jiang %B Proceedings of GRaM: the Second Edition of the Workshop on Geometry-grounded Representation Learning and Generative Modeling %C Proceedings of Machine Learning Research %D 2026 %E Alison Pouplin %E Sharvaree Vadgama %E Erik Bekkers %E Sékou-Oumar Kaba %E Hannah Lawrence %E Manuel Lecha %E Elizabeth Baker %E Julian Suk %E Robin Walters %E Jakub Tomczak %E Stefanie Jegelka %F pmlr-v326-gu26a %I PMLR %P 227--261 %U https://proceedings.mlr.press/v326/gu26a.html %V 326 %X Many learning and decision-making systems output conditional distributions rather than point predictions, yet are trained via locally “reasonable” myopic gradient updates that implicitly assume their composition remains globally stable and feasible. In RL (policy gradient/actor–critic) and LLMs (cross-entropy), drift is typically controlled by KL/Fisher trust regions, which need not reflect the true behavioral scale of policy change, so small per-step moves can accumulate into large transport-scale shifts that break stability, long-horizon evidence integration, and robustness (like a millimeter map error causing a catastrophic fall in physical space). We propose the Policy Laplacian Trace (PLT): retrieved historical policies define an OT-induced local graph, and each update solves a variational OT+KL proximal step coupling a Wasserstein barycenter term with KL regularization, yielding experience-induced Laplacian smoothing of task-gradient drift. Geometrically, PLT connects to Laplace learning in Wasserstein space: its discrete graph energy approximates a $p$-Dirichlet/Laplace–Beltrami energy on the realizable policy subset. Empirically, PLT is plug-and-play and improves PPO/MAPPO stability, sample efficiency, and robustness under controlled shifts, and strengthens LLM-as-policy performance on counterfactual trust, long-range factual recall, and few-shot novel-category learning across GPT-family models, while maintaining or improving base performance and calibration.
APA
Gu, X. & Jiang, C.. (2026). Laplacian Flows for Policy Learning from Experience. Proceedings of GRaM: the Second Edition of the Workshop on Geometry-grounded Representation Learning and Generative Modeling, in Proceedings of Machine Learning Research 326:227-261 Available from https://proceedings.mlr.press/v326/gu26a.html.

Related Material