Policy-Based Trajectory Clustering in Offline Reinforcement Learning

Xinqi Wang, Simon Shaolei Du, Hao Hu
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2196-2222, 2026.

Abstract

We introduce the task of clustering trajectories in offline reinforcement learning ({RL}) datasets to address the multi-modal nature of offline data. Such datasets often contain trajectories from diverse policies, and treating them as a single distribution can obscure structure and increase distributional shift. We formalize trajectory clustering by linking the KL-divergence of offline trajectory distributions with mixtures of policy-induced distributions. To solve this, we propose Policy-Guided K-means (PG-Kmeans) and Centroid-Attracted Autoencoder (CAAE). PG-Kmeans iteratively trains behavior cloning policies and assigns trajectories based on generation probabilities, while CAAE learns continuous latent representations regularized by a learnable codebook to achieve end-to-end clustering. We prove finite-step convergence of PG-Kmeans and analyze the ambiguity of optimal solutions caused by policy-induced conflicts. Experiments on D4RL and GridWorld show that PG-Kmeans and CAAE partition trajectories into coherent clusters and offer a framework for structuring offline data, with applications in data selection, curriculum learning, and policy transfer.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-wang26a, title = {Policy-Based Trajectory Clustering in Offline Reinforcement Learning}, author = {Wang, Xinqi and Du, Simon Shaolei and Hu, Hao}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {2196--2222}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/wang26a/wang26a.pdf}, url = {https://proceedings.mlr.press/v337/wang26a.html}, abstract = {We introduce the task of clustering trajectories in offline reinforcement learning ({RL}) datasets to address the multi-modal nature of offline data. Such datasets often contain trajectories from diverse policies, and treating them as a single distribution can obscure structure and increase distributional shift. We formalize trajectory clustering by linking the KL-divergence of offline trajectory distributions with mixtures of policy-induced distributions. To solve this, we propose Policy-Guided K-means (PG-Kmeans) and Centroid-Attracted Autoencoder (CAAE). PG-Kmeans iteratively trains behavior cloning policies and assigns trajectories based on generation probabilities, while CAAE learns continuous latent representations regularized by a learnable codebook to achieve end-to-end clustering. We prove finite-step convergence of PG-Kmeans and analyze the ambiguity of optimal solutions caused by policy-induced conflicts. Experiments on D4RL and GridWorld show that PG-Kmeans and CAAE partition trajectories into coherent clusters and offer a framework for structuring offline data, with applications in data selection, curriculum learning, and policy transfer.} }
Endnote
%0 Conference Paper %T Policy-Based Trajectory Clustering in Offline Reinforcement Learning %A Xinqi Wang %A Simon Shaolei Du %A Hao Hu %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-wang26a %I PMLR %P 2196--2222 %U https://proceedings.mlr.press/v337/wang26a.html %V 337 %X We introduce the task of clustering trajectories in offline reinforcement learning ({RL}) datasets to address the multi-modal nature of offline data. Such datasets often contain trajectories from diverse policies, and treating them as a single distribution can obscure structure and increase distributional shift. We formalize trajectory clustering by linking the KL-divergence of offline trajectory distributions with mixtures of policy-induced distributions. To solve this, we propose Policy-Guided K-means (PG-Kmeans) and Centroid-Attracted Autoencoder (CAAE). PG-Kmeans iteratively trains behavior cloning policies and assigns trajectories based on generation probabilities, while CAAE learns continuous latent representations regularized by a learnable codebook to achieve end-to-end clustering. We prove finite-step convergence of PG-Kmeans and analyze the ambiguity of optimal solutions caused by policy-induced conflicts. Experiments on D4RL and GridWorld show that PG-Kmeans and CAAE partition trajectories into coherent clusters and offer a framework for structuring offline data, with applications in data selection, curriculum learning, and policy transfer.
APA
Wang, X., Du, S.S. & Hu, H.. (2026). Policy-Based Trajectory Clustering in Offline Reinforcement Learning. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:2196-2222 Available from https://proceedings.mlr.press/v337/wang26a.html.

Related Material