D-ARL: A Distribution-Matched Asynchronous Reinforcement Learning Framework for Language Reasoning

Yinqi Bai, Tong Xialiang, Jie Wang, Hongyu Liu, Longdi Pan, Jiashuo Li, Zehao Wang, Jianye Hao, Mingxuan Yuan, Feng Wu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:5460-5478, 2026.

Abstract

Asynchronous reinforcement learning (RL) has shown notable success in accelerating the post-training of large language models (LLMs). However, its decoupled data generation and training paradigm introduces a fundamental distributional mismatch between data generated by stale behavior policies and current policy, leading to unstable training and degraded performance. To address this challenge, we propose D-ARL, a Distribution-matched Asynchronous Reinforcement Learning framework that selects high-quality asynchronous samples whose distributions are well aligned with the current policy for policy optimization. Specifically, D-ARL maintains a replay buffer that collects samples from the most recent $K$ behavior policies and proposes a variance-guided metric to select distribution-matched data. During training, D-ARL introduces a multi-behavior policy optimization algorithm to leverage the multi-source nature of the selected samples for policy update. Experiments on six widely used reasoning benchmarks show that D-ARL outperforms state-of-the-art asynchronous methods, achieving an average improvement of 6.4% in reasoning performance and 34.7% in sample efficiency.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bai26k, title = {D-{ARL}: A Distribution-Matched Asynchronous Reinforcement Learning Framework for Language Reasoning}, author = {Bai, Yinqi and Xialiang, Tong and Wang, Jie and Liu, Hongyu and Pan, Longdi and Li, Jiashuo and Wang, Zehao and Hao, Jianye and Yuan, Mingxuan and Wu, Feng}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {5460--5478}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bai26k/bai26k.pdf}, url = {https://proceedings.mlr.press/v306/bai26k.html}, abstract = {Asynchronous reinforcement learning (RL) has shown notable success in accelerating the post-training of large language models (LLMs). However, its decoupled data generation and training paradigm introduces a fundamental distributional mismatch between data generated by stale behavior policies and current policy, leading to unstable training and degraded performance. To address this challenge, we propose D-ARL, a Distribution-matched Asynchronous Reinforcement Learning framework that selects high-quality asynchronous samples whose distributions are well aligned with the current policy for policy optimization. Specifically, D-ARL maintains a replay buffer that collects samples from the most recent $K$ behavior policies and proposes a variance-guided metric to select distribution-matched data. During training, D-ARL introduces a multi-behavior policy optimization algorithm to leverage the multi-source nature of the selected samples for policy update. Experiments on six widely used reasoning benchmarks show that D-ARL outperforms state-of-the-art asynchronous methods, achieving an average improvement of 6.4% in reasoning performance and 34.7% in sample efficiency.} }
Endnote
%0 Conference Paper %T D-ARL: A Distribution-Matched Asynchronous Reinforcement Learning Framework for Language Reasoning %A Yinqi Bai %A Tong Xialiang %A Jie Wang %A Hongyu Liu %A Longdi Pan %A Jiashuo Li %A Zehao Wang %A Jianye Hao %A Mingxuan Yuan %A Feng Wu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bai26k %I PMLR %P 5460--5478 %U https://proceedings.mlr.press/v306/bai26k.html %V 306 %X Asynchronous reinforcement learning (RL) has shown notable success in accelerating the post-training of large language models (LLMs). However, its decoupled data generation and training paradigm introduces a fundamental distributional mismatch between data generated by stale behavior policies and current policy, leading to unstable training and degraded performance. To address this challenge, we propose D-ARL, a Distribution-matched Asynchronous Reinforcement Learning framework that selects high-quality asynchronous samples whose distributions are well aligned with the current policy for policy optimization. Specifically, D-ARL maintains a replay buffer that collects samples from the most recent $K$ behavior policies and proposes a variance-guided metric to select distribution-matched data. During training, D-ARL introduces a multi-behavior policy optimization algorithm to leverage the multi-source nature of the selected samples for policy update. Experiments on six widely used reasoning benchmarks show that D-ARL outperforms state-of-the-art asynchronous methods, achieving an average improvement of 6.4% in reasoning performance and 34.7% in sample efficiency.
APA
Bai, Y., Xialiang, T., Wang, J., Liu, H., Pan, L., Li, J., Wang, Z., Hao, J., Yuan, M. & Wu, F.. (2026). D-ARL: A Distribution-Matched Asynchronous Reinforcement Learning Framework for Language Reasoning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:5460-5478 Available from https://proceedings.mlr.press/v306/bai26k.html.

Related Material