AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training

Ling Chen, Houming Wu, Wenjie Yu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:17586-17602, 2026.

Abstract

Pipeline parallelism is essential for large-scale model training, but existing asynchronous approaches often degrade convergence due to parameter mismatch between forward and backward passes. We propose Asynchronous Multi-Directional Pipeline parallelism (AMDP) to mitigate this issue while sustaining high utilization. AMDP limits the first stage of each pipeline to process at most two minibatches before backpropagation, bounding the number of parameter updates between forward and backward passes. To alleviate the resulting pipeline bubbles, AMDP launches multiple concurrent pipelines and adapts their number according to pipeline depth. In addition, AMDP accumulates gradients across minibatches and applies them in a single update, ensuring that only a bounded number of minibatches experience parameter mismatch, limited to within one optimization step. Experiments on GPT- and BERT-style models demonstrate that AMDP significantly accelerates training while preserving convergence.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26ff, title = {{AMDP}: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training}, author = {Chen, Ling and Wu, Houming and Yu, Wenjie}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {17586--17602}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26ff/chen26ff.pdf}, url = {https://proceedings.mlr.press/v306/chen26ff.html}, abstract = {Pipeline parallelism is essential for large-scale model training, but existing asynchronous approaches often degrade convergence due to parameter mismatch between forward and backward passes. We propose Asynchronous Multi-Directional Pipeline parallelism (AMDP) to mitigate this issue while sustaining high utilization. AMDP limits the first stage of each pipeline to process at most two minibatches before backpropagation, bounding the number of parameter updates between forward and backward passes. To alleviate the resulting pipeline bubbles, AMDP launches multiple concurrent pipelines and adapts their number according to pipeline depth. In addition, AMDP accumulates gradients across minibatches and applies them in a single update, ensuring that only a bounded number of minibatches experience parameter mismatch, limited to within one optimization step. Experiments on GPT- and BERT-style models demonstrate that AMDP significantly accelerates training while preserving convergence.} }
Endnote
%0 Conference Paper %T AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training %A Ling Chen %A Houming Wu %A Wenjie Yu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26ff %I PMLR %P 17586--17602 %U https://proceedings.mlr.press/v306/chen26ff.html %V 306 %X Pipeline parallelism is essential for large-scale model training, but existing asynchronous approaches often degrade convergence due to parameter mismatch between forward and backward passes. We propose Asynchronous Multi-Directional Pipeline parallelism (AMDP) to mitigate this issue while sustaining high utilization. AMDP limits the first stage of each pipeline to process at most two minibatches before backpropagation, bounding the number of parameter updates between forward and backward passes. To alleviate the resulting pipeline bubbles, AMDP launches multiple concurrent pipelines and adapts their number according to pipeline depth. In addition, AMDP accumulates gradients across minibatches and applies them in a single update, ensuring that only a bounded number of minibatches experience parameter mismatch, limited to within one optimization step. Experiments on GPT- and BERT-style models demonstrate that AMDP significantly accelerates training while preserving convergence.
APA
Chen, L., Wu, H. & Yu, W.. (2026). AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:17586-17602 Available from https://proceedings.mlr.press/v306/chen26ff.html.

Related Material