On the Convergence of Self-Improving Online LLM Alignment

Xudong Wu, Pangpang Liu, Vaneet Aggarwal, Jiayu Chen
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7433-7467, 2026.

Abstract

The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel formulation of the problem to an efficient, single-level method. Empirically, SAIL has demonstrated strong performance on this task. However, a formal analysis of its convergence properties has been lacking. We identify a key theoretical challenge: the standard SAIL objective function is not guaranteed to be strongly concave due to unfavorable properties of its Hessian. To address this limitation, we propose a regularized objective, SAIL-RevKL, which incorporates a reverse Kullback-Leibler (KL) divergence penalty to improve the optimization landscape. Our central theoretical contribution is to prove that this regularized objective satisfies the {Polyak-Lojasiewicz} (PL) condition within a bounded parameter space. We establish global convergence guarantees, achieving a near-linear sample complexity. We further validate the effectiveness and stability of SAIL-RevKL through empirical evaluations, demonstrating that it outperforms the vanilla SAIL on both MuJoCo benchmarks and {LLM} alignment tasks.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-wu26c, title = {On the Convergence of Self-Improving Online {LLM} Alignment}, author = {Wu, Xudong and Liu, Pangpang and Aggarwal, Vaneet and Chen, Jiayu}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {7433--7467}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/wu26c/wu26c.pdf}, url = {https://proceedings.mlr.press/v337/wu26c.html}, abstract = {The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel formulation of the problem to an efficient, single-level method. Empirically, SAIL has demonstrated strong performance on this task. However, a formal analysis of its convergence properties has been lacking. We identify a key theoretical challenge: the standard SAIL objective function is not guaranteed to be strongly concave due to unfavorable properties of its Hessian. To address this limitation, we propose a regularized objective, SAIL-RevKL, which incorporates a reverse Kullback-Leibler (KL) divergence penalty to improve the optimization landscape. Our central theoretical contribution is to prove that this regularized objective satisfies the {Polyak-Lojasiewicz} (PL) condition within a bounded parameter space. We establish global convergence guarantees, achieving a near-linear sample complexity. We further validate the effectiveness and stability of SAIL-RevKL through empirical evaluations, demonstrating that it outperforms the vanilla SAIL on both MuJoCo benchmarks and {LLM} alignment tasks.} }
Endnote
%0 Conference Paper %T On the Convergence of Self-Improving Online LLM Alignment %A Xudong Wu %A Pangpang Liu %A Vaneet Aggarwal %A Jiayu Chen %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-wu26c %I PMLR %P 7433--7467 %U https://proceedings.mlr.press/v337/wu26c.html %V 337 %X The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel formulation of the problem to an efficient, single-level method. Empirically, SAIL has demonstrated strong performance on this task. However, a formal analysis of its convergence properties has been lacking. We identify a key theoretical challenge: the standard SAIL objective function is not guaranteed to be strongly concave due to unfavorable properties of its Hessian. To address this limitation, we propose a regularized objective, SAIL-RevKL, which incorporates a reverse Kullback-Leibler (KL) divergence penalty to improve the optimization landscape. Our central theoretical contribution is to prove that this regularized objective satisfies the {Polyak-Lojasiewicz} (PL) condition within a bounded parameter space. We establish global convergence guarantees, achieving a near-linear sample complexity. We further validate the effectiveness and stability of SAIL-RevKL through empirical evaluations, demonstrating that it outperforms the vanilla SAIL on both MuJoCo benchmarks and {LLM} alignment tasks.
APA
Wu, X., Liu, P., Aggarwal, V. & Chen, J.. (2026). On the Convergence of Self-Improving Online LLM Alignment. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:7433-7467 Available from https://proceedings.mlr.press/v337/wu26c.html.

Related Material