Convex Optimization for Alignment and Preference Learning on a Single GPU

Miria Feng, Mert Pilanci
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:30290-30320, 2026.

Abstract

Fine-tuning large language models (LLMs) to align with human preferences has driven the success of systems such as Gemini and ChatGPT. However, approaches like Reinforcement Learning from Human Feedback (RLHF) remain computationally expensive and complex. Direct Preference Optimization (DPO) offers a simpler alternative but has limitations such as inconsistent ranking accuracy, high dependence on GPU resources, and expensive hyperparameter tuning. We propose the Convex Optimization for Alignment and Preference Learning Algorithm (COALA): a novel lightweight strategy with strong theoretical guarantees. By leveraging the convex optimization reformulation of neural networks, COALA eliminates the need for a reference model and obtains significant reduction in both training time and VRAM consumption, thus enabling efficient training on a single GPU. Experiments across four datasets—including a 26621-sample synthetic Educational Feedback dataset—and six models (including Llama-3.1-8B) demonstrate COALA’s competitive performance and efficiency while utilizing as little as ${\sim}17.6%$ of DPO’s total TFLOPs. COALA exhibits stable, monotonically increasing rewards and reaches peak margins in significantly shorter time in comparison to traditional methods such as DPO and ORPO. To the best of our knowledge, this is the first time convex optimization has been effectively applied to preference fine-tuning of LLMs.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-feng26l, title = {Convex Optimization for Alignment and Preference Learning on a Single {GPU}}, author = {Feng, Miria and Pilanci, Mert}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {30290--30320}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/feng26l/feng26l.pdf}, url = {https://proceedings.mlr.press/v306/feng26l.html}, abstract = {Fine-tuning large language models (LLMs) to align with human preferences has driven the success of systems such as Gemini and ChatGPT. However, approaches like Reinforcement Learning from Human Feedback (RLHF) remain computationally expensive and complex. Direct Preference Optimization (DPO) offers a simpler alternative but has limitations such as inconsistent ranking accuracy, high dependence on GPU resources, and expensive hyperparameter tuning. We propose the Convex Optimization for Alignment and Preference Learning Algorithm (COALA): a novel lightweight strategy with strong theoretical guarantees. By leveraging the convex optimization reformulation of neural networks, COALA eliminates the need for a reference model and obtains significant reduction in both training time and VRAM consumption, thus enabling efficient training on a single GPU. Experiments across four datasets—including a 26621-sample synthetic Educational Feedback dataset—and six models (including Llama-3.1-8B) demonstrate COALA’s competitive performance and efficiency while utilizing as little as ${\sim}17.6%$ of DPO’s total TFLOPs. COALA exhibits stable, monotonically increasing rewards and reaches peak margins in significantly shorter time in comparison to traditional methods such as DPO and ORPO. To the best of our knowledge, this is the first time convex optimization has been effectively applied to preference fine-tuning of LLMs.} }
Endnote
%0 Conference Paper %T Convex Optimization for Alignment and Preference Learning on a Single GPU %A Miria Feng %A Mert Pilanci %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-feng26l %I PMLR %P 30290--30320 %U https://proceedings.mlr.press/v306/feng26l.html %V 306 %X Fine-tuning large language models (LLMs) to align with human preferences has driven the success of systems such as Gemini and ChatGPT. However, approaches like Reinforcement Learning from Human Feedback (RLHF) remain computationally expensive and complex. Direct Preference Optimization (DPO) offers a simpler alternative but has limitations such as inconsistent ranking accuracy, high dependence on GPU resources, and expensive hyperparameter tuning. We propose the Convex Optimization for Alignment and Preference Learning Algorithm (COALA): a novel lightweight strategy with strong theoretical guarantees. By leveraging the convex optimization reformulation of neural networks, COALA eliminates the need for a reference model and obtains significant reduction in both training time and VRAM consumption, thus enabling efficient training on a single GPU. Experiments across four datasets—including a 26621-sample synthetic Educational Feedback dataset—and six models (including Llama-3.1-8B) demonstrate COALA’s competitive performance and efficiency while utilizing as little as ${\sim}17.6%$ of DPO’s total TFLOPs. COALA exhibits stable, monotonically increasing rewards and reaches peak margins in significantly shorter time in comparison to traditional methods such as DPO and ORPO. To the best of our knowledge, this is the first time convex optimization has been effectively applied to preference fine-tuning of LLMs.
APA
Feng, M. & Pilanci, M.. (2026). Convex Optimization for Alignment and Preference Learning on a Single GPU. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:30290-30320 Available from https://proceedings.mlr.press/v306/feng26l.html.

Related Material