Leveraging Machine Unlearning for Cost-Efficient Preference Alignment

Xiaohua Feng, Yuyuan Li, Huwei Ji, Li Zhang, Jiaming Zhang, Tianyu Du, Chaochao Chen
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:30088-30114, 2026.

Abstract

Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream methods like reinforcement learning with human feedback face notable challenges. These approaches require high-quality datasets of positive preference examples, which are costly to obtain and computationally intensive. The LLM unlearning technique presents a promising alternative by directly removing the influence of negative examples. However, current research has primarily focused on empirical validation, lacking systematic quantitative analysis. To bridge this gap, we propose a framework linking PA with LLM unlearning. Through bi-level optimization, we first quantify how unlearning specific negative examples impacts PA performance. Our analysis reveals that these effects vary substantially across negative examples. Building on this insight, we pose a crucial question: how can we optimally select and weight negative examples for unlearning to maximize PA performance? To answer this, we propose Unlearning to Align (U2A), which leverages bi-level optimization to efficiently select and unlearn examples for optimal PA performance. We validate the proposed method through extensive experiments, with results confirming its effectiveness. Our code is available at https://anonymous.4open.science/r/U2A-9E75.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-feng26f, title = {Leveraging Machine Unlearning for Cost-Efficient Preference Alignment}, author = {Feng, Xiaohua and Li, Yuyuan and Ji, Huwei and Zhang, Li and Zhang, Jiaming and Du, Tianyu and Chen, Chaochao}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {30088--30114}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/feng26f/feng26f.pdf}, url = {https://proceedings.mlr.press/v306/feng26f.html}, abstract = {Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream methods like reinforcement learning with human feedback face notable challenges. These approaches require high-quality datasets of positive preference examples, which are costly to obtain and computationally intensive. The LLM unlearning technique presents a promising alternative by directly removing the influence of negative examples. However, current research has primarily focused on empirical validation, lacking systematic quantitative analysis. To bridge this gap, we propose a framework linking PA with LLM unlearning. Through bi-level optimization, we first quantify how unlearning specific negative examples impacts PA performance. Our analysis reveals that these effects vary substantially across negative examples. Building on this insight, we pose a crucial question: how can we optimally select and weight negative examples for unlearning to maximize PA performance? To answer this, we propose Unlearning to Align (U2A), which leverages bi-level optimization to efficiently select and unlearn examples for optimal PA performance. We validate the proposed method through extensive experiments, with results confirming its effectiveness. Our code is available at https://anonymous.4open.science/r/U2A-9E75.} }
Endnote
%0 Conference Paper %T Leveraging Machine Unlearning for Cost-Efficient Preference Alignment %A Xiaohua Feng %A Yuyuan Li %A Huwei Ji %A Li Zhang %A Jiaming Zhang %A Tianyu Du %A Chaochao Chen %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-feng26f %I PMLR %P 30088--30114 %U https://proceedings.mlr.press/v306/feng26f.html %V 306 %X Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream methods like reinforcement learning with human feedback face notable challenges. These approaches require high-quality datasets of positive preference examples, which are costly to obtain and computationally intensive. The LLM unlearning technique presents a promising alternative by directly removing the influence of negative examples. However, current research has primarily focused on empirical validation, lacking systematic quantitative analysis. To bridge this gap, we propose a framework linking PA with LLM unlearning. Through bi-level optimization, we first quantify how unlearning specific negative examples impacts PA performance. Our analysis reveals that these effects vary substantially across negative examples. Building on this insight, we pose a crucial question: how can we optimally select and weight negative examples for unlearning to maximize PA performance? To answer this, we propose Unlearning to Align (U2A), which leverages bi-level optimization to efficiently select and unlearn examples for optimal PA performance. We validate the proposed method through extensive experiments, with results confirming its effectiveness. Our code is available at https://anonymous.4open.science/r/U2A-9E75.
APA
Feng, X., Li, Y., Ji, H., Zhang, L., Zhang, J., Du, T. & Chen, C.. (2026). Leveraging Machine Unlearning for Cost-Efficient Preference Alignment. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:30088-30114 Available from https://proceedings.mlr.press/v306/feng26f.html.

Related Material