Alignment-Aware Decoding

Frédéric Berdoz, Luca A Lanzendörfer, René Caky, Roger Wattenhofer
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:7593-7612, 2026.

Abstract

Alignment of large language models remains a central challenge in natural language processing. Preference optimization has emerged as a popular and effective method for improving alignment, typically through training-time or prompt-based interventions. In this paper, we introduce alignment-aware decoding (AAD), a method to enhance model alignment directly at inference. Theoretically, AAD can be interpreted as implicit reward optimization, yet it requires no specialized training beyond the standard DPO setup. Empirically, AAD consistently outperforms strong baselines across diverse alignment benchmarks and model scales. Moreover, in data-constrained settings, AAD can produce high-quality synthetic data to improve alignment under standard decoding, providing a practical solution when labeled data is limited.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-berdoz26a, title = {Alignment-Aware Decoding}, author = {Berdoz, Fr\'{e}d\'{e}ric and Lanzend\"{o}rfer, Luca A and Caky, Ren\'{e} and Wattenhofer, Roger}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {7593--7612}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/berdoz26a/berdoz26a.pdf}, url = {https://proceedings.mlr.press/v306/berdoz26a.html}, abstract = {Alignment of large language models remains a central challenge in natural language processing. Preference optimization has emerged as a popular and effective method for improving alignment, typically through training-time or prompt-based interventions. In this paper, we introduce alignment-aware decoding (AAD), a method to enhance model alignment directly at inference. Theoretically, AAD can be interpreted as implicit reward optimization, yet it requires no specialized training beyond the standard DPO setup. Empirically, AAD consistently outperforms strong baselines across diverse alignment benchmarks and model scales. Moreover, in data-constrained settings, AAD can produce high-quality synthetic data to improve alignment under standard decoding, providing a practical solution when labeled data is limited.} }
Endnote
%0 Conference Paper %T Alignment-Aware Decoding %A Frédéric Berdoz %A Luca A Lanzendörfer %A René Caky %A Roger Wattenhofer %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-berdoz26a %I PMLR %P 7593--7612 %U https://proceedings.mlr.press/v306/berdoz26a.html %V 306 %X Alignment of large language models remains a central challenge in natural language processing. Preference optimization has emerged as a popular and effective method for improving alignment, typically through training-time or prompt-based interventions. In this paper, we introduce alignment-aware decoding (AAD), a method to enhance model alignment directly at inference. Theoretically, AAD can be interpreted as implicit reward optimization, yet it requires no specialized training beyond the standard DPO setup. Empirically, AAD consistently outperforms strong baselines across diverse alignment benchmarks and model scales. Moreover, in data-constrained settings, AAD can produce high-quality synthetic data to improve alignment under standard decoding, providing a practical solution when labeled data is limited.
APA
Berdoz, F., Lanzendörfer, L.A., Caky, R. & Wattenhofer, R.. (2026). Alignment-Aware Decoding. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:7593-7612 Available from https://proceedings.mlr.press/v306/berdoz26a.html.

Related Material