One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models

Daniel Fein, Max Lamparth, Violet Xiang, Mykel Kochenderfer, Nick Haber
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:29862-29896, 2026.

Abstract

Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences. However, RM-based preference-tuning is vulnerable to reward hacking, whereby LM policies learn undesirable behaviors from flawed RMs. By systematically measuring biases in five high-quality RMs, including the state-of-the-art, we find that issues persist despite prior work with respect to length, sycophancy, and overconfidence. We also discover new issues related to bias toward model-specific “styles” and answer-order. We categorize RM failures as tractable or resistant to linear intervention and propose a simple post-hoc intervention to mitigate low-complexity biases that arise from spurious correlations. Our proposed mechanistic reward shaping reduces targeted biases without degrading reward quality and while using minimal labeled data. The method is extensible to new biases, model-internal, and generalizes out-of-distribution.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-fein26a, title = {One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models}, author = {Fein, Daniel and Lamparth, Max and Xiang, Violet and Kochenderfer, Mykel and Haber, Nick}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {29862--29896}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/fein26a/fein26a.pdf}, url = {https://proceedings.mlr.press/v306/fein26a.html}, abstract = {Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences. However, RM-based preference-tuning is vulnerable to reward hacking, whereby LM policies learn undesirable behaviors from flawed RMs. By systematically measuring biases in five high-quality RMs, including the state-of-the-art, we find that issues persist despite prior work with respect to length, sycophancy, and overconfidence. We also discover new issues related to bias toward model-specific “styles” and answer-order. We categorize RM failures as tractable or resistant to linear intervention and propose a simple post-hoc intervention to mitigate low-complexity biases that arise from spurious correlations. Our proposed mechanistic reward shaping reduces targeted biases without degrading reward quality and while using minimal labeled data. The method is extensible to new biases, model-internal, and generalizes out-of-distribution.} }
Endnote
%0 Conference Paper %T One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models %A Daniel Fein %A Max Lamparth %A Violet Xiang %A Mykel Kochenderfer %A Nick Haber %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-fein26a %I PMLR %P 29862--29896 %U https://proceedings.mlr.press/v306/fein26a.html %V 306 %X Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences. However, RM-based preference-tuning is vulnerable to reward hacking, whereby LM policies learn undesirable behaviors from flawed RMs. By systematically measuring biases in five high-quality RMs, including the state-of-the-art, we find that issues persist despite prior work with respect to length, sycophancy, and overconfidence. We also discover new issues related to bias toward model-specific “styles” and answer-order. We categorize RM failures as tractable or resistant to linear intervention and propose a simple post-hoc intervention to mitigate low-complexity biases that arise from spurious correlations. Our proposed mechanistic reward shaping reduces targeted biases without degrading reward quality and while using minimal labeled data. The method is extensible to new biases, model-internal, and generalizes out-of-distribution.
APA
Fein, D., Lamparth, M., Xiang, V., Kochenderfer, M. & Haber, N.. (2026). One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:29862-29896 Available from https://proceedings.mlr.press/v306/fein26a.html.

Related Material