Recontextualization Mitigates Specification Gaming Without Modifying the Specification

Ariana Azarbal, Victor Gillioz, Vladimir Ivanov, Bryce Woodworth, Jacob Drori, Nevan Wichers, Aram Ebtekar, Alex Cloud, Alexander Matt Turner
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:4469-4516, 2026.

Abstract

Developers often struggle to specify correct training labels and rewards. Perhaps they don’t need to. We propose recontextualization, which reduces how often language models "game" training signals, performing misbehaviors those signals mistakenly reinforce. We show recontextualization prevents models from learning to 1) overfit evaluation criteria at the expense of chat response quality; 2) special-case code to pass incorrect tests; 3) overwrite evaluation functions rather than write correct code; and 4) become sycophantic. Our method works by generating completions from prompts discouraging misbehavior and then recontextualizing them as though they were in response to prompts permitting misbehavior. Recontextualization trains language models to resist misbehavior even when instructions permit it. This mitigates the reinforcement of misbehavior from misspecified training signals, reducing specification gaming without improving the supervision signal.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-azarbal26a, title = {Recontextualization Mitigates Specification Gaming Without Modifying the Specification}, author = {Azarbal, Ariana and Gillioz, Victor and Ivanov, Vladimir and Woodworth, Bryce and Drori, Jacob and Wichers, Nevan and Ebtekar, Aram and Cloud, Alex and Turner, Alexander Matt}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {4469--4516}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/azarbal26a/azarbal26a.pdf}, url = {https://proceedings.mlr.press/v306/azarbal26a.html}, abstract = {Developers often struggle to specify correct training labels and rewards. Perhaps they don’t need to. We propose recontextualization, which reduces how often language models "game" training signals, performing misbehaviors those signals mistakenly reinforce. We show recontextualization prevents models from learning to 1) overfit evaluation criteria at the expense of chat response quality; 2) special-case code to pass incorrect tests; 3) overwrite evaluation functions rather than write correct code; and 4) become sycophantic. Our method works by generating completions from prompts discouraging misbehavior and then recontextualizing them as though they were in response to prompts permitting misbehavior. Recontextualization trains language models to resist misbehavior even when instructions permit it. This mitigates the reinforcement of misbehavior from misspecified training signals, reducing specification gaming without improving the supervision signal.} }
Endnote
%0 Conference Paper %T Recontextualization Mitigates Specification Gaming Without Modifying the Specification %A Ariana Azarbal %A Victor Gillioz %A Vladimir Ivanov %A Bryce Woodworth %A Jacob Drori %A Nevan Wichers %A Aram Ebtekar %A Alex Cloud %A Alexander Matt Turner %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-azarbal26a %I PMLR %P 4469--4516 %U https://proceedings.mlr.press/v306/azarbal26a.html %V 306 %X Developers often struggle to specify correct training labels and rewards. Perhaps they don’t need to. We propose recontextualization, which reduces how often language models "game" training signals, performing misbehaviors those signals mistakenly reinforce. We show recontextualization prevents models from learning to 1) overfit evaluation criteria at the expense of chat response quality; 2) special-case code to pass incorrect tests; 3) overwrite evaluation functions rather than write correct code; and 4) become sycophantic. Our method works by generating completions from prompts discouraging misbehavior and then recontextualizing them as though they were in response to prompts permitting misbehavior. Recontextualization trains language models to resist misbehavior even when instructions permit it. This mitigates the reinforcement of misbehavior from misspecified training signals, reducing specification gaming without improving the supervision signal.
APA
Azarbal, A., Gillioz, V., Ivanov, V., Woodworth, B., Drori, J., Wichers, N., Ebtekar, A., Cloud, A. & Turner, A.M.. (2026). Recontextualization Mitigates Specification Gaming Without Modifying the Specification. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:4469-4516 Available from https://proceedings.mlr.press/v306/azarbal26a.html.

Related Material