Reasoning Models Struggle to Control their Chains of Thought

Chen Yueh-Han, Robert Mccarthy, Bruce W. Lee, He He, Micah Carroll, Tomek Korbak
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:152359-152387, 2026.

Abstract

Instruction following in LLMs captures models’ ability to change their visible behaviors as requested by users. Instead, we study models’ ability to control their chain-of-thought (CoT). This capability – CoT controllability – is undesirable because it could allow models to suppress signs of misbehavior in their CoT, thereby undermining our ability to monitor them. To measure this, we introduce the CoT-Control evaluation suite. We show that reasoning models are less able to follow instructions in their CoT than in their outputs: on instructions like reasoning about a genetics problem without mentioning the word “chromosome", Claude-Sonnet-4.5 complies only 5% of the time. We also find that CoT controllability is higher for larger models and decreases with more RL training, test-time compute, and increased problem difficulty. CoT controllability failures extend even to situations in which models are given incentives (as opposed to direct requests) to evade CoT monitors, although models that are told they’re being monitored exhibit slightly higher controllability. Similarly, eliciting controllability by adversarially optimizing prompts doesn’t meaningfully increase controllability. Our results leave us cautiously optimistic: reasoning models generally seem characterized by low CoT controllability. However, the mechanism behind this phenomenon is not well understood. Given its importance for maintaining CoT monitorability, we recommend that frontier labs keep tracking controllability for future models.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-yueh-han26a, title = {Reasoning Models Struggle to Control their Chains of Thought}, author = {Yueh-Han, Chen and Mccarthy, Robert and Lee, Bruce W. and He, He and Carroll, Micah and Korbak, Tomek}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {152359--152387}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/yueh-han26a/yueh-han26a.pdf}, url = {https://proceedings.mlr.press/v306/yueh-han26a.html}, abstract = {Instruction following in LLMs captures models’ ability to change their visible behaviors as requested by users. Instead, we study models’ ability to control their chain-of-thought (CoT). This capability – CoT controllability – is undesirable because it could allow models to suppress signs of misbehavior in their CoT, thereby undermining our ability to monitor them. To measure this, we introduce the CoT-Control evaluation suite. We show that reasoning models are less able to follow instructions in their CoT than in their outputs: on instructions like reasoning about a genetics problem without mentioning the word “chromosome", Claude-Sonnet-4.5 complies only 5% of the time. We also find that CoT controllability is higher for larger models and decreases with more RL training, test-time compute, and increased problem difficulty. CoT controllability failures extend even to situations in which models are given incentives (as opposed to direct requests) to evade CoT monitors, although models that are told they’re being monitored exhibit slightly higher controllability. Similarly, eliciting controllability by adversarially optimizing prompts doesn’t meaningfully increase controllability. Our results leave us cautiously optimistic: reasoning models generally seem characterized by low CoT controllability. However, the mechanism behind this phenomenon is not well understood. Given its importance for maintaining CoT monitorability, we recommend that frontier labs keep tracking controllability for future models.} }
Endnote
%0 Conference Paper %T Reasoning Models Struggle to Control their Chains of Thought %A Chen Yueh-Han %A Robert Mccarthy %A Bruce W. Lee %A He He %A Micah Carroll %A Tomek Korbak %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-yueh-han26a %I PMLR %P 152359--152387 %U https://proceedings.mlr.press/v306/yueh-han26a.html %V 306 %X Instruction following in LLMs captures models’ ability to change their visible behaviors as requested by users. Instead, we study models’ ability to control their chain-of-thought (CoT). This capability – CoT controllability – is undesirable because it could allow models to suppress signs of misbehavior in their CoT, thereby undermining our ability to monitor them. To measure this, we introduce the CoT-Control evaluation suite. We show that reasoning models are less able to follow instructions in their CoT than in their outputs: on instructions like reasoning about a genetics problem without mentioning the word “chromosome", Claude-Sonnet-4.5 complies only 5% of the time. We also find that CoT controllability is higher for larger models and decreases with more RL training, test-time compute, and increased problem difficulty. CoT controllability failures extend even to situations in which models are given incentives (as opposed to direct requests) to evade CoT monitors, although models that are told they’re being monitored exhibit slightly higher controllability. Similarly, eliciting controllability by adversarially optimizing prompts doesn’t meaningfully increase controllability. Our results leave us cautiously optimistic: reasoning models generally seem characterized by low CoT controllability. However, the mechanism behind this phenomenon is not well understood. Given its importance for maintaining CoT monitorability, we recommend that frontier labs keep tracking controllability for future models.
APA
Yueh-Han, C., Mccarthy, R., Lee, B.W., He, H., Carroll, M. & Korbak, T.. (2026). Reasoning Models Struggle to Control their Chains of Thought. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:152359-152387 Available from https://proceedings.mlr.press/v306/yueh-han26a.html.

Related Material