Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

Guanxu Chen, Jing Shao, Tao Luo, Lijie Hu, Qihao Lin, Dongrui Liu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:16553-16580, 2026.

Abstract

Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain unclear. Chain-of-thoughts (CoTs) have been commonly utilized to externalize LLMs’ thinking, but this strategy fails to accurately reflect LLMs’ thinking process. Techniques based on LLMs’ hidden representations provide an inner perspective to improve the monitorability of their latent thinking. However, previous methods only try to develop external modules instead of making LLMs themselves easier to monitor. In this paper, we propose a novel method, TELLME, improving the transparency of LLMs and helping monitors identify unsuitable and sensitive behaviors. Furthermore, we showcase the effectiveness of TELLME on detoxification tasks, where LLMs achieve consistent improvement among multimodal test sets, distinct architectures, and varying parameter scales. We further analyze TELLME’s improvement on LLMs’ generalization ability from both optimal transport theory and empirical perspectives.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26dr, title = {Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring}, author = {Chen, Guanxu and Shao, Jing and Luo, Tao and Hu, Lijie and Lin, Qihao and Liu, Dongrui}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {16553--16580}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26dr/chen26dr.pdf}, url = {https://proceedings.mlr.press/v306/chen26dr.html}, abstract = {Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain unclear. Chain-of-thoughts (CoTs) have been commonly utilized to externalize LLMs’ thinking, but this strategy fails to accurately reflect LLMs’ thinking process. Techniques based on LLMs’ hidden representations provide an inner perspective to improve the monitorability of their latent thinking. However, previous methods only try to develop external modules instead of making LLMs themselves easier to monitor. In this paper, we propose a novel method, TELLME, improving the transparency of LLMs and helping monitors identify unsuitable and sensitive behaviors. Furthermore, we showcase the effectiveness of TELLME on detoxification tasks, where LLMs achieve consistent improvement among multimodal test sets, distinct architectures, and varying parameter scales. We further analyze TELLME’s improvement on LLMs’ generalization ability from both optimal transport theory and empirical perspectives.} }
Endnote
%0 Conference Paper %T Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring %A Guanxu Chen %A Jing Shao %A Tao Luo %A Lijie Hu %A Qihao Lin %A Dongrui Liu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26dr %I PMLR %P 16553--16580 %U https://proceedings.mlr.press/v306/chen26dr.html %V 306 %X Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain unclear. Chain-of-thoughts (CoTs) have been commonly utilized to externalize LLMs’ thinking, but this strategy fails to accurately reflect LLMs’ thinking process. Techniques based on LLMs’ hidden representations provide an inner perspective to improve the monitorability of their latent thinking. However, previous methods only try to develop external modules instead of making LLMs themselves easier to monitor. In this paper, we propose a novel method, TELLME, improving the transparency of LLMs and helping monitors identify unsuitable and sensitive behaviors. Furthermore, we showcase the effectiveness of TELLME on detoxification tasks, where LLMs achieve consistent improvement among multimodal test sets, distinct architectures, and varying parameter scales. We further analyze TELLME’s improvement on LLMs’ generalization ability from both optimal transport theory and empirical perspectives.
APA
Chen, G., Shao, J., Luo, T., Hu, L., Lin, Q. & Liu, D.. (2026). Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:16553-16580 Available from https://proceedings.mlr.press/v306/chen26dr.html.

Related Material