Confidence is Not Universal: Task-Dependent Calibration and Emergent Behavior in LLMs

Chaeyun Jang, Moonseok Choi, Yegon Kim, Seungyoo Lee, Juho Lee, Hyungi Lee
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:50548-50571, 2026.

Abstract

Large language models (LLMs) increasingly support human decision-making, rendering human-interpretable confidence essential. However, it remains unclear whether verbalized confidence calibration generalizes across heterogeneous tasks without degrading accuracy. We show that universal confidence calibration fails. Across diverse benchmarks, we identify two incompatible task families with distinct confidence semantics. In reasoning-centric tasks, confidence supervision transfers within the family, often improving calibration while preserving or even improving accuracy, and induces emergent behaviors such as confidence-dependent reasoning length and self-verification. Retrieval- and copy-oriented tasks also exhibit within-family transfer, but fail to generalize to reasoning tasks, with cross-family supervision degrading both calibration and accuracy. Motivated by this finding, we disentangle confidence into reasoning uncertainty and evidence localization uncertainty. This simple decomposition restores cross-family generalization using supervised fine-tuning alone, suggesting that effective confidence alignment requires task-aware semantics rather than a universal scalar notion.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-jang26a, title = {Confidence is Not Universal: Task-Dependent Calibration and Emergent Behavior in {LLM}s}, author = {Jang, Chaeyun and Choi, Moonseok and Kim, Yegon and Lee, Seungyoo and Lee, Juho and Lee, Hyungi}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {50548--50571}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/jang26a/jang26a.pdf}, url = {https://proceedings.mlr.press/v306/jang26a.html}, abstract = {Large language models (LLMs) increasingly support human decision-making, rendering human-interpretable confidence essential. However, it remains unclear whether verbalized confidence calibration generalizes across heterogeneous tasks without degrading accuracy. We show that universal confidence calibration fails. Across diverse benchmarks, we identify two incompatible task families with distinct confidence semantics. In reasoning-centric tasks, confidence supervision transfers within the family, often improving calibration while preserving or even improving accuracy, and induces emergent behaviors such as confidence-dependent reasoning length and self-verification. Retrieval- and copy-oriented tasks also exhibit within-family transfer, but fail to generalize to reasoning tasks, with cross-family supervision degrading both calibration and accuracy. Motivated by this finding, we disentangle confidence into reasoning uncertainty and evidence localization uncertainty. This simple decomposition restores cross-family generalization using supervised fine-tuning alone, suggesting that effective confidence alignment requires task-aware semantics rather than a universal scalar notion.} }
Endnote
%0 Conference Paper %T Confidence is Not Universal: Task-Dependent Calibration and Emergent Behavior in LLMs %A Chaeyun Jang %A Moonseok Choi %A Yegon Kim %A Seungyoo Lee %A Juho Lee %A Hyungi Lee %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-jang26a %I PMLR %P 50548--50571 %U https://proceedings.mlr.press/v306/jang26a.html %V 306 %X Large language models (LLMs) increasingly support human decision-making, rendering human-interpretable confidence essential. However, it remains unclear whether verbalized confidence calibration generalizes across heterogeneous tasks without degrading accuracy. We show that universal confidence calibration fails. Across diverse benchmarks, we identify two incompatible task families with distinct confidence semantics. In reasoning-centric tasks, confidence supervision transfers within the family, often improving calibration while preserving or even improving accuracy, and induces emergent behaviors such as confidence-dependent reasoning length and self-verification. Retrieval- and copy-oriented tasks also exhibit within-family transfer, but fail to generalize to reasoning tasks, with cross-family supervision degrading both calibration and accuracy. Motivated by this finding, we disentangle confidence into reasoning uncertainty and evidence localization uncertainty. This simple decomposition restores cross-family generalization using supervised fine-tuning alone, suggesting that effective confidence alignment requires task-aware semantics rather than a universal scalar notion.
APA
Jang, C., Choi, M., Kim, Y., Lee, S., Lee, J. & Lee, H.. (2026). Confidence is Not Universal: Task-Dependent Calibration and Emergent Behavior in LLMs. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:50548-50571 Available from https://proceedings.mlr.press/v306/jang26a.html.

Related Material