Sample Margin-Aware Recalibration of Temperature Scaling

Haolan Guo, Linwei Tao, Haoyang Luo, Minjing Dong, Chang Xu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:38283-38315, 2026.

Abstract

Deep neural networks frequently exhibit overconfidence, undermining reliability in safety-critical applications. Existing adaptive methods rely on indirectly learned proxies of sample difficulty. We establish the logit margin as a direct and principled hardness indicator. We prove that margin constrains the feasible temperature range for a target confidence. Empirically, margin strongly correlates with decision boundary proximity and reveals systematic calibration patterns across difficulty levels. We further identify a fundamental flaw in NLL-based optimization: minimizing NLL can paradoxically worsen calibration. To address this, we introduce Charbonnier-SoftECE, a smooth objective that provably upper-bounds the smooth calibration error (smCE). Building on these insights, we propose SMART (Sample Margin-Aware Recalibration of Temperature), a lightweight method that learns a sample-wise margin-to-temperature mapping guided by our calibration-centric objective. Experiments demonstrate state-of-the-art calibration across CNNs and ViTs on standard, long-tailed, and distribution-shifted benchmarks, with minimal inference-time overhead. Code is available at: https://github.com/Misakaaaaaz/ICML2026-SMART.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-guo26s, title = {Sample Margin-Aware Recalibration of Temperature Scaling}, author = {Guo, Haolan and Tao, Linwei and Luo, Haoyang and Dong, Minjing and Xu, Chang}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {38283--38315}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/guo26s/guo26s.pdf}, url = {https://proceedings.mlr.press/v306/guo26s.html}, abstract = {Deep neural networks frequently exhibit overconfidence, undermining reliability in safety-critical applications. Existing adaptive methods rely on indirectly learned proxies of sample difficulty. We establish the logit margin as a direct and principled hardness indicator. We prove that margin constrains the feasible temperature range for a target confidence. Empirically, margin strongly correlates with decision boundary proximity and reveals systematic calibration patterns across difficulty levels. We further identify a fundamental flaw in NLL-based optimization: minimizing NLL can paradoxically worsen calibration. To address this, we introduce Charbonnier-SoftECE, a smooth objective that provably upper-bounds the smooth calibration error (smCE). Building on these insights, we propose SMART (Sample Margin-Aware Recalibration of Temperature), a lightweight method that learns a sample-wise margin-to-temperature mapping guided by our calibration-centric objective. Experiments demonstrate state-of-the-art calibration across CNNs and ViTs on standard, long-tailed, and distribution-shifted benchmarks, with minimal inference-time overhead. Code is available at: https://github.com/Misakaaaaaz/ICML2026-SMART.} }
Endnote
%0 Conference Paper %T Sample Margin-Aware Recalibration of Temperature Scaling %A Haolan Guo %A Linwei Tao %A Haoyang Luo %A Minjing Dong %A Chang Xu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-guo26s %I PMLR %P 38283--38315 %U https://proceedings.mlr.press/v306/guo26s.html %V 306 %X Deep neural networks frequently exhibit overconfidence, undermining reliability in safety-critical applications. Existing adaptive methods rely on indirectly learned proxies of sample difficulty. We establish the logit margin as a direct and principled hardness indicator. We prove that margin constrains the feasible temperature range for a target confidence. Empirically, margin strongly correlates with decision boundary proximity and reveals systematic calibration patterns across difficulty levels. We further identify a fundamental flaw in NLL-based optimization: minimizing NLL can paradoxically worsen calibration. To address this, we introduce Charbonnier-SoftECE, a smooth objective that provably upper-bounds the smooth calibration error (smCE). Building on these insights, we propose SMART (Sample Margin-Aware Recalibration of Temperature), a lightweight method that learns a sample-wise margin-to-temperature mapping guided by our calibration-centric objective. Experiments demonstrate state-of-the-art calibration across CNNs and ViTs on standard, long-tailed, and distribution-shifted benchmarks, with minimal inference-time overhead. Code is available at: https://github.com/Misakaaaaaz/ICML2026-SMART.
APA
Guo, H., Tao, L., Luo, H., Dong, M. & Xu, C.. (2026). Sample Margin-Aware Recalibration of Temperature Scaling. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:38283-38315 Available from https://proceedings.mlr.press/v306/guo26s.html.

Related Material