Eigenvalue Calibration for Semantic Embeddings of Large Language Models

Sebastian G. Gruber, Nassim Walha, Francis Bach, Florian Buettner
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:1769-1789, 2026.

Abstract

Uncertainty quantification is central to the reliable deployment of large language models ({LLMs}), and eigenvalues of semantic embeddings have recently emerged as a key tool in state-of-the-art methods. However, conventional calibration results developed for classification probabilities cannot be directly transferred to eigenvalues. We address this gap by proposing a novel framework for calibrating the eigenvalues of semantic embeddings. We interpret {LLMs} combined with semantic embeddings of their generated answers as density matrix predictors, and we propose a novel approach to calibrate density matrix predictors by applying temperature scaling to their eigenvalues. We establish entropy–risk equivalence under calibration, derive a central calibration inequality specific to eigenvalues, and prove that temperature-scaled eigenvalues optimize calibration when minimizing proper score risks. Experiments on a variety of real-world settings show that current {LLMs} are systematically overconfident, and validate our theoretical findings. Together, these results advance the foundations and practice of uncertainty quantification for semantic embeddings.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-gruber26a, title = {Eigenvalue Calibration for Semantic Embeddings of Large Language Models}, author = {Gruber, Sebastian G. and Walha, Nassim and Bach, Francis and Buettner, Florian}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {1769--1789}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/gruber26a/gruber26a.pdf}, url = {https://proceedings.mlr.press/v337/gruber26a.html}, abstract = {Uncertainty quantification is central to the reliable deployment of large language models ({LLMs}), and eigenvalues of semantic embeddings have recently emerged as a key tool in state-of-the-art methods. However, conventional calibration results developed for classification probabilities cannot be directly transferred to eigenvalues. We address this gap by proposing a novel framework for calibrating the eigenvalues of semantic embeddings. We interpret {LLMs} combined with semantic embeddings of their generated answers as density matrix predictors, and we propose a novel approach to calibrate density matrix predictors by applying temperature scaling to their eigenvalues. We establish entropy–risk equivalence under calibration, derive a central calibration inequality specific to eigenvalues, and prove that temperature-scaled eigenvalues optimize calibration when minimizing proper score risks. Experiments on a variety of real-world settings show that current {LLMs} are systematically overconfident, and validate our theoretical findings. Together, these results advance the foundations and practice of uncertainty quantification for semantic embeddings.} }
Endnote
%0 Conference Paper %T Eigenvalue Calibration for Semantic Embeddings of Large Language Models %A Sebastian G. Gruber %A Nassim Walha %A Francis Bach %A Florian Buettner %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-gruber26a %I PMLR %P 1769--1789 %U https://proceedings.mlr.press/v337/gruber26a.html %V 337 %X Uncertainty quantification is central to the reliable deployment of large language models ({LLMs}), and eigenvalues of semantic embeddings have recently emerged as a key tool in state-of-the-art methods. However, conventional calibration results developed for classification probabilities cannot be directly transferred to eigenvalues. We address this gap by proposing a novel framework for calibrating the eigenvalues of semantic embeddings. We interpret {LLMs} combined with semantic embeddings of their generated answers as density matrix predictors, and we propose a novel approach to calibrate density matrix predictors by applying temperature scaling to their eigenvalues. We establish entropy–risk equivalence under calibration, derive a central calibration inequality specific to eigenvalues, and prove that temperature-scaled eigenvalues optimize calibration when minimizing proper score risks. Experiments on a variety of real-world settings show that current {LLMs} are systematically overconfident, and validate our theoretical findings. Together, these results advance the foundations and practice of uncertainty quantification for semantic embeddings.
APA
Gruber, S.G., Walha, N., Bach, F. & Buettner, F.. (2026). Eigenvalue Calibration for Semantic Embeddings of Large Language Models. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:1769-1789 Available from https://proceedings.mlr.press/v337/gruber26a.html.

Related Material