E-Scores for (In)Correctness Assessment of Generative Model Outputs

Guneet S. Dhillon, Javier Gonzalez, Teodora Pandeva, Alicia Curth
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:190-198, 2026.

Abstract

While generative models, especially large language models (LLMs), are ubiquitous in today’s world, principled mechanisms to assess their (in)correctness are limited. Using the conformal prediction framework, previous works construct sets of LLM responses where the probability of including an incorrect response, or error, is capped at a user-defined tolerance level. However, since these methods are based on p-values, they are susceptible to p-hacking, i.e., choosing the tolerance level post-hoc can invalidate the guarantees. We therefore leverage e-values to complement generative model outputs with e-scores as measures of incorrectness. In addition to achieving the guarantees as before, e-scores further provide users with the flexibility of choosing data-dependent tolerance levels while upper bounding size distortion, a post-hoc notion of error. We experimentally demonstrate their efficacy in assessing LLM outputs under different forms of correctness: mathematical factuality and property constraints satisfaction.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-dhillon26a, title = { E-Scores for (In)Correctness Assessment of Generative Model Outputs }, author = {Dhillon, Guneet S. and Gonzalez, Javier and Pandeva, Teodora and Curth, Alicia}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {190--198}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/dhillon26a/dhillon26a.pdf}, url = {https://proceedings.mlr.press/v300/dhillon26a.html}, abstract = { While generative models, especially large language models (LLMs), are ubiquitous in today’s world, principled mechanisms to assess their (in)correctness are limited. Using the conformal prediction framework, previous works construct sets of LLM responses where the probability of including an incorrect response, or error, is capped at a user-defined tolerance level. However, since these methods are based on p-values, they are susceptible to p-hacking, i.e., choosing the tolerance level post-hoc can invalidate the guarantees. We therefore leverage e-values to complement generative model outputs with e-scores as measures of incorrectness. In addition to achieving the guarantees as before, e-scores further provide users with the flexibility of choosing data-dependent tolerance levels while upper bounding size distortion, a post-hoc notion of error. We experimentally demonstrate their efficacy in assessing LLM outputs under different forms of correctness: mathematical factuality and property constraints satisfaction. } }
Endnote
%0 Conference Paper %T E-Scores for (In)Correctness Assessment of Generative Model Outputs %A Guneet S. Dhillon %A Javier Gonzalez %A Teodora Pandeva %A Alicia Curth %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-dhillon26a %I PMLR %P 190--198 %U https://proceedings.mlr.press/v300/dhillon26a.html %V 300 %X While generative models, especially large language models (LLMs), are ubiquitous in today’s world, principled mechanisms to assess their (in)correctness are limited. Using the conformal prediction framework, previous works construct sets of LLM responses where the probability of including an incorrect response, or error, is capped at a user-defined tolerance level. However, since these methods are based on p-values, they are susceptible to p-hacking, i.e., choosing the tolerance level post-hoc can invalidate the guarantees. We therefore leverage e-values to complement generative model outputs with e-scores as measures of incorrectness. In addition to achieving the guarantees as before, e-scores further provide users with the flexibility of choosing data-dependent tolerance levels while upper bounding size distortion, a post-hoc notion of error. We experimentally demonstrate their efficacy in assessing LLM outputs under different forms of correctness: mathematical factuality and property constraints satisfaction.
APA
Dhillon, G.S., Gonzalez, J., Pandeva, T. & Curth, A.. (2026). E-Scores for (In)Correctness Assessment of Generative Model Outputs . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:190-198 Available from https://proceedings.mlr.press/v300/dhillon26a.html.

Related Material