Detecting Label Shift for Large Language Models with Conformal Test Martingales

Henrik Boström
Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:592-612, 2026.

Abstract

We study the problem of detecting label shift in deployed Large Language Models (LLMs), using model outputs, in the form of class labels, confidence estimates, and justifications, as proxies for true labels. We propose the use of conformal test martingales, which can detect shifts in both mean and dispersion across multiple streams while controlling the false alarm rate. We evaluate the approach on three binary and three multiclass text classification datasets using LLMs ranging from 3B to 32B parameters. We consider three martingale algorithms: Simple Jumper, Sleeper/Stayer, and Sleeper/Drifter, as well as their composite. The observed false alarm rates are consistently below the nominal level. Detection-time rankings are stable across datasets, models, and signals, with Sleeper/Drifter and the composite martingale outperforming the others, while Simple Jumper is least competitive. The composite provides only marginal gains over Sleeper/Drifter alone. Larger models tend to enable earlier detection of label shift, although the effect is uneven and non-monotonic. Among the signals, confidence is the most reliable proxy for binary datasets, whereas combining all three signals yields the best performance on two multiclass datasets. Justification is highly dataset-dependent: it fails for several datasets and models, yet is the only effective signal for the smallest model in one case.

Cite this Paper


BibTeX
@InProceedings{pmlr-v329-bostrom26b, title = {Detecting Label Shift for Large Language Models with Conformal Test Martingales}, author = {Bostr{\"o}m, Henrik}, booktitle = {Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications}, pages = {592--612}, year = {2026}, editor = {Ahlberg, Ernst and Johansson, Ulf and Boström, Henrik and Carlevaro, Alberto and Hallberg Szabadváry, Johan and Carlsson, Lars}, volume = {329}, series = {Proceedings of Machine Learning Research}, month = {02--04 Sep}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v329/main/assets/bostrom26b/bostrom26b.pdf}, url = {https://proceedings.mlr.press/v329/bostrom26b.html}, abstract = {We study the problem of detecting label shift in deployed Large Language Models (LLMs), using model outputs, in the form of class labels, confidence estimates, and justifications, as proxies for true labels. We propose the use of conformal test martingales, which can detect shifts in both mean and dispersion across multiple streams while controlling the false alarm rate. We evaluate the approach on three binary and three multiclass text classification datasets using LLMs ranging from 3B to 32B parameters. We consider three martingale algorithms: Simple Jumper, Sleeper/Stayer, and Sleeper/Drifter, as well as their composite. The observed false alarm rates are consistently below the nominal level. Detection-time rankings are stable across datasets, models, and signals, with Sleeper/Drifter and the composite martingale outperforming the others, while Simple Jumper is least competitive. The composite provides only marginal gains over Sleeper/Drifter alone. Larger models tend to enable earlier detection of label shift, although the effect is uneven and non-monotonic. Among the signals, confidence is the most reliable proxy for binary datasets, whereas combining all three signals yields the best performance on two multiclass datasets. Justification is highly dataset-dependent: it fails for several datasets and models, yet is the only effective signal for the smallest model in one case.} }
Endnote
%0 Conference Paper %T Detecting Label Shift for Large Language Models with Conformal Test Martingales %A Henrik Boström %B Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications %C Proceedings of Machine Learning Research %D 2026 %E Ernst Ahlberg %E Ulf Johansson %E Henrik Boström %E Alberto Carlevaro %E Johan Hallberg Szabadváry %E Lars Carlsson %F pmlr-v329-bostrom26b %I PMLR %P 592--612 %U https://proceedings.mlr.press/v329/bostrom26b.html %V 329 %X We study the problem of detecting label shift in deployed Large Language Models (LLMs), using model outputs, in the form of class labels, confidence estimates, and justifications, as proxies for true labels. We propose the use of conformal test martingales, which can detect shifts in both mean and dispersion across multiple streams while controlling the false alarm rate. We evaluate the approach on three binary and three multiclass text classification datasets using LLMs ranging from 3B to 32B parameters. We consider three martingale algorithms: Simple Jumper, Sleeper/Stayer, and Sleeper/Drifter, as well as their composite. The observed false alarm rates are consistently below the nominal level. Detection-time rankings are stable across datasets, models, and signals, with Sleeper/Drifter and the composite martingale outperforming the others, while Simple Jumper is least competitive. The composite provides only marginal gains over Sleeper/Drifter alone. Larger models tend to enable earlier detection of label shift, although the effect is uneven and non-monotonic. Among the signals, confidence is the most reliable proxy for binary datasets, whereas combining all three signals yields the best performance on two multiclass datasets. Justification is highly dataset-dependent: it fails for several datasets and models, yet is the only effective signal for the smallest model in one case.
APA
Boström, H.. (2026). Detecting Label Shift for Large Language Models with Conformal Test Martingales. Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, in Proceedings of Machine Learning Research 329:592-612 Available from https://proceedings.mlr.press/v329/bostrom26b.html.

Related Material