Understanding Measures of Uncertainty for Adversarial Example Detection

Lewis Smith, Yarin Gal
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:559-568, 2018.

Abstract

Measuring uncertainty is a promising technique for detecting adversarial examples, crafted in- puts on which the model predicts an incorrect class with high confidence. There are various measures of uncertainty, including predictive entropy and mutual information, each capturing distinct types of uncertainty. We study these measures, and shed light on why mutual infor- mation seems to be effective at the task of adver- sarial example detection. We highlight failure modes for MC dropout, a widely used approach for estimating uncertainty in deep models. This leads to an improved understanding of the draw- backs of current methods, and a proposal to im- prove the quality of uncertainty estimates using probabilistic model ensembles. We give illustra- tive experiments using MNIST to demonstrate the intuition underlying the different measures of uncertainty, as well as experiments on a real- world Kaggle dogs vs cats classification dataset.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-smith18a, title = {Understanding Measures of Uncertainty for Adversarial Example Detection}, author = {Smith, Lewis and Gal, Yarin}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {559--568}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/smith18a/smith18a.pdf}, url = {https://proceedings.mlr.press/r16/smith18a.html}, abstract = {Measuring uncertainty is a promising technique for detecting adversarial examples, crafted in- puts on which the model predicts an incorrect class with high confidence. There are various measures of uncertainty, including predictive entropy and mutual information, each capturing distinct types of uncertainty. We study these measures, and shed light on why mutual infor- mation seems to be effective at the task of adver- sarial example detection. We highlight failure modes for MC dropout, a widely used approach for estimating uncertainty in deep models. This leads to an improved understanding of the draw- backs of current methods, and a proposal to im- prove the quality of uncertainty estimates using probabilistic model ensembles. We give illustra- tive experiments using MNIST to demonstrate the intuition underlying the different measures of uncertainty, as well as experiments on a real- world Kaggle dogs vs cats classification dataset.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Understanding Measures of Uncertainty for Adversarial Example Detection %A Lewis Smith %A Yarin Gal %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-smith18a %I PMLR %P 559--568 %U https://proceedings.mlr.press/r16/smith18a.html %V R16 %X Measuring uncertainty is a promising technique for detecting adversarial examples, crafted in- puts on which the model predicts an incorrect class with high confidence. There are various measures of uncertainty, including predictive entropy and mutual information, each capturing distinct types of uncertainty. We study these measures, and shed light on why mutual infor- mation seems to be effective at the task of adver- sarial example detection. We highlight failure modes for MC dropout, a widely used approach for estimating uncertainty in deep models. This leads to an improved understanding of the draw- backs of current methods, and a proposal to im- prove the quality of uncertainty estimates using probabilistic model ensembles. We give illustra- tive experiments using MNIST to demonstrate the intuition underlying the different measures of uncertainty, as well as experiments on a real- world Kaggle dogs vs cats classification dataset. %Z Reissued by PMLR on 04 October 2026.
APA
Smith, L. & Gal, Y.. (2018). Understanding Measures of Uncertainty for Adversarial Example Detection. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:559-568 Available from https://proceedings.mlr.press/r16/smith18a.html. Reissued by PMLR on 04 October 2026.

Related Material