Statistically Undetectable Backdoors in Deep Neural Networks

Andrej Bogdanov, Alon Rosen, Neekon Vafa
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:8794-8823, 2026.

Abstract

We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the full descriptions of the models (e.g., all of the weights). The backdoor provides access to invariance-based adversarial examples for every input, mapping distant inputs to unusually close outputs. However, without the backdoor, it is provably impossible (under standard cryptographic assumptions) to generate any such adversarial examples in polynomial time. Our theoretical and preliminary empirical findings demonstrate a fundamental power asymmetry between model trainers and model users.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bogdanov26a, title = {Statistically Undetectable Backdoors in Deep Neural Networks}, author = {Bogdanov, Andrej and Rosen, Alon and Vafa, Neekon}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {8794--8823}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bogdanov26a/bogdanov26a.pdf}, url = {https://proceedings.mlr.press/v306/bogdanov26a.html}, abstract = {We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the full descriptions of the models (e.g., all of the weights). The backdoor provides access to invariance-based adversarial examples for every input, mapping distant inputs to unusually close outputs. However, without the backdoor, it is provably impossible (under standard cryptographic assumptions) to generate any such adversarial examples in polynomial time. Our theoretical and preliminary empirical findings demonstrate a fundamental power asymmetry between model trainers and model users.} }
Endnote
%0 Conference Paper %T Statistically Undetectable Backdoors in Deep Neural Networks %A Andrej Bogdanov %A Alon Rosen %A Neekon Vafa %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bogdanov26a %I PMLR %P 8794--8823 %U https://proceedings.mlr.press/v306/bogdanov26a.html %V 306 %X We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the full descriptions of the models (e.g., all of the weights). The backdoor provides access to invariance-based adversarial examples for every input, mapping distant inputs to unusually close outputs. However, without the backdoor, it is provably impossible (under standard cryptographic assumptions) to generate any such adversarial examples in polynomial time. Our theoretical and preliminary empirical findings demonstrate a fundamental power asymmetry between model trainers and model users.
APA
Bogdanov, A., Rosen, A. & Vafa, N.. (2026). Statistically Undetectable Backdoors in Deep Neural Networks. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:8794-8823 Available from https://proceedings.mlr.press/v306/bogdanov26a.html.

Related Material