Detecting Adversarial Attacks in Image Sequences with Conformal Test Martingales

Genghua Dong, Ziyun Li, Henrik Boström
Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:652-675, 2026.

Abstract

Deep neural networks used for image classification remain vulnerable to adversarial examples (AEs), especially under white-box attacks where the attacker can exploit information about the target model. Detecting whether a model is under attack is therefore an important step toward mitigating adversarial manipulation. Most existing AE detectors focus on instance-level or batch-level decisions, i.e., determining whether a single image or a fixed batch contains AEs. In this paper, we study a sequence-level detection setting using conformal test martingales. We formulate several nonconformity scores based on image and target-model quantities, convert them into conformal p-values, and monitor the resulting sequences with conformal test martingale algorithms. The resulting detection procedure is evaluated on ImageNet-100 image sequences with a pretrained ResNet-50 target model and three representative white-box attacks. The experiments show that conformal test martingales can detect attacks even when the monitoring segment contains a relatively small number of adversarial examples, but that detection performance depends strongly on both the nonconformity score and the martingale algorithm. Empirically, no single nonconformity score and algorithm pair works well for all attacks. These findings suggest that applying conformal test martingales to detect attacks may benefit from considering multiple image and target-model quantities and designing nonconformity scores according to the observed attack-induced changes.

Cite this Paper


BibTeX
@InProceedings{pmlr-v329-dong26a, title = {Detecting Adversarial Attacks in Image Sequences with Conformal Test Martingales}, author = {Dong, Genghua and Li, Ziyun and Bostr{\"o}m, Henrik}, booktitle = {Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications}, pages = {652--675}, year = {2026}, editor = {Ahlberg, Ernst and Johansson, Ulf and Boström, Henrik and Carlevaro, Alberto and Hallberg Szabadváry, Johan and Carlsson, Lars}, volume = {329}, series = {Proceedings of Machine Learning Research}, month = {02--04 Sep}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v329/main/assets/dong26a/dong26a.pdf}, url = {https://proceedings.mlr.press/v329/dong26a.html}, abstract = {Deep neural networks used for image classification remain vulnerable to adversarial examples (AEs), especially under white-box attacks where the attacker can exploit information about the target model. Detecting whether a model is under attack is therefore an important step toward mitigating adversarial manipulation. Most existing AE detectors focus on instance-level or batch-level decisions, i.e., determining whether a single image or a fixed batch contains AEs. In this paper, we study a sequence-level detection setting using conformal test martingales. We formulate several nonconformity scores based on image and target-model quantities, convert them into conformal p-values, and monitor the resulting sequences with conformal test martingale algorithms. The resulting detection procedure is evaluated on ImageNet-100 image sequences with a pretrained ResNet-50 target model and three representative white-box attacks. The experiments show that conformal test martingales can detect attacks even when the monitoring segment contains a relatively small number of adversarial examples, but that detection performance depends strongly on both the nonconformity score and the martingale algorithm. Empirically, no single nonconformity score and algorithm pair works well for all attacks. These findings suggest that applying conformal test martingales to detect attacks may benefit from considering multiple image and target-model quantities and designing nonconformity scores according to the observed attack-induced changes.} }
Endnote
%0 Conference Paper %T Detecting Adversarial Attacks in Image Sequences with Conformal Test Martingales %A Genghua Dong %A Ziyun Li %A Henrik Boström %B Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications %C Proceedings of Machine Learning Research %D 2026 %E Ernst Ahlberg %E Ulf Johansson %E Henrik Boström %E Alberto Carlevaro %E Johan Hallberg Szabadváry %E Lars Carlsson %F pmlr-v329-dong26a %I PMLR %P 652--675 %U https://proceedings.mlr.press/v329/dong26a.html %V 329 %X Deep neural networks used for image classification remain vulnerable to adversarial examples (AEs), especially under white-box attacks where the attacker can exploit information about the target model. Detecting whether a model is under attack is therefore an important step toward mitigating adversarial manipulation. Most existing AE detectors focus on instance-level or batch-level decisions, i.e., determining whether a single image or a fixed batch contains AEs. In this paper, we study a sequence-level detection setting using conformal test martingales. We formulate several nonconformity scores based on image and target-model quantities, convert them into conformal p-values, and monitor the resulting sequences with conformal test martingale algorithms. The resulting detection procedure is evaluated on ImageNet-100 image sequences with a pretrained ResNet-50 target model and three representative white-box attacks. The experiments show that conformal test martingales can detect attacks even when the monitoring segment contains a relatively small number of adversarial examples, but that detection performance depends strongly on both the nonconformity score and the martingale algorithm. Empirically, no single nonconformity score and algorithm pair works well for all attacks. These findings suggest that applying conformal test martingales to detect attacks may benefit from considering multiple image and target-model quantities and designing nonconformity scores according to the observed attack-induced changes.
APA
Dong, G., Li, Z. & Boström, H.. (2026). Detecting Adversarial Attacks in Image Sequences with Conformal Test Martingales. Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, in Proceedings of Machine Learning Research 329:652-675 Available from https://proceedings.mlr.press/v329/dong26a.html.

Related Material