Two Modalities Are Better Than One: Efficient Adversarial Purification via Multimodal Diffusion Models

Mingyuan Bai, Wei Huang, Tenghui Li, Andong Wang, Chao Li, Cesar F Caiafa, Junbin Gao, Qibin Zhao
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:5265-5279, 2026.

Abstract

Adversarial purification uses generative models to restore clean data distributions from unseen attacks without retraining classifiers. However, unimodal diffusion-based approaches struggle to preserve semantic consistency, while recent multimodal variants rely on computationally expensive adversarial training or distillation. Both approaches often lack theoretical guarantees. In this work, we propose MultiDAP, a novel framework leveraging multimodal diffusion models for efficient adversarial purification. MultiDAP first learns continuous class-agnostic prompts from clean data to capture rich semantic priors, replacing rigid hand-crafted templates. Guided by these prompts, MultiDAP purifies adversarial inputs by minimizing a regularized DDPM loss for only a few steps (e.g., 5-20). We provide theoretical guarantees for both the likelihood improvement via prompt learning and the convergence of the purification process. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet-1K demonstrate that MultiDAP matches the robustness of state-of-the-art baselines but with improved efficiency.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bai26b, title = {Two Modalities Are Better Than One: Efficient Adversarial Purification via Multimodal Diffusion Models}, author = {Bai, Mingyuan and Huang, Wei and Li, Tenghui and Wang, Andong and Li, Chao and Caiafa, Cesar F and Gao, Junbin and Zhao, Qibin}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {5265--5279}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bai26b/bai26b.pdf}, url = {https://proceedings.mlr.press/v306/bai26b.html}, abstract = {Adversarial purification uses generative models to restore clean data distributions from unseen attacks without retraining classifiers. However, unimodal diffusion-based approaches struggle to preserve semantic consistency, while recent multimodal variants rely on computationally expensive adversarial training or distillation. Both approaches often lack theoretical guarantees. In this work, we propose MultiDAP, a novel framework leveraging multimodal diffusion models for efficient adversarial purification. MultiDAP first learns continuous class-agnostic prompts from clean data to capture rich semantic priors, replacing rigid hand-crafted templates. Guided by these prompts, MultiDAP purifies adversarial inputs by minimizing a regularized DDPM loss for only a few steps (e.g., 5-20). We provide theoretical guarantees for both the likelihood improvement via prompt learning and the convergence of the purification process. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet-1K demonstrate that MultiDAP matches the robustness of state-of-the-art baselines but with improved efficiency.} }
Endnote
%0 Conference Paper %T Two Modalities Are Better Than One: Efficient Adversarial Purification via Multimodal Diffusion Models %A Mingyuan Bai %A Wei Huang %A Tenghui Li %A Andong Wang %A Chao Li %A Cesar F Caiafa %A Junbin Gao %A Qibin Zhao %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bai26b %I PMLR %P 5265--5279 %U https://proceedings.mlr.press/v306/bai26b.html %V 306 %X Adversarial purification uses generative models to restore clean data distributions from unseen attacks without retraining classifiers. However, unimodal diffusion-based approaches struggle to preserve semantic consistency, while recent multimodal variants rely on computationally expensive adversarial training or distillation. Both approaches often lack theoretical guarantees. In this work, we propose MultiDAP, a novel framework leveraging multimodal diffusion models for efficient adversarial purification. MultiDAP first learns continuous class-agnostic prompts from clean data to capture rich semantic priors, replacing rigid hand-crafted templates. Guided by these prompts, MultiDAP purifies adversarial inputs by minimizing a regularized DDPM loss for only a few steps (e.g., 5-20). We provide theoretical guarantees for both the likelihood improvement via prompt learning and the convergence of the purification process. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet-1K demonstrate that MultiDAP matches the robustness of state-of-the-art baselines but with improved efficiency.
APA
Bai, M., Huang, W., Li, T., Wang, A., Li, C., Caiafa, C.F., Gao, J. & Zhao, Q.. (2026). Two Modalities Are Better Than One: Efficient Adversarial Purification via Multimodal Diffusion Models. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:5265-5279 Available from https://proceedings.mlr.press/v306/bai26b.html.

Related Material