Do-Prompt: Causal Interventions Meet Variational Prompt Bottlenecks

Xueting Chen, Jun-Jie Huang, Yan Yan, Long Lan, Yuhua Tang, Wenjing Yang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:14651-14664, 2026.

Abstract

Multi-modal prompt learning is a parameter-efficient approach to adapting large vision-language models to downstream classification tasks. However, prompts can inadvertently evolve into a high-capacity pathway that encodes environment-dependent spurious correlations, which are predictive only in the source domain and thereby undermine transferability. To address this issue, this paper introduces Do-Prompt, a compress-and-intervene framework that brings together variational bottlenecks and causal interventions for robust prompt tuning. We model prompts as stochastic latent variables and impose a variational prompt bottleneck to explicitly regulate the information transmitted through prompts, effectively mitigating their tendency to memorize spurious nuisance cues. Building on this capacity constraint, we propose lightweight prompt-level interventions by perturbing the environment-related prompt components and enforcing prediction consistency under these do-style perturbations. This synergistic integration encourages reliance on task-stable, invariant semantics rather than spurious prompt content. Notably, Do-Prompt is plug-and-play compatible with existing multi-modal prompt tuning pipelines and introduces negligible computational overhead. Extensive experiments on base-to-novel generalization, cross-dataset transfer, and ImageNet distribution shifts demonstrate consistent performance gains, with particularly notable improvements on datasets exhibiting pronounced domain or texture biases.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26au, title = {Do-Prompt: Causal Interventions Meet Variational Prompt Bottlenecks}, author = {Chen, Xueting and Huang, Jun-Jie and Yan, Yan and Lan, Long and Tang, Yuhua and Yang, Wenjing}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {14651--14664}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26au/chen26au.pdf}, url = {https://proceedings.mlr.press/v306/chen26au.html}, abstract = {Multi-modal prompt learning is a parameter-efficient approach to adapting large vision-language models to downstream classification tasks. However, prompts can inadvertently evolve into a high-capacity pathway that encodes environment-dependent spurious correlations, which are predictive only in the source domain and thereby undermine transferability. To address this issue, this paper introduces Do-Prompt, a compress-and-intervene framework that brings together variational bottlenecks and causal interventions for robust prompt tuning. We model prompts as stochastic latent variables and impose a variational prompt bottleneck to explicitly regulate the information transmitted through prompts, effectively mitigating their tendency to memorize spurious nuisance cues. Building on this capacity constraint, we propose lightweight prompt-level interventions by perturbing the environment-related prompt components and enforcing prediction consistency under these do-style perturbations. This synergistic integration encourages reliance on task-stable, invariant semantics rather than spurious prompt content. Notably, Do-Prompt is plug-and-play compatible with existing multi-modal prompt tuning pipelines and introduces negligible computational overhead. Extensive experiments on base-to-novel generalization, cross-dataset transfer, and ImageNet distribution shifts demonstrate consistent performance gains, with particularly notable improvements on datasets exhibiting pronounced domain or texture biases.} }
Endnote
%0 Conference Paper %T Do-Prompt: Causal Interventions Meet Variational Prompt Bottlenecks %A Xueting Chen %A Jun-Jie Huang %A Yan Yan %A Long Lan %A Yuhua Tang %A Wenjing Yang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26au %I PMLR %P 14651--14664 %U https://proceedings.mlr.press/v306/chen26au.html %V 306 %X Multi-modal prompt learning is a parameter-efficient approach to adapting large vision-language models to downstream classification tasks. However, prompts can inadvertently evolve into a high-capacity pathway that encodes environment-dependent spurious correlations, which are predictive only in the source domain and thereby undermine transferability. To address this issue, this paper introduces Do-Prompt, a compress-and-intervene framework that brings together variational bottlenecks and causal interventions for robust prompt tuning. We model prompts as stochastic latent variables and impose a variational prompt bottleneck to explicitly regulate the information transmitted through prompts, effectively mitigating their tendency to memorize spurious nuisance cues. Building on this capacity constraint, we propose lightweight prompt-level interventions by perturbing the environment-related prompt components and enforcing prediction consistency under these do-style perturbations. This synergistic integration encourages reliance on task-stable, invariant semantics rather than spurious prompt content. Notably, Do-Prompt is plug-and-play compatible with existing multi-modal prompt tuning pipelines and introduces negligible computational overhead. Extensive experiments on base-to-novel generalization, cross-dataset transfer, and ImageNet distribution shifts demonstrate consistent performance gains, with particularly notable improvements on datasets exhibiting pronounced domain or texture biases.
APA
Chen, X., Huang, J., Yan, Y., Lan, L., Tang, Y. & Yang, W.. (2026). Do-Prompt: Causal Interventions Meet Variational Prompt Bottlenecks. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:14651-14664 Available from https://proceedings.mlr.press/v306/chen26au.html.

Related Material