Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG

Sarthak Choudhary, Nils Palumbo, Ashish Hooda, Krishnamurthy Dj Dvijotham, Somesh Jha
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:20543-20565, 2026.

Abstract

Retrieval-augmented generation (RAG) systems are vulnerable to attacks that inject poisoned passages into the retrieved context, even at low corruption rates. We show that existing attacks are not designed to be stealthy, allowing reliable detection and mitigation. We formalize a distinguishability-based security game to quantify stealth for such attacks. If a few poisoned passages control the response, they must bias the inference process more than the benign ones, inherently compromising stealth. This motivates analyzing intermediate signals of LLMs, such as attention weights, to approximate the influence of different passages on the response. Leveraging attention weights, we introduce the Normalized Passage Attention Score (NPAS) and a lightweight Attention-Variance Filter (AV Filter) that flags anomalous passages. Our method improves robustness, yielding up to 20% higher accuracy than baseline defenses. We also develop adaptive attacks that attempt to conceal such anomalies, achieving up to 35% success rate and underscoring the challenges of achieving true stealth in poisoning RAG systems.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-choudhary26b, title = {Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in {RAG}}, author = {Choudhary, Sarthak and Palumbo, Nils and Hooda, Ashish and Dvijotham, Krishnamurthy Dj and Jha, Somesh}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {20543--20565}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/choudhary26b/choudhary26b.pdf}, url = {https://proceedings.mlr.press/v306/choudhary26b.html}, abstract = {Retrieval-augmented generation (RAG) systems are vulnerable to attacks that inject poisoned passages into the retrieved context, even at low corruption rates. We show that existing attacks are not designed to be stealthy, allowing reliable detection and mitigation. We formalize a distinguishability-based security game to quantify stealth for such attacks. If a few poisoned passages control the response, they must bias the inference process more than the benign ones, inherently compromising stealth. This motivates analyzing intermediate signals of LLMs, such as attention weights, to approximate the influence of different passages on the response. Leveraging attention weights, we introduce the Normalized Passage Attention Score (NPAS) and a lightweight Attention-Variance Filter (AV Filter) that flags anomalous passages. Our method improves robustness, yielding up to 20% higher accuracy than baseline defenses. We also develop adaptive attacks that attempt to conceal such anomalies, achieving up to 35% success rate and underscoring the challenges of achieving true stealth in poisoning RAG systems.} }
Endnote
%0 Conference Paper %T Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG %A Sarthak Choudhary %A Nils Palumbo %A Ashish Hooda %A Krishnamurthy Dj Dvijotham %A Somesh Jha %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-choudhary26b %I PMLR %P 20543--20565 %U https://proceedings.mlr.press/v306/choudhary26b.html %V 306 %X Retrieval-augmented generation (RAG) systems are vulnerable to attacks that inject poisoned passages into the retrieved context, even at low corruption rates. We show that existing attacks are not designed to be stealthy, allowing reliable detection and mitigation. We formalize a distinguishability-based security game to quantify stealth for such attacks. If a few poisoned passages control the response, they must bias the inference process more than the benign ones, inherently compromising stealth. This motivates analyzing intermediate signals of LLMs, such as attention weights, to approximate the influence of different passages on the response. Leveraging attention weights, we introduce the Normalized Passage Attention Score (NPAS) and a lightweight Attention-Variance Filter (AV Filter) that flags anomalous passages. Our method improves robustness, yielding up to 20% higher accuracy than baseline defenses. We also develop adaptive attacks that attempt to conceal such anomalies, achieving up to 35% success rate and underscoring the challenges of achieving true stealth in poisoning RAG systems.
APA
Choudhary, S., Palumbo, N., Hooda, A., Dvijotham, K.D. & Jha, S.. (2026). Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:20543-20565 Available from https://proceedings.mlr.press/v306/choudhary26b.html.

Related Material