Erased but Not Forgotten: How Backdoors Compromise Concept Erasure

Tobias Braun, Jonas Henry Grebe, Patrick Mohr Gordillo, Marcus Rohrbach, Anna Rohrbach
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:9734-9756, 2026.

Abstract

The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to sexually explicit imagery. To mitigate such risks, prior work has proposed concept erasure methods that aim to sever unwanted concepts from the model via fine-tuning, yet it remains unclear whether these approaches truly remove all links to the harmful concept or merely conceal superficial connections. In this work, we reveal a critical vulnerability, the Erasure Evasion Backdoor (EEB): an adversary binds a backdoor trigger to a concept slated for removal, and this malicious link survives subsequent erasure. We show that both black-box and white-box adversaries can instantiate this threat. Across six state-of-the-art erasure methods, including robust ones that explicitly search for alternative representations of the target concept, EEB consistently exposes harmful content: up to 82% success against celebrity-identity unlearning, up to 94% for object erasure, and up to 16$\times$ amplification of explicit-content exposure. While EEB uncovers a blind spot in current erasure methods, it also provides a diagnostic tool for stress-testing future concept erasure techniques.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-braun26b, title = {Erased but Not Forgotten: How Backdoors Compromise Concept Erasure}, author = {Braun, Tobias and Grebe, Jonas Henry and Gordillo, Patrick Mohr and Rohrbach, Marcus and Rohrbach, Anna}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {9734--9756}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/braun26b/braun26b.pdf}, url = {https://proceedings.mlr.press/v306/braun26b.html}, abstract = {The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to sexually explicit imagery. To mitigate such risks, prior work has proposed concept erasure methods that aim to sever unwanted concepts from the model via fine-tuning, yet it remains unclear whether these approaches truly remove all links to the harmful concept or merely conceal superficial connections. In this work, we reveal a critical vulnerability, the Erasure Evasion Backdoor (EEB): an adversary binds a backdoor trigger to a concept slated for removal, and this malicious link survives subsequent erasure. We show that both black-box and white-box adversaries can instantiate this threat. Across six state-of-the-art erasure methods, including robust ones that explicitly search for alternative representations of the target concept, EEB consistently exposes harmful content: up to 82% success against celebrity-identity unlearning, up to 94% for object erasure, and up to 16$\times$ amplification of explicit-content exposure. While EEB uncovers a blind spot in current erasure methods, it also provides a diagnostic tool for stress-testing future concept erasure techniques.} }
Endnote
%0 Conference Paper %T Erased but Not Forgotten: How Backdoors Compromise Concept Erasure %A Tobias Braun %A Jonas Henry Grebe %A Patrick Mohr Gordillo %A Marcus Rohrbach %A Anna Rohrbach %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-braun26b %I PMLR %P 9734--9756 %U https://proceedings.mlr.press/v306/braun26b.html %V 306 %X The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to sexually explicit imagery. To mitigate such risks, prior work has proposed concept erasure methods that aim to sever unwanted concepts from the model via fine-tuning, yet it remains unclear whether these approaches truly remove all links to the harmful concept or merely conceal superficial connections. In this work, we reveal a critical vulnerability, the Erasure Evasion Backdoor (EEB): an adversary binds a backdoor trigger to a concept slated for removal, and this malicious link survives subsequent erasure. We show that both black-box and white-box adversaries can instantiate this threat. Across six state-of-the-art erasure methods, including robust ones that explicitly search for alternative representations of the target concept, EEB consistently exposes harmful content: up to 82% success against celebrity-identity unlearning, up to 94% for object erasure, and up to 16$\times$ amplification of explicit-content exposure. While EEB uncovers a blind spot in current erasure methods, it also provides a diagnostic tool for stress-testing future concept erasure techniques.
APA
Braun, T., Grebe, J.H., Gordillo, P.M., Rohrbach, M. & Rohrbach, A.. (2026). Erased but Not Forgotten: How Backdoors Compromise Concept Erasure. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:9734-9756 Available from https://proceedings.mlr.press/v306/braun26b.html.

Related Material