UnGuide: Learning to Forget with LoRA-Guided Diffusion Models

Alicja Polowczyk, Agnieszka Polowczyk, Dawid Malarz, Artur Kasymov, Jacek Tabor, Marcin Mazur, Przemysław Spurek
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:5471-5503, 2026.

Abstract

Large-scale text-to-image diffusion models pose safety and compliance risks due to their ability to generate harmful or undesirable content. Machine unlearning seeks to remove specific concepts from pretrained models while preserving overall generative quality. Low-Rank Adaptation ({LoRA}) enables parameter-efficient targeted forgetting, but naive {LoRA}-based unlearning often induces distributional shift, degrading fidelity and destabilizing sampling. In this paper, we propose UnGuide, a {LoRA}-based unlearning framework with stability-aware adaptive guidance. Our method dynamically modulates classifier-free guidance ({CFG}) based on early denoising trajectory variance. By estimating discrepancies between base and {LoRA}-adapted noise predictions during initial diffusion steps, UnGuide interpolates per prompt between the frozen and adapted models. High-variance trajectories receive stronger adapted-model guidance to enforce forgetting, while low-variance trajectories rely on the base model to preserve fidelity. This mechanism enables controlled concept suppression without prompt embedding modification or external segmentation. Experiments on object erasure and explicit content removal demonstrate improved forgetting efficacy and superior fidelity retention compared to existing {LoRA}-based approaches.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-polowczyk26a, title = {UnGuide: Learning to Forget with {LoRA}-Guided Diffusion Models}, author = {Polowczyk, Alicja and Polowczyk, Agnieszka and Malarz, Dawid and Kasymov, Artur and Tabor, Jacek and Mazur, Marcin and Spurek, Przemys{\l}aw}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {5471--5503}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/polowczyk26a/polowczyk26a.pdf}, url = {https://proceedings.mlr.press/v337/polowczyk26a.html}, abstract = {Large-scale text-to-image diffusion models pose safety and compliance risks due to their ability to generate harmful or undesirable content. Machine unlearning seeks to remove specific concepts from pretrained models while preserving overall generative quality. Low-Rank Adaptation ({LoRA}) enables parameter-efficient targeted forgetting, but naive {LoRA}-based unlearning often induces distributional shift, degrading fidelity and destabilizing sampling. In this paper, we propose UnGuide, a {LoRA}-based unlearning framework with stability-aware adaptive guidance. Our method dynamically modulates classifier-free guidance ({CFG}) based on early denoising trajectory variance. By estimating discrepancies between base and {LoRA}-adapted noise predictions during initial diffusion steps, UnGuide interpolates per prompt between the frozen and adapted models. High-variance trajectories receive stronger adapted-model guidance to enforce forgetting, while low-variance trajectories rely on the base model to preserve fidelity. This mechanism enables controlled concept suppression without prompt embedding modification or external segmentation. Experiments on object erasure and explicit content removal demonstrate improved forgetting efficacy and superior fidelity retention compared to existing {LoRA}-based approaches.} }
Endnote
%0 Conference Paper %T UnGuide: Learning to Forget with LoRA-Guided Diffusion Models %A Alicja Polowczyk %A Agnieszka Polowczyk %A Dawid Malarz %A Artur Kasymov %A Jacek Tabor %A Marcin Mazur %A Przemysław Spurek %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-polowczyk26a %I PMLR %P 5471--5503 %U https://proceedings.mlr.press/v337/polowczyk26a.html %V 337 %X Large-scale text-to-image diffusion models pose safety and compliance risks due to their ability to generate harmful or undesirable content. Machine unlearning seeks to remove specific concepts from pretrained models while preserving overall generative quality. Low-Rank Adaptation ({LoRA}) enables parameter-efficient targeted forgetting, but naive {LoRA}-based unlearning often induces distributional shift, degrading fidelity and destabilizing sampling. In this paper, we propose UnGuide, a {LoRA}-based unlearning framework with stability-aware adaptive guidance. Our method dynamically modulates classifier-free guidance ({CFG}) based on early denoising trajectory variance. By estimating discrepancies between base and {LoRA}-adapted noise predictions during initial diffusion steps, UnGuide interpolates per prompt between the frozen and adapted models. High-variance trajectories receive stronger adapted-model guidance to enforce forgetting, while low-variance trajectories rely on the base model to preserve fidelity. This mechanism enables controlled concept suppression without prompt embedding modification or external segmentation. Experiments on object erasure and explicit content removal demonstrate improved forgetting efficacy and superior fidelity retention compared to existing {LoRA}-based approaches.
APA
Polowczyk, A., Polowczyk, A., Malarz, D., Kasymov, A., Tabor, J., Mazur, M. & Spurek, P.. (2026). UnGuide: Learning to Forget with LoRA-Guided Diffusion Models. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:5471-5503 Available from https://proceedings.mlr.press/v337/polowczyk26a.html.

Related Material