Unlearning Isn’t Forgetting: Revealing Hidden Leakage in Class Unlearning Evaluations

Ali Ebrahimpour-Boroojeny, Yian Wang, Hari Sundaram
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:27569-27593, 2026.

Abstract

In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause information leakage about the forgotten class. We further propose a simple unlearning strategy to mitigate this issue. We introduce Class Membership Inference Attack (CMIA) that uses the probabilities the model assigns to neighboring classes to detect unlearned samples. We find that existing unlearning methods are vulnerable to CMIA across multiple datasets. We then propose a new fine-tuning objective that mitigates this privacy leakage by approximating, for forget-class inputs, the distribution over the remaining classes that a retrained-from-scratch model would produce. To construct this approximation, we estimate inter-class similarity and tilt the target model’s distribution accordingly. The resulting Tilted REWeighting (TREW) distribution serves as the desired distribution during fine-tuning. We also show that across multiple benchmarks, TREW matches or surpasses existing unlearning methods on prior unlearning metrics. More specifically, on CIFAR-10, it reduces the gap with retrained models by $19%$ and $46%$ for U-LiRA and CMIA scores, accordingly, compared to the SOTA method for each category.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-ebrahimpour-boroojeny26a, title = {Unlearning Isn’t Forgetting: Revealing Hidden Leakage in Class Unlearning Evaluations}, author = {Ebrahimpour-Boroojeny, Ali and Wang, Yian and Sundaram, Hari}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {27569--27593}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/ebrahimpour-boroojeny26a/ebrahimpour-boroojeny26a.pdf}, url = {https://proceedings.mlr.press/v306/ebrahimpour-boroojeny26a.html}, abstract = {In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause information leakage about the forgotten class. We further propose a simple unlearning strategy to mitigate this issue. We introduce Class Membership Inference Attack (CMIA) that uses the probabilities the model assigns to neighboring classes to detect unlearned samples. We find that existing unlearning methods are vulnerable to CMIA across multiple datasets. We then propose a new fine-tuning objective that mitigates this privacy leakage by approximating, for forget-class inputs, the distribution over the remaining classes that a retrained-from-scratch model would produce. To construct this approximation, we estimate inter-class similarity and tilt the target model’s distribution accordingly. The resulting Tilted REWeighting (TREW) distribution serves as the desired distribution during fine-tuning. We also show that across multiple benchmarks, TREW matches or surpasses existing unlearning methods on prior unlearning metrics. More specifically, on CIFAR-10, it reduces the gap with retrained models by $19%$ and $46%$ for U-LiRA and CMIA scores, accordingly, compared to the SOTA method for each category.} }
Endnote
%0 Conference Paper %T Unlearning Isn’t Forgetting: Revealing Hidden Leakage in Class Unlearning Evaluations %A Ali Ebrahimpour-Boroojeny %A Yian Wang %A Hari Sundaram %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-ebrahimpour-boroojeny26a %I PMLR %P 27569--27593 %U https://proceedings.mlr.press/v306/ebrahimpour-boroojeny26a.html %V 306 %X In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause information leakage about the forgotten class. We further propose a simple unlearning strategy to mitigate this issue. We introduce Class Membership Inference Attack (CMIA) that uses the probabilities the model assigns to neighboring classes to detect unlearned samples. We find that existing unlearning methods are vulnerable to CMIA across multiple datasets. We then propose a new fine-tuning objective that mitigates this privacy leakage by approximating, for forget-class inputs, the distribution over the remaining classes that a retrained-from-scratch model would produce. To construct this approximation, we estimate inter-class similarity and tilt the target model’s distribution accordingly. The resulting Tilted REWeighting (TREW) distribution serves as the desired distribution during fine-tuning. We also show that across multiple benchmarks, TREW matches or surpasses existing unlearning methods on prior unlearning metrics. More specifically, on CIFAR-10, it reduces the gap with retrained models by $19%$ and $46%$ for U-LiRA and CMIA scores, accordingly, compared to the SOTA method for each category.
APA
Ebrahimpour-Boroojeny, A., Wang, Y. & Sundaram, H.. (2026). Unlearning Isn’t Forgetting: Revealing Hidden Leakage in Class Unlearning Evaluations. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:27569-27593 Available from https://proceedings.mlr.press/v306/ebrahimpour-boroojeny26a.html.

Related Material