Understanding SAM’s Robustness to Noisy Labels Through Gradient Down-weighting

Hoang-Chau Luong, Thuc Nguyen-Quang, Dat Ba Tran, Minh-Triet Tran
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4159-4167, 2026.

Abstract

Sharpness-Aware Minimization (SAM) was introduced to improve generalization by seeking flat minima, yet it also exhibits robustness to label noise, a phenomenon that remains only partially understood. Prior work has mainly attributed this effect to SAM’s tendency to prolong the learning of clean samples. In this work, we provide a complementary explanation by analyzing SAM at the element-wise level. We show that when noisy gradients dominate a parameter direction, their influence is reduced by the stronger amplification of clean gradients. This slows the memorization of noisy labels while sustaining clean learning, offering a more complete account of SAM’s robustness. Building on this insight, we propose SANER (Sharpness-Aware Noise-Explicit Reweighting), a simple variant of SAM that explicitly magnifies this down-weighting effect. Experiments on benchmark image classification tasks with noisy labels demonstrate that SANER significantly mitigates noisy-label memorization and improves generalization over both SAM and SGD. Moreover, since SANER is designed from the mechanism of SAM, it can also be seamlessly integrated into SAM-like variants, further boosting their robustness.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-luong26a, title = { Understanding SAM’s Robustness to Noisy Labels Through Gradient Down-weighting }, author = {Luong, Hoang-Chau and Nguyen-Quang, Thuc and Tran, Dat Ba and Tran, Minh-Triet}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4159--4167}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/luong26a/luong26a.pdf}, url = {https://proceedings.mlr.press/v300/luong26a.html}, abstract = { Sharpness-Aware Minimization (SAM) was introduced to improve generalization by seeking flat minima, yet it also exhibits robustness to label noise, a phenomenon that remains only partially understood. Prior work has mainly attributed this effect to SAM’s tendency to prolong the learning of clean samples. In this work, we provide a complementary explanation by analyzing SAM at the element-wise level. We show that when noisy gradients dominate a parameter direction, their influence is reduced by the stronger amplification of clean gradients. This slows the memorization of noisy labels while sustaining clean learning, offering a more complete account of SAM’s robustness. Building on this insight, we propose SANER (Sharpness-Aware Noise-Explicit Reweighting), a simple variant of SAM that explicitly magnifies this down-weighting effect. Experiments on benchmark image classification tasks with noisy labels demonstrate that SANER significantly mitigates noisy-label memorization and improves generalization over both SAM and SGD. Moreover, since SANER is designed from the mechanism of SAM, it can also be seamlessly integrated into SAM-like variants, further boosting their robustness. } }
Endnote
%0 Conference Paper %T Understanding SAM’s Robustness to Noisy Labels Through Gradient Down-weighting %A Hoang-Chau Luong %A Thuc Nguyen-Quang %A Dat Ba Tran %A Minh-Triet Tran %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-luong26a %I PMLR %P 4159--4167 %U https://proceedings.mlr.press/v300/luong26a.html %V 300 %X Sharpness-Aware Minimization (SAM) was introduced to improve generalization by seeking flat minima, yet it also exhibits robustness to label noise, a phenomenon that remains only partially understood. Prior work has mainly attributed this effect to SAM’s tendency to prolong the learning of clean samples. In this work, we provide a complementary explanation by analyzing SAM at the element-wise level. We show that when noisy gradients dominate a parameter direction, their influence is reduced by the stronger amplification of clean gradients. This slows the memorization of noisy labels while sustaining clean learning, offering a more complete account of SAM’s robustness. Building on this insight, we propose SANER (Sharpness-Aware Noise-Explicit Reweighting), a simple variant of SAM that explicitly magnifies this down-weighting effect. Experiments on benchmark image classification tasks with noisy labels demonstrate that SANER significantly mitigates noisy-label memorization and improves generalization over both SAM and SGD. Moreover, since SANER is designed from the mechanism of SAM, it can also be seamlessly integrated into SAM-like variants, further boosting their robustness.
APA
Luong, H., Nguyen-Quang, T., Tran, D.B. & Tran, M.. (2026). Understanding SAM’s Robustness to Noisy Labels Through Gradient Down-weighting . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4159-4167 Available from https://proceedings.mlr.press/v300/luong26a.html.

Related Material