Robust Learning of A Group DRO Neuron

Guyang Cao, Shuyao Li, Sushrut Karmalkar, Jelena Diakonikolas
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4006-4014, 2026.

Abstract

We study the problem of learning a single neuron under standard squared loss in the presence of arbitrary label noise and group-level distributional shifts, for a broad family of covariate distributions. Our goal is to identify a "best-fit" neuron parameterized by ${\boldsymbol w}_{*}$ that performs well under the most challenging reweighting of the groups. Specifically, we address a Group Distributionally Robust Optimization problem: given sample access to $K$ distinct distributions ${\mathcal p_{[1]}},…, {\mathcal p_{[K]}}$, we seek to approximate ${\boldsymbol w}_{*}$ that minimizes the worst-case objective over convex combinations of group distributions ${\boldsymbol \lambda} \in \Delta_K$, where the objective is $\sum_{i \in [K]}\lambda_{[i]},\mathbb E_{(\mathbf x,y)\sim{\mathcal p_{[i]}}}(\sigma(\boldsymbol w\cdot\boldsymbol x)-y)^2 - \nu d_f(\boldsymbol\lambda,\tfrac1K\boldsymbol1)$ and $d_f$ is an $f$-divergence that imposes (optional) penalty on deviations from uniform group weights, scaled by a parameter $\nu \geq 0$. We develop a computationally efficient primal-dual algorithm that outputs a vector $\widehat{\boldsymbol w}$ that is constant-factor competitive with ${\boldsymbol w}_{*}$ under the worst-case group weighting. Our analytical framework directly confronts the inherent nonconvexity of the loss function, providing robust learning guarantees in the face of arbitrary label corruptions and group-specific distributional shifts. The implementation of the dual extrapolation update motivated by our algorithmic framework shows promise on LLM pre-training benchmarks.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-cao26c, title = { Robust Learning of A Group DRO Neuron }, author = {Cao, Guyang and Li, Shuyao and Karmalkar, Sushrut and Diakonikolas, Jelena}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4006--4014}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/cao26c/cao26c.pdf}, url = {https://proceedings.mlr.press/v300/cao26c.html}, abstract = { We study the problem of learning a single neuron under standard squared loss in the presence of arbitrary label noise and group-level distributional shifts, for a broad family of covariate distributions. Our goal is to identify a "best-fit" neuron parameterized by ${\boldsymbol w}_{*}$ that performs well under the most challenging reweighting of the groups. Specifically, we address a Group Distributionally Robust Optimization problem: given sample access to $K$ distinct distributions ${\mathcal p_{[1]}},…, {\mathcal p_{[K]}}$, we seek to approximate ${\boldsymbol w}_{*}$ that minimizes the worst-case objective over convex combinations of group distributions ${\boldsymbol \lambda} \in \Delta_K$, where the objective is $\sum_{i \in [K]}\lambda_{[i]},\mathbb E_{(\mathbf x,y)\sim{\mathcal p_{[i]}}}(\sigma(\boldsymbol w\cdot\boldsymbol x)-y)^2 - \nu d_f(\boldsymbol\lambda,\tfrac1K\boldsymbol1)$ and $d_f$ is an $f$-divergence that imposes (optional) penalty on deviations from uniform group weights, scaled by a parameter $\nu \geq 0$. We develop a computationally efficient primal-dual algorithm that outputs a vector $\widehat{\boldsymbol w}$ that is constant-factor competitive with ${\boldsymbol w}_{*}$ under the worst-case group weighting. Our analytical framework directly confronts the inherent nonconvexity of the loss function, providing robust learning guarantees in the face of arbitrary label corruptions and group-specific distributional shifts. The implementation of the dual extrapolation update motivated by our algorithmic framework shows promise on LLM pre-training benchmarks. } }
Endnote
%0 Conference Paper %T Robust Learning of A Group DRO Neuron %A Guyang Cao %A Shuyao Li %A Sushrut Karmalkar %A Jelena Diakonikolas %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-cao26c %I PMLR %P 4006--4014 %U https://proceedings.mlr.press/v300/cao26c.html %V 300 %X We study the problem of learning a single neuron under standard squared loss in the presence of arbitrary label noise and group-level distributional shifts, for a broad family of covariate distributions. Our goal is to identify a "best-fit" neuron parameterized by ${\boldsymbol w}_{*}$ that performs well under the most challenging reweighting of the groups. Specifically, we address a Group Distributionally Robust Optimization problem: given sample access to $K$ distinct distributions ${\mathcal p_{[1]}},…, {\mathcal p_{[K]}}$, we seek to approximate ${\boldsymbol w}_{*}$ that minimizes the worst-case objective over convex combinations of group distributions ${\boldsymbol \lambda} \in \Delta_K$, where the objective is $\sum_{i \in [K]}\lambda_{[i]},\mathbb E_{(\mathbf x,y)\sim{\mathcal p_{[i]}}}(\sigma(\boldsymbol w\cdot\boldsymbol x)-y)^2 - \nu d_f(\boldsymbol\lambda,\tfrac1K\boldsymbol1)$ and $d_f$ is an $f$-divergence that imposes (optional) penalty on deviations from uniform group weights, scaled by a parameter $\nu \geq 0$. We develop a computationally efficient primal-dual algorithm that outputs a vector $\widehat{\boldsymbol w}$ that is constant-factor competitive with ${\boldsymbol w}_{*}$ under the worst-case group weighting. Our analytical framework directly confronts the inherent nonconvexity of the loss function, providing robust learning guarantees in the face of arbitrary label corruptions and group-specific distributional shifts. The implementation of the dual extrapolation update motivated by our algorithmic framework shows promise on LLM pre-training benchmarks.
APA
Cao, G., Li, S., Karmalkar, S. & Diakonikolas, J.. (2026). Robust Learning of A Group DRO Neuron . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4006-4014 Available from https://proceedings.mlr.press/v300/cao26c.html.

Related Material