Linear Regression with Heteroskedastic Errors

Siddhant Chaudhary, Aditya Bhaskara
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:1054-1080, 2026.

Abstract

We study the classic linear regression problem under the challenging setting of heteroskedastic noise. We study a model in which we have $n$ sources that each observe linear measurements of an unknown $d$-dimensional vector $\beta$. Each source has an unknown and distinct variance for the error, and the goal is to estimate the parameter vector $\beta$, under the assumption that there exist \emph{sufficiently many} sources with small error (specifically, $m$ sources have variance at most $1$). Our results show that $\beta$ can be estimated to sub-constant error, as long as the number of small-error sources, i.e., $m$, is large enough. We prove two main results. First, we show that even with just one observation per source, under minimal assumptions on the linear measurements, $\beta$ can be estimated to sub-constant error when $m \ge n^{1- \frac{1}{4d}}$. Second, we show that if we have access to \emph{two} observations per source, under similar assumptions on the linear measurements, $\beta$ can be estimated to sub-constant error when $m \ge n^{5/6}$. Our results are related to the recent line of work on mean estimation with heteroskedastic variances, and more specifically, the \emph{subset of signals} model used in this literature. To the best of our knowledge, our results provide the first sub-constant recovery guarantees for regression in this model.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-chaudhary26a, title = {Linear Regression with Heteroskedastic Errors}, author = {Chaudhary, Siddhant and Bhaskara, Aditya}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {1054--1080}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/chaudhary26a/chaudhary26a.pdf}, url = {https://proceedings.mlr.press/v337/chaudhary26a.html}, abstract = {We study the classic linear regression problem under the challenging setting of heteroskedastic noise. We study a model in which we have $n$ sources that each observe linear measurements of an unknown $d$-dimensional vector $\beta$. Each source has an unknown and distinct variance for the error, and the goal is to estimate the parameter vector $\beta$, under the assumption that there exist \emph{sufficiently many} sources with small error (specifically, $m$ sources have variance at most $1$). Our results show that $\beta$ can be estimated to sub-constant error, as long as the number of small-error sources, i.e., $m$, is large enough. We prove two main results. First, we show that even with just one observation per source, under minimal assumptions on the linear measurements, $\beta$ can be estimated to sub-constant error when $m \ge n^{1- \frac{1}{4d}}$. Second, we show that if we have access to \emph{two} observations per source, under similar assumptions on the linear measurements, $\beta$ can be estimated to sub-constant error when $m \ge n^{5/6}$. Our results are related to the recent line of work on mean estimation with heteroskedastic variances, and more specifically, the \emph{subset of signals} model used in this literature. To the best of our knowledge, our results provide the first sub-constant recovery guarantees for regression in this model.} }
Endnote
%0 Conference Paper %T Linear Regression with Heteroskedastic Errors %A Siddhant Chaudhary %A Aditya Bhaskara %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-chaudhary26a %I PMLR %P 1054--1080 %U https://proceedings.mlr.press/v337/chaudhary26a.html %V 337 %X We study the classic linear regression problem under the challenging setting of heteroskedastic noise. We study a model in which we have $n$ sources that each observe linear measurements of an unknown $d$-dimensional vector $\beta$. Each source has an unknown and distinct variance for the error, and the goal is to estimate the parameter vector $\beta$, under the assumption that there exist \emph{sufficiently many} sources with small error (specifically, $m$ sources have variance at most $1$). Our results show that $\beta$ can be estimated to sub-constant error, as long as the number of small-error sources, i.e., $m$, is large enough. We prove two main results. First, we show that even with just one observation per source, under minimal assumptions on the linear measurements, $\beta$ can be estimated to sub-constant error when $m \ge n^{1- \frac{1}{4d}}$. Second, we show that if we have access to \emph{two} observations per source, under similar assumptions on the linear measurements, $\beta$ can be estimated to sub-constant error when $m \ge n^{5/6}$. Our results are related to the recent line of work on mean estimation with heteroskedastic variances, and more specifically, the \emph{subset of signals} model used in this literature. To the best of our knowledge, our results provide the first sub-constant recovery guarantees for regression in this model.
APA
Chaudhary, S. & Bhaskara, A.. (2026). Linear Regression with Heteroskedastic Errors. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:1054-1080 Available from https://proceedings.mlr.press/v337/chaudhary26a.html.

Related Material