Active learning for stochastic contextual linear bandits

Emma Brunskill, Ishani Karmarkar, Zhaoqi Li
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:577-585, 2026.

Abstract

A key goal in stochastic contextual linear bandits is to efficiently learn a near-optimal policy. Prior algorithms for this problem learn a policy by strategically sampling actions but naively (passively) sampling contexts from the underlying context distribution. However, in many practical scenarios—including online content recommendation, survey research, and clinical trials—practitioners can actively sample or recruit contexts based on prior knowledge of the context distribution. Despite this potential for \emph{active learning}, the role of strategic context sampling in stochastic contextual linear bandits is underexplored. We propose an algorithm that learns a near-optimal policy by strategically sampling rewards of context-action pairs. We prove \emph{instance-dependent} theoretical guarantees demonstrating that our active context sampling strategy can improve over the minimax rate by up to a factor of $\sqrt{d}$, where $d$ is the linear dimension. We show empirically that our algorithm reduces the number of samples needed to learn a near-optimal policy, in tasks such as warfarin dose prediction and joke recommendation.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-brunskill26a, title = { Active learning for stochastic contextual linear bandits }, author = {Brunskill, Emma and Karmarkar, Ishani and Li, Zhaoqi}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {577--585}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/brunskill26a/brunskill26a.pdf}, url = {https://proceedings.mlr.press/v300/brunskill26a.html}, abstract = { A key goal in stochastic contextual linear bandits is to efficiently learn a near-optimal policy. Prior algorithms for this problem learn a policy by strategically sampling actions but naively (passively) sampling contexts from the underlying context distribution. However, in many practical scenarios—including online content recommendation, survey research, and clinical trials—practitioners can actively sample or recruit contexts based on prior knowledge of the context distribution. Despite this potential for \emph{active learning}, the role of strategic context sampling in stochastic contextual linear bandits is underexplored. We propose an algorithm that learns a near-optimal policy by strategically sampling rewards of context-action pairs. We prove \emph{instance-dependent} theoretical guarantees demonstrating that our active context sampling strategy can improve over the minimax rate by up to a factor of $\sqrt{d}$, where $d$ is the linear dimension. We show empirically that our algorithm reduces the number of samples needed to learn a near-optimal policy, in tasks such as warfarin dose prediction and joke recommendation. } }
Endnote
%0 Conference Paper %T Active learning for stochastic contextual linear bandits %A Emma Brunskill %A Ishani Karmarkar %A Zhaoqi Li %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-brunskill26a %I PMLR %P 577--585 %U https://proceedings.mlr.press/v300/brunskill26a.html %V 300 %X A key goal in stochastic contextual linear bandits is to efficiently learn a near-optimal policy. Prior algorithms for this problem learn a policy by strategically sampling actions but naively (passively) sampling contexts from the underlying context distribution. However, in many practical scenarios—including online content recommendation, survey research, and clinical trials—practitioners can actively sample or recruit contexts based on prior knowledge of the context distribution. Despite this potential for \emph{active learning}, the role of strategic context sampling in stochastic contextual linear bandits is underexplored. We propose an algorithm that learns a near-optimal policy by strategically sampling rewards of context-action pairs. We prove \emph{instance-dependent} theoretical guarantees demonstrating that our active context sampling strategy can improve over the minimax rate by up to a factor of $\sqrt{d}$, where $d$ is the linear dimension. We show empirically that our algorithm reduces the number of samples needed to learn a near-optimal policy, in tasks such as warfarin dose prediction and joke recommendation.
APA
Brunskill, E., Karmarkar, I. & Li, Z.. (2026). Active learning for stochastic contextual linear bandits . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:577-585 Available from https://proceedings.mlr.press/v300/brunskill26a.html.

Related Material