Weighted Sequential Bayesian Inference for Non-Stationary Linear Contextual Bandits

Nicklas Werge, Yi-Shan Wu, Abdullah Akgül, Melih Kandemir
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7314-7340, 2026.

Abstract

In non-stationary linear contextual bandits, existing efficient algorithms typically rely on the Weighted Regularized Least-Squares (WRLS) estimator. Because WRLS only provides point estimates, previous methods typically construct surrogate distributions when aiming to perform {Bayesian}-like randomized exploration. To more properly establish the {Bayesian} principles, we introduce *Weighted Sequential {Bayesian}* (WSB) inference, which forms a sequence of posteriors over a sequence of non-stationary reward parameters. This {Bayesian} take allows us to isolate the influence of initial beliefs into a dynamic prior penalty evaluated through the posterior covariance, which typically decreases over time. Building on this framework, we instantiate three WSB-based algorithms for exploration: *WSB-LinUCB*, *WSB-RandLinUCB*, and *WSB-LinTS*. By extending a refined drift analysis to randomized exploration without requiring local norms, we establish frequentist regret guarantees that match state-of-the-art WRLS-based baselines. Empirically, WSB’s dynamic prior penalty reduces over-conservatism, allowing our algorithms to consistently match or exceed their WRLS-based counterparts. Lastly, we also provide a simplified proof for the time-uniform concentration of vector-valued martingales, a critical subroutine used throughout the literature, that might be of independent interest.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-werge26a, title = {Weighted Sequential {Bayesian} Inference for Non-Stationary Linear Contextual Bandits}, author = {Werge, Nicklas and Wu, Yi-Shan and Akg\"{u}l, Abdullah and Kandemir, Melih}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {7314--7340}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/werge26a/werge26a.pdf}, url = {https://proceedings.mlr.press/v337/werge26a.html}, abstract = {In non-stationary linear contextual bandits, existing efficient algorithms typically rely on the Weighted Regularized Least-Squares (WRLS) estimator. Because WRLS only provides point estimates, previous methods typically construct surrogate distributions when aiming to perform {Bayesian}-like randomized exploration. To more properly establish the {Bayesian} principles, we introduce *Weighted Sequential {Bayesian}* (WSB) inference, which forms a sequence of posteriors over a sequence of non-stationary reward parameters. This {Bayesian} take allows us to isolate the influence of initial beliefs into a dynamic prior penalty evaluated through the posterior covariance, which typically decreases over time. Building on this framework, we instantiate three WSB-based algorithms for exploration: *WSB-LinUCB*, *WSB-RandLinUCB*, and *WSB-LinTS*. By extending a refined drift analysis to randomized exploration without requiring local norms, we establish frequentist regret guarantees that match state-of-the-art WRLS-based baselines. Empirically, WSB’s dynamic prior penalty reduces over-conservatism, allowing our algorithms to consistently match or exceed their WRLS-based counterparts. Lastly, we also provide a simplified proof for the time-uniform concentration of vector-valued martingales, a critical subroutine used throughout the literature, that might be of independent interest.} }
Endnote
%0 Conference Paper %T Weighted Sequential Bayesian Inference for Non-Stationary Linear Contextual Bandits %A Nicklas Werge %A Yi-Shan Wu %A Abdullah Akgül %A Melih Kandemir %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-werge26a %I PMLR %P 7314--7340 %U https://proceedings.mlr.press/v337/werge26a.html %V 337 %X In non-stationary linear contextual bandits, existing efficient algorithms typically rely on the Weighted Regularized Least-Squares (WRLS) estimator. Because WRLS only provides point estimates, previous methods typically construct surrogate distributions when aiming to perform {Bayesian}-like randomized exploration. To more properly establish the {Bayesian} principles, we introduce *Weighted Sequential {Bayesian}* (WSB) inference, which forms a sequence of posteriors over a sequence of non-stationary reward parameters. This {Bayesian} take allows us to isolate the influence of initial beliefs into a dynamic prior penalty evaluated through the posterior covariance, which typically decreases over time. Building on this framework, we instantiate three WSB-based algorithms for exploration: *WSB-LinUCB*, *WSB-RandLinUCB*, and *WSB-LinTS*. By extending a refined drift analysis to randomized exploration without requiring local norms, we establish frequentist regret guarantees that match state-of-the-art WRLS-based baselines. Empirically, WSB’s dynamic prior penalty reduces over-conservatism, allowing our algorithms to consistently match or exceed their WRLS-based counterparts. Lastly, we also provide a simplified proof for the time-uniform concentration of vector-valued martingales, a critical subroutine used throughout the literature, that might be of independent interest.
APA
Werge, N., Wu, Y., Akgül, A. & Kandemir, M.. (2026). Weighted Sequential Bayesian Inference for Non-Stationary Linear Contextual Bandits. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:7314-7340 Available from https://proceedings.mlr.press/v337/werge26a.html.

Related Material