[edit]
Weighted Sequential Bayesian Inference for Non-Stationary Linear Contextual Bandits
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7314-7340, 2026.
Abstract
In non-stationary linear contextual bandits, existing efficient algorithms typically rely on the Weighted Regularized Least-Squares (WRLS) estimator. Because WRLS only provides point estimates, previous methods typically construct surrogate distributions when aiming to perform {Bayesian}-like randomized exploration. To more properly establish the {Bayesian} principles, we introduce *Weighted Sequential {Bayesian}* (WSB) inference, which forms a sequence of posteriors over a sequence of non-stationary reward parameters. This {Bayesian} take allows us to isolate the influence of initial beliefs into a dynamic prior penalty evaluated through the posterior covariance, which typically decreases over time. Building on this framework, we instantiate three WSB-based algorithms for exploration: *WSB-LinUCB*, *WSB-RandLinUCB*, and *WSB-LinTS*. By extending a refined drift analysis to randomized exploration without requiring local norms, we establish frequentist regret guarantees that match state-of-the-art WRLS-based baselines. Empirically, WSB’s dynamic prior penalty reduces over-conservatism, allowing our algorithms to consistently match or exceed their WRLS-based counterparts. Lastly, we also provide a simplified proof for the time-uniform concentration of vector-valued martingales, a critical subroutine used throughout the literature, that might be of independent interest.