[edit]
Exploiting Concavity Information in Contextual Bandit Optimization
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3762-3784, 2026.
Abstract
The contextual bandit models sequential decision-making problems in which rewards depend on both the chosen action and observed context. In many domains such as medicine, business, and engineering, prior knowledge provides structural information about the reward function that can be exploited to improve optimization efficiency. This paper studies settings where, for each observable context, the conditional mean reward is known to be concave with respect to the action variable. To leverage this structure, we develop a contextual bandit algorithm that conditions a {Bayesian} {Gaussian} Process posterior on concavity constraints. We propose a novel reward model that combines a concavity-preserving regression spline basis with a constrained GP posterior, yielding a tractable shape-constrained estimator. Building on this model, we construct a {UCB} algorithm and establish new posterior concentration inequalities for the constrained posterior, leading to regret guarantees that are never worse than applying standard GP-{UCB} without concavity information and can be strictly tighter when concavity is informative. Experiments on benchmark problems and a Warfarin dosing test application demonstrate substantial reductions in cumulative regret relative to state-of-the-art baselines.