Exploiting Concavity Information in Contextual Bandit Optimization

Kevin Li, Eric Laber
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3762-3784, 2026.

Abstract

The contextual bandit models sequential decision-making problems in which rewards depend on both the chosen action and observed context. In many domains such as medicine, business, and engineering, prior knowledge provides structural information about the reward function that can be exploited to improve optimization efficiency. This paper studies settings where, for each observable context, the conditional mean reward is known to be concave with respect to the action variable. To leverage this structure, we develop a contextual bandit algorithm that conditions a {Bayesian} {Gaussian} Process posterior on concavity constraints. We propose a novel reward model that combines a concavity-preserving regression spline basis with a constrained GP posterior, yielding a tractable shape-constrained estimator. Building on this model, we construct a {UCB} algorithm and establish new posterior concentration inequalities for the constrained posterior, leading to regret guarantees that are never worse than applying standard GP-{UCB} without concavity information and can be strictly tighter when concavity is informative. Experiments on benchmark problems and a Warfarin dosing test application demonstrate substantial reductions in cumulative regret relative to state-of-the-art baselines.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-li26k, title = {Exploiting Concavity Information in Contextual Bandit Optimization}, author = {Li, Kevin and Laber, Eric}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {3762--3784}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/li26k/li26k.pdf}, url = {https://proceedings.mlr.press/v337/li26k.html}, abstract = {The contextual bandit models sequential decision-making problems in which rewards depend on both the chosen action and observed context. In many domains such as medicine, business, and engineering, prior knowledge provides structural information about the reward function that can be exploited to improve optimization efficiency. This paper studies settings where, for each observable context, the conditional mean reward is known to be concave with respect to the action variable. To leverage this structure, we develop a contextual bandit algorithm that conditions a {Bayesian} {Gaussian} Process posterior on concavity constraints. We propose a novel reward model that combines a concavity-preserving regression spline basis with a constrained GP posterior, yielding a tractable shape-constrained estimator. Building on this model, we construct a {UCB} algorithm and establish new posterior concentration inequalities for the constrained posterior, leading to regret guarantees that are never worse than applying standard GP-{UCB} without concavity information and can be strictly tighter when concavity is informative. Experiments on benchmark problems and a Warfarin dosing test application demonstrate substantial reductions in cumulative regret relative to state-of-the-art baselines.} }
Endnote
%0 Conference Paper %T Exploiting Concavity Information in Contextual Bandit Optimization %A Kevin Li %A Eric Laber %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-li26k %I PMLR %P 3762--3784 %U https://proceedings.mlr.press/v337/li26k.html %V 337 %X The contextual bandit models sequential decision-making problems in which rewards depend on both the chosen action and observed context. In many domains such as medicine, business, and engineering, prior knowledge provides structural information about the reward function that can be exploited to improve optimization efficiency. This paper studies settings where, for each observable context, the conditional mean reward is known to be concave with respect to the action variable. To leverage this structure, we develop a contextual bandit algorithm that conditions a {Bayesian} {Gaussian} Process posterior on concavity constraints. We propose a novel reward model that combines a concavity-preserving regression spline basis with a constrained GP posterior, yielding a tractable shape-constrained estimator. Building on this model, we construct a {UCB} algorithm and establish new posterior concentration inequalities for the constrained posterior, leading to regret guarantees that are never worse than applying standard GP-{UCB} without concavity information and can be strictly tighter when concavity is informative. Experiments on benchmark problems and a Warfarin dosing test application demonstrate substantial reductions in cumulative regret relative to state-of-the-art baselines.
APA
Li, K. & Laber, E.. (2026). Exploiting Concavity Information in Contextual Bandit Optimization. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:3762-3784 Available from https://proceedings.mlr.press/v337/li26k.html.

Related Material