Adaptively Grouped Contextual Bandits for Heterogeneous Human-AI Decision Making with Conformal Prediction Sets

Yanchen Wu, Bo Li
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:136127-136192, 2026.

Abstract

Personalizing AI decision support for heterogeneous human decision-makers remains a key challenge. We study a collaboration workflow where AI provides a reduced prediction set via conformal prediction and the human makes the final decision based on the set. We formulate this personalization problem as a contextual bandit, where individual and task features form the context, candidate significance levels $\alpha$ serve as arms, and the optimal prediction-set size varies across contexts. To address large arm spaces and high-dimensional contexts, we introduce the Adaptively Grouped Contextual Bandit (AGCB) framework, which avoids global function approximation by exploiting two Human-AI structural assumptions: continuity and monotonicity. Continuity enables information sharing across nearby contexts and decisions, and drives a data-driven Zooming Mechanism that balances intra-group estimation error against inter-group approximation bias. Monotonicity converts each observation into directional counterfactual information over the $K$ candidate $\alpha$ values, reducing the arm-dependence factor from polynomial to logarithmic in $K$. Together, these mechanisms yield minimax-optimal dependence on the learning horizon $T$ for both cumulative and simple regret objectives. Empirical results confirm that AGCB achieves the strongest overall performance across most heterogeneous, data-scarce settings.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-wu26aa, title = {Adaptively Grouped Contextual Bandits for Heterogeneous Human-{AI} Decision Making with Conformal Prediction Sets}, author = {Wu, Yanchen and Li, Bo}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {136127--136192}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/wu26aa/wu26aa.pdf}, url = {https://proceedings.mlr.press/v306/wu26aa.html}, abstract = {Personalizing AI decision support for heterogeneous human decision-makers remains a key challenge. We study a collaboration workflow where AI provides a reduced prediction set via conformal prediction and the human makes the final decision based on the set. We formulate this personalization problem as a contextual bandit, where individual and task features form the context, candidate significance levels $\alpha$ serve as arms, and the optimal prediction-set size varies across contexts. To address large arm spaces and high-dimensional contexts, we introduce the Adaptively Grouped Contextual Bandit (AGCB) framework, which avoids global function approximation by exploiting two Human-AI structural assumptions: continuity and monotonicity. Continuity enables information sharing across nearby contexts and decisions, and drives a data-driven Zooming Mechanism that balances intra-group estimation error against inter-group approximation bias. Monotonicity converts each observation into directional counterfactual information over the $K$ candidate $\alpha$ values, reducing the arm-dependence factor from polynomial to logarithmic in $K$. Together, these mechanisms yield minimax-optimal dependence on the learning horizon $T$ for both cumulative and simple regret objectives. Empirical results confirm that AGCB achieves the strongest overall performance across most heterogeneous, data-scarce settings.} }
Endnote
%0 Conference Paper %T Adaptively Grouped Contextual Bandits for Heterogeneous Human-AI Decision Making with Conformal Prediction Sets %A Yanchen Wu %A Bo Li %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-wu26aa %I PMLR %P 136127--136192 %U https://proceedings.mlr.press/v306/wu26aa.html %V 306 %X Personalizing AI decision support for heterogeneous human decision-makers remains a key challenge. We study a collaboration workflow where AI provides a reduced prediction set via conformal prediction and the human makes the final decision based on the set. We formulate this personalization problem as a contextual bandit, where individual and task features form the context, candidate significance levels $\alpha$ serve as arms, and the optimal prediction-set size varies across contexts. To address large arm spaces and high-dimensional contexts, we introduce the Adaptively Grouped Contextual Bandit (AGCB) framework, which avoids global function approximation by exploiting two Human-AI structural assumptions: continuity and monotonicity. Continuity enables information sharing across nearby contexts and decisions, and drives a data-driven Zooming Mechanism that balances intra-group estimation error against inter-group approximation bias. Monotonicity converts each observation into directional counterfactual information over the $K$ candidate $\alpha$ values, reducing the arm-dependence factor from polynomial to logarithmic in $K$. Together, these mechanisms yield minimax-optimal dependence on the learning horizon $T$ for both cumulative and simple regret objectives. Empirical results confirm that AGCB achieves the strongest overall performance across most heterogeneous, data-scarce settings.
APA
Wu, Y. & Li, B.. (2026). Adaptively Grouped Contextual Bandits for Heterogeneous Human-AI Decision Making with Conformal Prediction Sets. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:136127-136192 Available from https://proceedings.mlr.press/v306/wu26aa.html.

Related Material