Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries

Linfeng Cao, Ming Shi, Ness Shroff
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:900-942, 2026.

Abstract

Personalized decision-making in multi-objective bandits requires learning user-specific trade-offs among competing objectives. Since arm utility depends on both unknown rewards and unknown preferences, existing methods infer preferences only from utility feedback, entangling preference learning with reward exploration. In practice, however, users often reveal their priorities through proactive conversational queries (e.g., “cheap and clean hotel”), yet this structured signal is not leveraged. We formalize a proactive query-based framework in which user queries provide structured preference signals. Modeling these signals via a Plackett–Luce subset choice model, we show that query-only learning is insufficient due to a fundamental shift-invariance barrier. To resolve this, we introduce MO-PQUCB, a hybrid algorithm that integrates query-based preference anchoring with bandit feedback through shift-invariant regularization and dual-exploration {UCB}. We prove that proactive queries accelerate preference estimation and yield improved regret scaling over prior preference-aware MO-MAB methods. Under corrupted queries, we further characterize statistical limits and design a robust estimator achieving near-optimal performance when corruption is sparse. Experiments validate both theoretical and practical gains.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-cao26a, title = {Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries}, author = {Cao, Linfeng and Shi, Ming and Shroff, Ness}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {900--942}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/cao26a/cao26a.pdf}, url = {https://proceedings.mlr.press/v337/cao26a.html}, abstract = {Personalized decision-making in multi-objective bandits requires learning user-specific trade-offs among competing objectives. Since arm utility depends on both unknown rewards and unknown preferences, existing methods infer preferences only from utility feedback, entangling preference learning with reward exploration. In practice, however, users often reveal their priorities through proactive conversational queries (e.g., “cheap and clean hotel”), yet this structured signal is not leveraged. We formalize a proactive query-based framework in which user queries provide structured preference signals. Modeling these signals via a Plackett–Luce subset choice model, we show that query-only learning is insufficient due to a fundamental shift-invariance barrier. To resolve this, we introduce MO-PQUCB, a hybrid algorithm that integrates query-based preference anchoring with bandit feedback through shift-invariant regularization and dual-exploration {UCB}. We prove that proactive queries accelerate preference estimation and yield improved regret scaling over prior preference-aware MO-MAB methods. Under corrupted queries, we further characterize statistical limits and design a robust estimator achieving near-optimal performance when corruption is sparse. Experiments validate both theoretical and practical gains.} }
Endnote
%0 Conference Paper %T Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries %A Linfeng Cao %A Ming Shi %A Ness Shroff %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-cao26a %I PMLR %P 900--942 %U https://proceedings.mlr.press/v337/cao26a.html %V 337 %X Personalized decision-making in multi-objective bandits requires learning user-specific trade-offs among competing objectives. Since arm utility depends on both unknown rewards and unknown preferences, existing methods infer preferences only from utility feedback, entangling preference learning with reward exploration. In practice, however, users often reveal their priorities through proactive conversational queries (e.g., “cheap and clean hotel”), yet this structured signal is not leveraged. We formalize a proactive query-based framework in which user queries provide structured preference signals. Modeling these signals via a Plackett–Luce subset choice model, we show that query-only learning is insufficient due to a fundamental shift-invariance barrier. To resolve this, we introduce MO-PQUCB, a hybrid algorithm that integrates query-based preference anchoring with bandit feedback through shift-invariant regularization and dual-exploration {UCB}. We prove that proactive queries accelerate preference estimation and yield improved regret scaling over prior preference-aware MO-MAB methods. Under corrupted queries, we further characterize statistical limits and design a robust estimator achieving near-optimal performance when corruption is sparse. Experiments validate both theoretical and practical gains.
APA
Cao, L., Shi, M. & Shroff, N.. (2026). Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:900-942 Available from https://proceedings.mlr.press/v337/cao26a.html.

Related Material