Expert Advice with Costly Observations

Lev Reyzin, Aadirupa Saha, Shuo Wu
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:5700-5715, 2026.

Abstract

Querying expert predictions or human feedback incurs explicit, often heterogeneous costs—a challenge central to crowdsourcing and {LLM} training. We study online learning with expert advice under observation costs, where the learner must adaptively select which experts to query ($p_i$ is expert-dependent) to minimize cumulative cost (losses plus query costs). While classical bandits and recent paid observation models have advanced the field, they fail to capture the rich cost-adaptive selection landscape and do not provide tight, problem-dependent performance guarantees for heterogeneous costs. This paper provides the first comprehensive characterization. For adversarial losses, we prove tight minimax regret $\Theta(\underset{m\in\{1,…,K\}}{\min} \{\sqrt{\frac{K}{m}\,T\log K} + (m-1)\bar pT\}),$ interpolating between bandit and full-information regimes and revealing the optimal observation budget, where $\bar p$ is a tight upper bound of $p_i$s. For stochastic losses, our novel $m$-ELIM algorithm achieves instance-dependent regret of $O\left(\frac{\log T}{m}\sum_{i \neq i^*} \frac{1}{\Delta_i} + (m-1)\bar{p}T\right)$, $\Delta_i$ being the suboptimality gap of the $i$-th arm, showing how gap structure and costs interact. Matching lower bounds (adversarial and stochastic) establishes minimax optimality. These results provide the first rigorous framework for optimally allocating annotation budgets across heterogeneous workers in crowdsourcing systems, directly informing cost-effective strategies for collecting human feedback in adaptive learning applications.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-reyzin26a, title = {Expert Advice with Costly Observations}, author = {Reyzin, Lev and Saha, Aadirupa and Wu, Shuo}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {5700--5715}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/reyzin26a/reyzin26a.pdf}, url = {https://proceedings.mlr.press/v337/reyzin26a.html}, abstract = {Querying expert predictions or human feedback incurs explicit, often heterogeneous costs—a challenge central to crowdsourcing and {LLM} training. We study online learning with expert advice under observation costs, where the learner must adaptively select which experts to query ($p_i$ is expert-dependent) to minimize cumulative cost (losses plus query costs). While classical bandits and recent paid observation models have advanced the field, they fail to capture the rich cost-adaptive selection landscape and do not provide tight, problem-dependent performance guarantees for heterogeneous costs. This paper provides the first comprehensive characterization. For adversarial losses, we prove tight minimax regret $\Theta(\underset{m\in\{1,…,K\}}{\min} \{\sqrt{\frac{K}{m}\,T\log K} + (m-1)\bar pT\}),$ interpolating between bandit and full-information regimes and revealing the optimal observation budget, where $\bar p$ is a tight upper bound of $p_i$s. For stochastic losses, our novel $m$-ELIM algorithm achieves instance-dependent regret of $O\left(\frac{\log T}{m}\sum_{i \neq i^*} \frac{1}{\Delta_i} + (m-1)\bar{p}T\right)$, $\Delta_i$ being the suboptimality gap of the $i$-th arm, showing how gap structure and costs interact. Matching lower bounds (adversarial and stochastic) establishes minimax optimality. These results provide the first rigorous framework for optimally allocating annotation budgets across heterogeneous workers in crowdsourcing systems, directly informing cost-effective strategies for collecting human feedback in adaptive learning applications.} }
Endnote
%0 Conference Paper %T Expert Advice with Costly Observations %A Lev Reyzin %A Aadirupa Saha %A Shuo Wu %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-reyzin26a %I PMLR %P 5700--5715 %U https://proceedings.mlr.press/v337/reyzin26a.html %V 337 %X Querying expert predictions or human feedback incurs explicit, often heterogeneous costs—a challenge central to crowdsourcing and {LLM} training. We study online learning with expert advice under observation costs, where the learner must adaptively select which experts to query ($p_i$ is expert-dependent) to minimize cumulative cost (losses plus query costs). While classical bandits and recent paid observation models have advanced the field, they fail to capture the rich cost-adaptive selection landscape and do not provide tight, problem-dependent performance guarantees for heterogeneous costs. This paper provides the first comprehensive characterization. For adversarial losses, we prove tight minimax regret $\Theta(\underset{m\in\{1,…,K\}}{\min} \{\sqrt{\frac{K}{m}\,T\log K} + (m-1)\bar pT\}),$ interpolating between bandit and full-information regimes and revealing the optimal observation budget, where $\bar p$ is a tight upper bound of $p_i$s. For stochastic losses, our novel $m$-ELIM algorithm achieves instance-dependent regret of $O\left(\frac{\log T}{m}\sum_{i \neq i^*} \frac{1}{\Delta_i} + (m-1)\bar{p}T\right)$, $\Delta_i$ being the suboptimality gap of the $i$-th arm, showing how gap structure and costs interact. Matching lower bounds (adversarial and stochastic) establishes minimax optimality. These results provide the first rigorous framework for optimally allocating annotation budgets across heterogeneous workers in crowdsourcing systems, directly informing cost-effective strategies for collecting human feedback in adaptive learning applications.
APA
Reyzin, L., Saha, A. & Wu, S.. (2026). Expert Advice with Costly Observations. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:5700-5715 Available from https://proceedings.mlr.press/v337/reyzin26a.html.

Related Material