A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback

Joseph Lazzaro, Davide Buffelli, Da-shan Shiu, Sattar Vakili
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:991-999, 2026.

Abstract

Preference feedback, in the form of pairwise comparisons rather than scalar scores, has seen increasing use in applications such as human-, laboratory-, and expert-in-the-loop design, as well as scientific discovery. We propose a Thompson Sampling (TS) approach to Bayesian optimization with preferential feedback that models comparisons using a monotone link on latent utility differences and leverages the dueling kernel induced by a base kernel. We provide a finite-time analysis showing that the performance of the proposed method matches that of standard TS for conventional Bayesian optimization with scalar feedback. The analysis exploits the anchor invariance of TS for challenger selection and introduces a double-TS pairing variant. We also demonstrate the performance of the method on both synthetic and real-world examples.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-lazzaro26a, title = { A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback }, author = {Lazzaro, Joseph and Buffelli, Davide and Shiu, Da-shan and Vakili, Sattar}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {991--999}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/lazzaro26a/lazzaro26a.pdf}, url = {https://proceedings.mlr.press/v300/lazzaro26a.html}, abstract = { Preference feedback, in the form of pairwise comparisons rather than scalar scores, has seen increasing use in applications such as human-, laboratory-, and expert-in-the-loop design, as well as scientific discovery. We propose a Thompson Sampling (TS) approach to Bayesian optimization with preferential feedback that models comparisons using a monotone link on latent utility differences and leverages the dueling kernel induced by a base kernel. We provide a finite-time analysis showing that the performance of the proposed method matches that of standard TS for conventional Bayesian optimization with scalar feedback. The analysis exploits the anchor invariance of TS for challenger selection and introduces a double-TS pairing variant. We also demonstrate the performance of the method on both synthetic and real-world examples. } }
Endnote
%0 Conference Paper %T A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback %A Joseph Lazzaro %A Davide Buffelli %A Da-shan Shiu %A Sattar Vakili %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-lazzaro26a %I PMLR %P 991--999 %U https://proceedings.mlr.press/v300/lazzaro26a.html %V 300 %X Preference feedback, in the form of pairwise comparisons rather than scalar scores, has seen increasing use in applications such as human-, laboratory-, and expert-in-the-loop design, as well as scientific discovery. We propose a Thompson Sampling (TS) approach to Bayesian optimization with preferential feedback that models comparisons using a monotone link on latent utility differences and leverages the dueling kernel induced by a base kernel. We provide a finite-time analysis showing that the performance of the proposed method matches that of standard TS for conventional Bayesian optimization with scalar feedback. The analysis exploits the anchor invariance of TS for challenger selection and introduces a double-TS pairing variant. We also demonstrate the performance of the method on both synthetic and real-world examples.
APA
Lazzaro, J., Buffelli, D., Shiu, D. & Vakili, S.. (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:991-999 Available from https://proceedings.mlr.press/v300/lazzaro26a.html.

Related Material