When and How Often is Weighted Majority Vote Optimal Under Log Loss?

Steven An, Sanjoy Dasgupta
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:87-132, 2026.

Abstract

Modern machine learning methods often require large, high quality labeled datasets, whose labels are potentially expensive and time consuming to obtain. One solution, seen in Programmatic Weak Supervision, crowdsourcing, and semi-supervised learning, is to cheaply obtain an ensemble of noisy labeling functions (LFs) and combine their predictions. Specifically, weighted majority vote ({WMV}) is a simple but well studied method to perform such a combination. Weighting strategies for {WMV} can be derived from a wide ranging set of assumptions (e.g. probabilistic, adversarial, etc.). However, existing analyses often suppose that the LF predictions are fixed, and characterize the conditions when said weighting strategies are optimal (among all weighting strategies). We take a different approach and show that all weighting strategies which only depend on LF accuracies, e.g. majority vote, are optimal (w.r.t. log loss) on a measure zero set of problems. A method to compute the proportion of problems where such aforementioned strategies are $\epsilon$ close to being optimal is presented and run. Other contributions include improved analysis of {WMV}’s excess error under log loss.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-an26a, title = {When and How Often is Weighted Majority Vote Optimal Under Log Loss?}, author = {An, Steven and Dasgupta, Sanjoy}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {87--132}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/an26a/an26a.pdf}, url = {https://proceedings.mlr.press/v337/an26a.html}, abstract = {Modern machine learning methods often require large, high quality labeled datasets, whose labels are potentially expensive and time consuming to obtain. One solution, seen in Programmatic Weak Supervision, crowdsourcing, and semi-supervised learning, is to cheaply obtain an ensemble of noisy labeling functions (LFs) and combine their predictions. Specifically, weighted majority vote ({WMV}) is a simple but well studied method to perform such a combination. Weighting strategies for {WMV} can be derived from a wide ranging set of assumptions (e.g. probabilistic, adversarial, etc.). However, existing analyses often suppose that the LF predictions are fixed, and characterize the conditions when said weighting strategies are optimal (among all weighting strategies). We take a different approach and show that all weighting strategies which only depend on LF accuracies, e.g. majority vote, are optimal (w.r.t. log loss) on a measure zero set of problems. A method to compute the proportion of problems where such aforementioned strategies are $\epsilon$ close to being optimal is presented and run. Other contributions include improved analysis of {WMV}’s excess error under log loss.} }
Endnote
%0 Conference Paper %T When and How Often is Weighted Majority Vote Optimal Under Log Loss? %A Steven An %A Sanjoy Dasgupta %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-an26a %I PMLR %P 87--132 %U https://proceedings.mlr.press/v337/an26a.html %V 337 %X Modern machine learning methods often require large, high quality labeled datasets, whose labels are potentially expensive and time consuming to obtain. One solution, seen in Programmatic Weak Supervision, crowdsourcing, and semi-supervised learning, is to cheaply obtain an ensemble of noisy labeling functions (LFs) and combine their predictions. Specifically, weighted majority vote ({WMV}) is a simple but well studied method to perform such a combination. Weighting strategies for {WMV} can be derived from a wide ranging set of assumptions (e.g. probabilistic, adversarial, etc.). However, existing analyses often suppose that the LF predictions are fixed, and characterize the conditions when said weighting strategies are optimal (among all weighting strategies). We take a different approach and show that all weighting strategies which only depend on LF accuracies, e.g. majority vote, are optimal (w.r.t. log loss) on a measure zero set of problems. A method to compute the proportion of problems where such aforementioned strategies are $\epsilon$ close to being optimal is presented and run. Other contributions include improved analysis of {WMV}’s excess error under log loss.
APA
An, S. & Dasgupta, S.. (2026). When and How Often is Weighted Majority Vote Optimal Under Log Loss?. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:87-132 Available from https://proceedings.mlr.press/v337/an26a.html.

Related Material