When Individually Calibrated Models Become Collectively Miscalibrated

Zhaohui Geoffrey Wang
Proceedings of The 1st Symposium on Probabilistic Machine Learning, PMLR 327:214-255, 2026.

Abstract

Probabilistic prediction systems often aggregate probability estimates from multiple models into a single decision. A natural assumption is that if each model is individually calibrated, the aggregate prediction will also be well calibrated. We show that this assumption fails in multi-agent settings: individually calibrated predictors can become collectively miscalibrated when their predictions interact strategically. This phenomenon arises in settings where predictions originate from multiple independent agents, including federated healthcare, multi-vendor intrusion detection, and crowdsourced forecasting, where agents optimize their own objectives. Specifically, we prove that under Brier-score-based aggregation with correlated beliefs, each agent’s individually optimal report systematically underestimates the positive-class probability, producing a Price of Anarchy (PoA) of 7.25x (mean aggregate bias -0.375). In contrast, VCG-based aggregation, which rewards each agent’s marginal contribution to aggregate accuracy, achieves the lowest PoA among all mechanisms studied (PoA $\approx$ 1.0x under dominant-strategy equilibrium). On three real-world datasets (NSL-KDD, UNSW-NB15, Credit Card Fraud) with feature-partitioned agents, VCG provides the strongest robustness guarantees among the aggregation methods we evaluate, while maintaining comparable accuracy. In data-sparse regimes ($n \leq 500$), VCG consistently outperforms stacking and majority voting; under adversarial agents, VCG maintains substantially lower false-negative rates than robust aggregation baselines. Adaptive weight updates further reduce false negatives by 20–22% under distribution shift, with $O(\sqrt{T})$ online regret guarantees. These results establish that how probabilistic predictions are aggregated matters as much as how well individual models are calibrated.

Cite this Paper


BibTeX
@InProceedings{pmlr-v327-wang26a, title = {When Individually Calibrated Models Become Collectively Miscalibrated }, author = {Wang, Zhaohui Geoffrey}, booktitle = {Proceedings of The 1st Symposium on Probabilistic Machine Learning}, pages = {214--255}, year = {2026}, editor = {Swaroop, Siddharth and Rügamer, David and Kristiadi, Agustinus}, volume = {327}, series = {Proceedings of Machine Learning Research}, month = {05 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v327/main/assets/wang26a/wang26a.pdf}, url = {https://proceedings.mlr.press/v327/wang26a.html}, abstract = { Probabilistic prediction systems often aggregate probability estimates from multiple models into a single decision. A natural assumption is that if each model is individually calibrated, the aggregate prediction will also be well calibrated. We show that this assumption fails in multi-agent settings: individually calibrated predictors can become collectively miscalibrated when their predictions interact strategically. This phenomenon arises in settings where predictions originate from multiple independent agents, including federated healthcare, multi-vendor intrusion detection, and crowdsourced forecasting, where agents optimize their own objectives. Specifically, we prove that under Brier-score-based aggregation with correlated beliefs, each agent’s individually optimal report systematically underestimates the positive-class probability, producing a Price of Anarchy (PoA) of 7.25x (mean aggregate bias -0.375). In contrast, VCG-based aggregation, which rewards each agent’s marginal contribution to aggregate accuracy, achieves the lowest PoA among all mechanisms studied (PoA $\approx$ 1.0x under dominant-strategy equilibrium). On three real-world datasets (NSL-KDD, UNSW-NB15, Credit Card Fraud) with feature-partitioned agents, VCG provides the strongest robustness guarantees among the aggregation methods we evaluate, while maintaining comparable accuracy. In data-sparse regimes ($n \leq 500$), VCG consistently outperforms stacking and majority voting; under adversarial agents, VCG maintains substantially lower false-negative rates than robust aggregation baselines. Adaptive weight updates further reduce false negatives by 20–22% under distribution shift, with $O(\sqrt{T})$ online regret guarantees. These results establish that how probabilistic predictions are aggregated matters as much as how well individual models are calibrated. } }
Endnote
%0 Conference Paper %T When Individually Calibrated Models Become Collectively Miscalibrated %A Zhaohui Geoffrey Wang %B Proceedings of The 1st Symposium on Probabilistic Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Siddharth Swaroop %E David Rügamer %E Agustinus Kristiadi %F pmlr-v327-wang26a %I PMLR %P 214--255 %U https://proceedings.mlr.press/v327/wang26a.html %V 327 %X Probabilistic prediction systems often aggregate probability estimates from multiple models into a single decision. A natural assumption is that if each model is individually calibrated, the aggregate prediction will also be well calibrated. We show that this assumption fails in multi-agent settings: individually calibrated predictors can become collectively miscalibrated when their predictions interact strategically. This phenomenon arises in settings where predictions originate from multiple independent agents, including federated healthcare, multi-vendor intrusion detection, and crowdsourced forecasting, where agents optimize their own objectives. Specifically, we prove that under Brier-score-based aggregation with correlated beliefs, each agent’s individually optimal report systematically underestimates the positive-class probability, producing a Price of Anarchy (PoA) of 7.25x (mean aggregate bias -0.375). In contrast, VCG-based aggregation, which rewards each agent’s marginal contribution to aggregate accuracy, achieves the lowest PoA among all mechanisms studied (PoA $\approx$ 1.0x under dominant-strategy equilibrium). On three real-world datasets (NSL-KDD, UNSW-NB15, Credit Card Fraud) with feature-partitioned agents, VCG provides the strongest robustness guarantees among the aggregation methods we evaluate, while maintaining comparable accuracy. In data-sparse regimes ($n \leq 500$), VCG consistently outperforms stacking and majority voting; under adversarial agents, VCG maintains substantially lower false-negative rates than robust aggregation baselines. Adaptive weight updates further reduce false negatives by 20–22% under distribution shift, with $O(\sqrt{T})$ online regret guarantees. These results establish that how probabilistic predictions are aggregated matters as much as how well individual models are calibrated.
APA
Wang, Z.G.. (2026). When Individually Calibrated Models Become Collectively Miscalibrated . Proceedings of The 1st Symposium on Probabilistic Machine Learning, in Proceedings of Machine Learning Research 327:214-255 Available from https://proceedings.mlr.press/v327/wang26a.html.

Related Material