Bayesian Adaptation Gym: A Benchmark for the Bayesian Low-Rank Adaptation of Multi-Modal Language Models

Colin Samplawski, Ramneet Kaur, Manoj Acharya, Anirban Roy, Adam D. Cobb
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:5919-5966, 2026.

Abstract

Large multi-modal language models are increasingly deployed in high-stakes domains, making well-calibrated uncertainty essential. Traditional {Bayesian} methods approximate posteriors over all model weights, which becomes intractable for modern large models. For this reason, recent work instead considers {Bayesian} low-rank adaptation to enable tractable posterior approximation. Due to a lack of a standardized benchmark to evaluate these approaches, it remains unclear where these methods provide meaningful benefits. To fill this gap, we introduce {Bayesian} Adaptation Gym (BAG), a benchmark for the {Bayesian} adaptation of multi-modal language models. BAG provides reference implementations of classic {Bayesian} baselines and state-of-the-art adaptation methods, along with a multi-modal dataset and task suite designed to probe calibration, robustness under distribution shift, and decision-making under uncertainty via active learning. Using BAG, we conduct and report extensive experiments across model sizes, datasets, and tasks to highlight the successes and failures of current {Bayesian} adaptation approaches. To enable further research, BAG is fully open source: https://github.com/SRI-CSL/BayesAdapt.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-samplawski26a, title = {{Bayesian} Adaptation Gym: A Benchmark for the {Bayesian} Low-Rank Adaptation of Multi-Modal Language Models}, author = {Samplawski, Colin and Kaur, Ramneet and Acharya, Manoj and Roy, Anirban and Cobb, {Adam} D.}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {5919--5966}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/samplawski26a/samplawski26a.pdf}, url = {https://proceedings.mlr.press/v337/samplawski26a.html}, abstract = {Large multi-modal language models are increasingly deployed in high-stakes domains, making well-calibrated uncertainty essential. Traditional {Bayesian} methods approximate posteriors over all model weights, which becomes intractable for modern large models. For this reason, recent work instead considers {Bayesian} low-rank adaptation to enable tractable posterior approximation. Due to a lack of a standardized benchmark to evaluate these approaches, it remains unclear where these methods provide meaningful benefits. To fill this gap, we introduce {Bayesian} Adaptation Gym (BAG), a benchmark for the {Bayesian} adaptation of multi-modal language models. BAG provides reference implementations of classic {Bayesian} baselines and state-of-the-art adaptation methods, along with a multi-modal dataset and task suite designed to probe calibration, robustness under distribution shift, and decision-making under uncertainty via active learning. Using BAG, we conduct and report extensive experiments across model sizes, datasets, and tasks to highlight the successes and failures of current {Bayesian} adaptation approaches. To enable further research, BAG is fully open source: https://github.com/SRI-CSL/BayesAdapt.} }
Endnote
%0 Conference Paper %T Bayesian Adaptation Gym: A Benchmark for the Bayesian Low-Rank Adaptation of Multi-Modal Language Models %A Colin Samplawski %A Ramneet Kaur %A Manoj Acharya %A Anirban Roy %A Adam D. Cobb %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-samplawski26a %I PMLR %P 5919--5966 %U https://proceedings.mlr.press/v337/samplawski26a.html %V 337 %X Large multi-modal language models are increasingly deployed in high-stakes domains, making well-calibrated uncertainty essential. Traditional {Bayesian} methods approximate posteriors over all model weights, which becomes intractable for modern large models. For this reason, recent work instead considers {Bayesian} low-rank adaptation to enable tractable posterior approximation. Due to a lack of a standardized benchmark to evaluate these approaches, it remains unclear where these methods provide meaningful benefits. To fill this gap, we introduce {Bayesian} Adaptation Gym (BAG), a benchmark for the {Bayesian} adaptation of multi-modal language models. BAG provides reference implementations of classic {Bayesian} baselines and state-of-the-art adaptation methods, along with a multi-modal dataset and task suite designed to probe calibration, robustness under distribution shift, and decision-making under uncertainty via active learning. Using BAG, we conduct and report extensive experiments across model sizes, datasets, and tasks to highlight the successes and failures of current {Bayesian} adaptation approaches. To enable further research, BAG is fully open source: https://github.com/SRI-CSL/BayesAdapt.
APA
Samplawski, C., Kaur, R., Acharya, M., Roy, A. & Cobb, A.D.. (2026). Bayesian Adaptation Gym: A Benchmark for the Bayesian Low-Rank Adaptation of Multi-Modal Language Models. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:5919-5966 Available from https://proceedings.mlr.press/v337/samplawski26a.html.

Related Material