On the Effect of Sampling Diversity in Scaling LLM Inference

Tianchun Wang, Yuanzhou Chen, Zichuan Liu, Jonathan Light, Weiyang Liu, Haifeng Chen, Xiang Zhang, Wei Cheng
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7137-7167, 2026.

Abstract

Large language model ({LLM}) scaling inference is key to unlocking greater performance, and leveraging diversity has proven an effective way to enhance it. Motivated by the observed relationship between solution accuracy and meaningful response diversity, we systematically study the effect of prompt diversity in scaling inference. We theoretically explain why diversified sampling improves Best-of-N scaling, showing that responses generated from diverse prompts after Best-of-N selection exhibit significantly lower error rates than those produced from stationary prompts. Building on this analysis, we derive a diversity-fidelity trade-off principle, that guides the design of sampling strategies introducing diversity. From this guidance, we instantiate a family of effective perturbation styles. We theoretically and empirically characterize when diversified exploration remains effective, demonstrating that it works under a variety of conditions, and we further show that under majority voting, diversity may vanish. We systematically evaluate the effectiveness of sampling diversity and show that, when applied appropriately in different contexts, meaningful perturbations yield stronger, task-dependent gains as diversity increases. Overall, this work provides a systematic analysis that offers a theoretical and empirical foundation for understanding the effect of diversity in {LLM} inference-time scaling.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-wang26f, title = {On the Effect of Sampling Diversity in Scaling {LLM} Inference}, author = {Wang, Tianchun and Chen, Yuanzhou and Liu, Zichuan and Light, Jonathan and Liu, Weiyang and Chen, Haifeng and Zhang, Xiang and Cheng, Wei}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {7137--7167}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/wang26f/wang26f.pdf}, url = {https://proceedings.mlr.press/v337/wang26f.html}, abstract = {Large language model ({LLM}) scaling inference is key to unlocking greater performance, and leveraging diversity has proven an effective way to enhance it. Motivated by the observed relationship between solution accuracy and meaningful response diversity, we systematically study the effect of prompt diversity in scaling inference. We theoretically explain why diversified sampling improves Best-of-N scaling, showing that responses generated from diverse prompts after Best-of-N selection exhibit significantly lower error rates than those produced from stationary prompts. Building on this analysis, we derive a diversity-fidelity trade-off principle, that guides the design of sampling strategies introducing diversity. From this guidance, we instantiate a family of effective perturbation styles. We theoretically and empirically characterize when diversified exploration remains effective, demonstrating that it works under a variety of conditions, and we further show that under majority voting, diversity may vanish. We systematically evaluate the effectiveness of sampling diversity and show that, when applied appropriately in different contexts, meaningful perturbations yield stronger, task-dependent gains as diversity increases. Overall, this work provides a systematic analysis that offers a theoretical and empirical foundation for understanding the effect of diversity in {LLM} inference-time scaling.} }
Endnote
%0 Conference Paper %T On the Effect of Sampling Diversity in Scaling LLM Inference %A Tianchun Wang %A Yuanzhou Chen %A Zichuan Liu %A Jonathan Light %A Weiyang Liu %A Haifeng Chen %A Xiang Zhang %A Wei Cheng %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-wang26f %I PMLR %P 7137--7167 %U https://proceedings.mlr.press/v337/wang26f.html %V 337 %X Large language model ({LLM}) scaling inference is key to unlocking greater performance, and leveraging diversity has proven an effective way to enhance it. Motivated by the observed relationship between solution accuracy and meaningful response diversity, we systematically study the effect of prompt diversity in scaling inference. We theoretically explain why diversified sampling improves Best-of-N scaling, showing that responses generated from diverse prompts after Best-of-N selection exhibit significantly lower error rates than those produced from stationary prompts. Building on this analysis, we derive a diversity-fidelity trade-off principle, that guides the design of sampling strategies introducing diversity. From this guidance, we instantiate a family of effective perturbation styles. We theoretically and empirically characterize when diversified exploration remains effective, demonstrating that it works under a variety of conditions, and we further show that under majority voting, diversity may vanish. We systematically evaluate the effectiveness of sampling diversity and show that, when applied appropriately in different contexts, meaningful perturbations yield stronger, task-dependent gains as diversity increases. Overall, this work provides a systematic analysis that offers a theoretical and empirical foundation for understanding the effect of diversity in {LLM} inference-time scaling.
APA
Wang, T., Chen, Y., Liu, Z., Light, J., Liu, W., Chen, H., Zhang, X. & Cheng, W.. (2026). On the Effect of Sampling Diversity in Scaling LLM Inference. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:7137-7167 Available from https://proceedings.mlr.press/v337/wang26f.html.

Related Material