Learning to Choose or Choosing to Learn: Best-of-N vs. Supervised Fine-Tuning for Bit String Generation

Seamus Somerstep, Vinod Raman, Unique Subedi, Yuekai Sun
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:1900-1908, 2026.

Abstract

Using the bit string generation problem as a case study, we theoretically compare two standard methods for adapting large language models to new tasks. The first, referred to as \emph{supervised fine-tuning}, involves training a new next token predictor on good generations. The second method, \emph{Best-of-N}, trains a reward model to select good responses from a collection generated by an unaltered base model. If the learning setting is realizable, we find that supervised fine-tuning outperforms BoN through a better dependence on the response length in its rate of convergence. If realizability fails, then depending on the failure mode, BoN can enjoy a better rate of convergence in either $n$ or a rate of convergence with better dependence on the response length.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-somerstep26a, title = { Learning to Choose or Choosing to Learn: Best-of-N vs. Supervised Fine-Tuning for Bit String Generation }, author = {Somerstep, Seamus and Raman, Vinod and Subedi, Unique and Sun, Yuekai}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {1900--1908}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/somerstep26a/somerstep26a.pdf}, url = {https://proceedings.mlr.press/v300/somerstep26a.html}, abstract = { Using the bit string generation problem as a case study, we theoretically compare two standard methods for adapting large language models to new tasks. The first, referred to as \emph{supervised fine-tuning}, involves training a new next token predictor on good generations. The second method, \emph{Best-of-N}, trains a reward model to select good responses from a collection generated by an unaltered base model. If the learning setting is realizable, we find that supervised fine-tuning outperforms BoN through a better dependence on the response length in its rate of convergence. If realizability fails, then depending on the failure mode, BoN can enjoy a better rate of convergence in either $n$ or a rate of convergence with better dependence on the response length. } }
Endnote
%0 Conference Paper %T Learning to Choose or Choosing to Learn: Best-of-N vs. Supervised Fine-Tuning for Bit String Generation %A Seamus Somerstep %A Vinod Raman %A Unique Subedi %A Yuekai Sun %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-somerstep26a %I PMLR %P 1900--1908 %U https://proceedings.mlr.press/v300/somerstep26a.html %V 300 %X Using the bit string generation problem as a case study, we theoretically compare two standard methods for adapting large language models to new tasks. The first, referred to as \emph{supervised fine-tuning}, involves training a new next token predictor on good generations. The second method, \emph{Best-of-N}, trains a reward model to select good responses from a collection generated by an unaltered base model. If the learning setting is realizable, we find that supervised fine-tuning outperforms BoN through a better dependence on the response length in its rate of convergence. If realizability fails, then depending on the failure mode, BoN can enjoy a better rate of convergence in either $n$ or a rate of convergence with better dependence on the response length.
APA
Somerstep, S., Raman, V., Subedi, U. & Sun, Y.. (2026). Learning to Choose or Choosing to Learn: Best-of-N vs. Supervised Fine-Tuning for Bit String Generation . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:1900-1908 Available from https://proceedings.mlr.press/v300/somerstep26a.html.

Related Material