Position: The Most Expensive Part of an LLM *should* be its Training Data

Nikhil Kandpal; Colin Raffel

Position: The Most Expensive Part of an LLM should be its Training Data

Nikhil Kandpal, Colin Raffel

Proceedings of the 42nd International Conference on Machine Learning, PMLR 267:81606-81616, 2025.

Abstract

Training a state-of-the-art Large Language Model (LLM) is an increasingly expensive endeavor due to growing computational, hardware, energy, and engineering demands. Yet, an often-overlooked (and seldom paid) expense is the human labor behind these models’ training data. Every LLM is built on an unfathomable amount of human effort: trillions of carefully written words sourced from books, academic papers, codebases, social media, and more. This position paper aims to assign a monetary value to this labor and argues that the most expensive part of producing an LLM should be the compensation provided to training data producers for their work. To support this position, we study 64 LLMs released between 2016 and 2024, estimating what it would cost to pay people to produce their training datasets from scratch. Even under highly conservative estimates of wage rates, the costs of these models’ training datasets are $10$-$1000$ times larger than the costs to train the models themselves, representing a significant financial liability for LLM providers. In the face of the massive gap between the value of training data and the lack of compensation for its creation, we highlight and discuss research directions that could enable fairer practices in the future.

Cite this Paper

BibTeX

@InProceedings{pmlr-v267-kandpal25a,
  title = 	 {Position: The Most Expensive Part of an {LLM} *should* be its Training Data},
  author =       {Kandpal, Nikhil and Raffel, Colin},
  booktitle = 	 {Proceedings of the 42nd International Conference on Machine Learning},
  pages = 	 {81606--81616},
  year = 	 {2025},
  editor = 	 {Singh, Aarti and Fazel, Maryam and Hsu, Daniel and Lacoste-Julien, Simon and Berkenkamp, Felix and Maharaj, Tegan and Wagstaff, Kiri and Zhu, Jerry},
  volume = 	 {267},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {13--19 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://raw.githubusercontent.com/mlresearch/v267/main/assets/kandpal25a/kandpal25a.pdf},
  url = 	 {https://proceedings.mlr.press/v267/kandpal25a.html},
  abstract = 	 {Training a state-of-the-art Large Language Model (LLM) is an increasingly expensive endeavor due to growing computational, hardware, energy, and engineering demands. Yet, an often-overlooked (and seldom paid) expense is the human labor behind these models’ training data. Every LLM is built on an unfathomable amount of human effort: trillions of carefully written words sourced from books, academic papers, codebases, social media, and more. This position paper aims to assign a monetary value to this labor and argues that the most expensive part of producing an LLM should be the compensation provided to training data producers for their work. To support this position, we study 64 LLMs released between 2016 and 2024, estimating what it would cost to pay people to produce their training datasets from scratch. Even under highly conservative estimates of wage rates, the costs of these models’ training datasets are $10$-$1000$ times larger than the costs to train the models themselves, representing a significant financial liability for LLM providers. In the face of the massive gap between the value of training data and the lack of compensation for its creation, we highlight and discuss research directions that could enable fairer practices in the future.}
}

Endnote

%0 Conference Paper
%T Position: The Most Expensive Part of an LLM *should* be its Training Data
%A Nikhil Kandpal
%A Colin Raffel
%B Proceedings of the 42nd International Conference on Machine Learning
%C Proceedings of Machine Learning Research
%D 2025
%E Aarti Singh
%E Maryam Fazel
%E Daniel Hsu
%E Simon Lacoste-Julien
%E Felix Berkenkamp
%E Tegan Maharaj
%E Kiri Wagstaff
%E Jerry Zhu	
%F pmlr-v267-kandpal25a
%I PMLR
%P 81606--81616
%U https://proceedings.mlr.press/v267/kandpal25a.html
%V 267
%X Training a state-of-the-art Large Language Model (LLM) is an increasingly expensive endeavor due to growing computational, hardware, energy, and engineering demands. Yet, an often-overlooked (and seldom paid) expense is the human labor behind these models’ training data. Every LLM is built on an unfathomable amount of human effort: trillions of carefully written words sourced from books, academic papers, codebases, social media, and more. This position paper aims to assign a monetary value to this labor and argues that the most expensive part of producing an LLM should be the compensation provided to training data producers for their work. To support this position, we study 64 LLMs released between 2016 and 2024, estimating what it would cost to pay people to produce their training datasets from scratch. Even under highly conservative estimates of wage rates, the costs of these models’ training datasets are $10$-$1000$ times larger than the costs to train the models themselves, representing a significant financial liability for LLM providers. In the face of the massive gap between the value of training data and the lack of compensation for its creation, we highlight and discuss research directions that could enable fairer practices in the future.

APA

Kandpal, N. & Raffel, C.. (2025). Position: The Most Expensive Part of an LLM *should* be its Training Data. Proceedings of the 42nd International Conference on Machine Learning, in Proceedings of Machine Learning Research 267:81606-81616 Available from https://proceedings.mlr.press/v267/kandpal25a.html.

Position: The Most Expensive Part of an LLM *should* be its Training Data

Abstract

Cite this Paper

Related Material

Position: The Most Expensive Part of an LLM should be its Training Data