scDataset: Scalable Data Loading for Deep Learning on Large-Scale Single-Cell Omics

Davide D’Ascenzo, Sebastiano Cultrera Di Montesano
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:22246-22262, 2026.

Abstract

Training deep learning models on single-cell datasets with hundreds of millions of cells requires loading data from disk, as these datasets exceed available memory. While random sampling provides the data diversity needed for effective training, it is prohibitively slow due to the random access pattern overhead, whereas sequential streaming achieves high throughput but introduces biases that degrade model performance. We present scDataset, a PyTorch data loader that enables efficient training from on-disk data with seamless integration across diverse storage formats. Our approach combines block sampling and batched fetching to achieve quasi-random sampling that balances I/O efficiency with minibatch diversity. On Tahoe-100M, a dataset of 100 million cells, scDataset achieves more than two orders of magnitude speedup compared to true random sampling while working directly with AnnData files. We provide theoretical bounds on minibatch diversity and empirically show that scDataset matches the performance of true random sampling across multiple classification tasks and model architectures.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-d-ascenzo26a, title = {sc{D}ataset: Scalable Data Loading for Deep Learning on Large-Scale Single-Cell Omics}, author = {D'Ascenzo, Davide and Montesano, Sebastiano Cultrera Di}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {22246--22262}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/d-ascenzo26a/d-ascenzo26a.pdf}, url = {https://proceedings.mlr.press/v306/d-ascenzo26a.html}, abstract = {Training deep learning models on single-cell datasets with hundreds of millions of cells requires loading data from disk, as these datasets exceed available memory. While random sampling provides the data diversity needed for effective training, it is prohibitively slow due to the random access pattern overhead, whereas sequential streaming achieves high throughput but introduces biases that degrade model performance. We present scDataset, a PyTorch data loader that enables efficient training from on-disk data with seamless integration across diverse storage formats. Our approach combines block sampling and batched fetching to achieve quasi-random sampling that balances I/O efficiency with minibatch diversity. On Tahoe-100M, a dataset of 100 million cells, scDataset achieves more than two orders of magnitude speedup compared to true random sampling while working directly with AnnData files. We provide theoretical bounds on minibatch diversity and empirically show that scDataset matches the performance of true random sampling across multiple classification tasks and model architectures.} }
Endnote
%0 Conference Paper %T scDataset: Scalable Data Loading for Deep Learning on Large-Scale Single-Cell Omics %A Davide D’Ascenzo %A Sebastiano Cultrera Di Montesano %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-d-ascenzo26a %I PMLR %P 22246--22262 %U https://proceedings.mlr.press/v306/d-ascenzo26a.html %V 306 %X Training deep learning models on single-cell datasets with hundreds of millions of cells requires loading data from disk, as these datasets exceed available memory. While random sampling provides the data diversity needed for effective training, it is prohibitively slow due to the random access pattern overhead, whereas sequential streaming achieves high throughput but introduces biases that degrade model performance. We present scDataset, a PyTorch data loader that enables efficient training from on-disk data with seamless integration across diverse storage formats. Our approach combines block sampling and batched fetching to achieve quasi-random sampling that balances I/O efficiency with minibatch diversity. On Tahoe-100M, a dataset of 100 million cells, scDataset achieves more than two orders of magnitude speedup compared to true random sampling while working directly with AnnData files. We provide theoretical bounds on minibatch diversity and empirically show that scDataset matches the performance of true random sampling across multiple classification tasks and model architectures.
APA
D’Ascenzo, D. & Montesano, S.C.D.. (2026). scDataset: Scalable Data Loading for Deep Learning on Large-Scale Single-Cell Omics. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:22246-22262 Available from https://proceedings.mlr.press/v306/d-ascenzo26a.html.

Related Material