NanoFLUX: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile Devices

Ruchika Chavhan, Malcolm Chadwick, Alberto Gil Couto Pimentel Ramos, Luca Morreale, Mehdi Noroozi, Abhinav Mehrotra
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:13278-13299, 2026.

Abstract

While large-scale text-to-image diffusion models continue to improve in visual quality, their increasing scale has widened the gap between state-of-the-art models and on-device solutions. To address this gap, we introduce NanoFLUX, a 2.4B text-to-image flow-matching model distilled from 17B FLUX.1-Schnell using a progressive compression pipeline designed to preserve generation quality. Our contributions include: (1) A model compression strategy driven by pruning redundant components in the diffusion transformer, reducing its size from 12B to 2B; (2) A ResNet-based token downsampling mechanism that reduces latency by allowing intermediate blocks to operate on lower-resolution tokens while preserving high-resolution processing elsewhere; (3) A novel text encoder distillation approach that leverages visual signals from early layers of the denoiser during sampling. Empirically, NanoFLUX generates $512 \times 512$ images in approximately 2.5 seconds on mobile devices, demonstrating the feasibility of high-quality on-device text-to-image generation.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chavhan26a, title = {{N}ano{FLUX}: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile Devices}, author = {Chavhan, Ruchika and Chadwick, Malcolm and Couto Pimentel Ramos, Alberto Gil and Morreale, Luca and Noroozi, Mehdi and Mehrotra, Abhinav}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {13278--13299}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chavhan26a/chavhan26a.pdf}, url = {https://proceedings.mlr.press/v306/chavhan26a.html}, abstract = {While large-scale text-to-image diffusion models continue to improve in visual quality, their increasing scale has widened the gap between state-of-the-art models and on-device solutions. To address this gap, we introduce NanoFLUX, a 2.4B text-to-image flow-matching model distilled from 17B FLUX.1-Schnell using a progressive compression pipeline designed to preserve generation quality. Our contributions include: (1) A model compression strategy driven by pruning redundant components in the diffusion transformer, reducing its size from 12B to 2B; (2) A ResNet-based token downsampling mechanism that reduces latency by allowing intermediate blocks to operate on lower-resolution tokens while preserving high-resolution processing elsewhere; (3) A novel text encoder distillation approach that leverages visual signals from early layers of the denoiser during sampling. Empirically, NanoFLUX generates $512 \times 512$ images in approximately 2.5 seconds on mobile devices, demonstrating the feasibility of high-quality on-device text-to-image generation.} }
Endnote
%0 Conference Paper %T NanoFLUX: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile Devices %A Ruchika Chavhan %A Malcolm Chadwick %A Alberto Gil Couto Pimentel Ramos %A Luca Morreale %A Mehdi Noroozi %A Abhinav Mehrotra %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chavhan26a %I PMLR %P 13278--13299 %U https://proceedings.mlr.press/v306/chavhan26a.html %V 306 %X While large-scale text-to-image diffusion models continue to improve in visual quality, their increasing scale has widened the gap between state-of-the-art models and on-device solutions. To address this gap, we introduce NanoFLUX, a 2.4B text-to-image flow-matching model distilled from 17B FLUX.1-Schnell using a progressive compression pipeline designed to preserve generation quality. Our contributions include: (1) A model compression strategy driven by pruning redundant components in the diffusion transformer, reducing its size from 12B to 2B; (2) A ResNet-based token downsampling mechanism that reduces latency by allowing intermediate blocks to operate on lower-resolution tokens while preserving high-resolution processing elsewhere; (3) A novel text encoder distillation approach that leverages visual signals from early layers of the denoiser during sampling. Empirically, NanoFLUX generates $512 \times 512$ images in approximately 2.5 seconds on mobile devices, demonstrating the feasibility of high-quality on-device text-to-image generation.
APA
Chavhan, R., Chadwick, M., Couto Pimentel Ramos, A.G., Morreale, L., Noroozi, M. & Mehrotra, A.. (2026). NanoFLUX: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile Devices. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:13278-13299 Available from https://proceedings.mlr.press/v306/chavhan26a.html.

Related Material