SFBD Flow: A Continuous-Optimization Framework for Training Diffusion Models with Noisy Samples

Haoye Lu, Darren Lo, Yaoliang Yu
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:388-396, 2026.

Abstract

Diffusion models achieve strong generative performance but often rely on large datasets that may include sensitive content. This challenge is compounded by the models’ tendency to memorize training data, raising privacy concerns. SFBD (Lu et al., 2025) addresses this by training on corrupted data and using limited clean samples to capture local structure and improve convergence. However, its iterative denoising and fine-tuning loop requires manual coordination, making it burdensome to implement. We reinterpret SFBD as an alternating projection algorithm and introduce a continuous variant, SFBD flow, that removes the need for alternating steps. We further show its connection to consistency constraint-based methods, and demonstrate that its practical instantiation, Online SFBD, consistently outperforms strong baselines across benchmarks.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-lu26a, title = { SFBD Flow: A Continuous-Optimization Framework for Training Diffusion Models with Noisy Samples }, author = {Lu, Haoye and Lo, Darren and Yu, Yaoliang}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {388--396}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/lu26a/lu26a.pdf}, url = {https://proceedings.mlr.press/v300/lu26a.html}, abstract = { Diffusion models achieve strong generative performance but often rely on large datasets that may include sensitive content. This challenge is compounded by the models’ tendency to memorize training data, raising privacy concerns. SFBD (Lu et al., 2025) addresses this by training on corrupted data and using limited clean samples to capture local structure and improve convergence. However, its iterative denoising and fine-tuning loop requires manual coordination, making it burdensome to implement. We reinterpret SFBD as an alternating projection algorithm and introduce a continuous variant, SFBD flow, that removes the need for alternating steps. We further show its connection to consistency constraint-based methods, and demonstrate that its practical instantiation, Online SFBD, consistently outperforms strong baselines across benchmarks. } }
Endnote
%0 Conference Paper %T SFBD Flow: A Continuous-Optimization Framework for Training Diffusion Models with Noisy Samples %A Haoye Lu %A Darren Lo %A Yaoliang Yu %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-lu26a %I PMLR %P 388--396 %U https://proceedings.mlr.press/v300/lu26a.html %V 300 %X Diffusion models achieve strong generative performance but often rely on large datasets that may include sensitive content. This challenge is compounded by the models’ tendency to memorize training data, raising privacy concerns. SFBD (Lu et al., 2025) addresses this by training on corrupted data and using limited clean samples to capture local structure and improve convergence. However, its iterative denoising and fine-tuning loop requires manual coordination, making it burdensome to implement. We reinterpret SFBD as an alternating projection algorithm and introduce a continuous variant, SFBD flow, that removes the need for alternating steps. We further show its connection to consistency constraint-based methods, and demonstrate that its practical instantiation, Online SFBD, consistently outperforms strong baselines across benchmarks.
APA
Lu, H., Lo, D. & Yu, Y.. (2026). SFBD Flow: A Continuous-Optimization Framework for Training Diffusion Models with Noisy Samples . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:388-396 Available from https://proceedings.mlr.press/v300/lu26a.html.

Related Material