Evaluating the Role of Great Pre-trained Diffusion Models in Few-shot Phase: Warm-up and Acceleration

Ruofeng Yang, Yongcan Li, Bo Jiang, Cheng Chen, Shuai Li
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7780-7817, 2026.

Abstract

Due to the customized requirements, few-shot diffusion models have attracted much attention. However, only a few works analyze few-shot models, and none involve the fast few-shot optimization process, which is important for quickly responding to users. In this work, we evaluate the role of each operation in the optimization process and prove the convergence guarantee for few-shot diffusion models. A standard operation for the few-shot model is only fine-tuning some key parameters to avoid overfitting the limited target dataset. We first show that this operation is insufficient from empirical and theoretical perspectives. Empirically, we conduct real-world few-shot fine-tuning experiments with underfitting and overfitting bad pre-trained models and show that the results are heavily influenced by these bad models. Theoretically, we also prove that the few-shot phase can not learn the ground-truth parameters and suffers from a small gradient when using a bad pre-trained model. Based on these results, we highlight the importance of a great pre-trained model by showing it can warm up few-shot models and lead to a strongly convex landscape for few-shot diffusion models. As a result, the few-shot model fast converges to the ground-truth parameters. In contrast, we show that with a bad initialization, the pretraining phase requires large optimization steps to converge.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-yang26d, title = {Evaluating the Role of Great Pre-trained Diffusion Models in Few-shot Phase: Warm-up and Acceleration}, author = {Yang, Ruofeng and Li, Yongcan and Jiang, Bo and Chen, Cheng and Li, Shuai}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {7780--7817}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/yang26d/yang26d.pdf}, url = {https://proceedings.mlr.press/v337/yang26d.html}, abstract = {Due to the customized requirements, few-shot diffusion models have attracted much attention. However, only a few works analyze few-shot models, and none involve the fast few-shot optimization process, which is important for quickly responding to users. In this work, we evaluate the role of each operation in the optimization process and prove the convergence guarantee for few-shot diffusion models. A standard operation for the few-shot model is only fine-tuning some key parameters to avoid overfitting the limited target dataset. We first show that this operation is insufficient from empirical and theoretical perspectives. Empirically, we conduct real-world few-shot fine-tuning experiments with underfitting and overfitting bad pre-trained models and show that the results are heavily influenced by these bad models. Theoretically, we also prove that the few-shot phase can not learn the ground-truth parameters and suffers from a small gradient when using a bad pre-trained model. Based on these results, we highlight the importance of a great pre-trained model by showing it can warm up few-shot models and lead to a strongly convex landscape for few-shot diffusion models. As a result, the few-shot model fast converges to the ground-truth parameters. In contrast, we show that with a bad initialization, the pretraining phase requires large optimization steps to converge.} }
Endnote
%0 Conference Paper %T Evaluating the Role of Great Pre-trained Diffusion Models in Few-shot Phase: Warm-up and Acceleration %A Ruofeng Yang %A Yongcan Li %A Bo Jiang %A Cheng Chen %A Shuai Li %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-yang26d %I PMLR %P 7780--7817 %U https://proceedings.mlr.press/v337/yang26d.html %V 337 %X Due to the customized requirements, few-shot diffusion models have attracted much attention. However, only a few works analyze few-shot models, and none involve the fast few-shot optimization process, which is important for quickly responding to users. In this work, we evaluate the role of each operation in the optimization process and prove the convergence guarantee for few-shot diffusion models. A standard operation for the few-shot model is only fine-tuning some key parameters to avoid overfitting the limited target dataset. We first show that this operation is insufficient from empirical and theoretical perspectives. Empirically, we conduct real-world few-shot fine-tuning experiments with underfitting and overfitting bad pre-trained models and show that the results are heavily influenced by these bad models. Theoretically, we also prove that the few-shot phase can not learn the ground-truth parameters and suffers from a small gradient when using a bad pre-trained model. Based on these results, we highlight the importance of a great pre-trained model by showing it can warm up few-shot models and lead to a strongly convex landscape for few-shot diffusion models. As a result, the few-shot model fast converges to the ground-truth parameters. In contrast, we show that with a bad initialization, the pretraining phase requires large optimization steps to converge.
APA
Yang, R., Li, Y., Jiang, B., Chen, C. & Li, S.. (2026). Evaluating the Role of Great Pre-trained Diffusion Models in Few-shot Phase: Warm-up and Acceleration. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:7780-7817 Available from https://proceedings.mlr.press/v337/yang26d.html.

Related Material