Summary
This paper proposes to speed up the training process of the diffusion model with an adaptive sampling strategy. The training process is
empirically divided into 3 stages based on increment: acceleration, deceleration, and convergence. Then the sampling strategy will reduce the frequency of steps from the convergence area. The importance of time steps is also considered. Five baselines are introduced compared with the proposed method on 2 datasets. Overall, this paper studies an important issue of diffusion models. However, there are major flaws and the experiment quality does not allow the acceptance of this paper.
Weaknesses
1. The rationale behind the motivation is not clearly stated and verified. The author states that the time steps can be empirically divided into 3 states. However, there is no empirical result to support the claim. All figures in the method part (i.e., Figure 1 and Figure 2) are pseudo figures. The value in real experiments should be provided.
2. It is not practical to decide the boundary of each state in real application. Also, it is not very clear how to decide the boundary of each state. If the depends on the convergent speed, the states could vary significantly based on the learning rate, model framework, and the quality of training data. Such a strategy is not practical. In fact, the analysis is not comprehensive. Is it possible that an adaptive learning rate will address the issue? The author should exclude other factors to verify it is the sampling quality, not another factor that results in the difference between the 3 stages.
3. The whole paper assumes that the diffusion model is DDPM. However, there are so many papers[1] that have already addressed the quality of sampling such as DDIM. The author should address the problem with a SOTA framework regarding efficiency.
4. The quality of the experiment is low. While the motivation is to speed up the training process, is it intuitive to report the real training time? The majority of the experiment is report FID with baselines. FID is not the metric to verify the efficiency. Figures 5 and 6 are used to report the convergent speed. However, it looks Log(FID) is not convergent yet. Also, what is the learning rate for each baseline in Figure 5? Why there are 3 sub-figures in Figure 5? Should it be a single figure including a comparison with all baselines?
5. Ablation study is missing. The author should remove the sampling strategy for each stage and vary the boundary. The presentation in the experiment could be improved.
6. Important baselines are missing including [2,3,4, 5] and many others.
[1] Shivam Gupta, Ajil Jalal, Aditya Parulekar, Eric Price, Zhiyang Xun:
Diffusion Posterior Sampling is Computationally Intractable.
[2] Tae Hong Moon, Moonseok Choi, EungGu Yun, Jongmin Yoon, Gayoung Lee, Jaewoong Cho, Juho Lee:
A Simple Early Exiting Framework for Accelerated Sampling in Diffusion Models. ICML 2024
[3] Zhiwei Tang, Jiasheng Tang, Hao Luo, Fan Wang, Tsung-Hui Chang:
Accelerating Parallel Sampling of Diffusion Models. ICML 2024
[4] Towards Faster Training of Diffusion Models: An Inspiration of A Consistency
Phenomenon
[5] Hongkai Zheng, Weili Nie, Arash Vahdat, Kamyar Azizzadenesheli, Anima Anandkumar:
Fast Sampling of Diffusion Models via Operator Learning. ICML 2023: 42390-42402
[6] Andy Shih, Suneel Belkhale, Stefano Ermon, Dorsa Sadigh, Nima Anari:
Parallel Sampling of Diffusion Models. NeurIPS 2023
Overall, there are major drawbacks in the proposed method and the quality of the experiment can be significantly improved.