Summary
Authors propose to use previous-stage generations as conditioning, thus re-generating improved images. They show via FID and linear probing that the generated images are higher quality than those generated with a single step. They also do extensive experiments to show that DoD can achieve lower FID faster and with less parameters.
Strengths
1) From a theoretical level, I find the paper very nice. It is a) simple / elegant and b) fast. From a novelty perspective, I have no concerns.
2) The paper includes multiple strong results concerning the practical usefulness of the method (num parameters, steps, GFLOPS)
3) This paper is well-written, and well-presented
Weaknesses
Points are in order of my perceived importance (most to least), indicating how heavily they weigh in my rating.
1) Table 5 (main results) seem to primarily compare against basic diffusion models, but miss comparisons against other post-hoc diffusion methods (e.g. MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer, Gao 2023 already cited in this paper; ReNO: Enhancing One-step Text-to-Image Models Through Reward-based Noise Optimization, Eyring 2024; ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation, Haji-Ali 2023). These should be the true competing methods.
2) LEM is a novel contribution, correct? There are other methods that use image features as conditioning (e.g. Diffusion Feedback Helps CLIP See Better, Wang 2024) and address the issue of making the task too easy if too much information is used--LEM should be justified by comparing to other methods of this type, both theoretically and quantitatively.
3) In Table 5 (main results), DoD is only the best on 3 / 5 metrics. This in itself is not convincing to me--however, I would be more convinced if other results were provided to show that DoD is much more efficient (e.g. the GFLOP or num parameter experiments) than the better performing methods, for relatively little performance drop. At the moment, these efficiency experiments are not done on the out-performing methods (StyleGAN-XL, BigGAN-deep, LDM-4-G, Flag-DiT-G).
4) Top-1 accuracy is included in Table 2--I believe it would be useful in the main results in Table 5 as well, as it is another way to show image generation quality.
5) The main results comparing to other SOTA (table 5) are very hidden, only at the end of the experimental section. They are also very slim on analysis. These need to be brought out more.
6) I have a hard time interpreting the qualitative results, as all images seem very similar. It would be helpful to provide the reader with some specific guidance as to what you perceive as the qualitative improvements.
Questions
In relation to my described weaknesses, the top areas that I see the most important for a strong submission and would need to be improved for me to raise my rating:
1) provide experiments against SOTA comparable models, meaning a) other post-hoc diffusion methods for DoD and b) other image feature conditioning methods for LEM. See W1 and W2 for specific suggestions of comparison methods. I am not particularly tied to these exact methods, but just giving them as a starting place to further understand what kinds of methods I mean.
2) Include better analysis within the main results (table 5), including a) more in-depth explanations in writing, b) better justification for why DoD is useful even though it is not the top-performing model on 3 / 5 metrics (without even including SOTA, as described in Q1)
Given convincing experiments, I would be included to greatly raise my rating, as I like the paper from a theoretical level.