Diffusion Models and Semi-Supervised Learners Benefit Mutually with Few Labels

In an effort to further advance semi-supervised generative and classification tasks, we propose a simple yet effective training strategy called dual pseudo training (DPT), built upon strong semi-supervised learners and diffusion models. DPT operates in three stages: training a classifier on partially labeled data to predict pseudo-labels; training a conditional generative model using these pseudo-labels to generate pseudo images; and retraining the classifier with a mix of real and pseudo images. Empirically, DPT consistently achieves SOTA performance of semi-supervised generation and classification across various settings. In particular, with one or two labels per class, DPT achieves a Fréchet Inception Distance (FID) score of 3.08 or 2.52 on ImageNet 256x256. Besides, DPT outperforms competitive semi-supervised baselines substantially on ImageNet classification tasks, achieving top-1 accuracies of 59.0 (+2.8), 69.5 (+3.0), and 74.4 (+2.0) with one, two, or five labels per class, respectively. Notably, our results demonstrate that diffusion can generate realistic images with only a few labels (e.g., <0.1%) and generative augmentation remains viable for semi-supervised classification. Our code is available at https://github.com/ML-GSAI/DPT.

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer JKn27/10 · confidence 4/52023-06-29

Summary

In this work, the authors propose a novel training strategy, dual pseudo training (DPT), that involves three steps: First, DPT uses state-of-the-art semi-supervised learning methods to predict pseudo-labels on a partially-labeled dataset and diffusion models. Second, DPT trains a conditional generative model with pseudo-labeled data for pseudo-images. Finally, the classifier is re-trained with a dataset which includes both real and pseudo images. DPT surpasses existing strong baseline diffusion models in terms of FID score and semi-supervised learning methods on image classification on CIFAR-10 and ImageNet.

Strengths

* The idea of using Semi-Supervised Learning (SSL) to train diffusion models and then improve SSL is novel. This is analogous to distillation methods that sometimes lead to better student performance than the teacher network. * The method boosts the performance of both the diffusion model and the discriminative model, becoming the new state-of-the-art method in both fields. * The writing is clear. I enjoy reading the paper. * The ablations in the appendix are comprehensive and answered several questions that I have.

Weaknesses

* The baseline diffusion method in this work that leverages full labels is inherently stronger than LDM. Since DPT does not outperform baseline with full supervision, the claim that "with one or two labels per class, DPT ... surpassing strong diffusion models with full labels, such as IDDPM, CDM, ADM, and LDM" is misleading to the readers that DPT, with pseudo-labels, leads to better performance than using full labels. * Since the authors use a different setting for diffusion method (e.g., with different backbone), the comparison with previous works (e.g., ADM and LDM) is not fair. The authors are encouraged to either use baseline settings from previous methods or migrate their setting to baseline methods to make sure the comparison is fair. * The discussions of computational efficiency (e.g., in terms of wall clock time) has been omitted. Diffusion models have been known to be slow in the inference process, especially when applied in the pixel space rather than the latent space. Note that this method is still valuable even if it's not particularly compute efficient, as long as the compute required is reasonable, as many semi-supervised learning methods are also less efficient compared to supervised learning.

Questions

Question: * The method seems to be sensitive to CFG. If I understand it correctly, a CFG of 0 means no guidance. While methods such as stable diffusion typically use a large CFG for an optimal trade-off between diversity and fidelity (7.5 by default, which is equivalent to 6.5 in the notation of this paper), the method uses CFG less than 1. Why is this large discrepancy? The authors are welcome to address the questions in the weakness section.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

3 good

Contribution

4 excellent

Limitations

As mentioned in the weakness section, the discussions of computational efficiency (e.g., in terms of wall clock time) has been omitted. The authors are encouraged to discuss their compute usage and wall clock time as well as the one for their baseline.

Reviewer hJXD6/10 · confidence 4/52023-07-05

Summary

To advance semi-supervised generative and classification tasks, this paper proposes a simple yet effective training strategy termed dual pseudo training (DPT). The proposed DPT operates in three stages: training a strong semi-supervised learner to predict pseudo labels, training a strong diffusion model using both real and pseudo labels to generate pseudo images, and retraining the semi-supervised learner with a mix of real and pseudo images. Experimental results show the effectiveness of the proposed DPT on semi-supervised generation and classification across various settings.

Strengths

1. The paper is well-written. 2. The proposed DPT bridging strong semi-supervised learners and diffusion models by pseudo labels and images is technically sound. 3. The idea that diffusion models and semi-supervised learners benefit mutually with few labels is interesting and of practical use. 4. It is impressive that the proposed DPT using 1% labels yields comparable results to the supervised baseline using 100% labels (cf. Table 2).

Weaknesses

1. Does this paper have any technical novelty? It should be highlighted in the abstract and introduction parts. 2. The proposed DPT improves semi-supervised learning models with generative augmentation via diffusion models. Despite substantial performance improvements, the computing overhead and time budget may be big. It would be better to provide a comparative analysis between the proposed method and existing ones. 3. At the second stage of the proposed DPT, pseudo labels may not always align with ground truth labels, making it hard to control semantics in samples. How does the pseudo label accuracy affect the pseudo image fidelity and semantics? Some illustrative examples are preferred. 4. Would the classifier retrained with a mix of real and pseudo images be used to predict pseudo labels as the first stage? The three stages may form a positive cycle to iteratively improve the pseudo label accuracy and pseudo image semantics. 5. Lines 57-58: The exploration on why classifiers can benefit generative models through class-level visualization and analysis is important for the method validation. It would be better to place it in the main text. 6. In Table 1, the semi-supervised generative model S^3GAN was proposed in 2019. Haven't there been any new methods in the last few years? 7. In tables of image generation results, different methods may differ in the model size, leading to unfair comparison. It would be better to show the number of model parameters in these tables. 8. Many recent related works on semi-supervised learning have not been included, discussed, and compared in this paper, e.g., [a-h]. [a] Zhang et al. Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling. NeurIPS, 2021. [b] Tang et al. Stochastic Consensus: Enhancing Semi-Supervised Learning with Consistency of Stochastic Classifiers. ECCV, 2022. [c] Wang et al. Unsupervised Selective Labeling for More Effective Semi-Supervised Learning. ECCV, 2022. [d] Tang et al. Towards Discovering the Effectiveness of Moderately Confident Samples for Semi-Supervised Learning. CVPR, 2022. [e] Lim et al. Class-Attentive Diffusion Network for Semi-Supervised Classification. AAAI, 2021. [f] Gong, S. et al. (2023). Diffusion Model Based Semi-supervised Learning on Brain Hemorrhage Images for Efficient Midline Shift Quantification. In: Frangi, A., de Bruijne, M., Wassermann, D., Navab, N. (eds) Information Processing in Medical Imaging. IPMI 2023. Lecture Notes in Computer Science, vol 13939. [g] Alshenoudy, A., Sabrowsky-Hirsch, B., Thumfart, S., Giretzlehner, M., Kobler, E. (2023). Semi-supervised Brain Tumor Segmentation Using Diffusion Models. In: Maglogiannis, I., Iliadis, L., MacIntyre, J., Dominguez, M. (eds) Artificial Intelligence Applications and Innovations. AIAI 2023. IFIP Advances in Information and Communication Technology, vol 675. [h] Chen et al. Debiased Self-Training for Semi-Supervised Learning. NeurIPS, 2022.

Questions

1. Does the "retraining" in the 3-rd stage mean fine-tuning or training from scratch? 2. Why do the strong diffusion models with full labels, such as IDDPM, CDM, ADM, and LDM, perform worse than the proposed DPL with few labels? How would the proposed DPT perform if full labels are used? It provides the uppper bound. 3. On Line 54, DiT-XL/2-G [7] (FID 2.27) is the SOTA supervised baseline, rather than [5] (FID 3.40). Please carefully check and correct. 5. In Fig. 2 (a), it is no need to indicate the label fraction by the bubble area, as the x-axis has already indicated the label fraction. Different shapes or colors are enough to show the advantage of the proposed DPT. 6. In Table 2, the "bold" and "underline" look a little messy. Some columns of evaluation metrics have no mark while others have many. A better marking scheme should be figured out to make the table easier to understand.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

Yes.

Reviewer wQJF6/10 · confidence 4/52023-07-06

Summary

This work presents an interesting interaction between a generative model and a semi-classifier. The pipeline can boost both the generative sample quality and discriminative performance of the semi-classifier in the few label settings. A three-stage pipeline is designed. First, the semi-classifier is trained on a few label data and the semi-label is acquired for all the unlabeled data. Second, the generative model is trained on both label and semi-label data to achieve comparable sample quality as a full-supervised one. At last, the semi-classifier can further improve its performance from the new samples generated by the learned generative model. DPT with two (i.e., < 0.2%) labels per class is comparable to SOTA supervised baseline (FID 2.52 vs. 2.29).

Strengths

- Well written and easy to follow. - Motivation is clear and the pipeline is simple. - Good results. DPT with two (i.e., < 0.2%) labels per class is comparable to SOTA supervised baseline (FID 2.52 vs. 2.29). - Interesting interaction between semi-classifier and diffusion model.

Weaknesses

- 61-63 Specifically, DPT with four labels per class achieves an error rate of 4.53% on 62 CIFAR-10, which even outperforms the performance of SOTA methods on CIFAR-10 63 with 25 labels per class. The Full-flex[1] algorithm has better performance than 4.53% for 4 labels per class on CIFAR-10, which might make this claim incorrect. - It's interesting to see if the framework is general to other semi-algorithm and diffusion model pairs. - lack of analysis of the key components of the interaction between the generative model and semi-classifier. For example, what is the key factor to make semi-classifier success for generative model in the design? And What is the key factor to make genertive model success for semi-classifier in the design? This should be discussed in the paper. [1] Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data

Questions

- Is the algorithm stable for different starting samples in the first stage? - what are the key components for the successful interaction between the generative model and semi-classifier? - How about the PSNR or other commonly used generation evaluation metrics? Since FID and IS are highly imagenet related. - Can it improve the upper bound of generation quality? If combining full supervised label data with this pipeline, will it surpass the 2.29 FID? - What about the improvement of unconditional generation? - The weaknesses mentioned above should be addressed.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

2 fair

Presentation

3 good

Contribution

2 fair

Limitations

yes

Reviewer 2c3F7/10 · confidence 3/52023-07-06

Summary

This paper proposes to exploit generative diffusion models to benefit semi-supervised learning and vice versa. Specifically, the proposed training strategy called dual pseudo training (DPT) consists of 3 stages: (1) Train a semi-supervised learning (SSL) model on labeled and unlabeled data (2) Train a conditional generative diffusion model using pseudo labels predicted by the SSL model in (1) (3) Use the generative model in (2) to generate more labeled images to train a classifier DPT achieves both SOTA performances on image generation and semi-supervised classification tasks on ImageNet and CIFAR-10, demonstrating "Diffusion Models and Semi-Supervised Learners Benefit Mutually with Few Labels".

Strengths

(1) The paper is well-written and easy to follow, (2) The proposed method is simple but effective. And it can leverage any SSL algorithms or generative models in a plug-and-play manner. (3) The results are very impressive, both in numbers (SOTA in both classification and generation tasks) and samples (Figure 1).

Weaknesses

(1) Naive combination of SSL and generative models is a straightforward idea; the novelty is limited.

Questions

I do not have questions for this paper.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

As stated by the author, filtering of pseudo labels in stage 1 and pseudo images in stage 2 should be investigated to reduce noise.

Reviewer kGQr7/10 · confidence 4/52023-07-13

Summary

This paper empirically studies the effectiveness of combining diffusion models and semi-supervised classification. First, an image encoder is trained using a self-supervised method, then finetuned with a small fraction of the ground-truth labels. Second, the image classifier assigns pseudo-labels to all unlabeled images in the training dataset which are then used to train a label-conditioned diffusion model. Finally, the diffusion model is used to generate image-label pairs that are used to further train the image classifier from the first step. Multiple experiments on CIFAR10 and ImageNet demonstrate the effectiveness of this approach, both in improving the diffusion model and the image classifier.

Strengths

- The paper demonstrates state-of-the-art results on both semi-supervised image classification and semi-supervised image generation - The presentation of both the model and the experiments is easy to follow and the text is well written - Exploring how diffusion models and discriminative models like image classifiers can mutually benefit from one another is interesting and relevant to both academics and practitioners

Weaknesses

1. As shown in [1], the FID can be decreased by aligning the histograms of Top-N classifications without improvements to the image quality. Albeit there exists no perfect method to quantify this effect, it would be insightful to see some investigation into whether the lower FID actually translates to improved image quality or merely better alignment in terms of Top-N classifications. 2. In the introduction, the paper claims that (unsupervised conditional) diffusion models trained with cluster assignments as pseudo-labels still underperform their ground-truth supervised counterparts. As shown in [2] this no longer holds true when the cluster number is set higher than the number of ground-truth classes. It would be interesting to explore how this insight could further benefit the semi-supervised setting, for example by further sub-clustering images with the same pseudo-labels. I consider the paper should be published if they fix their claim about unsupervised diffusion models. [1] T. Kynkäänniemi, T. Karras, M. Aittala, T. Aila, J. Lehtinen, "The Role of ImageNet Classes in Fréchet Inception Distance", ICLR 2023. [2] H. Tao, D. W. Zhang, Y. Asano, G. J. Burghouts, C. G. M. Snoek, "Self-Guided Diffusion Models", CVPR 2023.

Questions

The paper reports image generation results on ImageNet for multiple resolutions and differing numbers of ground-truth labels per class. What is the motivation behind the specific choices for the resolution and number of labels combinations? Does the choice of the resolution for the generative model affect the classification results?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

A potential negative impact of this line of work (semi-supervised generation/classification) is that it could amplify stereotypes where a few examples are taken as representative of the whole class (group). Some discussion on this would benefit the paper.

Reviewer kGQr2023-08-20

Thanks for the rebuttal. It addresses my initial concerns hence I raise my score from Weak Accept to Accept.

Authorsrebuttal2023-08-20

Thanks for the update!

Dear Reviewer kGQr, Thank you very much for your decision to update the rating to 'accept'. We highly appreciate it. Best regards, Authors

Reviewer JKn22023-08-10

The authors have posted an effective rebuttal that resolves my concerns. I still vote for acceptance of this work.

Authorsrebuttal2023-08-12

Additional experiments on the stability of DPT for different starting samples

Thank you for your question about DPT's stability for different starting samples. We add the experiment by randomly sampling labeled data in stage 1 three times on CIFAR-10 in Rebuttal Table 1. As shown in Rebuttal Table 1, our DPT achieves consistent improvement, which demonstrates that DPT is stable for different starting samples. **Rebuttal Table 1.** Error rates on CIFAR-10 32$\times$32 given 4 labels per class. | Method | labeled data 0 | labeled data 1 | labeled data 2 | | :------------------------: | :------------: | :------------: | :------------: | | FreeMatch | 4.95 | 5.07 | 4.76 | | DPT with EDM and FreeMatch | 4.53 | 4.91 | 4.60 |

Reviewer wQJF2023-08-20

Dear authors, Thanks for the answer from the rebuttal and the pointer from the appendix, most of my concerns are addressed. I raise my rating to 'weak accept'. Still, the explanation of why the classifier and generator could benefit each other is mostly empirical and results-driven. More detailed analysis or theoretical explanation should make this work more solid and potentially have a greater impact.

Authorsrebuttal2023-08-20

Thanks for the update!

Dear Reviewer wQJF, We are sincerely grateful for your insightful suggestion and your generous decision to update the rating to 'weak accept'. We will follow your suggestion, and try to add more detailed analysis and theoretical explanation in the future. Best regards, Authors

Reviewer hJXD2023-08-22

Thanks for the effective rebuttal that well addressed my concerns. Hence, I raise my rating from borderline accept to weak accept. I would be glad to see the authors include the posted revisions in the final version.

Authorsrebuttal2023-08-22

Thanks for the update!

Dear Reviewer hJXD, We are sincerely grateful for your insightful comments and your generous decision to update the rating to 'weak accept'. We highly appreciate it. Best regards, Authors

Program Chairsdecision2023-09-21

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC