Analyzing the Sample Complexity of Self-Supervised Image Reconstruction Methods

Supervised training of deep neural networks on pairs of clean image and noisy measurement achieves state-of-the-art performance for many image reconstruction tasks, but such training pairs are difficult to collect. Self-supervised methods enable training based on noisy measurements only, without clean images. In this work, we investigate the cost of self-supervised training in terms of sample complexity for a class of self-supervised methods that enable the computation of unbiased estimates of gradients of the supervised loss, including noise2noise methods. We analytically show that a model trained with such self-supervised training is as good as the same model trained in a supervised fashion, but self-supervised training requires more examples than supervised training. We then study self-supervised denoising and accelerated MRI empirically and characterize the cost of self-supervised training in terms of the number of additional samples required, and find that the performance gap between self-supervised and supervised training vanishes as a function of the training examples, at a problem-dependent rate, as predicted by our theory.

Paper

References (53)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer X8db5/10 · confidence 4/52023-07-04

Summary

The paper presents a theoretical analysis of the sample complexity of the problem of learning a linear denoiser with a self-supervised learning loss, and verifies the bounds on a series of experiments with linear denoising. The paper also presents an empirical study of the gap between self-supervised learning and supervised learning in the context of image reconstruction with neural networks (for denoising and compressive MRI problems).

Strengths

- A theoretical bound illustrates the role of sample complexity in the setting of self-supervised learning for image denoising. A bound scaling as 1/N where N is the dataset size is presented. - The gap between supervised and self-supervised learning is evaluated for various imaging tasks, showing that self-supervised losses which are unbiased estimators of the supervised loss can achieve a performance on par with supervised learning for large sample sizes.

Weaknesses

There is little link between the theoretical analysis of linear denoisers and the empirical results on non-linear denoising and reconstruction with deep networks. It is not clear whether the theory developed in the linear case can explain the non-linear setting. Moreover, some assumptions in the main theorem seem unrealistic in the context of deep learning, e.g. networks are trained on multiple epochs, not on a single pass as required by the theorem.

Questions

Does the dimension of the signal set play a significant role in Theorem 1? While the dimension d appears in eq. 5, it doesn't seem to impact strongly the final bound, which is somehow surprising. Why the paper doesn't analyze learning with a SURE-based loss? This should also be an unbiased estimator of the supervised loss for the denoising case.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

The paper discusses the limitations related to the choice of an specific neural network. However, I think it would be good to include a discussion of the limitations of using a linear denoiser analysis to understand the dynamics of learning highly non-linear denoisers.

Reviewer 8mky5/10 · confidence 4/52023-07-07

Summary

The work investigates the cost of self-supervised training by characterizing its sample complexity.

Strengths

1. The paper is based on the given theory and carries out corresponding empirical research on self-supervised denoising and accelerated MRI. 2. The paper shows that a model trained with such self-supervised training is as good as the same model trained in a supervised fashion, but self-supervised training requires more examples than supervised training. 3. The paper shows that the performance gap between self-supervised and supervised training vanishes as a function of the training examples, at a problem-dependent rate, as predicted by the theory.

Weaknesses

1. The main concern is that the theoretical approach of this paper seems to be similar to [1], just extending from supervised to self-supervised settings. The corresponding contribution of the theoretical approach should be further elucidated. 2. In the results reported by some previous self-supervised denoising works (e.g., Neighbor2Neighbor), Noise2Noise generally performed the same as supervised methods. But in this work, even with a lot of training data, there is still a gap between the two, what is the reason for this? 3. The paper is based on the setting of simple Gaussian noise. I would like to ask if the authors have done corresponding research or experiments on real-world RGB noise. Is this work still applicable to real-world situations? [1] Scaling laws for deep learning based image reconstruction. ICLR 2023.

Questions

Please see the weaknesses. I am willing to improve the score if the concerns are addressed well.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The limitations have been discussed in the paper.

Reviewer RTiS7/10 · confidence 3/52023-07-12

Summary

The paper is an study on the sample complexity for image reconstruction in two types of methods, self-supervised and supervised. The authors studies the risk bounds for the case of self-supervised methods. They then evaluate the convergance rates in numerical and empirical experiments for two problems, denoising and compressive sensing. The authors conclude that the convergence rates are similar in the two cases of unsupervised and supervised scenarios. However, the self-supervised approach requires more iterations to reach a similar performance.

Strengths

1) The authors theoretically study and find specific risk bounds for using the self-supervised method. 2) The authors study the sample complexity empiracally by considering a range of number of parameters for the network and a range of training set sizes.

Weaknesses

1) For many of the empirical experiments, the authors only report the best out of multiple runs. Illustrating the mean and some measure of variance among multiple runs gives a more complete picture rather than just the best case. Otherwise, the readers would wonder how reliable it is to do a single run of the self-supervised approach compared to the supervised approach.

Questions

1) In appendix 1, in the first line, I think the sign of the last term, $e$, should be negative, based on the given definition for $y\prime$. This results in changing the sign of a few terms in following lines. But the conclusion still holds. 2) Since there is an assumption that the noise distributions for training and inference to be the same in the case of compressed sensing, is it fair to call it self-supervised? 3) As the authors also point to in their limitations segment, the experiments are limited to U-net like design for the architectures. However, they mention they do not expect using different designs would change the qualitative results. Could they ellaborate on the intuition behind this expectation? - Typo: - Line 511: missing reference

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

2 fair

Limitations

The title and the claims in the paper point to image reconstruction in general. However, the experiments are only limited to denoising and compressed sensing. The behavior might be very different for some other image reconstruction tasks such as image inpainting or super-resolution. This limitation should be more pronounced in the claims.

Reviewer RTiS2023-08-15

Response to the authors

Thank you for your detailed response to the reviewers' comments on your paper.

Reviewer 8mky2023-08-18

Thanks for answering the questions. The concerns have been addressed. I have raised my rating.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC