Score-based generative models are provably robust: an uncertainty quantification perspective

Through an uncertainty quantification (UQ) perspective, we show that score-based generative models (SGMs) are provably robust to the multiple sources of error in practical implementation. Our primary tool is the Wasserstein uncertainty propagation (WUP) theorem, a model-form UQ bound that describes how the $L^2$ error from learning the score function propagates to a Wasserstein-1 ($\mathbf{d}_1$) ball around the true data distribution under the evolution of the Fokker-Planck equation. We show how errors due to (a) finite sample approximation, (b) early stopping, (c) score-matching objective choice, (d) score function parametrization expressiveness, and (e) reference distribution choice, impact the quality of the generative model in terms of a $\mathbf{d}_1$ bound of computable quantities. The WUP theorem relies on Bernstein estimates for Hamilton-Jacobi-Bellman partial differential equations (PDE) and the regularizing properties of diffusion processes. Specifically, PDE regularity theory shows that stochasticity is the key mechanism ensuring SGM algorithms are provably robust. The WUP theorem applies to integral probability metrics beyond $\mathbf{d}_1$, such as the total variation distance and the maximum mean discrepancy. Sample complexity and generalization bounds in $\mathbf{d}_1$ follow directly from the WUP theorem. Our approach requires minimal assumptions, is agnostic to the manifold hypothesis and avoids absolute continuity assumptions for the target distribution. Additionally, our results clarify the trade-offs among multiple error sources in SGMs.

Paper

References (35)

Scroll for more · 23 remaining

Similar papers

Peer review

Reviewer DZY46/10 · confidence 4/52024-07-11

Summary

This work studies the influence of different error terms for diffusion models from a continuous perspective under $W_1$ distance. They explain the reason why the early stopping parameter $\epsilon$ would lead to a memory phenomenon of diffusion models. To achieve these results, this work proposes a WUP theorem to explain the robustness of SGMs.

Strengths

1. The WUP theorem is novel since it can analyze generative flows instead of linear SDE and will arise independent interest. 2. The analysis of the early stopping parameter can deepen the understanding of the memory phenomenon. 3. The analysis of different objective functions is interesting.

Weaknesses

1. The abstract section mentions that stochasticity is the key mechanism ensuring SGM algorithms. However, this work does not discuss it in detail in the main content. It seems that the WUP theorem does not hold for the deterministic sampling process. It would be better to discuss it in detail.

Questions

Please see the Weakness part. Question 1: This work does not consider the influence of discretization error. It would be better to discuss the challenge when considering this error. Question 2: For me, one interesting point is the balance of different terms when considering $\epsilon$. As shown in Theorem 3.3, $e_5$ term has $\sqrt{\epsilon}$ dependence and $e_2$ term has $1/\sqrt{\epsilon}$. It seems that there exists a balance between these terms when $N$ is finite. It would be better to discuss this balance in detail.

Rating

6

Confidence

4

Soundness

3

Presentation

2

Contribution

3

Limitations

This work does not discuss the limitation and societal impact in a independent paragraph.

Reviewer gUwF7/10 · confidence 2/52024-07-11

Summary

The paper studies the robustness of score-based generative models (SGM) to different sources of error that are relevant in practice, such as the limited expressivity of the score function representation or the choice of reference distribution. Specifically, they upper bound the Wasserstein-1 and L1 distances between target and approximate distribution (given by an SGM) in terms of these different sources or error for both denoising and explicit score-matching objectives.

Strengths

- Score-based generative models are extremely relevant within the machine learning community, and quantifying their uncertainty in terms of how well they capture the target distribution is an important research problem. - The results are, to the best of my knowledge, new and it is impressive that they manage to isolate different sources of errors in their bounds while placing no restrictive assumptions on the target distribution.

Weaknesses

- The main weakness of the paper is its dense presentation. The authors did a good job of motivating and describing their contributions in the introduction, but the rest of the paper is not easy to follow. I am not sure, there is much room for improvement within the limited space of 9 pages, but I would suggest including a discussion on actionable insights one can get from these bounds. Throughout the paper, the authors comment on how important robustness is for the reliability of generative models in general, and I agree, but I could not get an intuition about when one can expect these bounds to be tight so as to ensure the resulting model is trustworthy. - While I respect the authors’ decision of going for an entirely theoretical paper, I cannot help but feel that some small scale experiments illustrating the tightness and usefulness of the bounds would be enlightening. ### Minor Issues - Line 35: “contributes” should probably be plural here. - Line 53: “recognizes” is misspelled. - Line 61: no need for “of” after “study”. - Line 100: if I’m not mistaken, the acronym SDE was not yet defined at this point in the text. - Line 102: Albeit clear from the context, $W$ and $\eta$ have not yet been defined by this point in the text. - Line 121: “however our results are generally apply” - Line 331: Word missing after “used”, probably “for”.

Questions

- In Section 6.2., the authors discuss an application of their bounds to likelihood-free inference. To that end, could the authors elaborate on how difficult it is to estimate their bounds in practice? Is accurately estimating the Lipschitz constant of the score function the main challenge?

Rating

7

Confidence

2

Soundness

3

Presentation

2

Contribution

3

Limitations

This is mainly a theoretical work and I cannot foresee any direct societal impact stemming from this research.

Reviewer PXMJ6/10 · confidence 4/52024-07-13

Summary

This paper studies the generalization error of diffusion models. The major tool is the Wasserstein uncertainty propagation theorem. With such a result and the regularity analysis in PDE, the authors establish robust analysis for diffusion models with respect to various errors.

Strengths

1. The authors examine the robustness of diffusion models from an uncertainty quantification perspective, that is not well-understood in prior work. 2. By leveraging the Wasserstein Uncertainty Propagation theorem, the authors provide the generalization error of diffusion models w.r.t. various error sources. 3. The paper is well-written and easy to follow. In particular, I appreciate the presentation of math derivation.

Weaknesses

1. The first major concern is the connection and comparison to the literature. The bounds in Theorems 3.2 and 3.3 are not explicit. I suggest the authors provide clear sample complexity results w.r.t. problem parameters. Furthermore, the authors should discuss how the bounds improved the existing results in the literature. 2. The analysis in this paper heavily relies on the PDE theory and regularity analysis. I suggest the authors discuss relevant prior work in Section 1.2. Also, I suggest the authors briefly introduce UQ used in other ML problems beyond diffusion models. Furthermore, the score approximation and estimation theory should be mentioned as that is one source of the errors.

Questions

I have some minor comments and questions: 1. The results established in this work are usually referred to sample complexity bounds or distribution estimation error. The authors use the name robust analysis while the meaning of "robustness" seems different from the one in robust optimization. I suggest the authors clarify this notion in the context. 2. The paper leverages the pathwise characterization of probability distribution generated by the forward process. I suggest the authors add more explanations and/or refer the readers to the literature when a PDE is introduced, e.g., Eq. (11). The same applies to all the FK and HJB equations.

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

NA

Reviewer PXMJ2024-08-12

Thank you for your detailed response. I have raised the score.

Reviewer gUwF2024-08-12

Thanks for the detailed response. I am satisfied with the answers and will maintain my positive score.

Reviewer DZY42024-08-12

Thanks for the detailed response. It would be better to add the discussion of W1 to the main content in the next version. I am satisfied with the answers and will maintain my positive score.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC