Provably Robust Score-Based Diffusion Posterior Sampling for Plug-and-Play Image Reconstruction

In a great number of tasks in science and engineering, the goal is to infer an unknown image from a small number of measurements collected from a known forward model describing certain sensing or imaging modality. Due to resource constraints, this task is often extremely ill-posed, which necessitates the adoption of expressive prior information to regularize the solution space. Score-based diffusion models, due to its impressive empirical success, have emerged as an appealing candidate of an expressive prior in image reconstruction. In order to accommodate diverse tasks at once, it is of great interest to develop efficient, consistent and robust algorithms that incorporate unconditional score functions of an image prior distribution in conjunction with flexible choices of forward models. This work develops an algorithmic framework for employing score-based diffusion models as an expressive data prior in general nonlinear inverse problems. Motivated by the plug-and-play framework in the imaging community, we introduce a diffusion plug-and-play method (DPnP) that alternatively calls two samplers, a proximal consistency sampler based solely on the likelihood function of the forward model, and a denoising diffusion sampler based solely on the score functions of the image prior. The key insight is that denoising under white Gaussian noise can be solved rigorously via both stochastic (i.e., DDPM-type) and deterministic (i.e., DDIM-type) samplers using the unconditional score functions. We establish both asymptotic and non-asymptotic performance guarantees of DPnP, and provide numerical experiments to illustrate its promise in solving both linear and nonlinear image reconstruction tasks. To the best of our knowledge, DPnP is the first provably-robust posterior sampling method for nonlinear inverse problems using unconditional diffusion priors.

Paper

References (84)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer juip7/10 · confidence 3/52024-07-01

Summary

The authors propose 'diffusion plug-and-play' (DPnP), a plug-and-play diffusion framework that alternatively calls what amounts to a consistency sampler based on the likelihood of the forward model, followed by (essentially) and unconditional diffusion step using the score function. The authors provide many subsequent theoretical results for this framework.

Strengths

1. The paper is well-written and easy to follow/understand (barring some notation issues, see Weaknesses). 2. Each step in the process of developing the DPnP algorithm is well-motivated and explained thoroughly. 3. The provided proofs for the main theorems in Appendix F are well-written and seem correct. 4. DPnP outperforms competitors in non-linear inverse problems in nearly all tested metrics in Section 5 and Appendix G. 5. The overall contribution is impactful with respect to non-linear inverse problems.

Weaknesses

1. This paper would benefit from improvements to the mathematical notation. Namely, it would be easier to read the math if scalar quantities were better differentiated from vector quantities (e.g., bold-faced vectors). 2. The formatting structure of the main paper is, at times, very granular with many lists. Personally, I appreciate this, but I feel that others may think that things could be ordered/structured better. This is a very minor issue, but worth keeping in mind. 3. I acknowledge that the primary contribution of this paper is the theoretical results associated with the DPnP framework, but I think that more robust experimental evaluation would have benefited this work. In particular, it would have been nice to see how the DPnP framework performs on more 'typical' linear inverse problems (e.g., inpainting, deblurring). This would have given a better sense of the overall performance of the approach, even if the focus is performance on non-linear inverse problems. I do not expect you to perform these experiments in the revision/for the rebuttal, I just wanted to note that the results would be more convincing if there were more of them. 4. It would be nice to discuss the connections between the proximal consistency sampler and projection-type approaches (e.g., that leveraged by Chung et al.'s 'Come-Closer-Diffuse-Faster) when the inverse problem is linear. In fact, I suspect that, for linear inverse problems, DPnP can be reduced to a simple projection step, followed by the reverse diffusion step. It would be worthwhile to discuss these connections, even if the discussion is relegated to an appendix. 5. In Algorithm 1, proximal consistency sampling is done before the denoising diffusion sampling. Why is this? The previously mentioned projection methods have the steps flipped, so I wonder why you have decided to structure the DPnP algorithm this way.

Questions

See Weaknesses.

Rating

7

Confidence

3

Soundness

4

Presentation

3

Contribution

3

Limitations

The authors adequately discussed the limitations of their method.

Reviewer ppwi5/10 · confidence 4/52024-07-04

Summary

This paper introduces a diffusion plug-and-play method (DPnP) that uses score-based diffusion models as expressive data priors for nonlinear inverse problems with general forward models. By combining a proximal consistency sampler and a denoising diffusion sampler, the method offers provably robust posterior sampling, with performance guarantees and demonstrated effectiveness across various tasks.

Strengths

This paper establishes both asymptotic and non-asymptotic performance guarantees for DPnP and provides numerical experiments to demonstrate its effectiveness across various tasks. The theoretical analysis presented is a valuable contribution to the field.

Weaknesses

Although I appreciate the theoretical aspect of this paper, as mentioned in the abstract, “this paper develops an algorithmic framework for employing score-based diffusion models as an expressive data prior in nonlinear inverse problems.” However, there are already existing works on embedding denoising diffusion models into plug-and-play frameworks as data priors, such as [1-2]. It would be helpful if the author could clearly highlight any new insights within this paper to distinguish it from previous work; otherwise, the novelty of this paper may appear somewhat incremental. Minor one: There are several typos, e.g., in the abstract, the sentence "Score-based diffusion models, thanks to its impressive empirical success, have emerged as an appealing candidate of an expressive prior in image reconstruction." The correct pronoun should be "their" instead of "its" to match the plural subject "Score-based diffusion models."

Questions

I wonder if the asymptotic consistency and non-asymptotic error analysis of DPnP, as established in this paper, demonstrate convergence. If so, it is necessary to verify this through numerical experiments. Could you please explain why the result of DPnPDDPM shown in Table 1 is smooth, while the one shown in Table 5 is noisy with a lot of noticeable noise? Thank you.

Rating

5

Confidence

4

Soundness

3

Presentation

3

Contribution

2

Limitations

As noted in the weaknesses section, the primary limitation is that diffusion-based PnP methods have already been proposed in [1-2] and applied to various inverse imaging problems.

Reviewer tiqf6/10 · confidence 3/52024-07-08

Summary

This paper introduces a diffusion-based sampling framework closely related to plug-and-play methods for solving general inverse problems. The technique alternates between two steps: calling a proximal consistency sampler that enforces data-fidelity, and regularization via a denoising diffusion sampler leveraging strong diffusion-based image priors. Theoretical results demonstrate asymptotic consistency and robustness to sampling errors. Numerical experiments show promising reconstruction quality.

Strengths

- The theoretical analysis is a valuable contribution. A lack of robustness to sampling errors and the resulting error accumulation has been a key challenge of diffusion-based solvers, especially in highly nonlinear tasks such as phase retrieval. - The paper is well-written overall and the structure is logical. - The experimental results are promising. In particular the proposed plug-and-play sampler achieves significant improvement over DPS, a well-established technique in the literature.

Weaknesses

- The experimental evaluation is somewhat lacking. It would be interesting to see comparison with more contemporary solvers such as ReSample [1], which has improved robustness against sampling errors due to a posterior mean correction scheme. Moreover, in-depth ablation studies on the multiple hyperparameters of the algorithm are missing. Thus, it is unclear how much hyperparameter tuning is necessary. - The proposed technique appears to have a very high compute cost (3000 NFEs). More discussion on the compute requirements and possible ways to accelerate the algorithm would be valuable. [1] Song, Bowen, et al. "Solving Inverse Problems with Latent Diffusion Models via Hard Data Consistency." The Twelfth International Conference on Learning Representations.

Questions

- Can the findings be extended to latent domain samplers? This would greatly improve the efficiency of the technique. - How does the method compare to other samplers such as ReSample? - How does performance scale with NFEs? - I would recommend changing DDS to some other abbreviation to avoid confusion with the Decomposed Diffusion Sampling method [2]. - What does G denote in line 136? [2] Chung, Hyungjin, Suhyeon Lee, and Jong Chul Ye. "Fast diffusion sampler for inverse problems by geometric decomposition." arXiv preprint arXiv:2303.05754 3.4 (2023).

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

Limitations are not clearly addressed in the paper.

Reviewer 43oG4/10 · confidence 4/52024-07-15

Summary

This paper focuses on developing a plug-and-play algorithm for using score-based diffusion models as an expressive prior for solving nonlinear inverse problems with general forward models. While going from current state $x_{k+1}$ to $x_{k}$, this paper makes two gradient updates: (1) from $x_{k+1}$ to $x_{k+\frac{1}{2}}$ using gradients from the measurement error and (2) from $x_{k+\frac{1}{2}}$ to $x_k$ using denoising diffusion score function. Theoretically, the authors prove robustness of the proposed algorithm in sampling from the posterior, and empirically, they show that the proposed algorithm outperforms one commonly used baseline DPS.

Strengths

1. This paper provides a robust algorithm for solving nonlinear inverse problems using unconditional score-based diffusion priors. 2. The paper is nicely written and theoretical analysis in a simple setting clearly demonstrates the major contributions of the paper.

Weaknesses

### **Weaknesses and comments** 1. Line 171: "Assumption on forward model is applicable to *many* applications of interest" what are the applications? 2. Assumption 1: what are some examples of $\mathcal{L(\cdot, y)}$ that are differentiable almost everywhere and used in practice? 3. Step 1 in **A stochastic DDPM-type sampler via heat flow**: How do you know the noise level $\eta$? This is typically unknown. 4. The paper is overloaded with lots of notations. To sample from the posterior, you need the measurement conditional score. By Bayes' theorem you can write this term as the combination of prior and likelihood. How do you get the unknown likelihood term exactly? Please explain the key idea without overloading with notations. 5. Inverting $\Sigma$ is difficult in high-dimension posterior sampling problems. How do you get around this issue unless you are making further approximations like diagonal covariance or scalar? In that case, how do you sample from the posterior exactly? 6. One of the interesting parts of the theoretical results in this line of research is the characterization of the discretization error. But this seems to be out of scope of this paper according to the authors (line 299). 7. Theoretical results are valid provided you get a tight TV bound for both the consistency and diffusion samplers. Any ideas how to obtain these bounds? 8. Experimental results are compared with DPS which the authors claim to be state-of-the-art. Many of the recently developed methods [1,2,3,4] outperform DPS and there is no comparison with any stronger baselines. See citations below and references therein. 9. For evaluation which dataset is used. Is it the same subset used in the original paper and other follow-up papers? The results differ quite a lot if you pick some smaller subset. 10. The authors are encouraged to cite the published version of the papers where applicable (see for instance CCL+22). References 1. Solving Linear Inverse Problems Provably via Posterior Sampling with Latent Diffusion Models. 2. Prompt-tuning latent diffusion models for inverse problems 3. Beyond First-Order Tweedie: Solving Inverse Problems using Latent Diffusion 4. Tweedie Moment Projected Diffusions for Inverse Problems

Questions

Please see the weakness section above.

Rating

4

Confidence

4

Soundness

3

Presentation

2

Contribution

2

Limitations

Yes, the authors have addressed the limitations.

Reviewer ppwi2024-08-09

Response to the author rebuttal

The reviewer has read the authors' rebuttal as well as the comments from other reviewers. Based on these, the reviewer prefers to maintain the initial score.

Reviewer ppwi2024-08-09

Further comments that may be helpful to improve the paper

The experiments presented are inadequate, and certain experimental outcomes are peculiar and fail to substantiate the claims made in the paper. For instance, the results for DPnPDDPM depicted in Table 1 are smooth, whereas those in Table 5 exhibit significant noise. The author did not address these issues during the rebuttal phase.

Authorsrebuttal2024-08-09

Thank you for your prompt reply, and clarification on our rebuttal

Thank you so much for responding to our rebuttal in a timely manner! We really appreciate it. We also appreciate your voicing of concerns regarding the experiments, and would like to clarify further on the performance DPnP-DDPM. In fact, we have addressed this under the headline of "Additional noise in DDS-DDPM" (we apologize the our algorithm name DPnP-DDPM was misspelled by DDS here due to auto-correction). We have to condense the review due to the character limit, and it is likely that you have missed it and therefore we want to repeat it below: > - Thank you for your sharp observations. As far as we can see, only the second row in Table 5 has noticeable noise for DPnP-DDPM. However, we also note that in this row, DPnP-DDPM also recovers visibly finer details and textures of the board and the text in the original image. This tradeoff between the capability of reconstructing finer details and the risk of introducing additional noise is indeed a general phenomenon that has been observed in previous works, e.g., in [SKZ+23, page 21]. In addition, we want to highlight that Table 1 is results for phase retrieval, while Table 5 is for quantized sensing, which are **different measurement forward models**. Therefore, their results are not directly comparable, and visually they may appear quite different. Again, we are happy to include new experiments to further substantiate our paper, if you are willing to provide further feedback. Thank you again for engaging with us!

Reviewer 43oG2024-08-12

Discussion with Authors

The reviewer thanks the authors for the detailed response. The reviewer is satisfied with the clarifications in Q1, Q2, Q3 and Q9. However, the major concerns still remain. Regarding Q4, the term $\log p(y|x_t)$ is typically not computed and existing methods (e.g. DPS) approximate this using $\log p(y|E[x_0|x_t])$, which is similar to the term $L(x;y)$ used in this paper except the additive score function and a cross-term from the Tweedie's formula. Since the noise level $\sigma$ is anyway not known and needs to be tuned in stepsize, the reviewer suspects that the proposed method doesn't offer any major advantages over existing methods at the cost of more compute. This observation is also supported by experiments in Tables 2 and 3. For Q5, how would you precompute $A^T\Sigma^{-1}y$ for inverse problems typically considered in practice, such as Gaussian deblur, motion deblur or super-resolution? What would be the storage space complexity in high-dimensional applications where images could be of size 1024x1024? The reviewer thinks that these issues have not been properly addressed in the paper. For Q6 and Q7, there is no discussion regarding this important piece of information in the main paper, which would essentially help the reader when to choose this algorithm over others. Especially, the dependence on $d$ in $\varepsilon_{DDS-DDPM}$ seems problematic for large-scale applications with $d$ of the order $10^6$ as in recent state-of-the-art inverse solvers. For Q8, the compared baselines are weak and do not adequately justify the claims of the paper. The reviewer was referring to more recently developed pixel space diffusion based inverse solvers such as TMPD or ReSample. The reviewer thanks the authors for providing some preliminary experiments on ReSample. The reviewer will follow the guidelines in revising the score if needed after all the questions have been addressed properly.

Authorsrebuttal2024-08-12

Thank you for your response, and further clarification

Thank you for engaging with us! We are happy to hear that many of your concerns have been addressed successfully, and appreciate your detailed comments that provide us an opportunity to clarify further the remaining points. Regarding your comments on Q4, we would like to clarify two points: - Our approach does not involve approximating $\log p(y|x_t)$ as in the previous algorithms. Our $\mathcal L(x; y)$ is not an approximation to $\log p(y|x_t)$; it is simply the likelihood $\log p(y|x_0)$ (where $x_0$ is the ground truth signal), using the notation in DPS paper, which is assumed known in most of the previous works. Our approach deviates significantly from DPS; the reason we were able to bypass this approximation is that we directly tackle the whole posterior distribution $\propto p^\star(x_0)p(y|x_0) = p^\star(x_0) \exp(\mathcal{L}(x_0; y))$ using a split Gibbs sampler, which is made practical by developing proximal samplers for both $p^\star(x_0)$ and $\exp(\mathcal{L}(x_0; y))$ with diffusion and MALA. - The measurement noise level $\sigma$ is assumed known or is a tunable parameter, which is consistent with most of the previous works (e.g. DDRM, DPS, LGD-MC, ReSample). As can be seen from the experimental results (including our rebuttal pdf), our algorithm demonstrates significant improvement for highly non-linear problems like phase retrieval, over recent works like LGD-MC and ReSample. Regarding your comments on Q5, we would like to note that (i) it is a common choice as in many previous works (e.g. the references above) that $\Sigma$ is chosen as a scalar identity, in which case it is not necessary to perform a matrix inversion; (ii) The storage cost of prefactorization has been found to be managable (of $O(n)$ order) with memory efficient SVD in many practical inverse problems, e.g. denoising, inpainting, super resolution, deblurring, and colorization, cf. DDRM [KEES22]; (iii) If memory is really of concern, there is also the option to simply use MALA for the proximal consistency step that avoids direct inversion, which is still theoretically sound as our theory tolerates errors for both subsamplers. Regarding your comments on Q6 and Q7, we would be happy to include more discussion in the final paper when space permits, which we suppressed in the submission and left a few references. Note that the dependence on $d$ in $\varepsilon_{\sf DDS-DDPM}$ can be further improved to $\tilde{O}(\sqrt{d/T}) + O(\varepsilon_{\sf score})$ using sharper results in [BDBDD24]. However, such dependency with $d$ is known to be tight and generally non-avoidable when plain diffusion models are used. Regarding your comments on Q8, since there is only one day left before the discussion deadline, it is challenging to provide more experiments result in time before the discussion period ends. Nonetheless, we are committed to include more algorithm evaluation in the final version. We also want to provide a bit more discussion regarding the additional experimental results regarding the full evaluation of ReSample on FFHQ dataset, which is already in the rebuttal pdf. Therein, it can be seen that LGD-MC, one of the baseline in our original submission, is a competitive baseline with performance close to that of ReSample. Our algorithm demonstrates significant advantages over both, especially on the phase retrieval task. We expect similar conclusions will hold when we compare ReSample on other datasets/tasks. Thank you again for your careful review. We appreciate your constructive feedback and are happy to discuss more.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC