Taming Generative Diffusion Prior for Universal Blind Image Restoration

Diffusion models have been widely utilized for image restoration. However, previous blind image restoration methods still need to assume the type of degradation model while leaving the parameters to be optimized, limiting their real-world applications. Therefore, we aim to tame generative diffusion prior for universal blind image restoration dubbed BIR-D, which utilizes an optimizable convolutional kernel to simulate the degradation model and dynamically update the parameters of the kernel in the diffusion steps, enabling it to achieve blind image restoration results even in various complex situations. Besides, based on mathematical reasoning, we have provided an empirical formula for the chosen of adaptive guidance scale, eliminating the need for a grid search for the optimal parameter. Experimentally, Our BIR-D has demonstrated superior practicality and versatility than off-the-shelf unsupervised methods across various tasks both on real-world and synthetic datasets, qualitatively and quantitatively. BIR-D is able to fulfill multi-guidance blind image restoration. Moreover, BIR-D can also restore images that undergo multiple and complicated degradations, demonstrating the practical applications.

Paper

References (55)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer rhDr3/10 · confidence 5/52024-07-06

Summary

This paper proposes a blind image restoration method by using a pre-trained diffusion model without additional prior knowledge. The proposed adaptive guidance scale is fancy which uses the loss function to judge its value while the degradation function design is confusing. The results on real-world benchmarks show great performance success, but the experiment to validate their proposed modules is fragile.

Strengths

1. The author proposes a non-training strategy to handle the blind image restoration tasks with no additional prior knowledge and achieve SOTA performance in the real world. 2. The way to control the guidance scale is fancy.

Weaknesses

1. The introduction of the method section is confusing, first, what is the “$\sum$” in Line 112 and Line 421? Second, the design of the degradation function D does not make sense, why does adding the M term estimate the noises and what does the noise mean here (Line 72)? 2. The proposed method is unreliable, the author said that they do not have additional prior information, but in my view, the usage of the pre-trained model (Diff-BIR) is the prior knowledge, this paper is just an extra refiner to refine a coarse-clean image to a better one. To validate the proposed method, do the extra ablation study on all the benchmarks without pre-trained model and show the quantitative and qualitative results w, w/o it. 3. The ablation study is fragile. i) Visualize the result of the degradation function D for degrading the clean image. ii) Visualize what M learned for different tasks. iii) Show the value changing of guidance scales during the denoising stage and propose the theoretical analysis of the changing trend.

Questions

See the weakness.

Rating

3

Confidence

5

Soundness

2

Presentation

2

Contribution

2

Limitations

Yes.

Authorsrebuttal2024-08-12

Looking forward to discussion

Dear Reviewer rhDr: We sincerely thank you for taking the time to this review and providing valuable comments. ___ Based on the reviewers' comments, we have made revisions to our manuscript to include the following changes. * We provide the trends of parameters in the adaptive guidance scale and optimizable convolution kernel in the sampling process to clarify how these designs contribute to the performance of BIR-D. * We clarify that we only use the first stage pre-training model for better initialization, rather than the pre-training diffusion model of DiffBIR. And we analyze the improvement of BIR-D using this first-stage pre-training model. * We have supplemented more details on the visualization, description and explanation of some symbols the reviewer mentioned in the review. ___ We hope our explanations have addressed your concerns. As we are in the discussion phase, we welcome any additional comments or questions regarding our response or the main paper. If further clarification is needed, please do not hesitate to mention it, and we will promptly address your inquiries. We look forward to receiving your feedback. Best wishes, The Authors

Authorsrebuttal2024-08-14

Dear reviewer rhDr, We sincerely thank you for your valuable time and feedback. We hope our existing rebuttal and official comments could address your previous concerns well. As the discussion phase is nearing its end, we remain open to addressing any remaining questions or concerns. If you have any further questions during the next discussion period, please let us know, and we will be happy to answer them. We look forward to receiving your feedback. Thank you once again! Sincerely, Best regards, The Authors

Reviewer CUUT5/10 · confidence 4/52024-07-12

Summary

This research introduces BIR-D, a novel approach to the universal challenge of blind image restoration. It leverages an adaptable convolutional kernel designed to emulate the degradation model, with the capability to refine its parameters progressively during the diffusion process. Furthermore, the work presents an empirical formula to guide the selection of the adaptive scale, a critical component in enhancing restoration accuracy. Extensive experiments substantiate the method's exceptional performance across a spectrum of restoration tasks, showcasing its robustness and efficacy.

Strengths

This research offers a novel perspective on Classifier-Guidance, highlighting the essential role of the guidance scale in the fidelity of image generation, and points out that applying a fixed guidance scale across all denoising steps is far from ideal. Therefore, it is necessary to innovate a method that enables the adaptive, real-time adjustment of the guidance scale at each stage of the diffusion process for degraded images in specific restoration tasks. The paper presents a robust validation of the BIR-D method through a comprehensive set of experiments across multiple image restoration tasks, such as deblurring, super-resolution enhancement, low light image enhancement, HDR image recovery, and multi-degradation image restoration.

Weaknesses

1. The contribution in question, which utilizes an optimizable convolutional kernel to simulate the degradation model and dynamically update the parameters of the kernel during the diffusion steps, may be perceived as lacking in novelty. You should provide a detailed comparison with the referenced [10], "Generative Diffusion Prior for Unified Image Restoration and Enhancement," about the different strategy for updating the degradation model. 2. The paper's exploration of the 'optimizable convolutional kernel' and 'adaptive guidance scale' could be enhanced by including an analysis of convergence trends or parameter behavior over time. Such analyses would clarify how these elements contribute to the method's performance. 3. The notations in this paper may lead to misunderstandings. Specifically, in Formula (6), the representation of $g$ lacks the subscript $xt=μ$, which is critical for clarity. Furthermore, the $N$ in Formula (19) should be distinguished from the $N$ used in Formula (20) to avoid ambiguity. Additionally, on line 418, the symbol $K$ should be replaced with $N$ for consistency. 4. The paper's explanation is not sufficiently clear, such as how BIR-D can accomplish multi-guidance blind image restoration.

Questions

1. Upon my review, the formula for calculating the guidance scale $s$ in equation (3), once combined and simplified with equation (1), yields an identity. This suggests that the information obtained from the current $xt$ sampling is independent of the update to $s$. Could you clarify this issue?

Rating

5

Confidence

4

Soundness

2

Presentation

3

Contribution

3

Limitations

None

Authorsrebuttal2024-08-12

Looking forward to discussion

Dear Reviewer CUUT: We sincerely thank you for devoting time to this review and providing valuable comments. ___ Based on the reviewers' comments, we have made revisions to our manuscript in the following areas. * We have supplemented the trends of parameters in the adaptive guidance scale and optimizable convolution kernel in the sampling process to better clarify how these designs contribute to the BIR-D's performance. * We have listed and analyzed the advantages and improvements of BIR-D compared to GDP from various perspectives. Importantly, we have re-clarified that the challenges previously associated with GDP are effectively addressed by BIR-D. * We have provided more details on multi-guidance blind image restoration, including diagram and pseudocode in Global PDF. * We have clarified the necessity of proposing empirical formulas with guidance scales and optimized the derivation approach and process to make it clearer to the readers. ___ We hope our explanations have addressed your concerns. As we are in the discussion phase, we welcome any additional comments or questions regarding our response or the main paper. If further clarification is needed, please do not hesitate to mention it, and we will promptly address your inquiries. We look forward to receiving your feedback. Best regards, The Authors

Reviewer xbfe7/10 · confidence 4/52024-07-14

Summary

The paper introduces BIR-D, a novel approach utilizing generative diffusion models for blind image restoration without requiring predefined degradation types. Traditional methods assume degradation models and optimize their parameters, limiting their applicability. BIR-D overcomes this by employing an optimizable convolutional kernel that simulates degradation dynamically during diffusion steps, allowing it to handle various complex degradations.

Strengths

The method stands out by integrating an optimizable convolutional kernel to dynamically adapt the degradation model during the diffusion steps, a concept not previously explored in the literature. The introduction of an empirical formula for adaptive guidance scale is innovative, eliminating the need for manual grid searches and enhancing the practicality of the approach across diverse image restoration tasks. The experimental results are robust, covering both qualitative and quantitative analyses on real-world and synthetic datasets. The superiority of BIR-D over existing methods is clearly demonstrated through comprehensive experimentation.

Weaknesses

The reviewer appreciates the innovative use of an optimizable convolutional kernel to dynamically adapt the degradation model during the diffusion steps. This is considered the most significant contribution of the work. However, this section lacks sufficient analysis and visualization. While the paper asserts that GDP [1] assumes specific degradation types and is not suitable for complex degradation models, the differences and improvements of the proposed degradation model compared to the one in GDP are not clearly explained. Additionally, the paper is missing some relevant references for blind IR [2] and earlier generative prior-based IR methods [3,4]. Overall, the reviewer appreciates the work and would be happy to adjust the rating if the aforementioned concerns are addressed. [1] Generative diffusion prior for unified image restoration and enhancement. CVPR'23 [2] AND: Adversarial neural degradation for learning blind image super-resolution. NeurIPS'23 [3] Image restoration with deep generative models. ICASSP'18 [4] Maximum a posteriori on a submanifold: a general image restoration method with gan. IJCNN'20

Questions

See weaknesses.

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors have addressed the limitations.

Authorsrebuttal2024-08-11

Official Comment by Authors

We sincerely appreciate your thought-provoking reviews and are pleased to see your upgrading decision. Following your valuable suggestions, we will carefully incorporate these revisions into our future version. To substantiate our results, we will release the code. Thank you once again for your positive rating and the time devoted to this review. Best regards, The Authors

Reviewer xbfe2024-08-11

Thank you for providing the additional details and clarifications in your response. I believes that the visualizations provided in the rebuttal file could help potential readers better understand the paper's contributions. The comparison to the GDP method also addresses my concerns. Therefore, I have increased my rating by one point.

Reviewer rhDr2024-08-12

The author only addresses partial of my concern. For the noise mask, the visualization is hard to understand and the explanation is not convincing. For pretrained weight, as you are doing the setting of universal, how can you add it to some of the tasks? This makes me question the veracity of the author's experiment in UNIVERSAL。 All in all, the original paper lacks too many experiments and I do not think it can be refined directly. I will keep my rating.

Authorsrebuttal2024-08-13

Dear Reviewer rhDr: We sincerely appreciate your time and effort in reviewing our paper. We would like to clarify the following issues you mentioned. 1. The setting of mask $\mathcal{M}$ is mainly used to solve the image restoration task of local regions with significant differences in brightness, as using an optimizable convolution kernel alone may not be able to effectively solve the brightness correction task of such regions. As shown in Global PDF Figures 4 and 5, a mask with the same dimension as the degraded image can learn the brightness and detail information of each local region of the image, which also assists the optimizable convolution kernel in simulating the degradation function. 2. The first stage pre-training model is only used to improve model performance in the two tasks of deblurring and motion blur reduction. Without the first stage pre-training model, our BIR-D can still achieve image restoration of deblurring and motion blur reduction tasks (see Table 6 in the main text and Figure 15 of Appendix G). It is worth noting that this first stage pre-training model is not used in the other blind image restoration tasks since other tasks can be effectively modeled by our devised optimizable convolution kernel. The universal capacity of our BIR-D comes from the design of an optimizable convolution kernel, which can effectively simulate any degradation models of most blind image restoration tasks. The experiment we conducted also proved this contribution. Here is a summary of the experiment we conducted. | Category | Task | Dataset | Figure | Table | |----------------|-----------------------------|-----------------------|-----------|-------| | | Deblurring | | 1,5(b),15 | 2 | | | Colorization | | 1,4,18 | 2 | | Linear Inverse | Super-resolution | ImageNet 1k | 1,5(a),17 | 2 | | | Inpainting | | 1,5(c),16 | 2 | | | Multi-task | | 1,9,10 | - | |----------------|-----------------------------|-----------------------|-----------|-------| | | BIR in Real-world Dataset | LFW,Wider | 1,3,11 | 1 | | | Low-light Enhancement | LOL,VE-LOL,LoLi-Phone | 1,6,12 | 3 | | Non-linear | Motion Blur Reduction | Gopro,HIDE | 1,8,13 | 4 | | | HDR Image Recovery | NTIRE 2021 | 1,7,14 | 4 | | | Realistic Image Restoration | Website | 1 | - | ___ We hope these clarifications will enhance your comprehension of our paper. If you have any further comments, please do not hesitate to mention it. We look forward to further communicating with you. Best wishes, The Authors

Reviewer CUUT2024-08-14

Most of the concerns have been addressed in the authors' response. I will raise the score.

Authorsrebuttal2024-08-14

Dear Reviewer CUUT: We sincerely appreciate your helpful and constructive review and are pleased to see your decision to raise your score. Based on your valuable suggestions, we will provide detailed explanations of the differences between BIR-D and GDP in the future version to highlight the strengths and improvements of BIR-D. Meanwhile, the parameter trends of the kernel and mask will be incorporated into our future version. We will also integrate the pipeline of multi-guidance BIR-D in the future version to make our multi-guidance method clearer. To further support our paper, we will carefully release our code. Thank you once again for your recognition of our work and the valuable time you have invested in this review. Best regards, The Authors

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC