Deep deraining networks consistently encounter substantial generalization issues when deployed in real-world applications, although they are successful in laboratory benchmarks. A prevailing perspective in deep learning encourages using highly complex data for training, with the expectation that richer image background content will facilitate overcoming the generalization problem. However, through comprehensive and systematic experimentation, we discover that this strategy does not enhance the generalization capability of these networks. On the contrary, it exacerbates the tendency of networks to overfit specific degradations. Our experiments reveal that better generalization in a deraining network can be achieved by simplifying the complexity of the training background images. This is because that the networks are ``slacking off'' during training, that is, learning the least complex elements in the image background and degradation to minimize training loss. When the background images are less complex than the rain streaks, the network will prioritize the background reconstruction, thereby suppressing overfitting the rain patterns and leading to improved generalization performance. Our research offers a valuable perspective and methodology for better understanding the generalization problem in low-level vision tasks and displays promising potential for practical application.
Paper
Similar papers
Peer review
Summary
This paper explores the generalization problem of image deraining. The authors claim that deep networks display a tendency to learn the less complex element in the task of separating image content from additive degradation. Overall, this paper contributes some new ideas and new insight to the community, but there are still some weaknesses that can be further improved.
Strengths
S1. The paper is well-organized and clearly written. S2. The authors point out the the generalization problem in the context of deraining tasks, which contributes some important ideas for the community. The key findings of this paper left a deep impression on me.
Weaknesses
W1. The author chose two CNN-based methods, RCDNet and SPDNet, to validate the main finding of this paper. I am curious whether the latest Transformer-based image deraining method still has the same conclusion. For example, Chen, et al. "Learning A Sparse Transformer Network for Effective Image Deraining." CVPR, 2023. In fact, the generalization performance of CNN-based and Transformer-based deraining methods is different. I suggest the author could add the experimental analysis of the Transformer-based deraining method. W2. The authors claim that the relative complexity between the background and rain determines the network behavior. In previous studies, Rain100L and Rain100H are two classic image deraining benchmark datasets that share the same background image. As far as I know, Rain100H (rain is more complex) has better generalization performance than Rain100L. I am curious about how to define the relative complexity between the background and rain. The expression here needs further improvement, which is a bit difficult to understand.
Questions
See the above Weaknesses part.
Rating
8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.
Confidence
5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
The author has clarified the limitations in the supplementary materials.
Call for Further Discussion
Dear reviewer FkGF: We thank you for the precious review time and valuable comments. And Thank you for recognizing our work We have provided corresponding responses and results, which we believe have covered your concerns. We hope to further discuss with you whether or not your concerns have been addressed. Please let us know if you still have any unclear parts of our work. Best, Paper 798 Authors.
Response to Authors
Thanks for the author's detailed responses. This paper contributes some new insight to low-level vision interpretability research, and the rebuttal also answered my concern clearly. Besides, it is better to add these analyses in the rebuttal to the released version.
Summary
This paper proposed to explore the problem of generalization problem in image deraining. The authors find that deep networks tend to overfit the degradation patterns present in the training set, which leads to poor generalization performance for unseen degradation patterns.
Strengths
1. Writing fairly well. 2. The problem is interesting. It is very interesting to explore such findings that better generalization in a deraining network can be achieved by simplifying the complexity of the training data. None of previous works have discussed these issues.
Weaknesses
1. Why did the author create a new dataset without conducting experiments on previous existing benchmarks. The author should explain this issue. 2. Is there any connection between the generalization problem proposed by the author and the interpretability of the low-level vision networks? I am confused about the conclusion of this paper.
Questions
Is the main findings of this paper applicable to other low-level vision tasks? Has the author conducted any extension experiments for other tasks?
Rating
8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
The authors adequately addressed the limitations.
Call for Further Discussion
Dear reviewer w1Er: We thank you for the precious review time and valuable comments. And Thank you for recognizing our work We have provided corresponding responses and results, which we believe have covered your concerns. We hope to further discuss with you whether or not your concerns have been addressed. Please let us know if you still have any unclear parts of our work. Best, Paper 798 Authors.
Response to Authors
Thanks for the author's thoughtful reply. The rebuttal addressed my concerns well. To my knowledge, there is a lack of interpretable research in the field of image rain removal, and this paper opens up a new perspective for the community. I believe that improving the interpretability of the solving process will be beneficial to improve the deraining performance and applications of image deraining methods, which will be important directions for future research. I was originally positive at the paper. When I checked other reviews and the rebuttal, I decided to raise my rating. I believe that the significance, key findings, experimental analysis, and discussion of this work meet the acceptance standards of NeurIPS, and I am inclined to accept this work. This will be an important pioneering work in the field of interpretable-based image deraining. Best, Reviewer w1Er
Summary
This paper is working on understanding generalization problem in image deraining. Several interesting findings are: 1) training with fewer background images leads to better deraining effect; 2) the relative complexity between the background and rain determines the network behavior; 3) a more complex background set makes it harder for the network to learn.
Strengths
there are some interesting findings in this paper for image deraining for example 1) training with fewer background images leads to better deraining effect. 2) the tradeoff between deraining and background reconstruction.
Weaknesses
1) although there are interesting finding, the overall contribution is limited, i.e., authors didn't provide a good solution to solve the problem. for example for figure 9, although results from author removed rain, but it also generated much more artifacts in the background which probably not good for image deraining task. I would expect more contribution from authors to make this paper accepted by NeurIPS. 2) My feeling is that there is fundamental limit in the current training for image derain, the simulated rain streaks is not real. I wondered wether the current findings from author because of this fundamental limit. 3) some related papers like [1] is not cited or compared. [1] Jie Xiao, Man Zhou, Xueyang Fu, Aiping Liu, Zheng-Jun Zha, Improving De-raining Generalization via Neural Reorganization, ICCV 2021. I have carefully read authors rebuttal and still thought the quality of this paper could be improved. So will keep my rating.
Questions
see the weakness part.
Rating
3: Reject: For instance, a paper with technical flaws, weak evaluation, inadequate reproducibility and incompletely addressed ethical considerations.
Confidence
5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.
Soundness
2 fair
Presentation
3 good
Contribution
2 fair
Limitations
yes, authors provided some limit of this paper, for example, this paper doesn't propose any new algorithms directly.
Call for Further Discussion
Dear reviewer e5xd: We thank you for the precious review time and valuable comments. We have provided corresponding responses and results, which we believe have covered your concerns. We hope to further discuss with you whether or not your concerns have been addressed. Please let us know if you still have any unclear parts of our work. Best, Paper 798 Authors.
Summary
This paper makes observations on the generalization problem of deraining networks by conducting extensive experiments, on the perspective of training data. This paper claims to find out some underexplored insights that less training data with less complex background can lead to better rain removal ability at the cost of the background reconstruction ability and that these conclusions can benefit the future research of deraining community. Two new quantitative metrics are proposed to isolate the evaluation of ability of rain removal and background restoration and experiments are based on them.
Strengths
1. It is interesting and valuable to investigate the reason why existing deraining methods fail to generalize to unseen rain streaks, from the perspective of training data. 2. It is reasonable to separately evaluate the ability of rain removal and background restoration of deraining models, for the two parts both have influence on the final restoration performance. 3. It is good to conduct extensive experiments to validate the assumptions and useful for those new researchers to deraining tasks. 4. The paper is easy to follow.
Weaknesses
Despite the interesting motivation and efforts on studying the generalization problem of deraining methods, there are some basic problems for this paper. 1. It is negative for the paper to hold the opinion that the background restoration quality can be sacrificed to get better rain removal ability. It is opposite to the practice of most image restoration tasks, i.e., keeping the degradation-free areas unchanged. Deraining, which consists of removing the rain streaks AND restoring those lost background details occluded by these rain streaks, is special that most part of the image is actually unaffected given only rain streaks exist and there is no rain accumulation. Thus, the objective of deraining is to remove as many rain streaks as possible when keeping the background unchanged. It agrees with the loss functions used to train models, that keeping the output more similar to the ground truth. The observation found in this paper, is actually not the deraining models overfit to the pattern of rain streaks, rather that models try to keep the background unchanged and do nothing to the image if rain streaks are not recognized. Training with less data indeed makes models to remove more rain, but also makes them to remove background details. That's why we need more powerful models to train on more diverse data to get better performance on real cases. 2. The higher value of proposed metric 'Rain Removal Performance' doesn't mean the better ability of rain removal. As mentioned above, the goal of deraining is to remove the rain streaks and restore the occluded background textures. Thus, simply calculating the difference between between the rain region of rainy and output images, cannot directly tell the ability of removing rain. For example, in the extreme case, the rain removal ability is 'the best' when the rain region is all set to zeros because the difference between the zeros and white rain is maximized, but it is unwanted for deraining. In addition, the difference between the rainy and ground truth in rain regions is different for different backgrounds. Therefore, it cannot directly tell the ability of rain removal by the relative value of the proposed metric. 3. The lower value of proposed metric 'Background reconstrution' doesn't mean the better ability of background reconstrution. Instead, it can only means that the models do less harm to background details when removing rain streaks. To be more specific, the reconstructed background only lies within the rain region rather than the unaffected background areas. Together with the above point, the metric of 'Rain Removal Performance' is a compoistion of the ability of removing rain and reconstructing background. Thus, the proposed two metrics acturally do not isolate the evaluation of removing rain and reconstructing background. Additionally, it can be found that only a certain amount of background images like 256-1k can help models to learn to keep unaffected areas unchanged, which may be good for future data synthsis. 4. The rain region is obtained by thresholding to calculate the new metrics. It is weird because the experiments are conducted on full synthetic data and the rain is manually superimposed on clean images, which means the accurate rain areas are already known in advance. The threshold value can also have negative effects on the correctness of observations. 5. It is infeasible to make more conclusions than that L1 loss for training is not the optimal choice when the whole training is based on L1 loss. Training with L1 loss aims to tell models to produce more similar outputs to ground truth pixel-wisely and evaluating the performance with PSNR and SSIM is persuing better visual quality. While this paper proposes new evaluation metrics and can only prove that L1 loss cannot help models to remove more rain at the cost of removing background details as expected by the paper. It is more important to make real solutions based on the proposed metrics to guide the training of models to get better results. But it is not considered in this paper. 6. It is simple to make seem-to-be-valuable observations via various experiments and such efforts should be respected. However, such conclusions can also be harmful to researchers who are new to this area if no absolute therories supports such insights. Authors also knows that better background preservation can lead to higher PSNR but remove less rain streaks as described in Sec 2.2. The marching direction of the community should be removing more rain streaks and handle more types of rain while keeping the background textures unchanged rather than removing more rain at the cost of background quality. It is recommended that authors also examine the generalization problem from the perspective of model complexity, and you may find that less complex models tends to remove more rain streaks at the cost of removing more backgrounds rather than doing nothing like more complex models. 7. Let's make a step backward, assuming these observations in this paper are insightful, the strategy proposed by authors that trying to find a balance of complexity of background and rain streaks should be applicable to all deraining models, no matter what architectures and complexities they have. However, the quantitative and qualitative results show that RCDNet may be the only benefited model of the proposed strategy. Thus, it is not a general solution that less training data brings better balance and generalization. 8. The selected SPDNet and RCDNet, etc. are actually not the state-of-the-art models in deraining. There are many more powerful deraining-oriented models or general backbone models in both CNN-based and Transformer-based architectures these years. It is more insightful to examine how these new powerful models behave when the training data of background images is decreasing. 9. For the second implication in Sec 4., is there any chance that larger rain range occludes more background areas and models need more data to learn how to keep unaffected regions unchanged and restore those occluded ones? Different perspectives can lead to various conclusions based on these experimental results. 10. As the LLM and SAM models in NLP and CV areas show up, it is more likely that surprisingly large amounts of training data is the key to generalization problem. However, it is hard to get pair-wise ground truth for so many low-level image restoration tasks, thus it is also valuable to study the possibility of training models to be better with proper data. But the insights claimed in this paper seem to be normal observations when conducting experiments. Most importantly, authors do not give any new solutions based on these observations like new loss functions. Actually, it is promising to design new loss functions in the spirit of isolation of removing rain and restoring background.
Questions
See the weaknesses above. It is interesting to see such study on low level vision which is underexplored, but still, the observations made in this paper have some problems. I disagree that the background quality can be the cost of removing more rain, which is against the purpose of image restoration, because usually a less complex method tends to remove both rain and background due to less capacity. The isolation of evaluation of ability of removing rain and restoring background also fails. Authors should reconsider their definition of reconstructing background. For deraining, the problematic areas are only rain regions rather than those unaffected background areas. The removal of rain and reconstruction of background take place in the same area while keeping the unaffected background unchanged is also important. Previous models years ago cannot handle them at the same time due to their inferior model capability and that's why recent models doing nothing to rain and keeping background tidy can get higher PSNR. For a better improvement of this paper, I recommend authors to try newly proposed powerful models and find out what the difference is. The perspective of model complexity is also important when discussing the generalization problem.
Rating
2: Strong Reject: For instance, a paper with major technical flaws, and/or poor evaluation, limited impact, poor reproducibility and mostly unaddressed ethical considerations.
Confidence
5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.
Soundness
3 good
Presentation
3 good
Contribution
2 fair
Limitations
NA.
Call for Further Discussion
Dear reviewer UxXt: We thank you for the precious review time and valuable comments. We have provided corresponding responses and results, which we believe have covered your concerns. We hope to further discuss with you whether or not your concerns have been addressed. Please let us know if you still have any unclear parts of our work. Best, Paper 798 Authors.
A Thanks Letter and Call for Discussion
Thanks to reviewers FkGF and w1Er for they have responded to our rebuttal. We thank the two reviewers for their recognition of our work. To avoid disturbing emails, we will not reply to your comments individually. Thanks for your time and effort! Reviewers UxXt and e5xd, For all your concerns, we have provided responses. Are there any deficiencies in our rebuttal? Whether the corresponding responses and results we provide cover your concerns? The discussion period ending date is August 21. Please let us know if you have any unsolved or other concerns. Once again we thank all reviewers and area chairs! Paper 798 author.
Last Chance for Discussion and A Summary
We thank all reviewers and area chairs for their valuable time and comments. The reviewer-author discussion deadline set by NeurIPS is drawing to a close. After discussing with reviewers and providing more clarifications/results/analyses, we would like to give a brief response. Reviewer FkGF, Reviewer w1Er all hold a positive side for our work. Reviewer w1Er commented "**this paper opens up a new perspective for the community**" and "**this will be an important pioneering work in the field of interpretable-based image deraining.**". Reviewer FkGF also commented "**This paper contributes some new insight to low-level vision interpretability research**". After the rebuttal, they raised the already positive paper score to “Strong Accept”. This speaks to their willingness to see the paper accepted. Reviewer e5xd holds negative rating at the first time. Reviewer e5xd contributed to the discussion, and we have also provided updated explanations. Hopefully our further discussion will address the reviewer's concerns. Reviewer UxXt also holds negative rating at the first time. But he\she lowered the rating **silently**, **without stating any reason**, and **without engaging in any discussion**. Not to mention we have sent several calls for discussions. We believe it is professional to participate in the discussion before the ending date and justify the rating. Once again we thank all reviewers and area chairs! Paper 798 author.
for figure 9, for the top image when check the wood board we could clear see it quite different from grouthtruth than other methods. for bottom images, when you check the top part of red bricks, middle of blue window, the eaves, it is also clearly worse than other methods. any reason why not compare with [32] and put a highly relevant paper to supplementary materials?
Response to Reviewer e5xd's Comment
Thanks for the reviewer's response. ```Figure 9```: We appreciate the astute observations by the reviewers about Figure 9. Upon revisiting, we do acknowledge the highlighted shortcomings. Given that our training utilized only 256 images without any specific techniques, the outcome, while expected, showcases our prowess in the removal of rain streaks. Indeed, while competing methodologies might sidestep these artifacts, they fall short in eliminating rain streaks effectively. However, we wish to emphasize that these results do not diminish the core contribution of our study. Our paper mainly focuses a counter-intuitive insight, shedding light on it from the lens of low-level vision interpretability. Such an understanding stands poised to reshape the community's grasp on the generalization challenge inherent to low-level vision. Our "implication" experiments provide a new possible way for real-world issues, revealing previously unknown results and potentials. We consciously chose not to present our method as a "novel method" but rather a deep dive into an intriguing phenomenon. We recognize various solutions exist to rectify the artifact issue -- from introducing post-processing modules to leveraging this outcome to steer another deraining network. However, introducing such augmentations might divert from our paper's central theme: a significant revelation about interpretability that holds paramount importance. We think whether the new model will perform better on the benchmark, and how much better it can be, is far less than the significance of our core interpretability findings contribution to the community. We still thank the reviewers for pointing out issues with our results. We will be honest about this flaw and highlight the issue of our results in this artifact in the section describing our results in the paper. Yet, we believe labeling our work as having "limited novelty" or worthy of "rejection" based on this alone might be an oversight. ```Compare with [32]```: We can, of course, provide comparison results with [32] if the reviewer insist. But we reiterate that our work's essence doesn't lie in benchmarking against state-of-the-art outcomes. As we have repeatedly emphasized, the main contribution of this paper is to demonstrate a counter-intuitive possibility and quantitatively analyze this counter-intuitive result from the perspective of interpretability. Our analysis will largely change the community's understanding of the generalization problem in low-level vision. We understand and respect the rain-removal community's pursuit of superior benchmark numbers and the significance of proposing novel and effective methods. But ignoring experiments that don't directly yield positive results or immediately inform algorithm development is concerning and worrisome. Although innovative works may face criticism from many sides, they are not inferior to methods-focused research. Our results show some unexpected findings. It wouldn't be good for our community to overlook them just because they're different from previous rain-removal studies that had top results. In response to the reviewer's specific request, current NIPS regulations prevent us from submitting a revised version with a comparison to [32]. But we'd love to include a direct comparison to [32] in the next version. We are confident in the enduring value of our work, no matter how many experiments are added. ```Thanks```: Once again, we're grateful for the feedback. Should there be further concerns needing clarification, please let us know. We're committed to ensuring our work's transparency and robustness.
Further Response to Reviewer e5xd's Comment
Dear Reviewer e5xd, We have responded to your question. We hope our latest response addresses your concerns. We realize that your main concern is paper contribution. But the other reviewers all acknowledged our contributions. Reviewer w1Er commented "**this paper opens up a new perspective for the community**" and "**this will be an important pioneering work in the field of interpretable-based image deraining**". Reviewer FkGF also commented "**This paper contributes some new insight to low-level vision interpretability research**". Even Reviewer UxXt acknowledges our contribution that "**It is interesting and valuable to investigate the reason**" and "**It is interesting to see such study on low level vision**". Hope you can consider our reply. Thanks for your efforts. Best, Paper 798 Authors.
Decision
Accept (poster)