Summary
This paper introduces the Reversible Decoupling Network (RDNet), a novel approach to single-image reflection removal that overcomes key limitations in existing methods. RDNet features a multi-column reversible encoder that preserves hierarchical semantic information by decoupling transmission- and reflection-related features, thus preventing information loss during feature interactions across scales. Additionally, an adaptive transmission-rate-aware prompt generator dynamically adjusts features by learning channel scaling factors from data, which enhances RDNet's generalization and robustness in real-world reflection scenarios.
Strengths
1.This work presents RDNet, which incorporates a multi-column reversible encoder to effectively preserve multi-scale semantic information. By decoupling transmission and reflection features, RDNet minimizes information loss during feature interactions, significantly enhancing reflection removal accuracy.
2.The proposed transmission-rate-aware prompt generator dynamically adjusts feature representations by learning channel scaling factors from the data. This design allows RDNet to achieve strong generalization and robustness across various real-world reflection scenarios.
3.The proposed method outperforms state-of-the-art methods in reflection removal both qualitatively and quantitatively.
Weaknesses
1.The Bidirectional Interaction Level in Figure 2 could benefit from clearer explanation, as the current description in the text is brief and may lead to misunderstandings.
2.The paper compares the proposed method with several existing reflection removal techniques. However, distinct datasets are used in the comparative experiments, which is unnecessary. Additionally, expanding the range of comparison methods would ensure a more comprehensive evaluation, as the current selection may not sufficiently demonstrate the method's effectiveness.
Questions
1.The methods section is not clear enough, for example, line 243 “Bidirectional Interaction Level”, but there is no further explanation and no reference to related work. What is the motivation for using the Bidirectional Interaction Level?
2.It would also be helpful to see results on a unified dataset, including visualizations. Could the authors provide results for the specified methods under consistent training and testing conditions? Such as:
a. Johnson, Justin, Alexandre Alahi, and Li Fei-Fei. "Perceptual losses for real-time style transfer and super-resolution." Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14. Springer International Publishing, 2016.
b.Wen, Qiang, et al. "Single image reflection removal beyond linearity." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019.
c.Kim, Soomin, Yuchi Huo, and Sung-Eui Yoon. "Single image reflection removal with physically-based training images." Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020.
c.Dong, Zheng, et al. "Location-aware single image reflection removal." Proceedings of the IEEE/CVF international conference on computer vision. 2021.
d.Song, Zhenbo, et al. "Robust single image reflection removal against adversarial attacks." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023.
e.Wang, Mengyi, et al. "Personalized single image reflection removal network through adaptive cascade refinement." Proceedings of the 31st ACM International Conference on Multimedia. 2023.