Cross-modal Active Complementary Learning with Self-refining Correspondence

Recently, image-text matching has attracted more and more attention from academia and industry, which is fundamental to understanding the latent correspondence across visual and textual modalities. However, most existing methods implicitly assume the training pairs are well-aligned while ignoring the ubiquitous annotation noise, a.k.a noisy correspondence (NC), thereby inevitably leading to a performance drop. Although some methods attempt to address such noise, they still face two challenging problems: excessive memorizing/overfitting and unreliable correction for NC, especially under high noise. To address the two problems, we propose a generalized Cross-modal Robust Complementary Learning framework (CRCL), which benefits from a novel Active Complementary Loss (ACL) and an efficient Self-refining Correspondence Correction (SCC) to improve the robustness of existing methods. Specifically, ACL exploits active and complementary learning losses to reduce the risk of providing erroneous supervision, leading to theoretically and experimentally demonstrated robustness against NC. SCC utilizes multiple self-refining processes with momentum correction to enlarge the receptive field for correcting correspondences, thereby alleviating error accumulation and achieving accurate and stable corrections. We carry out extensive experiments on three image-text benchmarks, i.e., Flickr30K, MS-COCO, and CC152K, to verify the superior robustness of our CRCL against synthetic and real-world noisy correspondences.

Paper

References (66)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer DZ4B7/10 · confidence 5/52023-07-02

Summary

This paper tackles a new challenge in image-text matching, namely, noisy correspondence, which refers to the mismatched image-text pairs that can mislead the model to learn incorrect cross-modal associations during training, resulting in a suboptimal cross-modal model that computes inaccurate similarities for retrieval. To address this challenge, the authors propose a general cross-modal robust complementary learning framework CRCL that can enhance the robustness and performance of existing image-text matching methods. The authors provide theoretical and empirical evidence that their method is effective and achieves state-of-the-art performance under the same settings compared with the existing robust baselines, such as NCR, DECL, and BiCro.

Strengths

1. This paper compares with a comprehensive and fair set of baselines, including the state-of-the-art robust method BiCro (CVPR’23). Moreover, the paper also includes CCR&CCS (WACV’23) and RCL (TPAMI’23) in the appendix, which is commendable. 2. This paper provides appropriate theoretical analysis in the appendix, which is satisfactory. The authors cleverly adapt the robust theory of the noisy label problem and views NC as an instance-level category noise label problem. The experimental results and the theoretical analysis are consistent and supportive of each other. 3. This paper shows promising performance on both synthetic NC and real NC datasets. Especially under high noise levels, CRCL demonstrates a strong robustness against NC. 4. From Figure 1, it is evident that SCC is effective in alleviating the accumulation of noise errors and improving the accuracy of rectified correspondence labels.

Weaknesses

1. There is a typo in lines 48-49: CSRL should be CRCL. 2. This paper verifies the generality of CRCL on three standard methods, i.e., VSE$\infty$, SGR, and SAF. How about the training cost of these methods, e.g., compared to the existing robust frameworks BiCro, NCR, and DECL? 3. The paper seems to lack a discussion of related works in the main text, although I found it in the appendix. I understand that it may be due to the space limitation. However, I think CRCL is a significant improvement over NCR (NeurIPS’2021) for the problem of noisy correspondence. Therefore, for the readers who are not familiar with NCR or noisy correspondence, it would be more helpful to discuss related work in the main text. 4. Is it possible to add Plausible-Match R-Precision (PMRP)[1] as an evaluation metric, since it may be more suitable for many-to-many matching evaluation? Although I believe that Recall can reflect most of the retrieval performance, PMRP may be more appealing. 5. I think SCC is a feasible and effective technique. What is the intuition behind it? Can the authors provide more explanation? [1] Chun S, Oh S J, De Rezende R S, et al. Probabilistic embeddings for cross-modal retrieval[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021: 8415-8424.

Questions

See Weaknesses.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

3 good

Contribution

4 excellent

Limitations

The authors discussed the limitations and Broader Impacts of CRCL in section 5.

Reviewer tALj6/10 · confidence 5/52023-07-05

Summary

This paper tackles a latent challenge in image-text matching, which is the presence of noisy correspondences between images and texts. The paper introduces a general framework that combines a robust loss function and a correspondence correction technique to enhance the existing models’ ability to cope with noisy correspondences. The paper shows that the proposed CRCL framework outperforms the state-of-the-art methods. The paper also provides both experimental and theoretical evidence to support the effectiveness of the proposed algorithms.

Strengths

Significance: Noisy correspondences are a common issue in image-text data, which can arise from alignment errors or weak cross-modal information. Studying how to deal with noisy correspondences can enable more applications of image-text learning, where the alignments are challenging and require domain expertise. This work makes significant contributions both theoretically and empirically, and offers a new perspective for future research on noisy correspondences and related fields. Clarity: This paper is well-written and clear, and the motivation and method are easy to follow. The technical proofs seem to be sound. However, there are some typos, undefined terms, and missing references in the paper. Please see Weaknesses and Questions below. Originality: The techniques used in this paper seem to be novel in the image-text matching field, and there are some technical innovations. Quality: I think the quality of this paper is high, and I did not find any major flaws. The authors conduct comprehensive experiments to demonstrate the effectiveness of the proposed method.

Weaknesses

- The related work section is incomplete. The authors should include and compare with some recent works on image-text matching, such as [1,2,3]. - Some typos: CRCL in line 48. - Eq.1 is confusing. I suggest rewriting it as $1 - y_{ik}$, with probability $\bar{\eta}_{ik}, \forall k\neq j$. - Eq.11 is not clear. The authors should explain the meaning and derivation of each term. [1] Huang Y, Wang Y, Zeng Y, et al. MACK: multimodal aligned conceptual knowledge for unpaired image-text matching[J]. Advances in Neural Information Processing Systems, 2022, 35: 7892-7904.\ [2] Goel S, Bansal H, Bhatia S, et al. Cyclip: Cyclic contrastive language-image pretraining[J]. Advances in Neural Information Processing Systems, 2022, 35: 6704-6719.\ [3] Pan Z, Wu F, Zhang B. Fine-Grained Image-Text Matching by Cross-Modal Hard Aligning Network[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023: 19275-19284.

Questions

- The authors should explain why their results are different from the previous works NCR and DECL. Are there any special settings or hyperparameters that affect the performance? The authors should also report the training efficiency of their method, such as the training time and the computational resources, and compare with the existing methods, such as DECL. - The authors should justify their choice of using MS-COCO and Flickr30K datasets, which have one-to-many correspondences (1 to 5), while their problem formulation assumes one-to-one correspondences. How does this affect the validity and applicability of their method? - The authors claim that ACL is a general framework that can be applied to any existing model. However, they do not provide any empirical evidence to support this claim under other tasks. For example, can ACL handle noisy labels? The authors should provide more experimental support for the generality of ACL. - For Eq.11, it is not clear how the initial labels are obtained. Is it based on the first epoch of the first SR? The authors should clarify this point and explain how they initialize the labels. - What does active loss mean? Is it the same as normalized loss [1]? The authors should define this term clearly. [1] Ma X, Huang H, Wang Y, et al. Normalized loss functions for deep learning with noisy labels[C]//International conference on machine learning. PMLR, 2020: 6543-6553.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

The authors have discussed the potential limitations and implications of their work.

Reviewer QLFm8/10 · confidence 5/52023-07-05

Summary

This paper presents a novel framework (CRCL) for cross-modal correspondence learning that can handle noisy image-text pairs. The key idea of CRCL is to use a complementary active loss (ACL) that balances between discriminative learning and robust learning. ACL leverages the rectified correspondence labels to adjust the loss function and assign active loss weights, which means that the potentially noisy pairs are more focused on robust learning. To obtain accurate rectified correspondence labels, the paper proposes a self-refining correspondence correction technique (SCC) that estimates the correlation between modalities and avoids error accumulation. The proposed framework is simple yet effective in mitigating the negative impact of noisy training pairs and achieving robust image-text matching. The paper also conducts extensive experiments on three benchmark datasets to demonstrate the effectiveness and rationality of the proposed method.

Strengths

1. The CRCL framework seems to be a general and flexible approach that can be easily integrated with existing methods to enhance the robustness of image-text matching. 2. This paper is well-written and easy to follow. The authors provide clear explanations and motivations for their method. 3. The authors simplify the training process by discarding the co-teaching scheme used in NCR and BiCro, and instead focus on the loss function and the correction technique to make the method concise and clear. 4. The experimental results are impressive and convincing. It is also noteworthy that CRCL can improve the performance of the pre-trained model, e.g., CLIP, as shown in Table 3.

Weaknesses

1. In Eq.3, there is a typo: $i_j$ should be $I_j$. 2. In Eq.11, some symbols are not clearly defined, e.g., $\hat{p}^(j,t-1)(*)$. Please explain their meanings and notations. 3. The NC problem can be viewed as a special case of the PMP problem. RCL[20] uses complementary contrastive learning (CCL) to deal with PMPs. Similarly, CRCL can also be regarded as a further extension and study of CCL. However, the paper does not provide a direct comparison between CRCL and RCL. It would be interesting to see how they differ in performance and insights. 4. In the supplementary material, there is a capitalization error: “For brevity” should be “for brevity”. 5. Why not include MSCN[A] as a baseline? MSCN seems to be from the same period as BiCro (CVPR’23) and also addresses the NC problem. 6. Some related works are missing: [B-C]. These papers also propose methods for image-text matching and could be relevant for comparison or discussion. [A] Noisy Correspondence Learning With Meta Similarity Correction, CVPR, 2023. [B] Learning Semantic Relationship Among Instances for Image-Text Matching, CVPR, 2023. [C] Fine-Grained Image-Text Matching by Cross-Modal Hard Aligning Network, CVPR, 2023.

Questions

This paper is well-written and easy to follow. It has all the essential elements of a good paper and I do not have any major objections to accept it. I appreciate the authors’ approach, which is more concise and theoretically sound than the existing robust image-text matching methods. However, there are some minor issues that need to be addressed, such as some typos and some recent related works that need to be further discussed. Noisy correspondence is an important research direction in the field of multimodality, as it is similar to noisy label learning. I hope the authors can not only improve the performance of their method, but also explore the potential impact of CRCL on more correspondence learning tasks.

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

4 excellent

Presentation

3 good

Contribution

3 good

Limitations

The authors should enrich further, as described in the Questions section.

Reviewer 8spS6/10 · confidence 4/52023-07-05

Summary

This manuscript focuses on image-text matching under the noisy correspondence setting. To achieve a noise robust multi-modal representation, the authors propose two components, including a Active Complementary Loss (ACL) and a Self-Refining Correspondence Correction (SRCC). In ACL, a complementary contrastive loss styled fomula is derived, in which a coefficient $q$ is set in seeking for a tighter bound between the divergence between risk of training with noisy correspondence and ideal setting. As for SRCC, the labeled matching score is relaxed by momentum updating to alleviate the noise. Finally, the authors conduct extensive experiments on image-text retrieval to show their performance.

Strengths

1. The motivation of the manuscript is novel, which is from noise tolerance loss function designing and noisy label correction simultaneously. 2. The proposed Active Complementary Loss has rigorous theoretical proof in evidenting the tighter bound to the divergence between noisy risk with ideal risk. 3. The experiments are adequate and extensive. What' s more, as the Tab. 1 shown, the proposed method is very stable under different noisy ratios, and the improvment is non-trivial under larger noisy ratio.

Weaknesses

1. Lacks necessary explaination in figures. Actually, the pure text claim to the proposed method could be harder to grab. 2. The Sec. 2.2 is complex and hard to follow. Beyond theoretical proof, how ACL works in intuition should be further discussed. 3. The existing organization to the proposed two components is poor. Actually, I didn' t see a clear connection between Sec. 2.2 with Sec. 2.3. Despite that they are all for noise correspondence, the author fail to discuss the relationship between ACL and SRCC. If not, I could treat this manuscript as a naive A+B technic.

Questions

None

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

None

Reviewer J1KY5/10 · confidence 4/52023-07-07

Summary

This paper focuses on the problem of noise correspondence in image-text matching tasks. To address this issue, this paper proposes a generalized cross-modal robust complementary learning framework, which not only reduces the risk of erroneous supervision from the active complementary loss but also obtains stable and accurate correspondence correction through a self-refining correspondence correction. Extensive experiments show that the developed model could significantly improve the effectiveness and robustness compared with the state-of-the-art approaches. The topic of this paper is of great practical interest and the motivation is clear.

Strengths

a. This paper proposes a novel generalized cross-modal robust complementary learning framework to address the noisy correspondence problem in image-text matching tasks, which enhances the effectiveness of existing methods through robust loss and correction techniques. b. This paper utilizes the active complementary loss that employs complementary pairs to conduct indirect cross-modal learning with exponential normalization to boost robustness against noisy correspondence. In addition, self-refining correspondence correction is proposed to obtain stable and accurate correspondence correction. c. Extensive experiments are provided to prove the effectiveness and robustness of the proposed model.

Weaknesses

a. Essentially, the authors propose a new loss function to solve the noise correspondence problem in the image-text matching task. The proposed active complementary loss and self-refining correspondence correction both serve the final loss. So I think the novelty may be limited. b. I noticed in the supplementary that when the framework proposed in this paper is applied to the VSE model, there is a leap in performance improvement (such as 60% noise, Flickr30K dataset, The bidirectional retrieval improved by 50.3% and 35.4% on R@1, respectively.) while the performance improvement was limited when applied to the SGRAF model. Therefore, I think the author should explain why the performance gap is so large. More extensive experiments should also be organized to demonstrate the robustness of the proposed framework. c. To my knowledge, in previous studies, the noise injected on the Flickr30K and MS COCO datasets was different. So how is the noise injected by the author generated?

Questions

Please refer to the item "weaknesses".

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

In addition to the effectiveness of the experimental results, I suggest that the authors can apply the proposed framework in more methods to enhance its robustness.

Reviewer DZ4B2023-08-19

I have carefully read the authors’ rebuttal and the other reviewers’ comments. I think the authors have satisfactorily addressed all my concerns and improved the quality of their paper. Therefore, I maintain my positive score for this paper.

Authorsrebuttal2023-08-19

We are very grateful for your feedback and positive recognition of our work. Your constructive comments have greatly enhanced the quality of our paper. We will revise our paper according to the reviews in the next version. Thank you again for your review and valuable time!

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC