Adapting the CLIPMem metric & Experiments on the large-scale YFCC100M dataset
We thank the reviewer for the detailed comments. We are glad that the reviewer recognizes our work as addressing “a gap in previous research” and appreciates that it “provides actionable insights,” “challenges established norms,” and “can benefit real-world multi-modal model training practices.” Below we address all of the points and questions raised by the reviewer one by one.
>**While tailored to CLIP, the metric and findings may need adaptation to apply effectively to other multi-modal models with different architectures. (...) Would you please comment on how the metric adapts to other multi-modal models besides CLIP?**
We tested our metric on the popular CLIP model, but many other multi-modal models follow the CLIP architecture, and our metric is immediately applicable to them as well. For example, multi-modal models with separate encoders and contrastive learning objectives can directly apply CLIPMem with minimal modifications (e.g., ALIGN [1], Florence [2], and LiT [3]), where memorization can be measured by evaluating the alignment scores between representations. For models with additional components other than contrastive alignment, CLIPMem can be applied after alignment before other operations like fusion (ALBEF [4]) or generative tasks (BLIP [5]). By doing so, CLIPMem can isolate and quantify memorization during alignment, being adaptable across different architectures.
>**The experiments focus on datasets like COCO and CC3M, so it’s unclear how well these findings generalize to other large-scale or domain-specific datasets.**
To address the reviewer’s comment, we added additional experiments with the larger YFCC100M dataset. To simulate the large data regime, we trained the model for one epoch, and then evaluate memorization. To ensure comparability to our initial experiments where we trained the model on 70,000 data points for 100 epochs (i.e., 7M samples seen during training), we trained with 7M samples from YFCC100M.
To fit our setup, we trained model f with 6950000 shared +50000 candidate samples and model g with 6950000 shared + 50000 independent samples. We observe that there is still significant memorization when training for one epoch.
| **Model, Epochs** | **CLIPMem** | **Lin. Prob. Acc. (ImageNet)** |
|:------------------------:|:-----------:|:------------------------------:|
| YFCC 7M, 1 Epoch | 0.425 | 64.83% ± 1.04% |
| Paper (Coco), 100 Epochs | 0.438 | 63.11% ± 0.91% |
We also assessed the memorized samples qualitatively, as shown in the new Figure 9 in the updated paper. They highlight that the insights remain the same even when training only for one epoch: atypical and miscaptioned samples are memorized.
**References:**
[1] “Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision” Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, Tom Duerig. ICML 2021.
[2] “Florence: A New Foundation Model for Computer Vision” Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, Ce Liu, Mengchen Liu, Zicheng Liu, Yumao Lu, Yu Shi, Lijuan Wang, Jianfeng Wang, Bin Xiao, Zhen Xiao, Jianwei Yang, Michael Zeng, Luowei Zhou, Pengchuan Zhang.
arXiv preprint arXiv:2111.11432 (2021).
[3] “LiT: Zero-Shot Transfer with Locked-image text Tuning” Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, Lucas Beyer. CVPR 2022.
[4] “Align before Fuse: Vision and Language Representation Learning with Momentum Distillation” Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare, Shafiq Joty, Caiming Xiong, Steven Hoi. NeurIPS 2021.
[5] “BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation” Junnan Li, Dongxu Li, Caiming Xiong, Steven Hoi. ICML 2022.