Author response to Reviewer dhjM
We would like to thank the reviewer for acknowledging our effort and we are encouraged that our responses have addressed most of your concerns. Regarding your last concern about the fairness of our experiments with traditional UDA methods, we have examined DANN and CDAN using a CLIP backbone to provide more insight into the limitations of domain-invariant feature learning in adapting pretrained CLIP to new domains. We will run more experiments on other datasets and incorporate obtained results to the revision. It is important to note that our baseline results, so far, are taken directly from prior work, as we adhere to the recent protocols for adapting CLIP from prior work.
| Method | Backbone | →C | →I | →P | Average | # Param |
|-|-|-|-|-|-|-|
| SOURCE ONLY | RN50-FC | 94.7 | 90.2 | 79.3 | 88.1 | 38M |
| DANN | RN50-FC | 96.0 | 92.3 | 80.2 | 89.5 | 38M |
| CDAN | RN50-FC | 96.2 | 92.0 | 80.4 | 89.5 | 38M |
| SOURCE ONLY | RN50 | 93.3 | 88.5 | 78.8 | 86.9 | 38M |
| DANN | RN50 | 94.5 | 91.5 | 79.0 | 88.3 | 38M |
| CDAN | RN50 | 94.7 | 92.0 | 79.0 | 88.6 | 38M |
| SOURCE ONLY | CLIP-RN50 | 93.0 | 90.7 | 78.7 | 87.4 | 102M |
| DANN | CLIP-RN50 | 95.0 | 91.7 | 79.2 | 88.6 | 102M |
| CDAN | CLIP-RN50 | 93.7 | 93.0 | 80.0 | 88.9 | 102M |
| DAN | ResNet50 | 93.3 | 92.2 | 77.6 | 87.7 | 48.9M |
| D-CORAL | ResNet50 | 93.6 | 91.7 | 77.1 | 87.5 | 47.5M |
| DANN | ResNet50 | 95.7 | 91.8 | 77.9 | 87.8 | 48.9M |
| PGA | Prompt-tuning| 96.8 | 95.7 | 84.6 | 92.4 | 114k |
| MPGA | Prompt-tuning| 97.4 | 96.5 | 84.7 | 92.9 | 131k |
Specifically, we applied DANN and CDAN on CLIP’s ResNet50 configured in three different ways: with a randomly initialized fully connected classifier (RN50-FC), with a frozen text encoder (RN50), and using the entire CLIP backbone (CLIP-RN50). The results, presented in the table below, demonstrate that appropriately adapting prior UDA methods to different parts of the CLIP model can yield better results compared to traditional methods on a ResNet backbone (e.g., DAN, CORAL, DANN). Eventhough, PGA and MPGA still exhibit superior performance, even with significantly fewer parameters finetuned. We hypothesize that relying solely on source classification loss and another objective for invariant feature learning can degrade CLIP’s rich semantic representation [R8, R9, R10, R11], which is crucial for predicting target domain data. To counteract this, utilizing target pseudo data (similar to self-training baseline in Table 1) or adopting a more carefully-designed optimization procedure that better leverages information from both source and target data—similar to our proposed method—could enhance performance.
We hope this response provides further insights into why traditional UDA methods may not perform as well as other prompt-based baselines. If the reviewer found any further unaddressed concerns, we are always happy to provide further clarifications and improve our work based on the constructive feedback from the reviewers.
[R8] Kumar, Ananya, et al. "Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution." ICLR 22.
[R9] Zheng, Zangwei, et al. "Preventing zero-shot transfer degradation in continual learning of vision-language models." ICCV 23.
[R10] Lai, Zhengfeng, et al. "Padclip: Pseudo-labeling with adaptive debiasing in clip for unsupervised domain adaptation." ICCV 23.
[R11] Ding, Yuxuan, et al. "Don't stop learning: Towards continual learning for the clip model." arXiv 22.