In multi-label classification, machine learning encounters the challenge of domain generalization when handling tasks with distributions differing from the training data. Existing approaches primarily focus on vision object recognition and neglect the integration of natural language. Recent advancements in vision-language pre-training leverages supervision from extensive visual-language pairs. This allows learning across diverse domains and enhances recognition in multi-modal scenarios, showcasing superior transfer learning capabilities in methods like CLIPood. However, CLIPood has several limitations: differences in the utilized loss, loss of generality in evaluating only a single backbone, and neglect of class-aware visual fusion.To address these, we propose this paper that infers the actual loss based on the implementation, broadens evaluations to larger vision-language backbones, and introduces Mixup-CLIPood with a novel mix-up loss for enhanced class-aware visual fusion.
Paper
References (25)
Scroll for more · 13 remaining