Robust Domain Generalization for Multi-modal Object Recognition

In multi-label classification, machine learning encounters the challenge of domain generalization when handling tasks with distributions differing from the training data. Existing approaches primarily focus on vision object recognition and neglect the integration of natural language. Recent advancements in vision-language pre-training leverages supervision from extensive visual-language pairs. This allows learning across diverse domains and enhances recognition in multi-modal scenarios, showcasing superior transfer learning capabilities in methods like CLIPood. However, CLIPood has several limitations: differences in the utilized loss, loss of generality in evaluating only a single backbone, and neglect of class-aware visual fusion.To address these, we propose this paper that infers the actual loss based on the implementation, broadens evaluations to larger vision-language backbones, and introduces Mixup-CLIPood with a novel mix-up loss for enhanced class-aware visual fusion.

Paper

References (25)

Scroll for more · 13 remaining

Similar papers

© 2026 NYSGPT2525 LLC