We propose a robust zero-shot medical diagnosis framework leveraging multimodal foundation models to identify novel diseases from text, images, and speech inputs. Using a pipeline based on CLIP and BioGPT, with evaluations on CheXpert and MIMIC-CXR datasets, we simulate unseen diseases and test the system's generalization ability. We perform a detailed performance comparison with three alternative models, conduct ablation studies, present embedding-based similarity modeling, and assess the system's scalability and efficiency. Results show that domain-specific prompt tuning and vision-language fusion significantly enhance accuracy, suggesting transformative potential in real-world low-resource healthcare environments.
Paper
Full text
Zero-Shot Medical Diagnosis Using Multimodal Foundation Models
Semantic Scholar · 2025
Abstract
We propose a robust zero-shot medical diagnosis framework leveraging multimodal foundation models to identify novel diseases from text, images, and speech inputs. Using a pipeline based on CLIP and BioGPT, with evaluations on CheXpert and MIMIC-CXR datasets, we simulate unseen diseases and test the system's generalization ability. We perform a detailed performance comparison with three alternative models, conduct ablation studies, present embedding-based similarity modeling, and assess the system's scalability and efficiency. Results show that domain-specific prompt tuning and vision-language fusion significantly enhance accuracy, suggesting transformative potential in real-world low-resource healthcare environments.