Zero-Shot Medical Diagnosis Using Multimodal Foundation Models

We propose a robust zero-shot medical diagnosis framework leveraging multimodal foundation models to identify novel diseases from text, images, and speech inputs. Using a pipeline based on CLIP and BioGPT, with evaluations on CheXpert and MIMIC-CXR datasets, we simulate unseen diseases and test the system's generalization ability. We perform a detailed performance comparison with three alternative models, conduct ablation studies, present embedding-based similarity modeling, and assess the system's scalability and efficiency. Results show that domain-specific prompt tuning and vision-language fusion significantly enhance accuracy, suggesting transformative potential in real-world low-resource healthcare environments.

Paper

Full text

PDF

Zero-Shot Medical Diagnosis Using Multimodal Foundation Models

Semantic Scholar · 2025

Abstract

We propose a robust zero-shot medical diagnosis framework leveraging multimodal foundation models to identify novel diseases from text, images, and speech inputs. Using a pipeline based on CLIP and BioGPT, with evaluations on CheXpert and MIMIC-CXR datasets, we simulate unseen diseases and test the system's generalization ability. We perform a detailed performance comparison with three alternative models, conduct ablation studies, present embedding-based similarity modeling, and assess the system's scalability and efficiency. Results show that domain-specific prompt tuning and vision-language fusion significantly enhance accuracy, suggesting transformative potential in real-world low-resource healthcare environments.

Similar papers

© 2026 NYSGPT2525 LLC