Multimodal Foundation Models for Early Disease Detection

Although healthcare data span EHRs, medical imaging, genomics, and wearable sensors, most diagnostic models continue to process these modalities in isolation, thereby limiting their ability to detect early cross-modal disease signatures. In this paper, we introduce a multimodal foundation model built on a transformer architecture that integrates heterogeneous clinical data through modality-specific encoders and cross-modal attention. In the model, each modality is mapped into a shared latent space and fused using multi-head attention with residual normalization. We implement the framework using a multimodal dataset that simulates early-stage disease patterns across EHR sequences, imaging patches, genomic profiles, and wearable signals, including missing-modality scenarios and label noise. The model is trained using supervised classification together with self-supervised reconstruction and contrastive alignment to improve robustness. Experimental evaluation demonstrates strong performance in early-detection settings, with stable classification metrics, reliable uncertainty estimates, and interpretable attention patterns. Our approach moves toward a foundation model that supports precision diagnostics, handles incomplete inputs, and improves early disease detection across oncology, cardiology, and neurology applications.

Paper

References (23)

Scroll for more · 11 remaining

Similar papers

© 2026 NYSGPT2525 LLC