Few-Shot Multimodal Medical Imaging: A Theoretical Framework

Medical imaging often faces the challenge of limited annotated data, especially when it comes to rare diseases in lowresource clinical settings. While different kinds of multimodal and meta-learners have been shown to improve performance in such scenarios, they often lack strict theoretical foundations that explain the conditions under which they work. In this work, we develop a comprehensive theoretical framework for few-shot multimodal medical imaging that simultaneously characterises the complexity of the sample, quantifies the uncertainty, as well as the interpretability. By leveraging PAC learning theory, Vapnik-Chervonenkis (VC) dimension analysis, and PACBayesian bounds, we derive lower bounds on the required number of labeled samples for reliable performance and show that different modalities of data complementing each other attenuate effective capacity through an information gain term. Furthermore, we suggest a formal metric of explanation stability and prove that the variance of explanations falls off at the rate of $1 / n$. To empirically validate the theoretical arguments above, we build a multimodal dataset and test an additive convolutional neural network with multilayer perceptron fusion in few-shot regimes. The results not only support the predicted multimodal performance improvements, but also show modality interference when the sample size is larger and corresponding reduction in predictive uncertainty. Collectively, our framework provides a principled foundation for the design of data-efficient, uncertainty-aware and interpretable diagnostic models for lowresource environments.

Paper

References (42)

Scroll for more · 30 remaining

Similar papers

© 2026 NYSGPT2525 LLC