Multimodal diagnosis of Parkinson's disease with an internet-based collaborative agent architecture of medical language models
Parkinson's disease (PD) remains one of the most prevalent neurodegenerative disorders, where delays in diagnosis compromise therapeutic outcomes and increase healthcare costs. Conventional unimodal approaches, based on voice, sensors, or imaging, face critical limitations, including small datasets, lack of reproducibility, and high infrastructure demands. To address these challenges, the proposed multimodal agent-based architecture integrates medical language models, audio signals, and neuroimaging, and is supported by data-machine learning pipelines and an edge-cloud infrastructure. The system leverages ensemble learning, large and vision language models, and Retrieval-Augmented Generation (RAG) to enhance clinical decision support. The transparency of the model was supported by explainability techniques (SHapley Additive exPlanations, permutation importance, partial dependence, and individual conditional expectation), which highlighted the main audio and sensor variables responsible for the predictions. Experimental evaluation confirmed the effectiveness of multimodal fusion. When integrated, the architecture achieved robust performance, with an accuracy of 0.86, an F1-score above 0.88, ROC-AUC greater than 0.93, and both sensitivity and specificity above 0.89. Calibration and hypothesis tests were validated by a low Brier score of 0.205 and an Expected Calibration Error of 0.151, while Decision Curve Analysis confirmed clinical relevance by minimizing false negatives, critical for early screening, and reducing redundant interventions. Multimodal fusion produced accurate, well-calibrated, and interpretable risk estimates for PD screening; larger prospective studies and cost-effectiveness analyses are needed to consolidate clinical applicability.
Paper
Full text
Multimodal diagnosis of Parkinson's disease with an internet-based collaborative agent architecture of medical language models
Semantic Scholar · Medicine · 2026
Abstract
Parkinson's disease (PD) remains one of the most prevalent neurodegenerative disorders, where delays in diagnosis compromise therapeutic outcomes and increase healthcare costs. Conventional unimodal approaches, based on voice, sensors, or imaging, face critical limitations, including small datasets, lack of reproducibility, and high infrastructure demands. To address these challenges, the proposed multimodal agent-based architecture integrates medical language models, audio signals, and neuroimaging, and is supported by data-machine learning pipelines and an edge-cloud infrastructure. The system leverages ensemble learning, large and vision language models, and Retrieval-Augmented Generation (RAG) to enhance clinical decision support. The transparency of the model was supported by explainability techniques (SHapley Additive exPlanations, permutation importance, partial dependence, and individual conditional expectation), which highlighted the main audio and sensor variables responsible for the predictions. Experimental evaluation confirmed the effectiveness of multimodal fusion. When integrated, the architecture achieved robust performance, with an accuracy of 0.86, an F1-score above 0.88, ROC-AUC greater than 0.93, and both sensitivity and specificity above 0.89. Calibration and hypothesis tests were validated by a low Brier score of 0.205 and an Expected Calibration Error of 0.151, while Decision Curve Analysis confirmed clinical relevance by minimizing false negatives, critical for early screening, and reducing redundant interventions. Multimodal fusion produced accurate, well-calibrated, and interpretable risk estimates for PD screening; larger prospective studies and cost-effectiveness analyses are needed to consolidate clinical applicability.