Automated diagnostic report generation lies at the core of clinical diagnosis and can alleviate clinician shortages. Existing diagnostic report generation methods have two major limitations: they rely on unimodal inputs (e.g., images), ignoring textual biomarkers like medical history, and lack proactive dialogue capabilities to elicit personalized clinical information. To address those issues, we propose a ProActive Multimodal Agentic (PAMA) system, which performs comprehensive disease analysis by examining biomarkers in diverse sources, including multi-view medical images, medical histories, and diagnostic conversations. Built upon a knowledge graph and recommendation-based dialogue architecture, PAMA actively initiates adaptive, multi-turn conversations with patients, which is integrated with visual data for robust and reliable report generation. Specifically, PAMA actively generates adaptive multi-turn questions to collect clinically relevant background information, and then fuses the resulting dialogue context with visual representations for robust diagnostic report generation. We validate our approach on two real-world benchmark datasets, MIMIC-CXR and IU-Xray, through extensive quantitative evaluations and comparisons with state-of-the-art baselines. Furthermore, we conduct user studies to assess the realism and clinical quality of the generated reports. Finally, we present real-world case studies to examine the performance of our system across diverse scenarios, demonstrating its robustness on scalability, complexity, and data variability.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex