Hierarchical Deep Multi-modal Network for Medical Visual Question Answering

Visual Question Answering in Medical domain (VQA-Med) plays an important role\nin providing medical assistance to the end-users. These users are expected to\nraise either a straightforward question with a Yes/No answer or a challenging\nquestion that requires a detailed and descriptive answer. The existing\ntechniques in VQA-Med fail to distinguish between the different question types\nsometimes complicates the simpler problems, or over-simplifies the complicated\nones. It is certainly true that for different question types, several distinct\nsystems can lead to confusion and discomfort for the end-users. To address this\nissue, we propose a hierarchical deep multi-modal network that analyzes and\nclassifies end-user questions/queries and then incorporates a query-specific\napproach for answer prediction. We refer our proposed approach as Hierarchical\nQuestion Segregation based Visual Question Answering, in short HQS-VQA. Our\ncontributions are three-fold, viz. firstly, we propose a question segregation\n(QS) technique for VQAMed; secondly, we integrate the QS model to the\nhierarchical deep multi-modal neural network to generate proper answers to the\nqueries related to medical images; and thirdly, we study the impact of QS in\nMedical-VQA by comparing the performance of the proposed model with QS and a\nmodel without QS. We evaluate the performance of our proposed model on two\nbenchmark datasets, viz. RAD and CLEF18. Experimental results show that our\nproposed HQS-VQA technique outperforms the baseline models with significant\nmargins. We also conduct a detailed quantitative and qualitative analysis of\nthe obtained results and discover potential causes of errors and their\nsolutions.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC