Multi-Agent Reinforcement Learning for Multi-Hop Reasoning in Multi-Modal Knowledge Graphs

Multi-hop reasoning over knowledge graphs (KGs) has become essential for various artificial intelligence applications, including intelligent question answering, medical diagnosis, and semantic search. Traditional methods, particularly embedding-based approaches like TransE and TransH, excel in simple, single-hop tasks but struggle in complex multi-hop reasoning scenarios. These approaches face significant challenges such as error accumulation across hops, limited representational capacity, and the inability to integrate multi-modal data effectively. To address these limitations, we propose a novel multi-agent reinforcement learning (MARL) framework that incorporates multi-modal knowledge graphs for improved multi-hop reasoning. The proposed system leverages three types of agents, each handling a different modality of data: structural, categorical, and descriptive. These agents collaboratively explore reasoning paths, using a reverse hyperplane projection technique to fuse multimodal data, ensuring structural information remains dominant while benefiting from the complementary insights of categorical and descriptive data. A collaborative reasoning reward function, incorporating path correctness, collaboration quality, reasoning efficiency, and multi-modal fusion quality, is designed to guide agents toward optimal reasoning paths. Experimental results demonstrate that the proposed framework through experiments on the Unified Medical Language System (UMLS) achieves notable performance improvements in multi-hop inference tasks and indicate that multimodal collaboration enables agents to narrow down plausible targets in dense biomedical neighborhoods. This approach provides a more accurate, interpretable, and efficient solution for complex knowledge graph reasoning, with potential applications in diverse domains such as healthcare, AI-driven question answering, and semantic search.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC