Multimodal graphs, which integrate diverse multimodal features and relations, are ubiquitous in real-world applications. However, existing multimodal graph learning methods are typically trained from scratch for specific graph data and tasks, failing to generalize across various multimodal graph data and tasks. To bridge this gap, we explore the potential of multimodal graph large language models (MG-LLM) to unify and generalize across diverse multimodal graph data and tasks. We propose a unified framework of multimodal graph data, tasks, and models, discovering the inherent multi-granularity and multi-scale characteristics in multimodal graphs. Specifically, we present five key desired characteristics for MG-LLM: (1) unified space for multimodal structures and attributes, (2) capability of handling diverse multimodal graph tasks, (3) multimodal graph in-context learning, (4) multimodal graph interaction with natural language, and (5) multimodal graph reasoning. We then elaborate on the key challenges, review existing literature, and highlight promising future research directions towards realizing these ambitious characteristics. Finally, we summarize existing multimodal graph datasets pertinent for model training. We believe this paper can contribute to the ongoing advancement of the research towards MG-LLM for generalization across multimodal graph data and tasks.
Paper
References (100)
Scroll for more · 38 remaining