VISAGE-X: Interpretable Graph-Based and Explainable Scene Reasoning for Vision-Language Understanding

Current developments in vision-language models such as CLIP, BLIP, large multimodal transformers have facilitated simultaneous perception of images and text but these models mostly act like black box predictors that do not reason through to some degree. This is a weakness that limits their application in situations that involve explainability and systematic decision-making. This work introduces a multimodal reasoning system that is interpretable and composes both a scene graph representation and a graph-based relational reasoner and an explainability system into the actual inference. The visual scenes are broken down to object-centric representations and are coded into scene graphs that describe the objects and the spatial and semantic relationship among them. The attentional mechanisms of Graph Neural Networks propagate information of a relational nature, whereas the inference explicitly over visual and graph and textual representations is performed by a multi-hop reasoning controller. Gradient-based visual attribution, graph-level importance estimation, and trackable reasoning paths are used to explain why the models made the decisions they did during the inference process, and provides visual and textual insights as to why the decision was taken. Visual Genome, GQA, and e-SNLI-VE are tested through the framework on the metrics of task accuracy, multi-hop reasoning accuracy and explanation faithfulness. Findings show structured graph-based reasoning to be more interpretable and provide better reasoning transparency than perception-only benchmarks and competitive on tasks, and thus integrated explainable reasoning is effective in providing trustworthy vision-language understanding.

Paper

Full text

PDF

VISAGE-X: Interpretable Graph-Based and Explainable Scene Reasoning for Vision-Language Understanding

Semantic Scholar · 2026

Abstract

Current developments in vision-language models such as CLIP, BLIP, large multimodal transformers have facilitated simultaneous perception of images and text but these models mostly act like black box predictors that do not reason through to some degree. This is a weakness that limits their application in situations that involve explainability and systematic decision-making. This work introduces a multimodal reasoning system that is interpretable and composes both a scene graph representation and a graph-based relational reasoner and an explainability system into the actual inference. The visual scenes are broken down to object-centric representations and are coded into scene graphs that describe the objects and the spatial and semantic relationship among them. The attentional mechanisms of Graph Neural Networks propagate information of a relational nature, whereas the inference explicitly over visual and graph and textual representations is performed by a multi-hop reasoning controller. Gradient-based visual attribution, graph-level importance estimation, and trackable reasoning paths are used to explain why the models made the decisions they did during the inference process, and provides visual and textual insights as to why the decision was taken. Visual Genome, GQA, and e-SNLI-VE are tested through the framework on the metrics of task accuracy, multi-hop reasoning accuracy and explanation faithfulness. Findings show structured graph-based reasoning to be more interpretable and provide better reasoning transparency than perception-only benchmarks and competitive on tasks, and thus integrated explainable reasoning is effective in providing trustworthy vision-language understanding.

Similar papers

© 2026 NYSGPT2525 LLC