Graphhopper: Multi-Hop Scene Graph Reasoning for Visual Question Answering

Visual Question Answering (VQA) is concerned with answering free-form\nquestions about an image. Since it requires a deep semantic and linguistic\nunderstanding of the question and the ability to associate it with various\nobjects that are present in the image, it is an ambitious task and requires\nmulti-modal reasoning from both computer vision and natural language\nprocessing. We propose Graphhopper, a novel method that approaches the task by\nintegrating knowledge graph reasoning, computer vision, and natural language\nprocessing techniques. Concretely, our method is based on performing\ncontext-driven, sequential reasoning based on the scene entities and their\nsemantic and spatial relationships. As a first step, we derive a scene graph\nthat describes the objects in the image, as well as their attributes and their\nmutual relationships. Subsequently, a reinforcement learning agent is trained\nto autonomously navigate in a multi-hop manner over the extracted scene graph\nto generate reasoning paths, which are the basis for deriving answers. We\nconduct an experimental study on the challenging dataset GQA, based on both\nmanually curated and automatically generated scene graphs. Our results show\nthat we keep up with a human performance on manually curated scene graphs.\nMoreover, we find that Graphhopper outperforms another state-of-the-art scene\ngraph reasoning model on both manually curated and automatically generated\nscene graphs by a significant margin.\n

Paper

References (40)

Scroll for more · 28 remaining

Similar papers

© 2026 NYSGPT2525 LLC