Found a Reason for me? Weakly-supervised Grounded Visual Question Answering using Capsules

The problem of grounding VQA tasks has seen an increased attention in the\nresearch community recently, with most attempts usually focusing on solving\nthis task by using pretrained object detectors. However, pre-trained object\ndetectors require bounding box annotations for detecting relevant objects in\nthe vocabulary, which may not always be feasible for real-life large-scale\napplications. In this paper, we focus on a more relaxed setting: the grounding\nof relevant visual entities in a weakly supervised manner by training on the\nVQA task alone. To address this problem, we propose a visual capsule module\nwith a query-based selection mechanism of capsule features, that allows the\nmodel to focus on relevant regions based on the textual cues about visual\ninformation in the question. We show that integrating the proposed capsule\nmodule in existing VQA systems significantly improves their performance on the\nweakly supervised grounding task. Overall, we demonstrate the effectiveness of\nour approach on two state-of-the-art VQA systems, stacked NMN and MAC, on the\nCLEVR-Answers benchmark, our new evaluation set based on CLEVR scenes with\nground truth bounding boxes for objects that are relevant for the correct\nanswer, as well as on GQA, a real world VQA dataset with compositional\nquestions. We show that the systems with the proposed capsule module\nconsistently outperform the respective baseline systems in terms of answer\ngrounding, while achieving comparable performance on VQA task.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC