Reducing Language Biases in Visual Question Answering with Visually-Grounded Question Encoder
Recent studies have shown that current VQA models are heavily biased on the\nlanguage priors in the train set to answer the question, irrespective of the\nimage. E.g., overwhelmingly answer "what sport is" as "tennis" or "what color\nbanana" as "yellow." This behavior restricts them from real-world application\nscenarios. In this work, we propose a novel model-agnostic question encoder,\nVisually-Grounded Question Encoder (VGQE), for VQA that reduces this effect.\nVGQE utilizes both visual and language modalities equally while encoding the\nquestion. Hence the question representation itself gets sufficient\nvisual-grounding, and thus reduces the dependency of the model on the language\npriors. We demonstrate the effect of VGQE on three recent VQA models and\nachieve state-of-the-art results on the bias-sensitive split of the VQAv2\ndataset; VQA-CPv2. Further, unlike the existing bias-reduction techniques, on\nthe standard VQAv2 benchmark, our approach does not drop the accuracy; instead,\nit improves the performance.\n
Paper
References (47)
Scroll for more · 35 remaining