Understanding Knowledge Gaps in Visual Question Answering: Implications for Gap Identification and Testing

Visual Question Answering (VQA) systems are tasked with answering natural\nlanguage questions corresponding to a presented image. Traditional VQA datasets\ntypically contain questions related to the spatial information of objects,\nobject attributes, or general scene questions. Recently, researchers have\nrecognized the need to improve the balance of such datasets to reduce the\nsystem's dependency on memorized linguistic features and statistical biases,\nwhile aiming for enhanced visual understanding. However, it is unclear whether\nany latent patterns exist to quantify and explain these failures. As an initial\nstep towards better quantifying our understanding of the performance of VQA\nmodels, we use a taxonomy of Knowledge Gaps (KGs) to tag questions with one or\nmore types of KGs. Each Knowledge Gap (KG) describes the reasoning abilities\nneeded to arrive at a resolution. After identifying KGs for each question, we\nexamine the skew in the distribution of questions for each KG. We then\nintroduce a targeted question generation model to reduce this skew, which\nallows us to generate new types of questions for an image. These new questions\ncan be added to existing VQA datasets to increase the diversity of questions\nand reduce the skew.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC