CLEVR_HYP: A Challenge Dataset and Baselines for Visual Question Answering with Hypothetical Actions over Images

Most existing research on visual question answering (VQA) is limited to\ninformation explicitly present in an image or a video. In this paper, we take\nvisual understanding to a higher level where systems are challenged to answer\nquestions that involve mentally simulating the hypothetical consequences of\nperforming specific actions in a given scenario. Towards that end, we formulate\na vision-language question answering task based on the CLEVR (Johnson et. al.,\n2017) dataset. We then modify the best existing VQA methods and propose\nbaseline solvers for this task. Finally, we motivate the development of better\nvision-language models by providing insights about the capability of diverse\narchitectures to perform joint reasoning over image-text modality. Our dataset\nsetup scripts and codes will be made publicly available at\nhttps://github.com/shailaja183/clevr_hyp.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC