Knowledge-Based Video Question Answering with Unsupervised Scene Descriptions

To understand movies, humans constantly reason over the dialogues and actions\nshown in specific scenes and relate them to the overall storyline already seen.\nInspired by this behaviour, we design ROLL, a model for knowledge-based video\nstory question answering that leverages three crucial aspects of movie\nunderstanding: dialog comprehension, scene reasoning, and storyline recalling.\nIn ROLL, each of these tasks is in charge of extracting rich and diverse\ninformation by 1) processing scene dialogues, 2) generating unsupervised video\nscene descriptions, and 3) obtaining external knowledge in a weakly supervised\nfashion. To answer a given question correctly, the information generated by\neach inspired-cognitive task is encoded via Transformers and fused through a\nmodality weighting mechanism, which balances the information from the different\nsources. Exhaustive evaluation demonstrates the effectiveness of our approach,\nwhich yields a new state-of-the-art on two challenging video question answering\ndatasets: KnowIT VQA and TVQA+.\n

Paper

References (59)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC