This paper presents a novel method, termed Bridge to Answer, to infer correct\nanswers for questions about a given video by leveraging adequate graph\ninteractions of heterogeneous crossmodal graphs. To realize this, we learn\nquestion conditioned visual graphs by exploiting the relation between video and\nquestion to enable each visual node using question-to-visual interactions to\nencompass both visual and linguistic cues. In addition, we propose bridged\nvisual-to-visual interactions to incorporate two complementary visual\ninformation on appearance and motion by placing the question graph as an\nintermediate bridge. This bridged architecture allows reliable message passing\nthrough compositional semantics of the question to generate an appropriate\nanswer. As a result, our method can learn the question conditioned visual\nrepresentations attributed to appearance and motion that show powerful\ncapability for video question answering. Extensive experiments prove that the\nproposed method provides effective and superior performance than\nstate-of-the-art methods on several benchmarks.\n
Paper
References (43)
Scroll for more · 31 remaining