Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries

Video Question Answering (Video QA) is a challenging video understanding task that requires models to compre-hend entire videos, identify the most relevant information based on contextual cues from a given question, and rea-son accurately to provide answers. Recent advancements in Multimodal Large Language Models (MLLMs) have trans-formed video QA by leveraging their exceptional common-sense reasoning capabilities. This progress is largely driven by the effective alignment between visual data and the language space of MLLMs. However, for video QA, an ad-ditional space-time alignment poses a considerable chal-lenge for extracting question-relevant information across frames. In this work, we investigate diverse temporal modeling techniques to integrate with MLLMs, aiming to achieve question-guided temporal modeling that leverages pre-trained visual and textual alignment in MLLMs. We propose T-Former, a novel temporal modeling method that creates a question-guided temporal bridge between frame-wise visual perception and the reasoning capabilities of LLMs. Our evaluation across multiple video QA bench-marks demonstrates that T-Former competes favorably with existing temporal modeling approaches and aligns with re-cent advancements in video QA.

Paper

Similar papers

© 2026 NYSGPT2525 LLC