Recent work has revealed that large language models (LLMs) can exhibit emergent theory-of-mind (ToM) capabilities—inferring human beliefs, desires, and intentions from text alone. Yet, everyday social reasoning often unfolds visually in dynamic contexts. This paper investigates whether multimodal LLMs can similarly demonstrate ToM skills in video-based tasks. Concretely, we propose a pipeline that fuses video and text signals, retrieves the most relevant frames for each query, and answers questions requiring spatio-temporal social understanding. We introduce a new frame localization benchmark, Theory of Mind Localization (ToMLoc), and show that finetuning a state-of-the-art Video-ChatGPT model on ToMLoc significantly improves performance on Social-Iq 2.0. Our results suggest that bridging textual and visual modalities is essential for capturing complex mental states in real-world scenarios. Moreover, retrieving key frames enhances interpretability by revealing how the model arrives at its inferences. These findings highlight the promise of video-based approaches for achieving more human-like social intelligence in LLMs.