General Response
We appreciate all reviewers for their detailed and constructive feedback. We have revised the paper accordingly. We upload the revised paper, with major updates highlighted in blue.
1. **Expanded Evaluation**: We have included **8 additional video benchmarks** for evaluation: ActivityNet-QA [1], EgoSchema [2], EventBench [3], LongVideoBench [4], PerceptionTest [5], MVBench [6], NExT-QA [7], and VNBench [8]. Alongside VideoMME [9], we now compare LongVILA against state-of-the-art methods on a total of **9 benchmarks**, demonstrating consistently strong performance. The results are shown in **Table 3 in the revision**.
2. **Needle in the Long Video Haystack Experiment**: We report stronger results, with the LongVILA model trained on 2048 frames achieving 99.8% accuracy on **6,000 frames (exceeding 1 million tokens)**. The results are illustrated in **Figure 2 in the revision**.
3. **VideoMME Update**: We have updated our results on **VideoMME** by adding the LongVILA-1.5B model and results for 256 frames. LongVILA-7B achieves **60.1% / 65.1%** accuracy for the settings without and with subtitles, respectively, as shown in **Table 4 in the revision**.
4. **Model Complexity Analysis**: We present a detailed analysis of **model complexity** across various factors including model size, components, number of frames, context length, latency, and FLOPs in **Table 10 in the revision**.
5. **Additional Baselines**: We compare our model against more baselines, including proprietary models (GPT-4V [10], GPT-4o [11], Gemini-1.5-Pro [12]) and open models (Flash-VStream [13], VideoLLaMA2.1 [14], PLLaVA [15], LLaVA One-Vision [16]), as presented in **Table 3 in the revision**.
6. **Additional Ablations**: We conducted further ablations on training schedule settings, as presented in **Table 1 in the revision**.
The following sections provide detailed responses to all reviewer comments.
[1] Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, Dacheng Tao: ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering. AAAI 2019
[2] Karttikeya Mangalam, Raiymbek Akshulakov, Jitendra Malik: EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding. NeurIPS 2023
[3] Yifan Du, Kun Zhou, Yuqi Huo, Yifan Li, Wayne Xin Zhao, Haoyu Lu, Zijia Zhao, Bingning Wang, Weipeng Chen, Ji-Rong Wen: Towards Event-oriented Long Video Understanding. 2024
[4] Haoning Wu, Dongxu Li, Bei Chen, Junnan Li: LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding. 2024
[5] Viorica Patraucean, et. al: Perception Test: A Diagnostic Benchmark for Multimodal Video Models. NeurIPS 2023
[6] Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Lou, Limin Wang, Yu Qiao: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark. CVPR 2024
[7] Junbin Xiao, Xindi Shang, Angela Yao, Tat-Seng Chua: NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions. CVPR 2021
[8] Zijia Zhao, Haoyu Lu, Yuqi Huo, Yifan Du, Tongtian Yue, Longteng Guo, Bingning Wang, Weipeng Chen, Jing Liu: Needle In A Video Haystack: A Scalable Synthetic Framework for Benchmarking Video MLLMs. 2024
[9] Chaoyou Fu, e t. al.: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. 2024
[10] OpenAI. Gpt-4v. 2023
[11] OpenAI. Hello gpt-4o. 2024
[12] Gemini Team. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024.
[13] Haoji Zhang, Yiqin Wang, Yansong Tang, Yong Liu, Jiashi Feng, Jifeng Dai, Xiaojie Jin: Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams. 2024
[14] Zesen Cheng, et. al.: VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs. 2024
[15] Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See-Kiong Ng, Jiashi Feng: PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning. 2024
[16] Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, Chunyuan Li: LLaVA-OneVision: Easy Visual Task Transfer. 2024