Summary
This paper presents Q-Bench-Video, a benchmark specifically designed to evaluate the video quality understanding capabilities of Large Multi-modal Models (LMMs). Recognizing a gap in existing benchmarks that focus on high-level video comprehension rather than quality assessment, Q-Bench-Video targets four key dimensions of video quality: technical, aesthetic, temporal, and AI-generated content (AIGC) distortions. The benchmark includes a diverse set of videos from natural scenes, AI-generated content, and computer graphics, ensuring a balanced distribution across quality levels. To comprehensively assess LMMs, Q-Bench-Video employs various question types—Yes-or-No, What-How, open-ended, and video pair comparisons—that capture the models’ performance across straightforward and complex tasks. Validated on 17 LMMs (12 open-source and 5 proprietary), the benchmark reveals that, while LMMs demonstrate foundational capabilities in video quality assessment, they fall significantly short of human-level performance, particularly in open-ended and AIGC-specific queries. Q-Bench-Video provides a new standard for evaluating video quality understanding and aims to drive future research on enhancing LMMs’ video quality perception.
Strengths
1. Q-Bench-Video is the first benchmark specifically focused on assessing video quality understanding in Large Multi-modal Models (LMMs), addressing a unique and underexplored aspect of LMMs that goes beyond typical video comprehension. By evaluating quality-related distortions—including technical, aesthetic, temporal, and AIGC-specific aspects—it provides a novel and comprehensive perspective on video quality assessment.
2. The benchmark is meticulously designed with diverse video sources (natural, AIGC, CG) and a balanced quality distribution through uniform sampling from quality-annotated datasets. The use of various question types (Yes-or-No, What-How, open-ended, and video pair comparisons) creates a robust framework that evaluates LMMs across both straightforward and complex scenarios, ensuring a thorough assessment of model capabilities.
3. Q-Bench-Video addresses an important gap in LMM evaluation by focusing on video quality understanding, which is essential for applications in video generation, content moderation, and quality control. The findings—highlighting substantial performance gaps between LMMs and human evaluators, especially in AIGC distortions—offer valuable insights for advancing LMMs and guide future research on improving their quality perception.
Weaknesses
1. While Q-Bench-Video includes diverse video types, there is limited exploration of how LMMs generalize across these domains. Testing LMMs on specific domains, such as medical or surveillance videos, would enhance the benchmark’s relevance by showing how well models handle domain-specific quality variations, which are critical in many real-world applications.
2. The benchmark results indicate that LMMs struggle significantly with open-ended questions, but the paper lacks a detailed breakdown of common error types or patterns in these responses. A more granular analysis could provide clearer insights into where models fail in nuanced video quality understanding and offer guidance on specific areas for model improvement.
3. The reliance on GPT-assisted evaluation for scoring open-ended responses may introduce subjectivity or bias, as the evaluation depends on the alignment of the assistant model’s judgments with human interpretation. Including human evaluations as a baseline for open-ended questions would provide a stronger reference and help validate the reliability of the GPT-assisted scores.
4. The benchmark primarily focuses on a balanced quality distribution but lacks an emphasis on evaluating LMMs in scenarios with challenging low-quality or noisy data, which are common in real-world settings. Incorporating more degraded video samples could better test the robustness of LMMs and highlight areas where models may require improvement for real-world deployment.
5. Some of the highly relevant video quality assessment works should be mentioned in the related work section, like [R1][R2][R3].
[R1] UGC-VQA: Benchmarking Blind Video Quality Assessment for User Generated Content, TIP 2021
[R2] RAPIQUE: Rapid and accurate video quality prediction of user generated content, OJSP 2021
[R3] FAVER: Blind Quality Prediction of Variable Frame Rate Videos, SPIC 202
Questions
1. Could you provide more detailed insights into the specific error patterns observed in LMMs’ open-ended responses?
2. Have you considered testing LMMs on additional specialized domains, such as gaming videos, animation, or screen content, to better assess cross-domain generalization?
3. Given the potential bias in GPT-assisted evaluation, did you consider incorporating human evaluations as a baseline for open-ended responses?
4. Would you consider adding a more challenging subset focused specifically on low-quality or noisy videos?