LVBench

An Extreme Long Video Understanding Benchmark

Models scored

23

evaluated

Modality

multimodal

Category

long context

+2 more

Published

2024

arxiv.org

Citations

348

Semantic Scholar

Influential

51

citations

References

53

cited works

Venue

IEEE International Conference on Computer Vision

published in

Abstract

Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, et al. (+7)

Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of meeting the demands of real-world applications such as embodied intelligence for long-term decision-making, in-depth movie reviews and discussions, and live sports commentary, all of which require comprehension of long videos spanning several hours. To address this gap, we introduce LVBench, a benchmark specifically designed for long video understanding. Our dataset comprises publicly sourced videos and encompasses a diverse set of tasks aimed at long video comprehension and information extraction. LVBench is designed to challenge multimodal models to demonstrate long-term memory and extended comprehension capabilities. Our extensive evaluations reveal that current multimodal models still underperform on these demanding long video understanding tasks. Through LVBench, we aim to spur the development of more advanced models capable of tackling the complexities of long video comprehension.

long contextmultimodalvision

Search

#ModelLabScore
01Seed 2.1 ProByteDance78
02Seed 2.1 TurboByteDance77
03Qwen3.7-PlusAlibaba Cloud / Qwen Team76
04Kimi K2.5Moonshot AI76
05Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team74
06Qwen3.5-27BAlibaba Cloud / Qwen Team74
07Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team71
08Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team71
09Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team68
10Qwen3 VL 32B InstructAlibaba Cloud / Qwen Team64
11Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team64
12Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team63
13Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team63
14Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team59
15Qwen3 VL 8B InstructAlibaba Cloud / Qwen Team58
16Qwen3 VL 4B InstructAlibaba Cloud / Qwen Team56
17Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team56
18Qwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team54
19Qwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team49
20Qwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team47
21Qwen2.5 VL 7B InstructAlibaba Cloud / Qwen Team45
22Nova ProAmazon42
23Nova LiteAmazon40

23 of 23 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC