Creator
Alibaba
Released
2025-10-14
Intelligence
8.4
Artificial Analysis Index
Coding
—
Artificial Analysis Index
In $/1M
$0.18
input tokens
Out $/1M
$0.70
output tokens
Blended $/1M
$0.31
3:1 blended
Speed
0
tokens / sec
Profile
License, openness, and modality — the governance layer.
License
Apache 2.0
commercial OK
Weights
Open
downloadable
Modalities
Image · Text · Video
Parameters
9B
total
Qwen3-VL is a large multimodal model that unifies vision, language, and reasoning to achieve human-level perception and cognition across text, images, and video. Built on a 235B-parameter architecture, it integrates early joint training of visual and textual modalities for strong language grounding. The model supports up to a 1 million-token context window and excels at visual understanding, spatial reasoning, long video comprehension, and tool-based interaction. It can generate code from images, perform precise 2D/3D object grounding, and operate digital interfaces like a visual agent. The “Instruct” version rivals Gemini 2.5 Pro in perception benchmarks, while the “Thinking” version leads in multimodal reasoning and STEM tasks. With multilingual OCR, creative writing, and fine-grained scene interpretation, Qwen3-VL establishes a new open-source frontier for integrated vision-language intelligence.
Capability profile
Category strength across 20 domains, via LLM Stats.
Benchmark breakdown
Independent evaluation scores, normalized to 0–100.