Qwen3 VL 30B A3B Instruct

Creator

Alibaba

Released

2025-10-03

Intelligence

10.0

Artificial Analysis Index

Coding

Artificial Analysis Index

In $/1M

$0.20

input tokens

Out $/1M

$0.80

output tokens

Blended $/1M

$0.35

3:1 blended

Speed

0

tokens / sec

Profile

License, openness, and modality — the governance layer.

License

Apache 2.0

commercial OK

Weights

Open

downloadable

Modalities

Image · Text · Video

Parameters

31B

total

Qwen3-VL is a large multimodal model that unifies vision, language, and reasoning to achieve human-level perception and cognition across text, images, and video. Built on a 235B-parameter architecture, it integrates early joint training of visual and textual modalities for strong language grounding. The model supports up to a 1 million-token context window and excels at visual understanding, spatial reasoning, long video comprehension, and tool-based interaction. It can generate code from images, perform precise 2D/3D object grounding, and operate digital interfaces like a visual agent. The “Instruct” version rivals Gemini 2.5 Pro in perception benchmarks, while the “Thinking” version leads in multimodal reasoning and STEM tasks. With multilingual OCR, creative writing, and fine-grained scene interpretation, Qwen3-VL establishes a new open-source frontier for integrated vision-language intelligence.

Capability profile

Category strength across 20 domains, via LLM Stats.

communication
100multimodal
100general
80writing
80language
80grounding
80creativity
80text-to-image
80structured output
80instruction following
803d
70math
70legal
70video
70
vision
70biology
70finance
70reasoning
70healthcare
70tool calling
70image to text
70spatial reasoning
70physics
60chemistry
60long context
60agents
50economics
50factuality
30
Via LLM Stats

Benchmark breakdown

Independent evaluation scores, normalized to 0–100.

GPQA Diamond
70
Humanity’s Last Exam
6
MMLU-Pro
76
SciCode
31
LiveCodeBench
48
MATH-500
AIME 2025
72
τ²-Bench (agentic)
19
Terminal-Bench Hard
6
IFBench
33
Via Artificial Analysis
© 2026 NYSGPT2525 LLC