Qwen3 VL 4B Instruct

Creator

Alibaba

Released

2025-10-14

Intelligence

4.1

Artificial Analysis Index

Coding

Artificial Analysis Index

In $/1M

$0.00

input tokens

Out $/1M

$0.00

output tokens

Blended $/1M

$0.00

3:1 blended

Speed

0

tokens / sec

Profile

License, openness, and modality — the governance layer.

License

Apache 2.0

commercial OK

Weights

Open

downloadable

Modalities

Image · Text

Parameters

4B

total

Qwen3-VL is a large multimodal model that unifies vision, language, and reasoning to achieve human-level perception and cognition across text, images, and video. Built on a 235B-parameter architecture, it integrates early joint training of visual and textual modalities for strong language grounding. The model supports up to a 1 million-token context window and excels at visual understanding, spatial reasoning, long video comprehension, and tool-based interaction. It can generate code from images, perform precise 2D/3D object grounding, and operate digital interfaces like a visual agent. The “Instruct” version rivals Gemini 2.5 Pro in perception benchmarks, while the “Thinking” version leads in multimodal reasoning and STEM tasks. With multilingual OCR, creative writing, and fine-grained scene interpretation, Qwen3-VL establishes a new open-source frontier for integrated vision-language intelligence.

Capability profile

Category strength across 20 domains, via LLM Stats.

communication
100multimodal
90general
80writing
80grounding
80creativity
80text-to-image
80instruction following
803d
70legal
70language
70image to text
70structured output
70math
60
video
60vision
60finance
60reasoning
60healthcare
60long context
60tool calling
60spatial reasoning
60factuality
50agents
40physics
40chemistry
40economics
40
Via LLM Stats

Benchmark breakdown

Independent evaluation scores, normalized to 0–100.

GPQA Diamond
37
Humanity’s Last Exam
4
MMLU-Pro
63
SciCode
14
LiveCodeBench
29
MATH-500
AIME 2025
37
τ²-Bench (agentic)
23
Terminal-Bench Hard
0
IFBench
32
Via Artificial Analysis

Providers

Inference hosts serving this model — their own pricing and measured performance, cheapest input first.

ProviderIn $/1MOut $/1MThroughputLatencyStatus
DeepInfra$0.10$0.60active
All providers

1 host · via LLM Stats

© 2026 NYSGPT2525 LLC