Phi-4 Multimodal Instruct

Creator

Microsoft AI

Released

2025-02-26

Intelligence

4.5

Artificial Analysis Index

Coding

Artificial Analysis Index

In $/1M

$0.00

input tokens

Out $/1M

$0.00

output tokens

Blended $/1M

$0.00

3:1 blended

Speed

26

tokens / sec

Profile

License, openness, and modality — the governance layer.

License

MIT

commercial OK

Weights

Open

downloadable

Modalities

Image · Text

Parameters

6B

total

Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a 128K token context length. Enhanced via SFT, DPO, and RLHF for instruction following and safety.

Capability profile

Category strength across 20 domains, via LLM Stats.

image to text
80vision
70general
70reasoning
70multimodal
70
3d
60math
60healthcare
60spatial reasoning
60
Via LLM Stats

Benchmark breakdown

Independent evaluation scores, normalized to 0–100.

GPQA Diamond
32
Humanity’s Last Exam
4
MMLU-Pro
49
SciCode
11
LiveCodeBench
13
MATH-500
69
AIME 2025
τ²-Bench (agentic)
Terminal-Bench Hard
IFBench
Via Artificial Analysis
© 2026 NYSGPT2525 LLC