IFBench Benchmark Leaderboard

Instruction & Long ContextLive

Category

Instruction & Long Context

Source

Artificial Analysis

evaluation of record

Models covered

449

in our data

Data status

Live

Top score

83.3%

best on record

Top model

Grok 4.3

SpaceXAI

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

A benchmark evaluating precise instruction-following generalization on 58 diverse, verifiable out-of-domain constraints that test models' ability to follow specific output requirements.

Leaderboard

Top 20 of 449 models we hold a score for.

1Grok 4.3mediumSpaceXAI
83.3%2Grok 4.20 0309ReasoningSpaceXAI
82.9%3MiniMax-M3MiniMax
82.9%4Nemotron 3 Ultra 550B A55BReasoningNVIDIA
81.4%5Grok 4.3highSpaceXAI
81.3%6Grok 4.20 0309 v2ReasoningSpaceXAI
81.2%7Grok 4.3lowSpaceXAI
81.0%8Qwen3.7 MaxAlibaba
80.5%9Nemotron Cascade 2 30B A3BNVIDIA
80.4%10MiMo-V2.5-ProXiaomi
79.9%11Nova 2.0 Pro PreviewlowAmazon
79.6%12DeepSeek V4 FlashReasoning, Max EffortDeepSeek
79.2%13Nova 2.0 Pro PreviewmediumAmazon
79.0%14Qwen3.5 397B A17BReasoningAlibaba
78.8%15Gemini 3 Flash PreviewReasoningGoogle
78.0%16Qwen3.7 PlusAlibaba
78.0%17GPT-5.2 CodexxhighOpenAI
77.6%18Gemini 3.1 Flash-LiteGoogle
77.2%19Gemini 3.1 Pro PreviewGoogle
77.1%20Qwen3.6 Max PreviewAlibaba
76.6%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Checks whether a model follows the letter of an instruction across 58 constraints it was not trained on, each one machine-checkable — answer in exactly three sentences, never use a particular word, end on a specific phrase. Higher is better; the score is the share of constraints satisfied, and the quality of the answer is not graded at all, only its compliance. That narrowness is both the limitation and the value: it isolates instruction-following from intelligence, and strong reasoning models routinely lose points here for ignoring a formatting rule they judged unimportant. Because the constraints are held out, it measures generalization rather than the memorized compliance that older instruction benchmarks reward.

© 2026 NYSGPT2525 LLC