IFEval

Instruction-Following Evaluation for Large Language Models

Models scored

65

evaluated

Modality

text

Category

general

+2 more

Published

2023

arxiv.org

Citations

917

Semantic Scholar

Influential

151

citations

References

0

cited works

Venue

arXiv.org

published in

Abstract

Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, et al. (+4)

One core capability of Large Language Models (LLMs) is to follow natural language instructions. However, the evaluation of such abilities is not standardized: Human evaluations are expensive, slow, and not objectively reproducible, while LLM-based auto-evaluation is potentially biased or limited by the ability of the evaluator LLM. To overcome these issues, we introduce Instruction-Following Eval (IFEval) for large language models. IFEval is a straightforward and easy-to-reproduce evaluation benchmark. It focuses on a set of"verifiable instructions"such as"write in more than 400 words"and"mention the keyword of AI at least 3 times". We identified 25 types of those verifiable instructions and constructed around 500 prompts, with each prompt containing one or more verifiable instructions. We show evaluation results of two widely available LLMs on the market. Our code and data can be found at https://github.com/google-research/google-research/tree/master/instruction_following_eval

generalinstruction followingstructured output

Search

#ModelLabScore
01Qwen3.5-27BAlibaba Cloud / Qwen Team95
02Qwen3.7-PlusAlibaba Cloud / Qwen Team95
03Qwen3.7 MaxAlibaba Cloud / Qwen Team94
04Qwen3.6 PlusAlibaba Cloud / Qwen Team94
05o3-miniOpenAI94
06Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team93
07Claude 3.7 SonnetAnthropic93
08Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team93
09Llama 3.3 70B InstructMeta92
10Nova ProAmazon92
11Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team92
12Qwen3.5-9BAlibaba Cloud / Qwen Team92
13Gemma 3 27BGoogle90
14Nemotron Nano 9B v2NVIDIA90
15Gemma 3 4BGoogle90
16Qwen3.5-4BAlibaba Cloud / Qwen Team90
17Kimi K2-Instruct-0905Moonshot AI90
18Kimi K2 InstructMoonshot AI90
19Nova LiteAmazon90
20LongCat-Flash-ChatMeituan90
21Llama 3.1 Nemotron Ultra 253B v1NVIDIA89
22Gemma 3 12BGoogle89
23Qwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team89
24Qwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team89
25Llama 3.1 405B InstructMeta89
26GPT-4.5OpenAI88
27Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team88
28Qwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team88
29Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team88
30Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team88
31Qwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team88
32Llama 3.1 70B InstructMeta88
33GPT-4.1OpenAI87
34Nova MicroAmazon87
35Kimi-k1.5Moonshot AI87
36DeepSeek-V3DeepSeek86
37Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team86
38Phi 4 Reasoning PlusMicrosoft85
39Sarvam-105BSarvam AI85
40Qwen3 VL 32B InstructAlibaba Cloud / Qwen Team85
41GPT-4.1 miniOpenAI84
42Qwen2.5 72B InstructAlibaba Cloud / Qwen Team84
43QwQ-32BAlibaba Cloud / Qwen Team84
44Qwen3 VL 8B InstructAlibaba Cloud / Qwen Team84
45Phi 4 ReasoningMicrosoft83
46Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team83
47Mistral Small 3 24B InstructMistral AI83
48Qwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team83
49Qwen3 VL 4B InstructAlibaba Cloud / Qwen Team82
50Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team82
51GPT-4oOpenAI81
52Llama 3.1 8B InstructMeta80
53Gemma 3 1BGoogle80
54Llama 3.1 Nemotron Nano 8B V1NVIDIA79
55Qwen3.5-2BAlibaba Cloud / Qwen Team79
56Llama 3.2 3B InstructMeta77
57MiniCPM-SALAOpenBMB76
58Granite 3.3 8B InstructIBM75
59Granite 3.3 8B BaseIBM75
60GPT-4.1 nanoOpenAI75

60 of 65 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC