SimpleQA

Measuring short-form factuality in large language models

Models scored

46

evaluated

Modality

text

Category

factuality

+2 more

Published

2024

arxiv.org

Citations

308

Semantic Scholar

Influential

61

citations

References

19

cited works

Venue

arXiv.org

published in

Abstract

Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, et al. (+4)

We present SimpleQA, a benchmark that evaluates the ability of language models to answer short, fact-seeking questions. We prioritized two properties in designing this eval. First, SimpleQA is challenging, as it is adversarially collected against GPT-4 responses. Second, responses are easy to grade, because questions are created such that there exists only a single, indisputable answer. Each answer in SimpleQA is graded as either correct, incorrect, or not attempted. A model with ideal behavior would get as many questions correct as possible while not attempting the questions for which it is not confident it knows the correct answer. SimpleQA is a simple, targeted evaluation for whether models"know what they know,"and our hope is that this benchmark will remain relevant for the next few generations of frontier models. SimpleQA can be found at https://github.com/openai/simple-evals.

factualitygeneralreasoning

Search

#ModelLabScore
01DeepSeek-V3.2-ExpDeepSeek97
02Grok 4 FastxAI95
03DeepSeek-V3.1DeepSeek93
04DeepSeek-R1-0528DeepSeek92
05ERNIE 5.0Baidu75
06Gemini 3 ProGoogle72
07Gemini 3 FlashGoogle69
08GPT-4.5OpenAI63
09DeepSeek-V4-Pro-MaxDeepSeek58
10Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team55
11Qwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team54
12Gemini 2.5 Pro Preview 06-05Google54
13Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team52
14Gemini 2.5 ProGoogle51
15Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team50
16Qwen3 VL 4B InstructAlibaba Cloud / Qwen Team48
17o1OpenAI47
18Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team44
19Gemini 3.1 Flash-LiteGoogle43
20o1-previewOpenAI42
21GPT-4oOpenAI38
22Kimi K2 BaseMoonshot AI35
23DeepSeek-V4-Flash-MaxDeepSeek34
24Kimi K2-Instruct-0905Moonshot AI31
25Kimi K2 InstructMoonshot AI31
26Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team27
27Gemini 2.5 FlashGoogle27
28DeepSeek-V3DeepSeek25
29Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team24
30Mistral Large 3 (675B Instruct 2512 NVFP4)Mistral AI24
31Mistral Large 3 (675B Instruct 2512 Eagle)Mistral AI24
32Mistral Large 3 (675B Instruct 2512)Mistral AI24
33Mistral Large 3 (675B Base)Mistral AI24
34Gemini 2.0 Flash-LiteGoogle22
35MiniMax M1 80KMiniMax19
36MiniMax M1 40KMiniMax18
37o3-miniOpenAI15
38Mistral Small 3.2 24B InstructMistral AI12
39Gemini 2.5 Flash-LiteGoogle11
40Mistral Small 3.1 24B InstructMistral AI10
41Gemma 3 27BGoogle10
42Gemma 3 12BGoogle6
43Gemma 3 4BGoogle4
44Phi 4Microsoft3
45Gemma 3 1BGoogle2
46ERNIE 4.5Baidu2

46 of 46 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC