ARC-C

Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Models scored

34

evaluated

Modality

text

Category

general

+1 more

Published

2018

arxiv.org

Citations

5,156

Semantic Scholar

Influential

643

citations

References

36

cited works

Venue

arXiv.org

published in

Abstract

Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, et al. (+3)

We present a new question set, text corpus, and baselines assembled to encourage AI research in advanced question answering. Together, these constitute the AI2 Reasoning Challenge (ARC), which requires far more powerful knowledge and reasoning than previous challenges such as SQuAD or SNLI. The ARC question set is partitioned into a Challenge Set and an Easy Set, where the Challenge Set contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurence algorithm. The dataset contains only natural, grade-school science questions (authored for human tests), and is the largest public-domain set of this kind (7,787 questions). We test several baselines on the Challenge Set, including leading neural models from the SQuAD and SNLI tasks, and find that none are able to significantly outperform a random baseline, reflecting the difficult nature of this task. We are also releasing the ARC Corpus, a corpus of 14M science sentences relevant to the task, and implementations of the three neural baseline models tested. Can your model perform better? We pose ARC as a challenge to the community.

generalreasoning

Search

#ModelLabScore
01MiMo-V2.5-ProXiaomi97
02Llama 3.1 405B InstructMeta97
03Claude 3 OpusAnthropic96
04Nova ProAmazon95
05Llama 3.1 70B InstructMeta95
06Claude 3 SonnetAnthropic93
07Jamba 1.5 LargeAI21 Labs93
08Nova LiteAmazon92
09Mistral Small 3 24B BaseMistral AI91
10Phi-3.5-MoE-instructMicrosoft91
11Nova MicroAmazon90
12Claude 3 HaikuAnthropic89
13Jamba 1.5 MiniAI21 Labs86
14Phi-3.5-mini-instructMicrosoft85
15Phi 4 MiniMicrosoft84
16Llama 3.1 8B InstructMeta83
17Llama 3.2 3B InstructMeta79
18Ministral 8B InstructMistral AI72
19Gemma 2 27BGoogle71
20Command R+Cohere71
21Qwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team71
22Qwen2.5 32B InstructAlibaba Cloud / Qwen Team70
23Llama 3.1 Nemotron 70B InstructNVIDIA69
24Qwen2 72B InstructAlibaba Cloud / Qwen Team69
25Gemma 2 9BGoogle68
26Qwen2.5 14B InstructAlibaba Cloud / Qwen Team67
27Hermes 3 70BNous Research66
28Gemma 3n E4B Instructed LiteRT PreviewGoogle62
29Gemma 3n E4BGoogle62
30Qwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team61
31Gemma 3n E2B Instructed LiteRT (Preview)Google52
32Gemma 3n E2BGoogle52
33Granite 3.3 8B BaseIBM51
34ERNIE 4.5Baidu41

34 of 34 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC