HellaSwag

Can a Machine Really Finish Your Sentence?

Models scored

27

evaluated

Modality

text

Category

reasoning

Published

2019

arxiv.org

Citations

4,301

Semantic Scholar

Influential

435

citations

References

22

cited works

Venue

Annual Meeting of the Association for Computational Linguistics

published in

Abstract

Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, et al. (+1)

Recent work by Zellers et al. (2018) introduced a new task of commonsense natural language inference: given an event description such as “A woman sits at a piano,” a machine must select the most likely followup: “She sets her fingers on the keys.” With the introduction of BERT, near human-level performance was reached. Does this mean that machines can perform human level commonsense inference? In this paper, we show that commonsense inference still proves difficult for even state-of-the-art models, by presenting HellaSwag, a new challenge dataset. Though its questions are trivial for humans (>95% accuracy), state-of-the-art models struggle (<48%). We achieve this via Adversarial Filtering (AF), a data collection paradigm wherein a series of discriminators iteratively select an adversarial set of machine-generated wrong answers. AF proves to be surprisingly robust. The key insight is to scale up the length and complexity of the dataset examples towards a critical ‘Goldilocks’ zone wherein generated text is ridiculous to humans, yet often misclassified by state-of-the-art models. Our construction of HellaSwag, and its resulting difficulty, sheds light on the inner workings of deep pretrained models. More broadly, it suggests a new path forward for NLP research, in which benchmarks co-evolve with the evolving state-of-the-art in an adversarial way, so as to present ever-harder challenges.

Search

#ModelLabScore
01Claude 3 OpusAnthropic95
02GPT-4OpenAI95
03Gemini 1.5 ProGoogle93
04MiMo-V2.5-ProXiaomi90
05Claude 3 SonnetAnthropic89
06Command R+Cohere89
07Hermes 3 70BNous Research88
08Qwen2 72B InstructAlibaba Cloud / Qwen Team88
09Gemini 1.5 FlashGoogle87
10Gemma 2 27BGoogle86
11Claude 3 HaikuAnthropic86
12Llama 3.1 Nemotron 70B InstructNVIDIA86
13Qwen2.5 32B InstructAlibaba Cloud / Qwen Team85
14Phi-3.5-MoE-instructMicrosoft84
15Mistral NeMo InstructMistral AI84
16Qwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team83
17Gemma 2 9BGoogle82
18Granite 3.3 8B BaseIBM80
19Gemma 3n E4B Instructed LiteRT PreviewGoogle79
20Gemma 3n E4BGoogle79
21Qwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team77
22Gemma 3n E2B Instructed LiteRT (Preview)Google72
23Gemma 3n E2BGoogle72
24Llama 3.2 3B InstructMeta70
25Phi-3.5-mini-instructMicrosoft69
26Phi 4 MiniMicrosoft69
27ERNIE 4.5Baidu33

27 of 27 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC