Models scored
27
evaluated
Modality
text
Category
reasoning
Published
2019
arxiv.org
Citations
4,301
Semantic Scholar
Influential
435
citations
References
22
cited works
Venue
Annual Meeting of the Association for Computational Linguistics
published in
Abstract
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, et al. (+1)
Recent work by Zellers et al. (2018) introduced a new task of commonsense natural language inference: given an event description such as “A woman sits at a piano,” a machine must select the most likely followup: “She sets her fingers on the keys.” With the introduction of BERT, near human-level performance was reached. Does this mean that machines can perform human level commonsense inference? In this paper, we show that commonsense inference still proves difficult for even state-of-the-art models, by presenting HellaSwag, a new challenge dataset. Though its questions are trivial for humans (>95% accuracy), state-of-the-art models struggle (<48%). We achieve this via Adversarial Filtering (AF), a data collection paradigm wherein a series of discriminators iteratively select an adversarial set of machine-generated wrong answers. AF proves to be surprisingly robust. The key insight is to scale up the length and complexity of the dataset examples towards a critical ‘Goldilocks’ zone wherein generated text is ridiculous to humans, yet often misclassified by state-of-the-art models. Our construction of HellaSwag, and its resulting difficulty, sheds light on the inner workings of deep pretrained models. More broadly, it suggests a new path forward for NLP research, in which benchmarks co-evolve with the evolving state-of-the-art in an adversarial way, so as to present ever-harder challenges.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Claude 3 Opus | Anthropic | 95 |
| 02 | GPT-4 | OpenAI | 95 |
| 03 | Gemini 1.5 Pro | 93 | |
| 04 | MiMo-V2.5-Pro | Xiaomi | 90 |
| 05 | Claude 3 Sonnet | Anthropic | 89 |
| 06 | Command R+ | Cohere | 89 |
| 07 | Hermes 3 70B | Nous Research | 88 |
| 08 | Qwen2 72B Instruct | Alibaba Cloud / Qwen Team | 88 |
| 09 | Gemini 1.5 Flash | 87 | |
| 10 | Gemma 2 27B | 86 | |
| 11 | Claude 3 Haiku | Anthropic | 86 |
| 12 | Llama 3.1 Nemotron 70B Instruct | NVIDIA | 86 |
| 13 | Qwen2.5 32B Instruct | Alibaba Cloud / Qwen Team | 85 |
| 14 | Phi-3.5-MoE-instruct | Microsoft | 84 |
| 15 | Mistral NeMo Instruct | Mistral AI | 84 |
| 16 | Qwen2.5-Coder 32B Instruct | Alibaba Cloud / Qwen Team | 83 |
| 17 | Gemma 2 9B | 82 | |
| 18 | Granite 3.3 8B Base | IBM | 80 |
| 19 | Gemma 3n E4B Instructed LiteRT Preview | 79 | |
| 20 | Gemma 3n E4B | 79 | |
| 21 | Qwen2.5-Coder 7B Instruct | Alibaba Cloud / Qwen Team | 77 |
| 22 | Gemma 3n E2B Instructed LiteRT (Preview) | 72 | |
| 23 | Gemma 3n E2B | 72 | |
| 24 | Llama 3.2 3B Instruct | Meta | 70 |
| 25 | Phi-3.5-mini-instruct | Microsoft | 69 |
| 26 | Phi 4 Mini | Microsoft | 69 |
| 27 | ERNIE 4.5 | Baidu | 33 |
27 of 27 models · score normalized 0–100 where available