Arena Hard
From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Models scored
26
evaluated
Modality
text
Category
creativity
+3 more
Published
2024
arxiv.org
Citations
490
Semantic Scholar
Influential
103
citations
References
66
cited works
Venue
International Conference on Machine Learning
published in
Abstract
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, et al. (+4)
The rapid evolution of Large Language Models (LLMs) has outpaced the development of model evaluation, highlighting the need for continuous curation of new, challenging benchmarks. However, manual curation of high-quality, human-aligned benchmarks is expensive and time-consuming. To address this, we introduce BenchBuilder, an automated pipeline that leverages LLMs to curate high-quality, open-ended prompts from large, crowd-sourced datasets, enabling continuous benchmark updates without human in the loop. We apply BenchBuilder to datasets such as Chatbot Arena and WildChat-1M, extracting challenging prompts and utilizing LLM-as-a-Judge for automatic model evaluation. To validate benchmark quality, we propose new metrics to measure a benchmark's alignment with human preferences and ability to separate models. We release Arena-Hard-Auto, a benchmark consisting 500 challenging prompts curated by BenchBuilder. Arena-Hard-Auto provides 3x higher separation of model performances compared to MT-Bench and achieves 98.6% correlation with human preference rankings, all at a cost of $20. Our work sets a new framework for the scalable curation of automated benchmarks from extensive data.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Qwen3 235B A22B | Alibaba Cloud / Qwen Team | 96 |
| 02 | Qwen3 32B | Alibaba Cloud / Qwen Team | 94 |
| 03 | Qwen3 30B A3B | Alibaba Cloud / Qwen Team | 91 |
| 04 | Llama-3.3 Nemotron Super 49B v1 | NVIDIA | 88 |
| 05 | Mistral Small 3 24B Instruct | Mistral AI | 88 |
| 06 | Qwen2.5 72B Instruct | Alibaba Cloud / Qwen Team | 81 |
| 07 | Phi 4 Reasoning Plus | Microsoft | 79 |
| 08 | DeepSeek-V2.5 | DeepSeek | 76 |
| 09 | Phi 4 | Microsoft | 75 |
| 10 | Phi 4 Reasoning | Microsoft | 73 |
| 11 | Ministral 8B Instruct | Mistral AI | 71 |
| 12 | Jamba 1.5 Large | AI21 Labs | 65 |
| 13 | Mistral Small 4 | Mistral AI | 58 |
| 14 | Granite 3.3 8B Instruct | IBM | 58 |
| 15 | Granite 3.3 8B Base | IBM | 58 |
| 16 | Mistral Large 3 | Mistral AI | 55 |
| 17 | MiniStral 3 (14B Instruct 2512) | Mistral AI | 55 |
| 18 | Qwen2.5 7B Instruct | Alibaba Cloud / Qwen Team | 52 |
| 19 | Ministral 3 (8B Instruct 2512) | Mistral AI | 51 |
| 20 | Jamba 1.5 Mini | AI21 Labs | 46 |
| 21 | Mistral Small 3.2 24B Instruct | Mistral AI | 43 |
| 22 | Phi-3.5-MoE-instruct | Microsoft | 38 |
| 23 | Phi-3.5-mini-instruct | Microsoft | 37 |
| 24 | Phi 4 Mini | Microsoft | 33 |
| 25 | Ministral 3 (3B Instruct 2512) | Mistral AI | 31 |
| 26 | IBM Granite 4.0 Tiny Preview | IBM | 27 |
26 of 26 models · score normalized 0–100 where available