Arena Hard

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Models scored

26

evaluated

Modality

text

Category

creativity

+3 more

Published

2024

arxiv.org

Citations

490

Semantic Scholar

Influential

103

citations

References

66

cited works

Venue

International Conference on Machine Learning

published in

Abstract

Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, et al. (+4)

The rapid evolution of Large Language Models (LLMs) has outpaced the development of model evaluation, highlighting the need for continuous curation of new, challenging benchmarks. However, manual curation of high-quality, human-aligned benchmarks is expensive and time-consuming. To address this, we introduce BenchBuilder, an automated pipeline that leverages LLMs to curate high-quality, open-ended prompts from large, crowd-sourced datasets, enabling continuous benchmark updates without human in the loop. We apply BenchBuilder to datasets such as Chatbot Arena and WildChat-1M, extracting challenging prompts and utilizing LLM-as-a-Judge for automatic model evaluation. To validate benchmark quality, we propose new metrics to measure a benchmark's alignment with human preferences and ability to separate models. We release Arena-Hard-Auto, a benchmark consisting 500 challenging prompts curated by BenchBuilder. Arena-Hard-Auto provides 3x higher separation of model performances compared to MT-Bench and achieves 98.6% correlation with human preference rankings, all at a cost of $20. Our work sets a new framework for the scalable curation of automated benchmarks from extensive data.

creativitygeneralreasoningwriting

Search

#ModelLabScore
01Qwen3 235B A22BAlibaba Cloud / Qwen Team96
02Qwen3 32BAlibaba Cloud / Qwen Team94
03Qwen3 30B A3BAlibaba Cloud / Qwen Team91
04Llama-3.3 Nemotron Super 49B v1NVIDIA88
05Mistral Small 3 24B InstructMistral AI88
06Qwen2.5 72B InstructAlibaba Cloud / Qwen Team81
07Phi 4 Reasoning PlusMicrosoft79
08DeepSeek-V2.5DeepSeek76
09Phi 4Microsoft75
10Phi 4 ReasoningMicrosoft73
11Ministral 8B InstructMistral AI71
12Jamba 1.5 LargeAI21 Labs65
13Mistral Small 4Mistral AI58
14Granite 3.3 8B InstructIBM58
15Granite 3.3 8B BaseIBM58
16Mistral Large 3Mistral AI55
17MiniStral 3 (14B Instruct 2512)Mistral AI55
18Qwen2.5 7B InstructAlibaba Cloud / Qwen Team52
19Ministral 3 (8B Instruct 2512)Mistral AI51
20Jamba 1.5 MiniAI21 Labs46
21Mistral Small 3.2 24B InstructMistral AI43
22Phi-3.5-MoE-instructMicrosoft38
23Phi-3.5-mini-instructMicrosoft37
24Phi 4 MiniMicrosoft33
25Ministral 3 (3B Instruct 2512)Mistral AI31
26IBM Granite 4.0 Tiny PreviewIBM27

26 of 26 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC