Multi-Challenge
A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Models scored
29
evaluated
Modality
text
Category
communication
+1 more
Published
2025
arxiv.org
Citations
138
Semantic Scholar
Influential
16
citations
References
0
cited works
Venue
Annual Meeting of the Association for Computational Linguistics
published in
Abstract
Ved Sirdeshmukh, Kaustubh Deshpande, Johannes Mols, Lifeng Jin, et al. (+6)
We present MultiChallenge, a pioneering benchmark evaluating large language models (LLMs) on conducting multi-turn conversations with human users, a crucial yet underexamined capability for their applications. MultiChallenge identifies four categories of challenges in multi-turn conversations that are not only common and realistic among current human-LLM interactions, but are also challenging to all current frontier LLMs. All 4 challenges require accurate instruction-following, context allocation, and in-context reasoning at the same time. We also develop LLM as judge with instance-level rubrics to facilitate an automatic evaluation method with fair agreement with experienced human raters. Despite achieving near-perfect scores on existing multi-turn evaluation benchmarks, all frontier models have less than 50% accuracy on MultiChallenge, with the top-performing Claude 3.5 Sonnet (June 2024) achieving just a 41.4% average accuracy.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Nova 2 Pro | Amazon | 78 |
| 02 | Nova 2 Lite | Amazon | 77 |
| 03 | Nova 2 Omni | Amazon | 76 |
| 04 | GPT-5 | OpenAI | 70 |
| 05 | Qwen3.5-397B-A17B | Alibaba Cloud / Qwen Team | 68 |
| 06 | Nemotron 3 Ultra (550B A55B) | NVIDIA | 64 |
| 07 | Step3-VL-10B | StepFun | 63 |
| 08 | Qwen3.5-122B-A10B | Alibaba Cloud / Qwen Team | 62 |
| 09 | Qwen3.5-27B | Alibaba Cloud / Qwen Team | 61 |
| 10 | o3 | OpenAI | 60 |
| 11 | Qwen3.5-35B-A3B | Alibaba Cloud / Qwen Team | 60 |
| 12 | Nemotron 3 Super (120B A12B) | NVIDIA | 55 |
| 13 | Qwen3.5-9B | Alibaba Cloud / Qwen Team | 55 |
| 14 | Kimi K2 Instruct | Moonshot AI | 54 |
| 15 | Kimi K2-Instruct-0905 | Moonshot AI | 54 |
| 16 | MAI-Thinking-1 | Microsoft | 53 |
| 17 | Qwen3.5-4B | Alibaba Cloud / Qwen Team | 49 |
| 18 | MiniMax M1 40K | MiniMax | 45 |
| 19 | MiniMax M1 80K | MiniMax | 45 |
| 20 | GPT-4.5 | OpenAI | 44 |
| 21 | o4-mini | OpenAI | 43 |
| 22 | GPT-4o | OpenAI | 40 |
| 23 | o3-mini | OpenAI | 40 |
| 24 | Nemotron 3 Nano (30B A3B) | NVIDIA | 39 |
| 25 | GPT-4.1 | OpenAI | 38 |
| 26 | GPT-4.1 mini | OpenAI | 36 |
| 27 | Qwen3.5-2B | Alibaba Cloud / Qwen Team | 34 |
| 28 | Qwen3.5-0.8B | Alibaba Cloud / Qwen Team | 19 |
| 29 | GPT-4.1 nano | OpenAI | 15 |
29 of 29 models · score normalized 0–100 where available