Multi-Challenge

A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Models scored

29

evaluated

Modality

text

Category

communication

+1 more

Published

2025

arxiv.org

Citations

138

Semantic Scholar

Influential

16

citations

References

0

cited works

Venue

Annual Meeting of the Association for Computational Linguistics

published in

Abstract

Ved Sirdeshmukh, Kaustubh Deshpande, Johannes Mols, Lifeng Jin, et al. (+6)

We present MultiChallenge, a pioneering benchmark evaluating large language models (LLMs) on conducting multi-turn conversations with human users, a crucial yet underexamined capability for their applications. MultiChallenge identifies four categories of challenges in multi-turn conversations that are not only common and realistic among current human-LLM interactions, but are also challenging to all current frontier LLMs. All 4 challenges require accurate instruction-following, context allocation, and in-context reasoning at the same time. We also develop LLM as judge with instance-level rubrics to facilitate an automatic evaluation method with fair agreement with experienced human raters. Despite achieving near-perfect scores on existing multi-turn evaluation benchmarks, all frontier models have less than 50% accuracy on MultiChallenge, with the top-performing Claude 3.5 Sonnet (June 2024) achieving just a 41.4% average accuracy.

communicationreasoning

Search

#ModelLabScore
01Nova 2 ProAmazon78
02Nova 2 LiteAmazon77
03Nova 2 OmniAmazon76
04GPT-5OpenAI70
05Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team68
06Nemotron 3 Ultra (550B A55B)NVIDIA64
07Step3-VL-10BStepFun63
08Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team62
09Qwen3.5-27BAlibaba Cloud / Qwen Team61
10o3OpenAI60
11Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team60
12Nemotron 3 Super (120B A12B)NVIDIA55
13Qwen3.5-9BAlibaba Cloud / Qwen Team55
14Kimi K2 InstructMoonshot AI54
15Kimi K2-Instruct-0905Moonshot AI54
16MAI-Thinking-1Microsoft53
17Qwen3.5-4BAlibaba Cloud / Qwen Team49
18MiniMax M1 40KMiniMax45
19MiniMax M1 80KMiniMax45
20GPT-4.5OpenAI44
21o4-miniOpenAI43
22GPT-4oOpenAI40
23o3-miniOpenAI40
24Nemotron 3 Nano (30B A3B)NVIDIA39
25GPT-4.1OpenAI38
26GPT-4.1 miniOpenAI36
27Qwen3.5-2BAlibaba Cloud / Qwen Team34
28Qwen3.5-0.8BAlibaba Cloud / Qwen Team19
29GPT-4.1 nanoOpenAI15

29 of 29 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC