Models scored
58
evaluated
Modality
text
Category
agents
+2 more
Published
2025
arxiv.org
Citations
511
Semantic Scholar
Influential
80
citations
References
28
cited works
Venue
arXiv.org
published in
Abstract
Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, et al. (+6)
We present BrowseComp, a simple yet challenging benchmark for measuring the ability for agents to browse the web. BrowseComp comprises 1,266 questions that require persistently navigating the internet in search of hard-to-find, entangled information. Despite the difficulty of the questions, BrowseComp is simple and easy-to-use, as predicted answers are short and easily verifiable against reference answers. BrowseComp for browsing agents can be seen as analogous to how programming competitions are an incomplete but useful benchmark for coding agents. While BrowseComp sidesteps challenges of a true user query distribution, like generating long answers or resolving ambiguity, it measures the important core capability of exercising persistence and creativity in finding information. BrowseComp can be found at https://github.com/openai/simple-evals.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Kimi K3 | Moonshot AI | 91 |
| 02 | Claude Opus 5 | Anthropic | 91 |
| 03 | GPT-5.6 Sol | OpenAI | 90 |
| 04 | GPT-5.5 Pro | OpenAI | 90 |
| 05 | GPT-5.6 Terra | OpenAI | 88 |
| 06 | Claude Mythos Preview | Anthropic | 87 |
| 07 | Kimi K2.6 | Moonshot AI | 86 |
| 08 | Seed 2.1 Pro | ByteDance | 86 |
| 09 | Gemini 3.1 Pro | 86 | |
| 10 | Seed 2.1 Turbo | ByteDance | 85 |
| 11 | Claude Sonnet 5 | Anthropic | 85 |
| 12 | GPT-5.5 | OpenAI | 84 |
| 13 | Claude Opus 4.8 | Anthropic | 84 |
| 14 | Hy3 | Tencent | 84 |
| 15 | Claude Opus 4.6 | Anthropic | 84 |
| 16 | MiniMax M3 | MiniMax | 84 |
| 17 | DeepSeek-V4-Pro-Max | DeepSeek | 83 |
| 18 | GPT-5.6 Luna | OpenAI | 83 |
| 19 | GPT-5.4 | OpenAI | 83 |
| 20 | Claude Opus 4.7 | Anthropic | 79 |
| 21 | GLM-5.1 | Zhipu AI | 79 |
| 22 | GPT-5.2 Pro | OpenAI | 78 |
| 23 | Seed 2.0 Pro | ByteDance | 77 |
| 24 | MiniMax M2.5 | MiniMax | 76 |
| 25 | GLM-5 | Zhipu AI | 76 |
| 26 | Kimi K2.5 | Moonshot AI | 75 |
| 27 | Claude Sonnet 4.6 | Anthropic | 75 |
| 28 | DeepSeek-V4-Flash-Max | DeepSeek | 73 |
| 29 | Step-3.5-Flash | StepFun | 69 |
| 30 | Qwen3.5-397B-A17B | Alibaba Cloud / Qwen Team | 69 |
| 31 | GPT-5.2 | OpenAI | 66 |
| 32 | Qwen3.5-122B-A10B | Alibaba Cloud / Qwen Team | 64 |
| 33 | MiniMax M2.1 | MiniMax | 62 |
| 34 | Qwen3.5-35B-A3B | Alibaba Cloud / Qwen Team | 61 |
| 35 | Qwen3.5-27B | Alibaba Cloud / Qwen Team | 61 |
| 36 | Kimi K2-Thinking-0905 | Moonshot AI | 60 |
| 37 | MiMo-V2-Flash | Xiaomi | 58 |
| 38 | LongCat-Flash-Thinking-2601 | Meituan | 57 |
| 39 | GPT-5 | OpenAI | 55 |
| 40 | GLM-4.7 | Zhipu AI | 52 |
| 41 | o4-mini | OpenAI | 52 |
| 42 | DeepSeek-V3.2 (Thinking) | DeepSeek | 51 |
| 43 | DeepSeek-V3.2 | DeepSeek | 51 |
| 44 | o3 | OpenAI | 50 |
| 45 | Sarvam-105B | Sarvam AI | 50 |
| 46 | Mistral Medium 3.5 | Mistral AI | 49 |
| 47 | GLM-4.6 | Zhipu AI | 45 |
| 48 | Grok 4 Fast | xAI | 45 |
| 49 | Nemotron 3 Ultra (550B A55B) | NVIDIA | 44 |
| 50 | MiniMax M2 | MiniMax | 44 |
| 51 | GLM-4.7-Flash | Zhipu AI | 43 |
| 52 | DeepSeek-V3.2-Exp | DeepSeek | 40 |
| 53 | Sarvam-30B | Sarvam AI | 36 |
| 54 | Nemotron 3 Super (120B A12B) | NVIDIA | 31 |
| 55 | DeepSeek-V3.1 | DeepSeek | 30 |
| 56 | GLM-4.5 | Zhipu AI | 26 |
| 57 | GLM-4.5-Air | Zhipu AI | 21 |
| 58 | DeepSeek-R1-0528 | DeepSeek | 9 |
58 of 58 models · score normalized 0–100 where available