BrowseComp

A Simple Yet Challenging Benchmark for Browsing Agents

Models scored

58

evaluated

Modality

text

Category

agents

+2 more

Published

2025

arxiv.org

Citations

511

Semantic Scholar

Influential

80

citations

References

28

cited works

Venue

arXiv.org

published in

Abstract

Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, et al. (+6)

We present BrowseComp, a simple yet challenging benchmark for measuring the ability for agents to browse the web. BrowseComp comprises 1,266 questions that require persistently navigating the internet in search of hard-to-find, entangled information. Despite the difficulty of the questions, BrowseComp is simple and easy-to-use, as predicted answers are short and easily verifiable against reference answers. BrowseComp for browsing agents can be seen as analogous to how programming competitions are an incomplete but useful benchmark for coding agents. While BrowseComp sidesteps challenges of a true user query distribution, like generating long answers or resolving ambiguity, it measures the important core capability of exercising persistence and creativity in finding information. BrowseComp can be found at https://github.com/openai/simple-evals.

agentsreasoningsearch

Search

#ModelLabScore
01Kimi K3Moonshot AI91
02Claude Opus 5Anthropic91
03GPT-5.6 SolOpenAI90
04GPT-5.5 ProOpenAI90
05GPT-5.6 TerraOpenAI88
06Claude Mythos PreviewAnthropic87
07Kimi K2.6Moonshot AI86
08Seed 2.1 ProByteDance86
09Gemini 3.1 ProGoogle86
10Seed 2.1 TurboByteDance85
11Claude Sonnet 5Anthropic85
12GPT-5.5OpenAI84
13Claude Opus 4.8Anthropic84
14Hy3Tencent84
15Claude Opus 4.6Anthropic84
16MiniMax M3MiniMax84
17DeepSeek-V4-Pro-MaxDeepSeek83
18GPT-5.6 LunaOpenAI83
19GPT-5.4OpenAI83
20Claude Opus 4.7Anthropic79
21GLM-5.1Zhipu AI79
22GPT-5.2 ProOpenAI78
23Seed 2.0 ProByteDance77
24MiniMax M2.5MiniMax76
25GLM-5Zhipu AI76
26Kimi K2.5Moonshot AI75
27Claude Sonnet 4.6Anthropic75
28DeepSeek-V4-Flash-MaxDeepSeek73
29Step-3.5-FlashStepFun69
30Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team69
31GPT-5.2OpenAI66
32Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team64
33MiniMax M2.1MiniMax62
34Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team61
35Qwen3.5-27BAlibaba Cloud / Qwen Team61
36Kimi K2-Thinking-0905Moonshot AI60
37MiMo-V2-FlashXiaomi58
38LongCat-Flash-Thinking-2601Meituan57
39GPT-5OpenAI55
40GLM-4.7Zhipu AI52
41o4-miniOpenAI52
42DeepSeek-V3.2 (Thinking)DeepSeek51
43DeepSeek-V3.2DeepSeek51
44o3OpenAI50
45Sarvam-105BSarvam AI50
46Mistral Medium 3.5Mistral AI49
47GLM-4.6Zhipu AI45
48Grok 4 FastxAI45
49Nemotron 3 Ultra (550B A55B)NVIDIA44
50MiniMax M2MiniMax44
51GLM-4.7-FlashZhipu AI43
52DeepSeek-V3.2-ExpDeepSeek40
53Sarvam-30BSarvam AI36
54Nemotron 3 Super (120B A12B)NVIDIA31
55DeepSeek-V3.1DeepSeek30
56GLM-4.5Zhipu AI26
57GLM-4.5-AirZhipu AI21
58DeepSeek-R1-0528DeepSeek9

58 of 58 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC