LLM Reference
activeAgents

BrowseComp

Metric: Score (higher is better)Introduced: 2025

OpenAI benchmark (released April 2025) measuring browsing agents' ability to locate hard-to-find, entangled facts on the open web. 1,266 questions with easy-to-verify short answers; widely reported across frontier model launches.

Models ranked

18

tracked on this benchmark

Score band

91.2 – 60.6

best → lowest tracked

Snapshot trend

+18.00

Jul 3 → Jul 16 · 1 models

Leaderboard

Tracked models ranked by Score (higher is better).

Compare candidates
#Model variant and provenanceScore
1
Kimi K3
Version: BrowseComp; 300K-token context compaction; max reasoning effortHarness: Not recordedEvaluator: Not recordedObserved: Jul 16, 2026Confidence: Not recordedSource

Notes: Model variant: Kimi K3 (max reasoning effort) with 300K-token context compaction. Benchmark variant: BrowseComp. Harness/evaluator not disclosed by Moonshot. Source posture: Moonshot/Kimi self-reported launch table; confidence: medium because the value is vendor-reported and not independently reproduced. Recommended seed value: 91.2. Kimi also reports 90.4 with 1M context and no context management; do not conflate the variants.

91.2
2
GPT-5.6 Sol
Version: BrowseComp; standard Sol; agentic browsing; OpenAI GA launch harnessHarness: Not recordedEvaluator: Not recordedObserved: Jul 9, 2026Confidence: Not recordedSource

Notes: Standard single-agent BrowseComp row; kept separate from Sol Ultra multi-agent BrowseComp.

90.4
3
Gemini 3.1 Pro Preview
Version: BrowseComp (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
85.9
4
Claude Sonnet 5
Version: BrowseComp single-agent; adaptive thinking at maximum effort, 10M-token limit with context compaction triggered at 200kHarness: Not recordedEvaluator: Not recordedObserved: Jun 30, 2026Confidence: Not recordedSource
84.7
5
GPT-5.5
Version: BrowseComp (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
84.4
6
Claude Opus 4.6
Version: BrowseComp (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
84.0
7
Nex-N2-Pro
Version: BrowseComp — self-reportedHarness: Not recordedEvaluator: Not recordedObserved: Jun 4, 2026Confidence: Not recordedSource

Notes: DAT-5798: Self-reported. Not independently verified.

83.7
8
MiniMax M3
Version: MiniMax-reportedHarness: Not recordedEvaluator: Not recordedObserved: Jun 1, 2026Confidence: Not recordedSource
83.5
9
DeepSeek V4 Pro
Version: BrowseComp (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
83.4
10
Kimi K2.6
Version: BrowseComp (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
83.2
11
LongCat-2.0
Version: BrowseComp; Search Agent chart; LongCat launch benchmarkHarness: Not recordedEvaluator: Not recordedObserved: Jun 30, 2026Confidence: Not recordedSource

Notes: Official LongCat-2.0 launch chart reports 79.86. Evaluator/harness details are not fully reproducible from public text. Confidence: confirmed for reported value, medium for reproducibility.

79.9
12
Claude Opus 4.7
Version: BrowseComp (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
79.3
13
Step 3.7 Flash
Version: From official StepFun blog post (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
75.8
14
Agents-A1
Version: BrowseComp accuracy, vendor-reported on Agents-A1 model cardHarness: Not recordedEvaluator: Not recordedObserved: Jun 26, 2026Confidence: Not recordedSource

Notes: Official Intern Science model card reports Agents-A1 at 75.51 on BrowseComp. Evaluator/source: Intern Science self-reported model-card/project benchmark table. Public harness details are not fully reproducible from the seed row. Confidence: medium for reported value, low-medium for reproducibility. Recommended seed value: 75.51.

75.5
15
DeepSeek V4 Flash
Version: BrowseComp pass@1, V4-Flash Think MaxHarness: Not recordedEvaluator: Not recordedObserved: Jul 3, 2026Confidence: Not recordedSource
73.2
16
Hunyuan Hy3 Preview
Version: BrowseComp (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
67.1
17
Qwen3.5-122B-A10B
Version: BrowseComp (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
63.8
18
Kimi K2.5
Version: BrowseComp (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
60.6

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

Related benchmarks

Last reviewed: Jun 7, 2026

Resources