BrowseComp
OpenAI benchmark (released April 2025) measuring browsing agents' ability to locate hard-to-find, entangled facts on the open web. 1,266 questions with easy-to-verify short answers; widely reported across frontier model launches.
Models ranked
18
tracked on this benchmark
Score band
91.2 – 60.6
best → lowest tracked
Snapshot trend
+18.00
Jul 3 → Jul 16 · 1 models
Leaderboard
Tracked models ranked by Score (higher is better).
Notes: Model variant: Kimi K3 (max reasoning effort) with 300K-token context compaction. Benchmark variant: BrowseComp. Harness/evaluator not disclosed by Moonshot. Source posture: Moonshot/Kimi self-reported launch table; confidence: medium because the value is vendor-reported and not independently reproduced. Recommended seed value: 91.2. Kimi also reports 90.4 with 1M context and no context management; do not conflate the variants.
Notes: Standard single-agent BrowseComp row; kept separate from Sol Ultra multi-agent BrowseComp.
Notes: DAT-5798: Self-reported. Not independently verified.
Notes: Official LongCat-2.0 launch chart reports 79.86. Evaluator/harness details are not fully reproducible from public text. Confidence: confirmed for reported value, medium for reproducibility.
Notes: Official Intern Science model card reports Agents-A1 at 75.51 on BrowseComp. Evaluator/source: Intern Science self-reported model-card/project benchmark table. Public harness details are not fully reproducible from the seed row. Confidence: medium for reported value, low-medium for reproducibility. Recommended seed value: 75.51.
How to read this benchmark
This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.
Trust this score when
- There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
- The model list covers the same version family you can actually deploy today.
- Top candidates overlap with your required routing and feature requirements.
Be cautious when
- There is only one benchmark snapshot or the dataset appears stale.
- The benchmark metric direction is opposite of your decision objective.
- The score difference between options is narrow and likely within implementation variance.
Related benchmarks
Last reviewed: Jun 7, 2026