activeReasoning

Humanity's Last Exam

Metric: Score (higher is better)

Multi-domain expert-level exam covering science, math, humanities, and more

Models ranked

34

tracked on this benchmark

Score band

64.5 – 16.2

best → lowest tracked

Snapshot trend

-6.70

Jul 16 → Sep 10 · 1 models

Leaderboard

Tracked models ranked by Score (higher is better).

Compare candidates
#Model variant and provenanceScore
1
Claude Mythos 5
Version: HLE (with tools)Harness: Not recordedEvaluator: Not recordedObserved: Jun 9, 2026Confidence: Not recordedSource

Notes: Official Anthropic launch table starred the 64.5% HLE with-tools score as Claude Mythos 5. The 59.0% no-tools variant is recorded in the handoff but not seeded because modelBenchmark rows are unique by modelSlug and benchmarkSlug.

64.5
2
GLM-5.3-Flash
Version: HLE with toolsHarness: Not recordedEvaluator: Not recordedObserved: Aug 26, 2026Confidence: Not recordedSource
55.3
3
Claude Opus 4.7
Version: HLE with tools (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
54.7
4
Claude Opus 4.6
Version: HLE with tools enabled (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
53.0
5
Gemini 3.1 Pro Preview
Version: HLE with tools (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
51.4
6
Kimi K2.5
Version: HLE-Full with tools (agentic) (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
50.2
7
Step 3.7 Flash
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: May 29, 2026Confidence: Not recordedSource
47.2
8
Kimi K3
Version: HLE-Full without tools; max reasoning effortHarness: Not recordedEvaluator: Not recordedObserved: Jul 16, 2026Confidence: Not recordedSource

Notes: Model variant: Kimi K3 (max reasoning effort). Benchmark variant: HLE-Full without tools. Harness/evaluator not disclosed by Moonshot. Source posture: Moonshot/Kimi self-reported launch table; confidence: medium because the value is vendor-reported and not independently reproduced. Recommended seed value: 43.5.

43.5
9
Claude Sonnet 5
Version: Humanity's Last Exam; no tools; reasoning-only; thinking auto; 1M total-token cap; Claude Opus 4.6 graderHarness: Not recordedEvaluator: Not recordedObserved: Jun 30, 2026Confidence: Not recordedSource
43.2
10
GPT-5.4
Version: HLE (xhigh setting) (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
41.6
11
Qwen3.7-Max
Version: HLE no tools (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
41.4
12
GPT-5.5
Version: HLE without tools (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
41.4
13
GLM-5.2
Version: HLE (accuracy, text-only assumed)Harness: Not recordedEvaluator: Not recordedObserved: Jun 13, 2026Confidence: Not recordedSource
40.5
14
Gemini 3.5 Flash
Version: HLE (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
40.2
15
ERNIE 5.0
Version: Humanity's Last Exam (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
39.0
16
DeepSeek V4 Pro
Version: HLE (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
37.7
17
DeepSeek V4.1 Flash
Version: HLE Pass@1; text-only; max reasoning effortHarness: Not recordedEvaluator: Not recordedObserved: Sep 10, 2026Confidence: Not recordedSource
36.8
18
DeepSeek V4 Flash
Version: Humanity's Last Exam pass@1, no tools, V4-Flash Think MaxHarness: Not recordedEvaluator: Not recordedObserved: Jul 3, 2026Confidence: Not recordedSource
34.8
19
Kimi K2.6
Version: HLE-Full without tools (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
34.7
20
Claude Sonnet 4.6
Version: HLE without tools (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
33.2
21
Grok 4.20
Version: HLE (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
32.2
22
GLM-5.1
Version: HLE text-only (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
31.0
23
GLM-5
Version: HLE text-only (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
30.5
24
Grok 4.20 Reasoning
Version: HLE for Grok 4 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
30.0
25
Qwen3.6 Max Preview
Version: HLE Thinking Level Max (no tools), from DataLearner (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
28.8
Showing the top 25 of 34 tracked models. Browse all models.

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

Related benchmarks

Last reviewed: May 20, 2026

Resources