LLM Reference
activeReasoning

Humanity's Last Exam

Metric: Score (higher is better)

Multi-domain expert-level exam covering science, math, humanities, and more High benchmark score alone doesn't make a model the right pick — weigh it against pricing, API availability, and release date.

Models ranked

32

tracked on this benchmark

Score band

64.5 – 16.2

best → lowest tracked

Snapshot trend

+18.10

Jul 1 → Jul 16 · 1 models

Leaderboard

Tracked models ranked by Score (higher is better).

Compare candidates
#Model variant and provenanceScore
1
Claude Mythos 5
Version: HLE (with tools)Harness: Not recordedEvaluator: Not recordedObserved: Jun 9, 2026Confidence: Not recordedSource

Notes: Official Anthropic launch table starred the 64.5% HLE with-tools score as Claude Mythos 5. The 59.0% no-tools variant is recorded in the handoff but not seeded because modelBenchmark rows are unique by modelSlug and benchmarkSlug.

64.5
2
Claude Opus 4.7
Version: HLE with tools (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
54.7
3
Claude Opus 4.6
Version: HLE with tools enabled (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
53.0
4
Gemini 3.1 Pro Preview
Version: HLE with tools (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
51.4
5
Kimi K2.5
Version: HLE-Full with tools (agentic) (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
50.2
6
Step 3.7 Flash
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: May 29, 2026Confidence: Not recordedSource
47.2
7
Kimi K3
Version: HLE-Full without tools; max reasoning effortHarness: Not recordedEvaluator: Not recordedObserved: Jul 16, 2026Confidence: Not recordedSource

Notes: Model variant: Kimi K3 (max reasoning effort). Benchmark variant: HLE-Full without tools. Harness/evaluator not disclosed by Moonshot. Source posture: Moonshot/Kimi self-reported launch table; confidence: medium because the value is vendor-reported and not independently reproduced. Recommended seed value: 43.5.

43.5
8
Claude Sonnet 5
Version: Humanity's Last Exam; no tools; reasoning-only; thinking auto; 1M total-token cap; Claude Opus 4.6 graderHarness: Not recordedEvaluator: Not recordedObserved: Jun 30, 2026Confidence: Not recordedSource
43.2
9
GPT-5.4
Version: HLE (xhigh setting) (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
41.6
10
Qwen3.7-Max
Version: HLE no tools (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
41.4
11
GPT-5.5
Version: HLE without tools (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
41.4
12
GLM-5.2
Version: HLE (accuracy, text-only assumed)Harness: Not recordedEvaluator: Not recordedObserved: Jun 13, 2026Confidence: Not recordedSource
40.5
13
Gemini 3.5 Flash
Version: HLE (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
40.2
14
ERNIE 5.0
Version: Humanity's Last Exam (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
39.0
15
DeepSeek V4 Pro
Version: HLE (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
37.7
16
DeepSeek V4 Flash
Version: Humanity's Last Exam pass@1, no tools, V4-Flash Think MaxHarness: Not recordedEvaluator: Not recordedObserved: Jul 3, 2026Confidence: Not recordedSource
34.8
17
Kimi K2.6
Version: HLE-Full without tools (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
34.7
18
Claude Sonnet 4.6
Version: HLE without tools (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
33.2
19
Grok 4.20
Version: HLE (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
32.2
20
GLM-5.1
Version: HLE text-only (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
31.0
21
GLM-5
Version: HLE text-only (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
30.5
22
Grok 4.20 Reasoning
Version: HLE for Grok 4 (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
30.0
23
Qwen3.6 Max Preview
Version: HLE Thinking Level Max (no tools), from DataLearner (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
28.8
24
Qwen3.5-397B-A17B
Version: HLE with CoT, no tools, from official model card (accuracy)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
28.7
25
Grok 4
Version: Humanity's Last Exam no tools, pass@1, base Grok 4Harness: Not recordedEvaluator: Not recordedObserved: Jul 1, 2026Confidence: Not recordedSource
25.4
Showing the top 25 of 32 tracked models. Browse all models.

How to read this benchmark

This benchmark scores models where higher is better. Scores are useful for directional filtering and shortlisting — not for universal quality ranking. Prefer benchmarks closest to your workload, then validate the linked model pages for pricing, context window, and provider availability.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

FAQ

What does the Humanity's Last Exam benchmark measure?

Multi-domain expert-level exam covering science, math, humanities, and more On this page it lists 32 tracked model variants where higher is better.

Is a higher Humanity's Last Exam score always better?

For this benchmark, higher is better. A high score helps you shortlist, but confirm pricing, context window, and provider availability on each model page before committing — the top scorer is not always the right pick for your workload or budget.

How current is this Humanity's Last Exam data?

This benchmark was last reviewed on May 20, 2026. The tracked score average moved +18.10 points across the last 3 snapshots.

Related benchmarks

Last reviewed: May 20, 2026

Resources