LLM Reference
active

Terminal-Bench 2.1

Metric: Score (higher is better)

Terminal-Bench 2.1 is an agentic terminal coding benchmark measuring model performance on complex terminal-based programming tasks requiring multi-step reasoning and tool use. An updated version of Terminal-Bench 2.0.

Models ranked

16

tracked on this benchmark

Score band

88.8 – 51.7

best → lowest tracked

Snapshot trend

-13.60

Aug 12 → Aug 14 · 1 models

Leaderboard

Tracked models ranked by Score (higher is better).

Compare candidates
#Model variant and provenanceScore
1
GPT-5.6 Sol
Version: Terminal-Bench 2.1; standard mode; OpenAI GA launch harnessHarness: Not recordedEvaluator: Not recordedObserved: Jul 9, 2026Confidence: Not recordedSource

Notes: OpenAI GA launch table reports 88.8% in standard mode. Sol Ultra multi-agent mode reports 91.9% on the same benchmark but is non-comparable to standard single-model rows and is documented here only in notes.

88.8
2
Kimi K3
Version: Terminal-Bench 2.1; KimiCode harness; max reasoning effortHarness: Not recordedEvaluator: Not recordedObserved: Jul 16, 2026Confidence: Not recordedSource

Notes: Model variant: Kimi K3 (max reasoning effort). Benchmark variant: Terminal-Bench 2.1. Harness/evaluator: KimiCode. Source posture: Moonshot/Kimi self-reported launch table; confidence: medium because the value is vendor-reported and not independently reproduced. Recommended seed value: 88.3.

88.3
3
Claude Mythos 5
Version: Terminal-Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Jun 9, 2026Confidence: Not recordedSource

Notes: Official Anthropic launch table starred this as a Claude Mythos 5 score, not a Claude Fable 5 score.

88.0
4
Qwen3.8-Max
Version: Terminal Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Aug 12, 2026Confidence: Not recordedSource
86.6
5
Gemini 3.7 Flash
Version: Terminal-bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Aug 13, 2026Confidence: Not recordedSource

Notes: Do not join Terminal-bench 3.0 14.9 onto this slug.

85.8
6
Grok 4.5
Version: Terminal Bench 2.1 (xAI first-party reported chart; harness not specified)Harness: Not recordedEvaluator: Not recordedObserved: Jul 8, 2026Confidence: Not recordedSource

Notes: xAI launch blog Terminal Bench 2.1 chart reports Grok 4.5 at 83.3%. Evaluator/source: xAI first-party launch chart only; blog footer cites competitor figures from developer system cards/leaderboards but does not name Grok 4.5's Terminal Bench 2.1 harness, agent scaffold, or reasoning level. Variant: Grok 4.5 (harness unspecified). Harness status: not specified in primary source — treat as directional vendor-reported chart, not reproducible from seed metadata. Confidence: medium for reported value, low for reproducibility/comparability. Recommended seed value: 83.3 pending harness documentation or independent leaderboard confirmation.

83.3
7
GLM-5.2
Version: Terminal-Bench 2.1 (% tasks completed)Harness: Not recordedEvaluator: Not recordedObserved: Jun 13, 2026Confidence: Not recordedSource
82.7
8
Claude Sonnet 5
Version: Terminal-Bench 2.1; mini-SWE-agent harness on GKE, 1x timeout rate, 3x memory ceiling, xhigh effort, 445 trialsHarness: Not recordedEvaluator: Not recordedObserved: Jun 30, 2026Confidence: Not recordedSource
80.4
9
GPT-5.5
Version: Terminal-Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Jun 18, 2026Confidence: Not recordedSource

Notes: DAT-6271: Separate from GPT-5.5 Terminal-Bench 2.0 row; do not collapse benchmark versions.

78.2
10
GPT-5.5 Pro
Version: Terminal-Bench 2.1; standard GPT-5.5 weightsHarness: Not recordedEvaluator: Not recordedObserved: May 28, 2026Confidence: Not recordedSource

Notes: Directly comparable with Claude Opus 4.8 Terminal-Bench 2.1. Do not collapse with Terminal-Bench 2.0.

78.2
11
Claude Opus 4.8
Version: Terminal-Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: May 28, 2026Confidence: Not recordedSource

Notes: Official Anthropic launch benchmark. Do not duplicate this score onto Terminal-Bench 2.0.

74.6
12
Qwen3.8-27B
Version: Terminal Bench 2.1 (Terminus)Harness: TerminusEvaluator: Not recordedObserved: Aug 14, 2026Confidence: Not recordedSource
73.0
13
LongCat-2.0
Version: Terminal-Bench 2.1; Code Agent chart; LongCat launch benchmarkHarness: Not recordedEvaluator: Not recordedObserved: Jun 30, 2026Confidence: Not recordedSource

Notes: Official LongCat-2.0 launch chart reports 70.8. Evaluator/harness details are not fully reproducible from public text; keep separate from third-party Terminal-Bench leaderboard rows. Confidence: confirmed for reported value, medium for reproducibility.

70.8
14
MiniMax M3
Version: MiniMax-reportedHarness: Not recordedEvaluator: Not recordedObserved: Jun 1, 2026Confidence: Not recordedSource
66.0
15
Step 3.7 Flash
Version: Comparison: Step 3.5 Flash 53.37%, DeepSeek V4 Flash 62.0%, Gemini 3.5 Flash 76.2%, GPT-5.5 82.7%, Claude Opus 4.7 69.4%Harness: Not recordedEvaluator: Not recordedObserved: May 29, 2026Confidence: Not recordedSource
59.5
16
Muse Glimmer

Configuration: High Reasoning

Version: TerminalBench 2.1 (with terminus2)Harness: Not recordedEvaluator: Not recordedObserved: Aug 10, 2026Confidence: Not recordedSource
51.7

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

Last reviewed: May 28, 2026

Resources