active

Terminal-Bench 2.1

Metric: Score (higher is better)

Terminal-Bench 2.1 is an agentic terminal coding benchmark measuring model performance on complex terminal-based programming tasks requiring multi-step reasoning and tool use. An updated version of Terminal-Bench 2.0.

Models ranked

27

tracked on this benchmark

Score band

90.6 – 17.4

best → lowest tracked

Snapshot trend

+48.73

Oct 7 → Oct 9 · 3 models

Leaderboard

Tracked models ranked by Score (higher is better).

Compare candidates
#Model variant and provenanceScore
1
DeepSeek V4.1 Flash
Version: Terminal-Bench 2.1 Pass@1; DeepSeek Harness Minimal mode; 1M contextHarness: Not recordedEvaluator: Not recordedObserved: Sep 10, 2026Confidence: Not recordedSource
90.6
2
Xiaomi MiMo-V2.6-Pro
Version: Terminal Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Oct 9, 2026Confidence: Not recordedSource
89.9
3
GPT-5.6 Sol
Version: Terminal-Bench 2.1; standard mode; OpenAI GA launch harnessHarness: Not recordedEvaluator: Not recordedObserved: Jul 9, 2026Confidence: Not recordedSource

Notes: OpenAI GA launch table reports 88.8% in standard mode. Sol Ultra multi-agent mode reports 91.9% on the same benchmark but is non-comparable to standard single-model rows and is documented here only in notes.

88.8
4
Kimi K3
Version: Terminal-Bench 2.1; KimiCode harness; max reasoning effortHarness: Not recordedEvaluator: Not recordedObserved: Jul 16, 2026Confidence: Not recordedSource

Notes: Model variant: Kimi K3 (max reasoning effort). Benchmark variant: Terminal-Bench 2.1. Harness/evaluator: KimiCode. Source posture: Moonshot/Kimi self-reported launch table; confidence: medium because the value is vendor-reported and not independently reproduced. Recommended seed value: 88.3.

88.3
5
Claude Mythos 5
Version: Terminal-Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Jun 9, 2026Confidence: Not recordedSource

Notes: Official Anthropic launch table starred this as a Claude Mythos 5 score, not a Claude Fable 5 score.

88.0
6
Xiaomi MiMo-V2.6-Flash
Version: Terminal Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Oct 9, 2026Confidence: Not recordedSource
87.6
7
Qwen3.8-Max
Version: Terminal Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Aug 12, 2026Confidence: Not recordedSource
86.6
8
Gemini 3.7 Flash
Version: Terminal-bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Aug 13, 2026Confidence: Not recordedSource

Notes: Do not join Terminal-bench 3.0 14.9 onto this slug.

85.8
9
GLM-5.3-Flash
Version: Terminal-Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Aug 26, 2026Confidence: Not recordedSource
84.3
10
Grok 4.5
Version: Terminal Bench 2.1 (xAI first-party reported chart; harness not specified)Harness: Not recordedEvaluator: Not recordedObserved: Jul 8, 2026Confidence: Not recordedSource

Notes: xAI launch blog Terminal Bench 2.1 chart reports Grok 4.5 at 83.3%. Evaluator/source: xAI first-party launch chart only; blog footer cites competitor figures from developer system cards/leaderboards but does not name Grok 4.5's Terminal Bench 2.1 harness, agent scaffold, or reasoning level. Variant: Grok 4.5 (harness unspecified). Harness status: not specified in primary source — treat as directional vendor-reported chart, not reproducible from seed metadata. Confidence: medium for reported value, low for reproducibility/comparability. Recommended seed value: 83.3 pending harness documentation or independent leaderboard confirmation.

83.3
11
GLM-5.2
Version: Terminal-Bench 2.1 (% tasks completed)Harness: Not recordedEvaluator: Not recordedObserved: Jun 13, 2026Confidence: Not recordedSource
82.7
12
Claude Sonnet 5
Version: Terminal-Bench 2.1; mini-SWE-agent harness on GKE, 1x timeout rate, 3x memory ceiling, xhigh effort, 445 trialsHarness: Not recordedEvaluator: Not recordedObserved: Jun 30, 2026Confidence: Not recordedSource
80.4
13
GPT-5.5
Version: Terminal-Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Jun 18, 2026Confidence: Not recordedSource

Notes: Separate from GPT-5.5 Terminal-Bench 2.0 row; do not collapse benchmark versions.

78.2
14
GPT-5.5 Pro
Version: Terminal-Bench 2.1; standard GPT-5.5 weightsHarness: Not recordedEvaluator: Not recordedObserved: May 28, 2026Confidence: Not recordedSource

Notes: Directly comparable with Claude Opus 4.8 Terminal-Bench 2.1. Do not collapse with Terminal-Bench 2.0.

78.2
15
Motif 3
Version: Terminal-Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Oct 9, 2026Confidence: Not recordedSource
74.9
16
Claude Opus 4.8
Version: Terminal-Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: May 28, 2026Confidence: Not recordedSource

Notes: Official Anthropic launch benchmark. Do not duplicate this score onto Terminal-Bench 2.0.

74.6
17
Qwen3.8-27B
Version: Terminal Bench 2.1 (Terminus)Harness: TerminusEvaluator: Not recordedObserved: Aug 14, 2026Confidence: Not recordedSource
73.0
18
LongCat-2.0
Version: Terminal-Bench 2.1; Code Agent chart; LongCat launch benchmarkHarness: Not recordedEvaluator: Not recordedObserved: Jun 30, 2026Confidence: Not recordedSource

Notes: Official LongCat-2.0 launch chart reports 70.8. Evaluator/harness details are not fully reproducible from public text; keep separate from third-party Terminal-Bench leaderboard rows. Confidence: confirmed for reported value, medium for reproducibility.

70.8
19
MiniMax M3
Version: MiniMax-reportedHarness: Not recordedEvaluator: Not recordedObserved: Jun 1, 2026Confidence: Not recordedSource
66.0
20
MAI-Code-1.1-Flash
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Oct 8, 2026Confidence: Not recordedSource

Notes: Self-reported; same Copilot harness. Pass rate 62.9; avg tokens 17.0K.

62.9
21
Step 3.7 Flash
Version: Comparison: Step 3.5 Flash 53.37%, DeepSeek V4 Flash 62.0%, Gemini 3.5 Flash 76.2%, GPT-5.5 82.7%, Claude Opus 4.7 69.4%Harness: Not recordedEvaluator: Not recordedObserved: May 29, 2026Confidence: Not recordedSource
59.5
22
Nemotron 3 Ultra
Version: Terminal-Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Oct 7, 2026Confidence: Not recordedSource
56.4
23
Muse Glimmer

Configuration: High Reasoning

Version: TerminalBench 2.1 (with terminus2)Harness: Not recordedEvaluator: Not recordedObserved: Aug 10, 2026Confidence: Not recordedSource
51.7
24
Granite 4.2 30B
Version: Terminal-Bench 2.1Harness: Not recordedEvaluator: Not recordedObserved: Oct 7, 2026Confidence: Not recordedSource
29.2
25
Kolibri 1
Version: TerminalBench 2.1 (Harbor)Harness: Not recordedEvaluator: Not recordedObserved: Oct 3, 2026Confidence: Not recordedSource

Notes: Aleph Alpha first-party HF model card Aleph-Alpha/Kolibri-1 (FP8 release checkpoint), Evaluation > Post-training table; same value in launch blog/landing page. Harness: Aleph-Alpha-Research eval-framework (Harbor for TerminalBench/SWE-Bench), Kolibri at reasoning effort high, each model at its documented context window. BF16 repo card reports 30.7. Self-reported by the lab.

27.7
Showing the top 25 of 27 tracked models. Browse all models.

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

Last reviewed: May 28, 2026

Resources