LLM Reference
activeCodingAgents

Terminal-Bench 2.0

Metric: % Tasks Completed (higher is better)Introduced: 2026

Second-generation terminal agent benchmark with 89 high-quality tasks spanning software engineering, machine learning, security, data science, and other real shell environments. High benchmark score alone doesn't make a model the right pick — weigh it against pricing, API availability, and release date.

Models ranked

28

tracked on this benchmark

Score band

82.7 – 10.1

best → lowest tracked

Snapshot trend

+20.90

Jun 9 → Jul 3 · 1 models

Leaderboard

Tracked models ranked by % Tasks Completed (higher is better).

Compare candidates
#Model variant and provenanceScore
1
GPT-5.5
Version: Terminal-Bench 2.0 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
82.7
2
GPT-5.5 Pro
Version: Terminal-Bench 2.0; standard GPT-5.5 weightsHarness: Not recordedEvaluator: Not recordedObserved: Apr 23, 2026Confidence: Not recordedSource

Notes: Not directly comparable to Claude Opus 4.8 in this datapack; no Opus 4.8 Terminal-Bench 2.0 row was found. Use Terminal-Bench 2.1 for the head-to-head comparison.

82.7
3
GPT-5.3-Codex
Version: Terminal-Bench 2.0 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
77.3
4
Gemini 3.5 Flash
Version: Terminal-Bench 2.0 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
76.2
5
GPT-5.4
Version: Terminal-Bench 2.0 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
75.1
6
Qwen3.7-Max
Version: TerminusHarness: Not recordedEvaluator: Not recordedObserved: May 20, 2026Confidence: Not recordedSource
69.7
7
Claude Opus 4.7
Version: Terminal-Bench 2.0 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
69.4
8
Composer 2.5
Version: Terminal-Bench 2.0Harness: Not recordedEvaluator: Not recordedObserved: May 21, 2026Confidence: Not recordedSource

Notes: Confidence: medium. DAT-4756 secondary-source score for the Cursor IDE-native Composer 2.5 agent; useful for comparison but still reflects the full Cursor agent system rather than a standalone base-model run.

69.3
9
Xiaomi MiMo-V2.5-Pro
Version: Terminal-Bench 2.0 (score)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
68.4
10
Kimi K2.6
Version: Terminal-Bench 2.0 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
66.7
11
MiniMax M3
Version: Terminal-Bench 2.0 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
66.0
12
Claude Opus 4.6
Version: Terminal-Bench 2.0 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
65.4
13
Qwen3.6-Max
Version: Terminal-Bench 2.0 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
65.4
14
GPT-5.2 Codex
Version: Terminal-Bench 2.0; OpenAI launch comparison; GPT-5.2-Codex xhighHarness: Not recordedEvaluator: Not recordedObserved: Dec 18, 2025Confidence: Not recordedSource

Notes: OpenAI GPT-5.3-Codex appendix reports GPT-5.2-Codex 64.0%, GPT-5.2 62.2%, GPT-5.1-Codex-Max 58.1%. Evaluator: OpenAI self-reported launch benchmark. Harness details are not fully reproducible from public text; compare separately from tbench.ai third-party agent rows. Confidence: confirmed for reported value, medium for reproducibility.

64.0
15
GLM-5.1
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Apr 7, 2026Confidence: Not recordedSource
63.5
16
Composer 2
Version: Official Harbor evaluation framework (score)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
61.7
17
Claude Sonnet 4.6
Version: Terminal-Bench 2.0Harness: Not recordedEvaluator: Not recordedObserved: Feb 17, 2026Confidence: Not recordedSource

Notes: Confidence: medium. DAT-4934 found this via DataCamp citing Anthropic release notes; terminal-bench.com leaderboard was not independently accessible, and Composer 2.5 comparisons remain cross-harness.

59.1
18
DeepSeek V4 Pro
Version: independent evaluation (BenchLM, June 2026)Harness: Not recordedEvaluator: Not recordedObserved: Jun 18, 2026Confidence: Not recordedSource
59.1
19
MiniMax M2.7
Version: OpenRouter Terminal Bench 2 listingHarness: Not recordedEvaluator: Not recordedObserved: Jun 1, 2026Confidence: Not recordedSource
57.0
20
DeepSeek V4 Flash
Version: Terminal Bench 2.0 accuracy, V4-Flash Think MaxHarness: Not recordedEvaluator: Not recordedObserved: Jul 3, 2026Confidence: Not recordedSource
56.9
21
MAI-Code-1-Flash
Version: Terminal-Bench 2.0 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
54.8
22
Hunyuan Hy3 Preview
Version: Terminal-Bench 2.0 (score)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
54.4
23
Qwen3.5-397B-A17B
Version: Terminal-Bench 2.0 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
52.5
24
Kimi K2.5
Version: Terminal-Bench 2.0 (accuracy%)Harness: Not recordedEvaluator: Not recordedObserved: Jun 7, 2026Confidence: Not recordedSource
50.8
25
MAI-Thinking-1
Version: Terminal-Bench 2.0Harness: Not recordedEvaluator: Not recordedObserved: Jun 2, 2026Confidence: Not recordedSource

Notes: Microsoft AI technical report Table 11 result for MAI-Thinking-1.

46.0
Showing the top 25 of 28 tracked models. Browse all models.

How to read this benchmark

This benchmark scores models where higher is better. Scores are useful for directional filtering and shortlisting — not for universal quality ranking. Prefer benchmarks closest to your workload, then validate the linked model pages for pricing, context window, and provider availability.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

FAQ

What does the Terminal-Bench 2.0 benchmark measure?

Second-generation terminal agent benchmark with 89 high-quality tasks spanning software engineering, machine learning, security, data science, and other real shell environments. On this page it lists 28 tracked model variants where higher is better.

Is a higher Terminal-Bench 2.0 score always better?

For this benchmark, higher is better. A high score helps you shortlist, but confirm pricing, context window, and provider availability on each model page before committing — the top scorer is not always the right pick for your workload or budget.

How current is this Terminal-Bench 2.0 data?

This benchmark was last reviewed on May 21, 2026. The tracked score average moved +20.90 points across the last 3 snapshots.

Related benchmarks

Last reviewed: May 21, 2026

Resources