Terminal-Bench 2.1
Terminal-Bench 2.1 is an agentic terminal coding benchmark measuring model performance on complex terminal-based programming tasks requiring multi-step reasoning and tool use. An updated version of Terminal-Bench 2.0.
Models ranked
16
tracked on this benchmark
Score band
88.8 – 51.7
best → lowest tracked
Snapshot trend
-13.60
Aug 12 → Aug 14 · 1 models
Leaderboard
Tracked models ranked by Score (higher is better).
Notes: OpenAI GA launch table reports 88.8% in standard mode. Sol Ultra multi-agent mode reports 91.9% on the same benchmark but is non-comparable to standard single-model rows and is documented here only in notes.
Notes: Model variant: Kimi K3 (max reasoning effort). Benchmark variant: Terminal-Bench 2.1. Harness/evaluator: KimiCode. Source posture: Moonshot/Kimi self-reported launch table; confidence: medium because the value is vendor-reported and not independently reproduced. Recommended seed value: 88.3.
Notes: Official Anthropic launch table starred this as a Claude Mythos 5 score, not a Claude Fable 5 score.
Notes: Do not join Terminal-bench 3.0 14.9 onto this slug.
Notes: xAI launch blog Terminal Bench 2.1 chart reports Grok 4.5 at 83.3%. Evaluator/source: xAI first-party launch chart only; blog footer cites competitor figures from developer system cards/leaderboards but does not name Grok 4.5's Terminal Bench 2.1 harness, agent scaffold, or reasoning level. Variant: Grok 4.5 (harness unspecified). Harness status: not specified in primary source — treat as directional vendor-reported chart, not reproducible from seed metadata. Confidence: medium for reported value, low for reproducibility/comparability. Recommended seed value: 83.3 pending harness documentation or independent leaderboard confirmation.
Notes: DAT-6271: Separate from GPT-5.5 Terminal-Bench 2.0 row; do not collapse benchmark versions.
Notes: Directly comparable with Claude Opus 4.8 Terminal-Bench 2.1. Do not collapse with Terminal-Bench 2.0.
Notes: Official Anthropic launch benchmark. Do not duplicate this score onto Terminal-Bench 2.0.
Notes: Official LongCat-2.0 launch chart reports 70.8. Evaluator/harness details are not fully reproducible from public text; keep separate from third-party Terminal-Bench leaderboard rows. Confidence: confirmed for reported value, medium for reproducibility.
Configuration: High Reasoning
How to read this benchmark
This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.
Trust this score when
- There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
- The model list covers the same version family you can actually deploy today.
- Top candidates overlap with your required routing and feature requirements.
Be cautious when
- There is only one benchmark snapshot or the dataset appears stale.
- The benchmark metric direction is opposite of your decision objective.
- The score difference between options is narrow and likely within implementation variance.
Last reviewed: May 28, 2026