Terminal-Bench 2.1
Terminal-Bench 2.1 is an agentic terminal coding benchmark measuring model performance on complex terminal-based programming tasks requiring multi-step reasoning and tool use. An updated version of Terminal-Bench 2.0.
Models ranked
27
tracked on this benchmark
Score band
90.6 – 17.4
best → lowest tracked
Snapshot trend
+48.73
Oct 7 → Oct 9 · 3 models
Leaderboard
Tracked models ranked by Score (higher is better).
Notes: OpenAI GA launch table reports 88.8% in standard mode. Sol Ultra multi-agent mode reports 91.9% on the same benchmark but is non-comparable to standard single-model rows and is documented here only in notes.
Notes: Model variant: Kimi K3 (max reasoning effort). Benchmark variant: Terminal-Bench 2.1. Harness/evaluator: KimiCode. Source posture: Moonshot/Kimi self-reported launch table; confidence: medium because the value is vendor-reported and not independently reproduced. Recommended seed value: 88.3.
Notes: Official Anthropic launch table starred this as a Claude Mythos 5 score, not a Claude Fable 5 score.
Notes: Do not join Terminal-bench 3.0 14.9 onto this slug.
Notes: xAI launch blog Terminal Bench 2.1 chart reports Grok 4.5 at 83.3%. Evaluator/source: xAI first-party launch chart only; blog footer cites competitor figures from developer system cards/leaderboards but does not name Grok 4.5's Terminal Bench 2.1 harness, agent scaffold, or reasoning level. Variant: Grok 4.5 (harness unspecified). Harness status: not specified in primary source — treat as directional vendor-reported chart, not reproducible from seed metadata. Confidence: medium for reported value, low for reproducibility/comparability. Recommended seed value: 83.3 pending harness documentation or independent leaderboard confirmation.
Notes: Separate from GPT-5.5 Terminal-Bench 2.0 row; do not collapse benchmark versions.
Notes: Directly comparable with Claude Opus 4.8 Terminal-Bench 2.1. Do not collapse with Terminal-Bench 2.0.
Notes: Official Anthropic launch benchmark. Do not duplicate this score onto Terminal-Bench 2.0.
Notes: Official LongCat-2.0 launch chart reports 70.8. Evaluator/harness details are not fully reproducible from public text; keep separate from third-party Terminal-Bench leaderboard rows. Confidence: confirmed for reported value, medium for reproducibility.
Notes: Self-reported; same Copilot harness. Pass rate 62.9; avg tokens 17.0K.
Configuration: High Reasoning
Notes: Aleph Alpha first-party HF model card Aleph-Alpha/Kolibri-1 (FP8 release checkpoint), Evaluation > Post-training table; same value in launch blog/landing page. Harness: Aleph-Alpha-Research eval-framework (Harbor for TerminalBench/SWE-Bench), Kolibri at reasoning effort high, each model at its documented context window. BF16 repo card reports 30.7. Self-reported by the lab.
How to read this benchmark
This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.
Trust this score when
- There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
- The model list covers the same version family you can actually deploy today.
- Top candidates overlap with your required routing and feature requirements.
Be cautious when
- There is only one benchmark snapshot or the dataset appears stale.
- The benchmark metric direction is opposite of your decision objective.
- The score difference between options is narrow and likely within implementation variance.
Last reviewed: May 28, 2026