What is Terminal-Bench 4.0?
Terminal-Bench 4.0 is a set of 66 tasks that an AI agent has to finish inside a real Linux terminal sandbox. Each task ships its own container, a solution written by a human, and tests that check the end state of the machine when the agent stops. The agent either leaves the system in the state the tests expect or it fails. Nobody grades the transcript.
The tasks are long and they are not all software. Version 4.0 spans software, science, machine learning, operations, hardware, security and media work. Harbor and the Laude Institute host the benchmark, and Snorkel AI supports it through its Open Benchmarks Grants program, per Snorkel's leaderboard page.
The design comes from the Terminal-Bench paper by Merrill et al., submitted on January 17, 2026, with more than 80 authors from Stanford, the Laude Institute, Anthropic, Snorkel AI and other groups, Alex Dimakis and Ludwig Schmidt among them. That paper describes version 2.0 and states the problem: "Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models." What sets Terminal-Bench apart, in the authors' words, is its "emphasis on diverse, long-horizon tasks collected from experts," run "inside a real terminal shell" rather than "the synthetic environments used by other benchmarks."
The scores on this page come from three places. Vals AI runs every model independently on one harness. Snorkel mirrors the official tbench.ai board, and Anthropic published its own numbers with Claude Opus 5.5. Several primary pieces were not retrievable on October 6, 2026: OpenAI's GPT-6 Astra launch post, Google's Gemini 4 evaluation methodology page and the data rows of the tbench.ai board itself. Nothing retrieved states the board's per-task timeout or step cap, OpenAI's harness and effort for Astra, Vals' per-row reasoning effort, Anthropic's harness for any row or effort for Fable 5.1 and Opus 5, who ran each mirrored board row, or the mirror's snapshot date.
How the tasks and harness work
Vals counts 66 community-contributed tasks across seven categories.
| Category | Tasks |
|---|---|
| Software | 18 |
| Science | 14 |
| ML | 11 |
| Operations | 9 |
| Hardware | 5 |
| Security | 5 |
| Media | 4 |
Software is the largest category and still covers barely more than a quarter of the set. A model that is strong at repository bug fixing and weak at, say, a science pipeline or a hardware task will not top this benchmark.
Tasks run on Harbor, which comes with container sandbox providers already integrated. One command runs the full set:
Harbor Hub recommends Modal or Daytona as the sandbox for 4.0, and lists GPU and multi-container tasks in the set. The paper's own runs used Daytona with 32 to 100 containers in parallel. The task source is public on GitHub under Apache-2.0.
The benchmark does not force a particular tool. Section 3.1 of the paper PDF says "we do not explicitly require agents to use a terminal as their sole tool," and agents "are free to manipulate the container however they please." That freedom makes model comparisons messy, because "agent and model performance are hard to decouple" and each lab tunes its own scaffold to its own models. The authors built Terminus 2, "a simple scaffold" with "a single tool, a headless terminal," as a neutral testbed. Vals takes the same approach with mini-swe-agent, "a minimal, model-agnostic agent harness" where bash is the only tool. The official board does the opposite and runs each vendor's own agent, Claude Code for Anthropic, Codex for OpenAI and Grok Build for Grok. That choice is why the board can't rank vendors against each other, which comes up again in the leaderboard section.
Agents have internet access so they can install packages and search the web. Under Vals' setup each task gets an 8-hour agent phase, and each verifier has its own task-specific timeout. Running past either one scores zero for the task. Vals uses Daytona as the sandbox.
For version 2.0, the paper's authors also screened tasks with an adversarial exploit agent that hunts for ways to cheat the verifiers, a form of red teaming aimed at the benchmark rather than the model. A pass on a screened task reflects the intended solution path. No retrieved source describes the same screening for the 4.0 set.
How scoring works
Each task is pass/fail. The verifier's full test suite has to pass, there is no partial credit, and the per-task metric is pass@1. No LLM judge is involved anywhere.
The three publishers aggregate differently:
- Vals reports avg@3, the mean of three runs per task, so each model has 198 task attempts. It ranks by a headline score and, for five Anthropic rows, also publishes a score that counts attempts served by an older model as failures. More on that below.
- The official board reports resolution rate with "a margin reflecting run-to-run variance," plus total tokens and total cost across all 66 tasks. Snorkel notes those totals are "not normalized for effort setting."
- Anthropic reports single scores with a standard error, ±2.6 points for Opus 5.5 and ±1.6 to 2 points for its other models.
The board's margins run from ±2.6 to ±3.8 points. Two models within three points of each other on the same table are not separated by this benchmark.
Versions
Version 2.0, the one in the paper, had 89 tasks. Vals still runs an 89-task Terminal-Bench 2.1, which it describes as using "89 tasks with unique categories ranging from model training to system administration." No retrieved source says what 2.1 changed from 2.0 or when it shipped.
Version 3.0.0 had 74 tasks. Version 4.0.0 removed 8 of them, revised 20 and added none, leaving 66. Snorkel's FAQ gives the reason: "Harbor's v4.0.0 release calibrates task resources, fixes broken tasks, and retires saturated ones, tasks frontier agents had already maxed out and that no longer separate models from each other. The eight removed tasks fall under that umbrella." Nobody has published a per-task reason for the 20 revisions.
Compared with the 89-task 2.x line, Vals says "all 66 tasks are new with no overlap from the previous 89-task version." Domain categories replace the old difficulty tiers, every task allows the 8-hour agent phase, and the scope reaches well past software engineering. Vals draws the conclusion directly: "Because the task set is entirely different, scores are not comparable across versions." A model's 2.0 score and its 4.0 score measure different things, and the gap between them says nothing about progress.
For a buyer, the 4.0 changes mean the set still separates current frontier models. The tasks that every top agent already solved are gone.
Current leaderboard
Vals AI independent run
Vals AI runs all 43 listed models through mini-swe-agent with bash as the only tool, three runs per task, updated October 1, 2026. It is the one source where every row shares a harness, which makes it the table to rank by.
The headline ranking hides a problem that Vals caught. When Vals sent requests to Anthropic's API, some were answered by an older model. Vals' footnote: "Five Anthropic rows include provider-side fallback. Opus 5.5 had 22 of 198 task attempts served by Opus 5 or Claude Opus 4.8; counting those as failures lowers it from 65.15% to 58.08%, behind Sonnet 5.5 and Astra." The table shows both numbers.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 65.15% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. Fallback-adjusted 58.08%, 22 of 198 attempts served by Opus 5 or Opus 4.8. $13.20 and 1h04m per task |
| Claude Sonnet 5.5 | Anthropic | 64.14% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. Fallback-adjusted 62.63%, 4 attempts served by Claude Sonnet 5. $16.51 and 1h22m per task |
| GPT-6 Astra | OpenAI | 59.60% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. No fallback note. $9.58 and 35m42s per task |
| Claude Fable 5.1 | Anthropic | 58.08% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. Fallback-adjusted 50.00%, 22 attempts served by Opus 5. $17.18 and 1h00m per task |
| Gemini 4 Argon | 57.58% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. No fallback note. $17.64 and 53m35s per task | |
| GPT-6.1 Sol | OpenAI | 55.05% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. No fallback note. $1.72 and 45m12s per task |
| Claude Opus 5 | Anthropic | 53.53% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. Fallback-adjusted 53.03%, 1 attempt served by Opus 4.8. $18.60 and 1h07m per task |
| GPT-6 Sol | OpenAI | 44.44% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. No fallback note. $5.79 and 34m56s per task |
| Claude Fable 5 | Anthropic | 41.41% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. Fallback-adjusted 33.84%, 42 attempts served by Opus 4.8. $30.34 and 1h08m per task |
| GLM 5.3 | Not listed | 38.89% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. No fallback note. $9.37 and 1h23m per task |
| GPT-5.6 Sol | OpenAI | 37.88% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. No fallback note. $7.98 and 33m46s per task |
| Qwen 3.8 Max | Not listed | 34.34% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. Nine provider refusals scored as failures. $10.63 and 2h28m per task |
| Step 5 Preview | Not listed | 31.82% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. No fallback note. $3.11 and 1h58m per task |
| MiMo V2.6 Pro | Not listed | 31.31% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. No fallback note. $0.50 and 2h42m per task |
| Grok 4.7 | Not listed | 28.79% | mini-swe-agent, bash only, avg@3 | 2026-10-01 | Vals. No fallback note. $18.09 and 50m21s per task |
One harness, one dataset and one scoring rule, so every row compares with every other row in this table and with nothing outside it. Score is Vals' headline. Organization is filled only where the sources name the vendor. Rows 16 to 43 are omitted; the best of them are Muse Spark 1.3 Max at 24.75%, Claude Opus 4.8 at 23.23%, GPT-5.6 Terra at 22.73% and Gemini 3.8 Flash at 19.19%. Vals also scored three refusals each on Sonnet 5 and Gemini 3.8 Flash as failures, and lists none for Opus 5.5.
Read the table with the adjusted column and the order changes. Sonnet 5.5 leads at 62.63%, 4.55 points ahead of Opus 5.5, which it beats at $16.51 per task against $13.20. GPT-6 Astra is second at 59.60% with no fallback note, then Opus 5.5 at 58.08% and Gemini 4 Argon at 57.58%. Fable 5.1 takes the largest hit, 8.08 points, from 58.08% to 50.00%. I rank by the adjusted score. It counts only the passes the requested model produced itself, so it is the floor for what that model did on its own, while the headline mixes in answers from older models.
On adjusted scores the cost frontier has three models. GPT-6.1 Sol scores 55.05% at $1.72 per task, GPT-6 Astra 59.60% at $9.58, and Claude Sonnet 5.5 62.63% at $16.51. Everything else is dominated. Astra beats Opus 5.5 by 1.5 points and costs $3.62 less per task, and GPT-6.1 Sol beats GPT-6 Sol by 10.6 points at under a third of the cost. On headline scores the frontier is GPT-6.1 Sol, Astra and Opus 5.5, and Opus 5.5 dominates Sonnet 5.5. The fallback correction alone decides which Claude model belongs on the frontier.
Official board, as mirrored by Snorkel
The tbench.ai leaderboard returned no data rows. Snorkel's page mirrors it, so the tables below are the official board as mirrored by Snorkel. Each vendor's model runs in its own agent at a stated effort level, and the board reports tokens and cost as totals over all 66 tasks. Those totals are on a different basis from Vals' per-task cost, which is why they stay out of the chart.
Claude Code, max effort:
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | 57.9% ±3.8 | Claude Code, max effort | 2026-09-01 | Snorkel mirror. 2.7B tokens, $6.2k over 66 tasks |
| Claude Fable 5 | Anthropic | 44.5% ±3.8 | Claude Code, max effort | 2026-06-09 | Snorkel mirror. 3.8B tokens, $7.3k over 66 tasks |
| GLM-5.3 | Not listed | 41.8% ±3.2 | Claude Code, max effort | 2026-08-14 | Snorkel mirror. 8.7B tokens, $2.7k over 66 tasks |
| Claude Opus 4.8 | Anthropic | 23.6% ±3.6 | Claude Code, max effort | 2026-05-28 | Snorkel mirror. 6.4B tokens, $6.5k over 66 tasks |
| Claude Sonnet 5 | Anthropic | 12.4% ±3.1 | Claude Code, max effort | 2026-06-30 | Snorkel mirror. 21.6B tokens, $9.6k over 66 tasks |
Same agent and effort, so rows compare within this table only. Date is the model release date the board lists.
Claude Code, xhigh effort:
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5 | Anthropic | 53.9% ±3.2 | Claude Code, xhigh effort | 2026-07-24 | Snorkel mirror. 6.9B tokens, $6.1k over 66 tasks. Anthropic cites 51.8% on the public board |
A single row at xhigh, so it does not compare with the max-effort Claude Code table above. Date is the model release date.
Codex, max effort:
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 58.2% ±2.8 | Codex, max effort | 2026-09-03 | Snorkel mirror. 1.5B tokens, $3.3k over 66 tasks |
| GPT-5.6 Sol | OpenAI | 37.3% ±3.8 | Codex, max effort | 2026-06-26 | Snorkel mirror. 4.4B tokens, $2.5k over 66 tasks |
| GPT-5.6 Terra | OpenAI | 21.5% ±3.3 | Codex, max effort | 2026-06-26 | Snorkel mirror. 5.7B tokens, $1.7k over 66 tasks |
| GPT-5.6 Luna | OpenAI | 17.3% ±2.8 | Codex, max effort | 2026-06-26 | Snorkel mirror. 11.6B tokens, $0.3k over 66 tasks |
Same agent and effort, so rows compare within this table only, and not with the Claude Code tables. Date is the model release date.
Grok Build, high effort:
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Grok 4.6 | Not listed | 20.3% ±3.1 | Grok Build, high effort | 2026-08-12 | Snorkel mirror. 4.0B tokens, $3.6k over 66 tasks |
| Grok 4.5 | Not listed | 12.4% ±2.6 | Grok Build, high effort | 2026-07-16 | Snorkel mirror. 3.4B tokens, $2.1k over 66 tasks |
Same agent and effort, so rows compare within this table only. Date is the model release date.
Grok Build, xhigh effort:
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Grok 4.7 | Not listed | 37.6% ±3.5 | Grok Build, xhigh effort | 2026-09-21 | Snorkel mirror. 5.5B tokens, $3.7k over 66 tasks |
A single row at xhigh, so it does not compare with the high-effort Grok table. Date is the model release date.
mini-SWE-agent, high effort:
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Gemini 3.8 Flash | 19.1% ±3.4 | mini-SWE-agent, high effort | 2026-09-02 | Snorkel mirror. 17.2B tokens, $1.8k over 66 tasks | |
| Gemini 3.7 Flash | 11.2% ±2.4 | mini-SWE-agent, high effort | 2026-08-13 | Snorkel mirror. 11.1B tokens, $1.3k over 66 tasks |
The only Gemini rows on the board. Same agent and effort, so rows compare within this table only, and not with Vals' mini-swe-agent run, which has its own setup and scoring. Date is the model release date.
The board answers a narrower question than Vals does. It tells you how much a newer model gains inside one vendor's agent. Fable 5.1 scores 57.9% in Claude Code at max effort and GPT-6 Astra 58.2% in Codex at max effort (both mirror, not primary), but those two numbers come from different agents and do not rank the models against each other.
Anthropic's launch numbers
Anthropic published Terminal-Bench 4.0 results with the Claude Opus 5.5 announcement on September 22, 2026. The page says "Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort."
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 66.4% | Anthropic's setup, adaptive thinking, max effort | 2026-09-22 | Vendor-reported. Standard error ±2.6 |
| Claude Fable 5.1 | Anthropic | 55.8% | Anthropic's setup | 2026-09-22 | Vendor-reported. Standard error ±1.6 to 2 |
| Claude Opus 5 | Anthropic | 52.3% | Anthropic's setup | 2026-09-22 | Vendor-reported. Standard error ±1.6 to 2 |
One vendor's own runs on one setup, so rows compare with each other and not with Vals or the board. Anthropic's footnote ties the setup to the board: "The public leaderboard (5 trials/task, Claude Code harness) reports Claude Opus 5 at 51.8%; our setup reproduces it at 52.3%, within noise."
Anthropic's table also carries two OpenAI figures, with the footnote "GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI."
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 57.9% | OpenAI's setup | 2026-09-22 | OpenAI-reported, quoted by Anthropic. No standard error |
| GPT-5.6 Sol | OpenAI | 37.3% | OpenAI's setup | 2026-09-22 | OpenAI-reported, quoted by Anthropic. No standard error |
Second-hand vendor figures from a different lab's setup, so they do not compare with the Anthropic rows above. Date is the publication date of Anthropic's post.
Anthropic's chart caption makes a cost claim: "Terminal-Bench 4.0 measures how well a model can complete complex, multi-step professional tasks within a command line interface. Opus 5.5 at default effort beats Opus 5 at max effort for about a fifth of the cost. It matches GPT-6 Astra at about 40% of the cost." That is Anthropic's claim about its own runs. On Vals' harness the cost order runs the other way, with Opus 5.5 at $13.20 per task and Astra at $9.58.
Historical progression
No model has a published score on both 4.0 and an earlier version, and the task sets do not overlap, so there is no cross-version trend line to draw. What exists is a v2.0 snapshot from the paper and same-harness progressions within model families on 4.0.
The paper ran 16 frontier models in six agents, at least five runs per combination, 32,155 trials in total, on the 89-task v2.0 set. Models with adjustable reasoning ran at the provider default, "medium" for Anthropic and OpenAI.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 4.5 | Anthropic | 58% | Terminus 2, medium effort, v2.0 | Jan 2026 | Paper, section 4 |
| Gemini 3 Pro | 57% | Terminus 2, provider default, v2.0 | Jan 2026 | Paper, section 4 | |
| Kimi K2 Thinking | Not listed | 36% | Terminus 2, provider default, v2.0 | Jan 2026 | Paper, section 4. Best open-weight model |
Version 2.0 with 89 tasks and no overlap with 4.0, so these rows compare only with each other. GPT-5.2 scored 63% on v2.0 on Codex CLI, a different agent, so it sits outside the table. The paper also found that "model selection is usually more important than agent scaffold." Codex CLI's resolution rate rose 52% moving from GPT-5-Nano to GPT-5.2, while Gemini 2.5 Pro gained 17% on Terminus 2 over OpenHands.
Within 4.0, family progressions on one harness:
| Series | Source | Scores | Change |
|---|---|---|---|
| Claude Opus, headline | Vals, mini-swe-agent | Opus 4.8 23.23%, Opus 5 53.53%, Opus 5.5 65.15% | +30.30, +11.62 |
| Claude Opus, fallback-adjusted | Vals, mini-swe-agent | Opus 4.8 23.23%, Opus 5 53.03%, Opus 5.5 58.08% | +29.80, +5.05 |
| Claude Fable | Vals, mini-swe-agent | Fable 5 41.41% to Fable 5.1 58.08%, adjusted 33.84% to 50.00% | +16.67, adjusted +16.16 |
| GPT Sol | Vals, mini-swe-agent | GPT-5.6 Sol 37.88%, GPT-6 Sol 44.44%, GPT-6.1 Sol 55.05% | +6.56, +10.61 |
| Claude, Claude Code max | Board, mirror, not primary | Opus 4.8 23.6%, Fable 5 44.5%, Fable 5.1 57.9% | Fable 5 to 5.1 +13.4 |
| Grok, Grok Build high | Board, mirror, not primary | Grok 4.5 12.4%, Grok 4.6 20.3% | +7.9 |
| OpenAI, Codex max | Board, mirror, not primary | GPT-5.6 Sol 37.3%, GPT-6 Astra 58.2% | Different tiers, no delta |
Each row holds one harness and effort, so changes compare within a row only. Opus 4.8 has no fallback note on Vals, so its adjusted and headline scores match. Grok 4.7 ran at xhigh and is not chained to the high-effort Grok series.
The Opus rows are the ones to look at. On the headline, Opus 5.5 gained 11.62 points over Opus 5. Counting only passes each model produced itself, the gain is 5.05 points. Opus 5 had one fallback attempt and Opus 5.5 had 22, so the two bases split almost entirely on the newer model. The 5.05-point gain is the floor. The 11.62-point headline gain includes passes that Opus 5 and Opus 4.8 produced on Opus 5.5's behalf.
Documented failure modes
No retrieved source publishes a failure analysis of the 4.0 task set. The first four items below are 4.0 evidence. The rest come from the v2.0 paper, which studied the 89-task set mostly on Terminus 2, and are not 4.0 findings.
Terminal-Bench 4.0 evidence
Provider-side fallback. Vals found Anthropic requests served by older models on five rows: 22 of 198 attempts for Opus 5.5, 22 for Fable 5.1, 42 for Fable 5, 4 for Sonnet 5.5 and 1 for Opus 5. It moved Opus 5.5 by 7.07 points and Fable 5.1 by 8.08. This is a harness-level risk for anyone calling a vendor API from their own agent, so record which model served each response.
Timeouts. Exceeding the 8-hour agent limit or a verifier's own timeout scores zero. Vals' latency per task runs from 33m46s for GPT-5.6 Sol to 2h42m for MiMo V2.6 Pro.
Broken and saturated tasks. Snorkel's FAQ says v4.0.0 "fixes broken tasks" and retires saturated ones, with 20 revised and 8 removed from v3.0.0. No source counts how many tasks were broken.
Noise. Board margins run from ±2.6 to ±3.8 points, and Anthropic reports ±2.6 for Opus 5.5.
Terminal-Bench 2.0 evidence
Execution errors dominate. The paper sorts failures using a taxonomy based on MAST. Execution errors cover disobeying the specification, repeating steps and missing termination conditions. Coherence errors cover reasoning that doesn't match the action, context loss and drifting off task. Verification errors cover stopping early and checking weakly or not at all. "Execution errors dominate for Opus 4.5 and GPT-5.2," and Qwen Coder 480B shows higher rates in every class. When you read a failed trajectory, look first for an ignored spec or a repeated step.
Missing executables. Of 3,800 sampled command-level failures, 24.1% were commands calling executables that were not installed or not on PATH. Agents that skip checking their environment burn steps on this.
Effort does not buy accuracy. The paper found "essentially no correlation" between average turns per trial and success (r=-0.028, p=0.916) and a weak negative correlation between average output tokens and success (r=-0.170, p=0.515). One run took up to two hours and almost 100 million tokens on a single task. Long runs are not a sign of capability, which is why cost per task is the number to compare, not runtime. See test-time compute for the broader trade.
Contamination. Every task is public on GitHub. The paper embeds the Big-Bench canary string in each file and says "We consider the development of a private test set to be outside the scope of this paper." Its own Agentic Benchmark Checklist review calls the canary "largely a symbolic safeguard" and notes "Even the tasks not selected to be included with this paper, which may serve as a 'held-out' set, are still publicly accessible." No retrieved source describes a private split for 4.0, so treat 4.0 scores as exposed to training-data leakage.
Cheating. The authors have "not observed this behavior in tens of thousands of agent trajectories" of agents reading oracle solutions, even with internet access, and tell users to "remain vigilant." A high score is not evidence of cheating. A vendor-reported score still deserves a trajectory audit before you rely on it.
Reproducibility. "Variability in machine resources and container runtime enforcement can lead to differences in effective task environments." The same model scores differently on different sandbox providers, so rerun on your own infrastructure before comparing against a published figure.
How it compares to related benchmarks
SWE-bench is the obvious neighbor. The Terminal-Bench paper lists it among benchmarks focused on software engineering. Only 18 of Terminal-Bench 4.0's 66 tasks are software; the other 48 cover science, ML, operations, hardware, security and media. A SWE-bench score tells you about patching a repository. Terminal-Bench tells you whether an agent can drive a whole machine to a goal.
DeepSWE goes the other way, with 113 original software-engineering tasks across 91 repositories from Datacurve. If your workload is code, DeepSWE's 113 software tasks give a denser signal than Terminal-Bench 4.0's 18.
Terminal-Bench Science is the sibling built on the same Harbor stack. Steven Dillmann at Stanford released its 70 tasks on August 27, 2026, with the Terminal-Bench and Harbor teams, covering research workflows in life, physical, earth, math and engineering sciences, and calibrated it harder than Terminal-Bench 3.0. Terminal-Bench 4.0 has a 14-task science category inside a broader set.
The paper groups OSWorld, WebArena, Visual Web Arena and AppWorld as computer-use benchmarks, a separate category. Those agents work through a GUI or a browser. Terminal-Bench's interface is a shell, and under Vals the only tool is bash. tau-bench and the Berkeley Function Calling Leaderboard measure tool use against an API, while Terminal-Bench tasks run in a real container. Older shell benchmarks, per the paper, "measure narrow aspects of the command line interface, like optimizing shell scripts, configuring software environments, or translating natural language to Bash commands." Terminal-Bench is "focused on general agentic manipulation of computers." GAIA is another agentic test worth reading alongside it. For the wider picture of how these fit together, see LLM benchmarks and LLM evaluation.
What Terminal-Bench 4.0 means for teams choosing a model
Sonnet 5.5 is the strongest Claude model once fallback is counted. On Vals' headline, Opus 5.5 leads Sonnet 5.5 by 1.01 points, 65.15% to 64.14%. With fallback attempts counted as failures, Sonnet 5.5 leads by 4.55 points, 62.63% to 58.08%, at $16.51 per task against $13.20. Anthropic's own run puts Opus 5.5 at 66.4%, a vendor number with no stated harness. If you are picking between the two for terminal work, the independent adjusted score favors Sonnet 5.5.
GPT-6 Astra is the efficient frontier pick. It scores 59.60% on Vals at $9.58 and 35m42s per task, the fastest of the top six, with no fallback note. If your agent runs inside a CI job or an on-call automation with a time budget, Astra's 35m42s per task against Opus 5.5's 1h04m decides it.
GPT-6.1 Sol is the budget pick. It scores 55.05% at $1.72 per task, 7.7 times cheaper than Opus 5.5. It trails Opus 5.5 by 10.1 points on the headline and by 3.0 points on the adjusted score, and it sits on the cost frontier on both bases. For high-volume work where a failed attempt gets retried, it is the obvious starting point.
Older tiers are a bad trade. Claude Opus 5 scores 53.53% headline and 53.03% adjusted at $18.60 per task. Fable 5 scores 41.41% and 33.84% at $30.34. Both cost more than Opus 5.5 and GPT-6.1 Sol and score below both.
Test your shortlist in your own agent architecture, and log which model answered. Fallback moved Opus 5.5 by 7.07 points and Fable 5.1 by 8.08 on Vals, and the official board runs each vendor in its own agent, so no board row compares across vendors. A 2-point gap at the top of any table here is inside the noise. A 7-point fallback effect is not.
The Klu LLM leaderboard tracks these models across other benchmarks. To run the same kind of pass/fail, cost-per-task eval on your own infrastructure and runbooks, see how engineering teams set it up in Klu on the engineering page.