What is AutomationBench?
AutomationBench is Zapier's benchmark for AI agents that carry out business workflows across several applications. The agent gets one trigger message and nothing else. It has to discover the API endpoints it needs, read and follow a policy document embedded in the task, and leave every simulated SaaS system in the correct final state. Daniel Shepard and Robin Salimans of Zapier published it in April 2026 as arXiv 2604.18934, and Zapier announced it on its blog with a public leaderboard at zapier.com/benchmarks.
Earlier agent benchmarks such as WebArena and OSWorld test long-horizon browsing or desktop control. None of them asks a model to coordinate several apps, find its own endpoints and obey business rules in the same task. That combination is what most real automation work looks like, and it is what AutomationBench grades. Grading is a set of deterministic assertions on end state. There is no LLM judge and no partial credit in the official score.
As of 2026-10-06, Gemini 4 Argon at High effort leads Zapier's v1.0.6 leaderboard at 51.29%. Claude Sonnet 5.5 at Max is the best non-Google entry at 44.75%, 6.54 points back, and GPT 6 Astra at Max sits at 41.4%. The top model still fails about half of these workflows outright.
How the tasks and harness work
The public set has 600 tasks, 100 in each of six business domains: Sales, Marketing, Operations, Support, Finance and HR. Zapier also keeps a private held-out set of more than 600 harder tasks, which has not been released. Zapier's own footnote on one entry ("260 of 657") puts the private set at 657 tasks. The official leaderboard runs on the private set, so a vendor can get a validated score without the option of training on the full task pool.
The environment simulates 47 SaaS apps behind about 500 API endpoints: CRM, inbox, calendar, project management, Google Sheets and others. Pydantic models hold the state of each app and keep the parts of real APIs that trip agents up, including pagination, required fields and common error cases. The tasks then layer on the problems that live business data has. There are duplicate records, stale data, contacts with similar names, inconsistent formats, and policies written into the task that the agent has to apply without being reminded.
The agent gets two tools. Search runs BM25 over the API schemas and returns the top five matches. Execute sends an HTTP-style request with a method, URL and body. Parallel tool calls are supported. Nobody hands the model a curated function list for each app, so a large share of every task is working out which endpoint does what before calling it. A model that guesses a path or assumes a record lives in the obvious place gets punished here, where a benchmark with hand-written tools would let it through.
Each task runs under a 50-step cap, adjustable with --max-steps. The median task finishes in single-digit steps and the cap is rarely hit. AutomationBench is a test of getting a short chain exactly right, not of stamina over hundreds of turns.
The code and public tasks are on GitHub. uv run auto-bench --model [name] runs the public set locally with up to 100 tasks in parallel, and the benchmark is also available on the Prime Intellect Environments Hub.
How scoring works
Every task carries a list of assertions about the final state of the simulated apps. Positive assertions check that the required actions happened. Negative assertions check that the agent avoided forbidden actions, which is where the embedded policies bite.
The official score is task_completed_correctly, a binary pass or fail per task, averaged over all tasks. An agent that gets nine of ten assertions right on a task scores zero for it. The harness also reports partial_credit, the fraction of assertions satisfied, but the leaderboard ranks on the binary number. I think that is the right call for this domain. A workflow that updates the CRM and forgets the follow-up is a broken workflow, and someone has to find and fix it by hand.
Because the grader reads program state, no LLM-as-a-judge sits between the agent and its score. Two runs that leave the same end state get the same grade.
Versions and variants
Zapier hardens AutomationBench periodically and states that scores across versions are not comparable. The current leaderboard runs v1.0.6, released 2026-07-31. That release added a Google Drive path for finding spreadsheet IDs, relaxed formatting requirements that were too rigid, stopped failing agents for side effects the task never specified, and moved to framework 0.2.0 with prompt caching and refusal tracking. v1.0.5, released 2026-07-16, added baselines to the docs, fixed fairness and task-contract problems, and made the runner more stable. Both are in the changelog. Those fairness edits move scores, which is why the tables below keep each version separate.
Artificial Analysis runs an independent variant, AutomationBench-AA, announced on 2026-07-06. It uses Zapier's private subset of 657 tasks across 40 simulated apps, with nearly 12,000 assertions, each classed as an objective or a guardrail. Every task gets one run under a 50-turn cap. The metric is the share of objectives completed without violating any guardrail. That is a different number from Zapier's binary task score, so AA results get their own tables.
Claude entries on Zapier's board run with "default fallbacks". When Anthropic's safeguards route a request away from the model, it goes to a fallback model instead of counting as a failure. This changes Claude scores by a lot, as the composite entries below show, and anyone reading a Claude number on this benchmark needs to know whether fallbacks were on.
Two things are not published. Zapier gives effort labels such as High, XHigh and Max, but no reasoning token budgets and no temperature. Google has published no AutomationBench methodology for Gemini 4 Argon in a fetchable location; its Gemini 4 evals methodology page returned a 404. Every Argon number on this page comes from Zapier.
Current leaderboard
Zapier runs every model on its board itself, on the private set, at v1.0.6. The table below is the top 10 of 121 models, retrieved 2026-10-06. Cost per task is the standard list price for every row, including Argon, which Google is currently selling at a promotional price.
Zapier official leaderboard, v1.0.6
Same version, private set, metric and pricing basis for every row; not comparable with the Anthropic, Artificial Analysis, paper or README tables further down.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Gemini 4 Argon | 51.29% | Zapier harness, private set, High effort | v1.0.6 (2026-07-31), fetched 2026-10-06 | Zapier, Zapier-run; $1.70/task list, $0.85 promo | |
| Gemini 4 Argon | 50.08% | Zapier harness, private set, Medium effort | v1.0.6 (2026-07-31), fetched 2026-10-06 | Zapier, Zapier-run; $1.54/task list, $0.77 promo | |
| Claude Sonnet 5.5 | Anthropic | 44.75% | Zapier harness, private set, Max, default fallbacks | v1.0.6 (2026-07-31), fetched 2026-10-06 | Zapier, Zapier-run; $1.14/task |
| Claude Opus 5.5 | Anthropic | 42.47% | Zapier harness, private set, Max, default fallbacks | v1.0.6 (2026-07-31), fetched 2026-10-06 | Zapier, Zapier-run; $1.44/task |
| GPT 6 Astra | OpenAI | 41.4% | Zapier harness, private set, Max effort | v1.0.6 (2026-07-31), fetched 2026-10-06 | Zapier, Zapier-run; $1.73/task |
| GPT 6 Astra | OpenAI | 38.96% | Zapier harness, private set, XHigh effort | v1.0.6 (2026-07-31), fetched 2026-10-06 | Zapier, Zapier-run; $1.50/task |
| GPT 6 Astra | OpenAI | 37.14% | Zapier harness, private set, High effort | v1.0.6 (2026-07-31), fetched 2026-10-06 | Zapier, Zapier-run; $1.44/task |
| Claude Sonnet 5.5 | Anthropic | 36.83% | Zapier harness, private set, XHigh, default fallbacks | v1.0.6 (2026-07-31), fetched 2026-10-06 | Zapier, Zapier-run; $0.48/task |
| Claude Opus 5.5 | Anthropic | 35.77% | Zapier harness, private set, XHigh, default fallbacks | v1.0.6 (2026-07-31), fetched 2026-10-06 | Zapier, Zapier-run; $0.89/task |
| GPT 6 Astra | OpenAI | 34.09% | Zapier harness, private set, Medium effort | v1.0.6 (2026-07-31), fetched 2026-10-06 | Zapier, Zapier-run; $1.27/task |
Zapier lists v1.0.6 as the board version and gives no per-row run date. No Google, OpenAI or Anthropic page publishes these Argon, Astra or Sonnet 5.5 v1.0.6 numbers.
Argon holds the top two spots. Going from Medium to High buys 1.21 points for $0.16 more per task. Argon Medium beats Sonnet 5.5 Max by 5.33 points for $0.40 more.
Below Argon, the surprise is that Anthropic's smaller model wins. Sonnet 5.5 beats Opus 5.5 at both effort levels: 44.75% to 42.47% at Max for $0.30 less per task, and 36.83% to 35.77% at XHigh for $0.41 less. GPT 6 Astra at Max costs the most of the top five, $1.73, and lands 9.89 points behind Argon High. Dropping Astra from Max to Medium loses 7.31 points and saves $0.46.
Four configurations make up the cost frontier: Sonnet 5.5 XHigh (36.83% at $0.48), Sonnet 5.5 Max (44.75% at $1.14), Argon Medium (50.08% at $1.54) and Argon High (51.29% at $1.70). Every GPT 6 Astra and Opus 5.5 configuration in the top 10 is beaten on both score and cost by a Sonnet 5.5 entry.
Domain leaders on the same board
The same Zapier page breaks scores out by domain. These rows share the v1.0.6 private-set setup of the table above.
| Domain | Leader | Score | Closest other model | Score |
|---|---|---|---|---|
| Sales | Gemini 4 Argon Medium | 58.12% | Claude Opus 5.5 | 45.3% |
| Marketing | Gemini 4 Argon High | 58.0% | Claude Sonnet 5.5 | 52.0% |
| Operations | Gemini 4 Argon Medium | 68.0% | Claude Opus 5.5 | 65.0% |
| Support | Gemini 4 Argon High | 45.0% | GPT 6 Astra | 39.0% |
| Finance | Claude Sonnet 5.5 Max | 50.0% | Gemini 4 Argon | 49.17% |
| HR | Gemini 4 Argon High | 39.17% | Claude Sonnet 5.5 | 31.67% |
Argon leads five of six domains. Its widest margin is Sales, 12.82 points over Opus 5.5. Operations has the highest top score at 68.0%, and HR the lowest at 39.17%. Finance is the one domain Argon does not win, where Sonnet 5.5 Max edges it by 0.83 points.
Anthropic's Claude Opus 5.5 release table
Zapier ran these during early access and Anthropic published them in its Claude Opus 5.5 announcement. No-fallback run with no stated version or set; not comparable with the Zapier board or any other table here.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 40.0% | Zapier harness, early access, no fallbacks; page default is adaptive thinking at max effort | Run date not stated | Anthropic; Zapier-run, Anthropic-reported; safeguard interventions counted as failures |
| Claude Fable 5.1 | Anthropic | 31.4% | Setting not stated in Anthropic's table | Run date not stated | Anthropic; Zapier-run, Anthropic-reported; Zapier's footnote says it includes Opus 5 fallback completions |
Anthropic's footnote says counting safeguard interventions as failures "resulted in a lower score than Claude Opus 5.5 would achieve in practice."
Zapier figures quoted by Anthropic
The same Anthropic page says these three rows "come from Zapier's public leaderboard." Only the Astra row traces to a stated version and effort; the Sol and Opus 5 rows rank against nothing else on this page.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 41.4% | Zapier harness; effort not stated by Anthropic | Not stated | Anthropic; matches the Astra Max row on Zapier's v1.0.6 board |
| GPT-5.6 Sol | OpenAI | 28.8% | Zapier harness; version and effort unknown | Not stated | Anthropic; Zapier-run, Anthropic-reported |
| Claude Opus 5 | Anthropic | 26.9% | Zapier harness; version and effort unknown | Not stated | Anthropic; Zapier-run, Anthropic-reported |
Fallback composite on Zapier's board
This entry is listed separately because a second model did part of the work and the cost covers only one of them. Not comparable with single-model rows on any table.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Fable 5.1 with Opus 5 fallback | Anthropic | 31.4% | Zapier harness, private set; Opus 5 handled 260 of 657 tasks; effort not stated | Fetched 2026-10-06 | Zapier, Zapier-run; cost per task covers Fable 5.1 only, combo cost not published |
Zapier's note reads: "Opus 5 completes steps Fable 5.1's safety classifier refuses, then Fable finishes the task." Opus 5 touched about 40% of the tasks, and those completions count toward the 31.4%.
AutomationBench-AA, independent
Artificial Analysis ran these on 657 private-subset tasks with its objectives-without-guardrail-violation metric, launched 2026-07-06. Different metric and older models than Zapier's board; compare these rows only with each other.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 4.8 | Anthropic | 48.5% | AA harness, max effort, one run, 50 turns | 2026-07-06 | Artificial Analysis; ~$1.50/task |
| Gemini 3.5 Flash | 42.6% | AA harness, one run, 50 turns | 2026-07-06 | Artificial Analysis; $0.49/task | |
| GPT-5.5 | OpenAI | 42.1% | AA harness, xhigh effort, one run, 50 turns | 2026-07-06 | Artificial Analysis; $1.32/task |
| GLM-5.2 | Z.ai | 27.8% | AA harness, max effort, one run, 50 turns | 2026-07-06 | Artificial Analysis; best open-weights model; cost not given |
Gemini 3.5 Flash edges GPT-5.5 by 0.5 points at 37% of its cost, and it has the best ratio on the board: 15.0 objectives completed per guardrail violation. DeepSeek V4, Gemini 3.1 Flash-Lite and Qwen3.7 Plus each cost under $0.05 per task. The article does not state that any of these rows used a fallback.
AutomationBench-AA fallback composite
Same source and metric as the AA table, listed apart because Opus 4.8 did part of the work.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Fable 5 with Opus 4.8 fallback | Anthropic | 48.6% | AA harness, max effort; Opus 4.8 fallback on ~18% of tasks | 2026-07-06 | Artificial Analysis; cost not given |
Artificial Analysis lists Fable 5 as the leader at 48.6%. With nearly a fifth of its tasks handed to Opus 4.8, the 0.1-point margin over Opus 4.8 alone says nothing about Fable 5 on its own.
Public repo leaderboard
The GitHub README reports pass rates on the 600-task public set at each model's highest reasoning effort. Public set, undated, version not stated; not comparable with the private-set board.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5 | Anthropic | 50.3% | Zapier harness, public set, max | Undated, fetched 2026-10-06 | GitHub, Zapier-run |
| Kimi K3 | Not stated in README | 46.67% | Zapier harness, public set, max | Undated, fetched 2026-10-06 | GitHub, Zapier-run |
| Claude Fable 5 | Anthropic | 46.17% | Zapier harness, public set, max | Undated, fetched 2026-10-06 | GitHub; scored after the July 2026 stricter classifier |
| GPT-5.6 Sol | OpenAI | 45.83% | Zapier harness, public set, max | Undated, fetched 2026-10-06 | GitHub, Zapier-run |
| Gemini 3.6 Flash | 45.00% | Zapier harness, public set, high | Undated, fetched 2026-10-06 | GitHub, Zapier-run | |
| Claude Opus 4.8 | Anthropic | 41.00% | Zapier harness, public set, max | Undated, fetched 2026-10-06 | GitHub, Zapier-run |
| Gemini 3.5 Flash | 38.33% | Zapier harness, public set, high | Undated, fetched 2026-10-06 | GitHub, Zapier-run | |
| GPT-5.6 Terra | OpenAI | 37.17% | Zapier harness, public set, max | Undated, fetched 2026-10-06 | GitHub, Zapier-run |
| Claude Sonnet 5 | Anthropic | 34.67% | Zapier harness, public set, max | Undated, fetched 2026-10-06 | GitHub, Zapier-run |
| GLM 5.2 | Z.ai | 26.17% | Zapier harness, public set, max | Undated, fetched 2026-10-06 | GitHub, Zapier-run |
Original paper, April 2026
The paper's own baseline runs, on the public set at release. Different set and version from everything above; kept as the launch record.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 4.7 | Anthropic | 9.9% | Paper harness, public set | 2026-04-21 | arXiv 2604.18934; $1.80/task |
| Gemini 3.1 Pro | 9.6% | Paper harness, public set | 2026-04-21 | arXiv 2604.18934; $0.54/task | |
| GPT 5.4 | OpenAI | 7.6% | Paper harness, public set, high | 2026-04-21 | arXiv 2604.18934; $1.93/task |
| Claude Sonnet 4.6 | Anthropic | 5.3% | Paper harness, public set | 2026-04-21 | arXiv 2604.18934; $1.81/task |
| Claude Haiku 4.5 | Anthropic | 1.5% | Paper harness, public set | 2026-04-21 | arXiv 2604.18934; $0.18/task |
At launch, every model scored under 10%. Gemini 3.1 Pro came within 0.3 points of Opus 4.7 at less than a third of the cost.
Dated snapshots
Three moments in AutomationBench's short history, each on a different set, metric or version. They are listed by date and do not form a trend line.
| Date | Set, metric and version | Leader | Score | Source |
|---|---|---|---|---|
| 2026-04-21 | Public set, binary task score, paper v1 | Claude Opus 4.7 | 9.9% | arXiv 2604.18934 |
| 2026-07-06 | 657-task private subset, AA metric | Claude Fable 5 with Opus 4.8 fallback | 48.6% | Artificial Analysis |
| v1.0.6 (2026-07-31), fetched 2026-10-06 | Zapier private set, binary task score | Gemini 4 Argon High | 51.29% | Zapier |
The jump from 9.9% to 51.29% looks like progress, but the two numbers come from different task sets and benchmark versions, so the gap between them measures nothing.
Failure modes and limitations
False confidence. This is the dominant failure in the paper. Among failed tasks, the agent declared success while the data was wrong in 72% of Opus failures, 91% of Gemini failures and 84% of GPT 5.4 failures. Those were April 2026 models, but the pattern is the reason AutomationBench grades end state instead of trusting the transcript. An agent's "Done" message carries no information about whether the work got done.
Giving up on search too early. Models assume records sit in a default location instead of exploring. With only a BM25 search over schemas to work from, a model that stops after one query writes to the wrong place.
Incomplete list processing. On tasks with several items, models handle some and drop the rest. Under binary scoring, one dropped item fails the task.
Instruction deviation. Models paraphrase or ignore precise requirements, which collides with assertions that check exact values and with the embedded policies.
Finance is hard. In Artificial Analysis's runs, models completed about half the share of objectives on Finance that they did on Support or Operations. Zapier's v1.0.6 domain table puts the Finance leader at 50.0%, against 68.0% for Operations.
Safeguards change Claude scores. Anthropic's footnote says safeguard interventions counted as failures in the no-fallback Opus 5.5 run. Fable 5 fell back to Opus 4.8 on about 18% of AA tasks, and Fable 5.1 handed 260 of 657 tasks to Opus 5 in Zapier's composite entry. A Claude score on this benchmark depends on whether fallbacks were on, and the same will be true in your deployment.
Synthetic data. The authors acknowledge a "risk of lack of realism and impossibility" in simulated apps, and say reasoning gaps will remain after the current challenges are solved.
Contamination. The 600 public tasks are released, so a model can be trained on them. That is why the official board runs on the private set, and why README numbers on the public set carry less weight than the board.
How AutomationBench compares to related benchmarks
The comparison points come from the AutomationBench paper, and the differences are about what the agent touches.
WebArena and Mind2Web, in the paper's words, "evaluate long-horizon web tasks in browsing environments." AutomationBench has no browser. The agent works entirely through REST API calls.
OSWorld "evaluates open-ended computer-use tasks spanning real web and desktop applications." AutomationBench swaps real apps for simulated ones behind REST APIs.
For tau3-bench, the paper's objection is that "each task still operates within a single application rather than requiring cross-application coordination." AutomationBench tasks span several apps by design.
AutomationBench grades programmatic end state only, so no model grades another model. LLM-as-a-judge covers the alternative approach. For a broader map of where agent evals fit, see LLM benchmarks and LLM evaluation.
What AutomationBench means for teams choosing a model
If you are automating business workflows across SaaS tools, this is the most direct public signal available, and it says a few specific things about current frontier models.
Gemini 4 Argon is the strongest model on Zapier's private set. High scores 51.29% at $1.70 per task at list price, and Medium gets within 1.21 points for $0.16 less. Argon is on promotional pricing right now, $0.85 per task at High and $0.77 at Medium, but the board ranks on list price. Budget from list price unless your contract locks in the promotion.
Among Anthropic models, pick Sonnet 5.5 over Opus 5.5 for this kind of work. Sonnet wins at both Max and XHigh and costs less at both. Sonnet 5.5 at XHigh is the value pick on the whole board: 36.83% at $0.48, which is 72% of Argon High's score at 28% of its list cost. If you run high volumes of simple workflows with a human checking exceptions, that trade is worth testing first.
GPT 6 Astra is hard to justify on this benchmark. At Max it is the most expensive configuration in the top five and still 9.89 points behind Argon High, and Sonnet 5.5 Max beats it for $0.59 less per task.
Two things apply whichever model you pick. The best model fails about half these tasks, and in the paper's runs most failures came with a confident success message. Put end-state verification outside the agent: check the CRM record, the sheet row, the sent message, rather than the agent's summary. And if you deploy Claude, evaluate with the same safeguard and fallback configuration you will ship, because Zapier's Claude numbers assume default fallbacks are on.
Public leaderboards like this one are a starting point. Compare current scores across benchmarks on the Klu LLM leaderboard, and see how teams run the same kind of end-state eval against their own workflows in Klu on the operations page.