Benchmarks / AutomationBench

AutomationBench

Built by Zapier · Daniel Shepard, Robin Salimans · released 21 Apr 2026

Zapier's test of whether an agent can carry out a cross-app business workflow on its own. Each task boots a small simulated company (CRM, inbox, calendar, sheets, helpdesk) in 47 simulated SaaS apps, hands the agent one trigger message, and checks the state it leaves behind. Tasks come from workflow patterns on Zapier's platform in sales, marketing, operations, support, finance, and HR.

Frontier

42.5%

Pass rate (task_completed_correctly)

Claude Opus 5.5 (default fallbacks, Max)

Sep 2026 · Source: Zapier (benchmark maintainers)

First on board version 1.0.6 at $1.44 per task. "Default fallbacks": tasks the model refused were rerun with Anthropic's default fallback routing; other task results are unchanged. Zapier ranks this entry, so we show it. Anthropic's Opus 5.5 post reports 40.0% for the same model without fallbacks (Zapier-run, early access). The board shows no date; read 2026-09-24.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at Zapier

AutomationBench: Pass rate (task_completed_correctly) over time, 3 recorded results. 0%20%40%60%80%100%Apr 2026Jul 2026Oct 2026 Opus 4.7 (max): 9.9% (21 Apr 2026) GPT 6 Astra (Max): 41.4% (Sep 2026) Claude Opus 5.5 (default fallbacks, Max): 42.5% (Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/automationbench/"><img src="https://canagentswork.com/og/benchmarks-automationbench.png" width="600" height="315" alt="AutomationBench: the best result is 42.5% (Claude Opus 5.5 (default fallbacks, Max), Sep 2026)." loading="lazy"></a>

Markdown:

[![AutomationBench: the best result is 42.5% (Claude Opus 5.5 (default fallbacks, Max), Sep 2026).](https://canagentswork.com/og/benchmarks-automationbench.png)](https://canagentswork.com/benchmarks/automationbench/)

What it measures

Strict pass rate: a task counts only when every deterministic assertion on the final state holds, including negative assertions (for example no email to the wrong team). The agent finds endpoints with a search tool and calls them with an execute tool, up to 50 steps, with no human in the loop. Official scores come from a held-out private set of about 657 tasks that mirrors the public set but is harder. A partial-credit score (share of assertions passed) is reported for diagnosis only.

Pass rate (task_completed_correctly): Share of private-set tasks where every scored assertion on the final environment state passes. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
api-tools, simulated-workplace
Grading
state-check
Tasks
657
Contamination
Leaderboard scores use a held-out private task set. The public set is for research and its scores differ from the board by design (for example Claude Opus 5 at 50.3% on the public set).
Reuse
The 600 public tasks (100 per domain) and the harness ship in the repository under the MIT license. The private leaderboard set is not released. (open-mit)

Limits to keep in mind

  • The leading entries are "default fallbacks" runs: tasks that a model refused were rerun through the provider's default fallback model, and the rerun result counts. Zapier ranks these entries on its board, so the atlas shows them, but a no-fallback run of the same model scores lower (Anthropic reports 40.0% for Opus 5.5 without fallbacks, run by Zapier during early access). Source
  • Version updates change the tasks. In 1.0.6 (2026-07-31) Zapier fixed task fairness, made private tasks harder, and reran the official entries. Scores before and after a version change are not directly comparable, and the launch paper's scores (best 9.9%) predate 1.0.5. Source
  • Tasks and environments are synthetic (generated with models, then hardened and sampled for manual review), so some tasks may be unrealistic or impossible. Zapier says bugs remain. Source
  • Scores are from a single run per model. Zapier says run-to-run variance is typically within 1%. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemPass rate (task_completed_correctly)DateSource
GPT 6 Astra (Max)41.4%Sep 2026Zapier · primary
Claude Opus 5.5 (default fallbacks, Max) · frontier42.5%Sep 2026Zapier · primary
Opus 4.7 (max)9.9%21 Apr 2026Zapier · primary

Timeline

  • Sep 2026 — Claude Opus 5.5 tops AutomationBench 1.0.6 at 42.47%, up from 9.9% at launch. Source
  • 31 Jul 2026 — AutomationBench 1.0.6 fixes task fairness, hardens private tasks, and reruns official entries. Source
  • 21 Apr 2026 — Zapier publishes AutomationBench; best model completes 9.9% of cross-app workflows. Source

Where it sits in the atlas

Work ladder: Direct evidence for Management and business operations. On the work ladder it counts as 42.5% of tasks completed. How the ladder works

Work it measures (O*NET work activities): Implement procedures or processes; Maintain operational records; Maintain sales or financial records; Communicate with others about operational plans or activities; Respond to customer problems or inquiries; Perform administrative or clerical activities.

Management and business operationsSales and marketing (partial)Customer support (partial)Finance and accounting (partial)Office and administrative support (partial)

Business workflow automationEnterprise workflowsCRM operationsCustomer service

Last checked 24 Sep 2026 against 9 primary sources. See an error? Tell us.

How to cite

Credit the original work first: AutomationBench by Zapier (https://arxiv.org/abs/2604.18934).

Then, if you used this page:

Can Agents Work. "AutomationBench: frontier results and sources." https://canagentswork.com/benchmarks/automationbench/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-automationbench,
  title        = {{AutomationBench: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/automationbench/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}