Benchmarks / Remote Labor Index

Remote Labor Index (RLI)

Built by CAIS and Scale AI · released Oct 2025

Real freelance projects, collected from professionals on Upwork, that an agent must deliver end to end. Trained human evaluators compare each AI deliverable with the deliverable that a paid professional made.

Frontier

15.8%

Automation rate

Claude Fable 5

1 Jul 2026 · Source: CAIS (benchmark maintainers)

CAIS says the three new models were paired with stronger agent scaffolding.

Open

The best result is below 20% of the ceiling, or below half of human parity.

See the full leaderboard at CAIS

Remote Labor Index: Automation rate over time, 5 recorded results. 0%20%40%60%80%100%Oct 2025Jan 2026Apr 2026Jul 2026 Manus: 2.5% (Oct 2025) Claude Opus 4.6 (Claude Cowork scaffold): 4.17% (2026) GPT-5.5: 6.3% (1 Jul 2026) Claude Opus 4.8: 8.3% (1 Jul 2026) Claude Fable 5: 15.8% (1 Jul 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/remote-labor-index/"><img src="https://canagentswork.com/og/benchmarks-remote-labor-index.png" width="600" height="315" alt="Remote Labor Index: the best result is 15.8% (Claude Fable 5, 1 Jul 2026)." loading="lazy"></a>

Markdown:

[![Remote Labor Index: the best result is 15.8% (Claude Fable 5, 1 Jul 2026).](https://canagentswork.com/og/benchmarks-remote-labor-index.png)](https://canagentswork.com/benchmarks/remote-labor-index/)

What it measures

Whether a reasonable client would accept the agent's deliverable, compared with a gold-standard deliverable from a professional freelancer. Projects come from 23 Upwork domains, for example 3D and CAD, architecture, graphic design, video and animation, audio, data analysis, and web apps.

Automation rate: Share of projects where the AI deliverable is judged at least as good as the professional deliverable. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Project level: whole projects judged by an acceptance standard
Environment
cli, computer-use
Grading
human-expert, pairwise-human
Tasks
240
Human reference
A professional freelancer's accepted deliverable. Mean human completion time 28.9 hours (median 11.5 hours). Mean project value $632.60 (median $200).
Contamination
Leaderboard scores use a private set of 230 projects.
Reuse
Scores come from 230 private projects. 10 public projects and the open-source evaluation platform are released for qualitative analysis. (cite-only)

Limits to keep in mind

  • Excludes projects that need direct client interaction, physical labor, or long-term evaluation (for example SEO). Source
  • Trained experts grade every deliverable by hand, so new agents are tested at a low frequency. Source
  • Agents had a maximum generation budget of $30 per project. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemAutomation rateDateSource
GPT-5.56.3%1 Jul 2026CAIS · primary
Claude Opus 4.88.3%1 Jul 2026CAIS · primary
Claude Fable 5 · frontier15.8%1 Jul 2026CAIS · primary
Claude Opus 4.6 (Claude Cowork scaffold)
harness: Claude Cowork
4.17%2026CAIS · primary
Manus2.5%Oct 2025Scale AI · primary

Where sources disagree

CAIS, a co-author of the benchmark, reports 15.8% for Claude Fable 5. Two aggregator pages report 16.1%. We did not find the cause. A later re-grade is one possible cause.

  • 15.8safe.ai (primary) · CAIS post dated 2026-07-01.
  • 16.1www.benchleader.com (third-party) · BenchLeader page, seen 2026-09-23.
  • 16.1ai-intensify.com (third-party) · News article, seen 2026-09-23.

We show 15.8. Status: open.

Timeline

  • 1 Jul 2026 — Best Remote Labor Index score rises to 15.8%. Source

Where it sits in the atlas

Design, media, and writingArchitecture and engineering (partial)Data and analytics (partial)Software engineering (partial)

Freelance projectsProfessional deliverables

Last checked 23 Sep 2026 against 2 primary sources, with a second independent check. See an error? Tell us.

How to cite

Credit the original work first: Remote Labor Index by CAIS and Scale AI (https://arxiv.org/abs/2510.26787).

Then, if you used this page:

Can Agents Work. "Remote Labor Index: frontier results and sources." https://canagentswork.com/benchmarks/remote-labor-index/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-remote-labor-index,
  title        = {{Remote Labor Index: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/remote-labor-index/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}