Benchmarks / SWE-Lancer

SWE-Lancer

Built by OpenAI · Samuel Miserendino, Michele Wang, Tejal Patwardhan, Johannes Heidecke · released 17 Feb 2025

Real freelance software jobs from Upwork, all from the Expensify codebase, worth $1 million in actual payouts. Independent contributor (IC) tasks ask the model to fix a bug or build a feature; manager tasks ask it to pick the best of several freelancer proposals. The public split, SWE-Lancer Diamond, is what labs report.

Frontier

80%

IC SWE Diamond pass@1

gpt-5.1-codex-max

18 Nov 2025 · Source: OpenAI (benchmark maintainers)

IC SWE Diamond, pass@1 averaged over 3 runs, on the 2025-07-17 offline dataset. The system card shows the value only in a chart labeled 80%. OpenAI is the maintainer and reports its own model. Later OpenAI system cards do not include SWE-Lancer.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

SWE-Lancer: IC SWE Diamond pass@1 over time, 4 recorded results. 0%20%40%60%80%100%Apr 2025Jul 2025Oct 2025 Claude 3.5 Sonnet: 26.2% (17 Feb 2025) gpt-5 (no browsing): 55% (18 Nov 2025) gpt-5-codex: 67% (18 Nov 2025) gpt-5.1-codex-max: 80% (18 Nov 2025)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/swe-lancer/"><img src="https://canagentswork.com/og/benchmarks-swe-lancer.png" width="600" height="315" alt="SWE-Lancer: the best result is 80% (gpt-5.1-codex-max, 18 Nov 2025)." loading="lazy"></a>

Markdown:

[![SWE-Lancer: the best result is 80% (gpt-5.1-codex-max, 18 Nov 2025).](https://canagentswork.com/og/benchmarks-swe-lancer.png)](https://canagentswork.com/benchmarks/swe-lancer/)

What it measures

For IC tasks, the model's patch must pass end-to-end browser tests that professional engineers wrote and reviewed three times. For manager tasks, the choice must match the proposal the real hiring manager picked. The paper reports dollars earned; OpenAI's later system cards report pass@1 on the IC SWE Diamond set, which is the metric here. The 2025-07-17 offline revision keeps 198 IC Diamond tasks and removes internet access during runs.

IC SWE Diamond pass@1: Share of Diamond-set independent contributor tasks whose end-to-end tests pass, one attempt per task. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
repo, cli
Grading
automated-tests
Tasks
1,488
Human reference
Each task's value is the real payout to the freelancer who did it, from $250 to $32,000 in the full set. IC tasks range from 15-minute bug fixes to multi-week features.
Contamination
Tasks come from a public open-source repository (Expensify). OpenAI keeps most tasks private and disables internet during runs.
Reuse
Diamond split is public in the openai/preparedness repository (MIT). The remaining tasks are a private holdout. (open-mit)

Limits to keep in mind

  • One codebase (the open-source Expensify repository), so results say little about other stacks or domains. OpenAI lists this as a limitation. Source
  • OpenAI is both maintainer and the only lab that reports results. No other model developer adopted the benchmark, and OpenAI system cards from December 2025 on no longer include it. Source
  • Dataset changed on 2025-07-17: 39 of 237 IC Diamond tasks were dropped and internet access was removed, so results before and after are not comparable. Source
  • pass@1 is one sample per task, and OpenAI notes significant variance between runs. Later system cards average three runs. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemIC SWE Diamond pass@1DateSource
gpt-5 (no browsing)55%18 Nov 2025OpenAI · primary
gpt-5-codex67%18 Nov 2025OpenAI · primary
gpt-5.1-codex-max · frontier80%18 Nov 2025OpenAI · primary
Claude 3.5 Sonnet26.2%17 Feb 2025OpenAI · primary

Timeline

  • 18 Nov 2025 — GPT-5.1-Codex-Max reaches 80% on SWE-Lancer IC Diamond. Source
  • 17 Feb 2025 — OpenAI releases SWE-Lancer with $1 million of Upwork tasks. Source

Where it sits in the atlas

Software engineering

Issue resolutionFeature developmentFreelance projects

Last checked 23 Sep 2026 against 6 primary sources. See an error? Tell us.

How to cite

Credit the original work first: SWE-Lancer by OpenAI (https://arxiv.org/abs/2502.12115).

Then, if you used this page:

Can Agents Work. "SWE-Lancer: frontier results and sources." https://canagentswork.com/benchmarks/swe-lancer/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-swe-lancer,
  title        = {{SWE-Lancer: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/swe-lancer/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}