Benchmarks / CyberGym

CyberGym

Built by UC Berkeley · released 3 Jun 2025

1,507 real vulnerabilities from 188 open-source projects, drawn from Google's OSS-Fuzz. Given a text description of a bug and the unpatched codebase, an agent must write a proof-of-concept input that triggers the vulnerability. UC Berkeley maintains a public leaderboard with team submissions.

Frontier

98.5%

Success rate (Level 1)

Creation (天工), multi-model · agent: Creation

7 Sep 2026 · Source: UC Berkeley (third party)

Team submission listed first on the official leaderboard. Uses a runnable vulnerable Docker image ("dynamic" label). Sangfor AI (GLM-5.3) follows at 97.21% and Alipay AI4SDL at 96.75%.

Saturated

The best result is at 90% or more of the ceiling.

See the full leaderboard at UC Berkeley

CyberGym: Success rate (Level 1) over time, 5 recorded results. 0%20%40%60%80%100%Jul 2025Oct 2025Jan 2026Apr 2026Jul 2026Oct 2026 OpenHands (Claude Sonnet 3.7): 11.9% (15 May 2025) OpenHands (GPT-5): 39.4% (5 Dec 2025) Claude Mythos Preview: 83.1% (7 Apr 2026) MDASH (multi-model): 91% (17 Jun 2026) Creation (天工), multi-model: 98.5% (7 Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/cybergym/"><img src="https://canagentswork.com/og/benchmarks-cybergym.png" width="600" height="315" alt="CyberGym: the best result is 98.5% (Creation (天工), multi-model, 7 Sep 2026)." loading="lazy"></a>

Markdown:

[![CyberGym: the best result is 98.5% (Creation (天工), multi-model, 7 Sep 2026).](https://canagentswork.com/og/benchmarks-cybergym.png)](https://canagentswork.com/benchmarks/cybergym/)

What it measures

Whether an agent can reproduce a known vulnerability in a large real codebase. A task counts as solved when the generated proof of concept crashes the pre-patch build but not the post-patch build. The headline (Level 1) gives the agent the vulnerability description and the source. Other levels give less or more information. The leaderboard notes when a team used a runnable vulnerable image or test-time memory.

Success rate (Level 1): Share of the 1,507 instances where the agent produces a working proof of concept for the target vulnerability. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
repo, cli
Grading
automated-tests
Tasks
1,507
Contamination
Vulnerabilities and their fixes are public in OSS-Fuzz and project histories, so training data may include them.
Reuse
Apache-2.0 (GitHub repository license). Dataset hosted on Hugging Face (about 240 GB). (open-apache)

Limits to keep in mind

  • Leaderboard results are run and submitted by the teams themselves. Agent runs are stochastic, and the maintainers say small score differences may not reflect real capability gaps now that leading systems score high. Source
  • Vulnerability descriptions can be ambiguous, which adds noise to grading. Source
  • Top entries use multiple models, orchestration, and test-time memory across instances. They measure agent systems, not single models. Source
  • The task is reproduction of known bugs, not discovery. The paper separately reports 34 zero-days found by agents in open-ended runs. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemSuccess rate (Level 1)DateSource
Creation (天工), multi-model · frontier98.5%7 Sep 2026UC Berkeley · third-party
MDASH (multi-model)91%17 Jun 2026Microsoft Research · lab-reported
Claude Mythos Preview83.1%7 Apr 2026Anthropic · lab-reported
OpenHands (GPT-5)39.4%5 Dec 2025UC Berkeley · primary
OpenHands (Claude Sonnet 3.7)11.9%15 May 2025UC Berkeley · primary

Timeline

  • 7 Sep 2026 — CyberGym leader reaches 98.5% reproduction rate. Source
  • 3 Jun 2025 — UC Berkeley releases CyberGym with 1,507 real vulnerabilities. Source

Where it sits in the atlas

SecuritySoftware engineering (partial)

Vulnerability research

Last checked 23 Sep 2026 against 5 primary sources. See an error? Tell us.

How to cite

Credit the original work first: CyberGym by UC Berkeley (https://arxiv.org/abs/2506.02548).

Then, if you used this page:

Can Agents Work. "CyberGym: frontier results and sources." https://canagentswork.com/benchmarks/cybergym/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-cybergym,
  title        = {{CyberGym: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/cybergym/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}