Benchmarks / Will It Survive? Agent code survival study
Will It Survive? Agent code survival study (Code survival)
Built by Concordia University · Musfiqur Rahman, Emad Shihab · released 23 Jan 2026
A survival analysis of code that coding agents merged into 201 open-source projects, drawn from the AIDev dataset. It tracks whether agent-written lines and files are later modified, compared with human-written code in the same projects over matched windows, and classifies why they changed.
Frontier
No result recorded yet. See the source for current results.
What it measures
How long merged agent-authored code lasts before someone changes it. The cohort is 201 repositories and 5,171 PRs (3,003 agent, 2,168 human), tracked at file and line level. At the line level, agent code had a 16% lower hazard of modification than human code (hazard ratio 0.842, p < 0.001) and a 15.8 percentage-point lower modification rate. Line-level "death rates": Cursor 38.7%, Claude Code 41.0%, OpenAI Codex 48.5%, GitHub Copilot 48.6%, Devin 71.7%, human baseline 69.3%. When agent code changed, corrective fixes were slightly more common (26.3% vs 23.0%).
Line-level modification (death) rate: Share of merged lines later modified or deleted within the observation window. Lower means the code survived longer. Lower is better.
No fixed reference point, so we do not rate its status. Field signal from one observational study. No results file; the paper's figures are in the description.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- field-data
- Grading
- field-signal
- Human reference
- Human-authored lines in the same repositories had a 69.3% line-level death rate.
- Reuse
- Paper CC BY 4.0. Underlying data comes from the AIDev dataset. (cite-only)
Limits to keep in mind
- Preprint under review at EASE 2026 (as of v1, 2026-01-23). Not yet peer reviewed. Source
- The authors suggest code ownership may explain the result: developers avoid touching code with no clear human owner. Longer survival therefore does not by itself show higher quality. Source
- The file-level hazard ratio was 1.038 (p = 0.052), so the survival advantage holds at the line level but not clearly at the file level. Effect sizes for intent differences are small. Source
Where it sits in the atlas
Go to the source
- Paper arxiv.org
Last checked 23 Sep 2026 against 2 primary sources. See an error? Tell us.
How to cite
Credit the original work first: Will It Survive? Agent code survival study by Concordia University (https://arxiv.org/abs/2601.16809).
Then, if you used this page:
Can Agents Work. "Will It Survive? Agent code survival study: frontier results and sources." https://canagentswork.com/benchmarks/code-survival-study/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-code-survival-study,
title = {{Will It Survive? Agent code survival study: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/code-survival-study/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}