Benchmarks / TheAgentCompany

TheAgentCompany

Built by Carnegie Mellon University · Frank F. Xu, Yufan Song, Boxuan Li, Yuxuan Tang, et al. · released 18 Dec 2024

A simulated software company with self-hosted GitLab, a project tracker (Plane), ownCloud file storage, and RocketChat. An agent gets work tasks like a digital employee and must browse, code, run programs, and talk to simulated coworkers to finish them.

Frontier

42.9%

Tasks resolved

TTE-MatrixAgent + DeepSeek-V3.2 · harness: TTE-MatrixAgent

10 Nov 2025 · Source: Carnegie Mellon University (benchmark maintainers)

Top entry on the leaderboard, marked checked by the maintainers. Partial-credit score 52.4, 29.91 steps and $0.40 per task on average. Environment model Qwen Plus. The agent code is not open source. Newest entry on the board as of 2026-09-23.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at Carnegie Mellon University

TheAgentCompany: Tasks resolved over time, 4 recorded results. 0%20%40%60%80%100%Jan 2025Apr 2025Jul 2025Oct 2025 OpenHands + Claude 3.5 Sonnet: 24% (17 Dec 2024) OpenHands + Gemini 2.5 Pro: 30.3% (10 May 2025) OpenHands-Versa + Claude Sonnet 4: 33.1% (14 Jun 2025) TTE-MatrixAgent + DeepSeek-V3.2: 42.9% (10 Nov 2025)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/theagentcompany/"><img src="https://canagentswork.com/og/benchmarks-theagentcompany.png" width="600" height="315" alt="TheAgentCompany: the best result is 42.9% (TTE-MatrixAgent + DeepSeek-V3.2, 10 Nov 2025)." loading="lazy"></a>

Markdown:

[![TheAgentCompany: the best result is 42.9% (TTE-MatrixAgent + DeepSeek-V3.2, 10 Nov 2025).](https://canagentswork.com/og/benchmarks-theagentcompany.png)](https://canagentswork.com/benchmarks/theagentcompany/)

What it measures

Whether an agent completes 175 work tasks in a sealed company environment by browsing the web, writing code, running programs, and messaging coworkers. Checkers inspect the final state of the systems (for example a merged change, a filed document, or a message sent) and award full credit only when every checkpoint passes. Simulated coworkers are language-model agents that the agent must ask for information.

Tasks resolved: Share of the 175 tasks fully completed. The leaderboard also shows a partial-credit score that rewards passed checkpoints. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
simulated-workplace, browser, cli
Grading
state-check
Tasks
175
Contamination
All tasks and evaluators are public, so models trained after December 2024 may have seen them. The maintainers have not published a contamination analysis.
Reuse
MIT license on the GitHub repository. Task images, evaluators, and company data are public. (open-mit)

Limits to keep in mind

  • The leaderboard's newest entry is from November 2025, so it does not reflect models released in 2026. Source
  • Coworkers are simulated by a language model, and results depend on which model plays the environment (the board lists it per entry). Source
  • The company is a small software firm, so tasks lean toward engineering and internal tooling rather than the full range of office work. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemTasks resolvedDateSource
TTE-MatrixAgent + DeepSeek-V3.2 · frontier
harness: TTE-MatrixAgent
42.9%10 Nov 2025Carnegie Mellon University · primary
OpenHands-Versa + Claude Sonnet 4
harness: OpenHands-Versa
33.1%14 Jun 2025Carnegie Mellon University · primary
OpenHands + Gemini 2.5 Pro
harness: OpenHands
30.3%10 May 2025Carnegie Mellon University · primary
OpenHands + Claude 3.5 Sonnet
harness: OpenHands
24%17 Dec 2024Carnegie Mellon University · primary

Timeline

  • 18 Dec 2024 — CMU releases TheAgentCompany. Source

Where it sits in the atlas

Software engineeringManagement and business operations (partial)Office and administrative support (partial)

Enterprise workflowsFeature developmentLong-horizon autonomy

Last checked 23 Sep 2026 against 3 primary sources. See an error? Tell us.

How to cite

Credit the original work first: TheAgentCompany by Carnegie Mellon University (https://arxiv.org/abs/2412.14161).

Then, if you used this page:

Can Agents Work. "TheAgentCompany: frontier results and sources." https://canagentswork.com/benchmarks/theagentcompany/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-theagentcompany,
  title        = {{TheAgentCompany: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/theagentcompany/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}