<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Can Agents Work? — timeline</title><link>https://canagentswork.com/timeline/</link><description>Launches, big jumps, saturation, and retirements across agent benchmarks.</description><language>en</language><item><title>Scale AI releases SWE-Bench Pro V2 with 642 validated tasks</title><link>https://canagentswork.com/benchmarks/swe-bench-pro/</link><guid isPermaLink="false">2026-09-22-swe-bench-pro-v2</guid><pubDate>Tue, 22 Sep 2026 00:00:00 GMT</pubDate><description>V2 drops 89 tasks from the 731-task public set, rewrites 529 problem statements, repairs verifiers, and adds a HARD-51 subset. It runs under a locked protocol: no network during the agent phase and re-grading on a fresh sandbox. The leaderboard still shows V1 results. Source: https://github.com/scaleapi/SWE-bench_Pro-os</description></item><item><title>APEX-Agents 1.1 stops rewarding hedged answers</title><link>https://canagentswork.com/benchmarks/apex-agents/</link><guid isPermaLink="false">2026-09-08-apex-agents-1-1</guid><pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate><description>Mercor found models raising scores by giving several answers at once ("scattergunning"). Version 1.1 trims the set to 240 audited tasks, adds a judge that scores hedged rubric items as zero, and fixes tools. Scores rise in aggregate and are not comparable with 1.0. Claude Fable 5.1 leads at 68.6% Pass@1. Source: https://www.mercor.com/blog/introducing-apex-agents-1-1</description></item><item><title>CyberGym leader reaches 98.5% reproduction rate</title><link>https://canagentswork.com/benchmarks/cybergym/</link><guid isPermaLink="false">2026-09-07-cybergym-frontier-98</guid><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate><description>The Creation multi-model agent tops the official leaderboard at 98.47%, up from under 20% at launch in June 2025 and 83.1% for Claude Mythos Preview in April 2026. Several team submissions now exceed 95%, and the maintainers warn that small differences may not be meaningful. Source: https://www.cybergym.io/cybergym/</description></item><item><title>GPT-6 Astra takes the Vending-Bench 2 lead at about $15,500</title><link>https://canagentswork.com/benchmarks/vending-bench-2/</link><guid isPermaLink="false">2026-09-07-vending-bench-2-gpt-6-astra</guid><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate><description>Andon Labs reports that GPT-6 Astra finished the simulated year with an average balance of $15,515, the first OpenAI model to lead the board and the largest gap over second place so far (Claude Opus 5, $11,181.87). Andon Labs credits steady negotiation all year and never prepaying suppliers that had gone out of business. Source: https://andonlabs.com/blog/gpt-6-astra-vending-bench</description></item><item><title>GDPval-AA v2.1 re-anchors the Elo scale</title><link>https://canagentswork.com/benchmarks/gdpval-aa/</link><guid isPermaLink="false">2026-09-gdpval-aa-v2-1-rescale</guid><pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate><description>Artificial Analysis pins DeepSeek V4.1 Flash (max) at 1600 Elo and fits ratings with a Crowd-BT model. Rank order is largely unchanged, but v2.1 scores are not comparable with v2 scores, which were anchored to human expert performance at 1000. Source: https://artificialanalysis.ai/methodology/intelligence-benchmarking</description></item><item><title>First verified OSWorld score above 90%</title><link>https://canagentswork.com/benchmarks/osworld-verified/</link><guid isPermaLink="false">2026-07-25-osworld-verified-passes-90</guid><pubDate>Sat, 25 Jul 2026 00:00:00 GMT</pubDate><description>The Intelligence-Indeed Agent completes 90.19% of OSWorld-Verified tasks at 100 steps in a run verified by the maintainers. A week later Claude Fable 5 posts 85.96% as the best general model. Both are well above the 72.36% human figure from the original study. Source: https://os-world.github.io/static/data/osworld_verified_results.xlsx</description></item><item><title>OpenAI audit estimates about 30% of SWE-Bench Pro tasks are broken</title><link>https://canagentswork.com/benchmarks/swe-bench-pro/</link><guid isPermaLink="false">2026-07-08-openai-audit-swe-bench-pro</guid><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate><description>OpenAI reviews the 731-task public split with an agent pipeline and five human engineers per flagged task. The pipeline flags 200 tasks (27.4%) and humans flag 249 (34.1%) as broken. OpenAI withdraws its earlier recommendation to adopt SWE-Bench Pro. Source: https://openai.com/index/separating-signal-from-noise-coding-evaluations/</description></item><item><title>Best Remote Labor Index score rises to 15.8%</title><link>https://canagentswork.com/benchmarks/remote-labor-index/</link><guid isPermaLink="false">2026-07-01-rli-frontier-quadruples</guid><pubDate>Wed, 01 Jul 2026 00:00:00 GMT</pubDate><description>CAIS reports that Claude Fable 5 completes 15.8% of real freelance projects at a client-acceptable standard. At launch in October 2025, the best agent completed 2.5%. Source: https://safe.ai/blog/significant-increase-in-digital-labor-automation</description></item><item><title>BrowseComp scores pass 90%</title><link>https://canagentswork.com/benchmarks/browsecomp/</link><guid isPermaLink="false">2026-07-browsecomp-passes-90</guid><pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate><description>OpenAI reports 90.4% for GPT-5.6 Sol and 92.2% for its four-agent "ultra" mode. By September 2026 GPT-6 Astra (91.5%) and Claude Opus 5 (90.8%) also sit above 90%, leaving little headroom. Source: https://openai.com/index/gpt-5-6/</description></item><item><title>OSWorld 2.0 succeeds OSWorld-Verified</title><link>https://canagentswork.com/benchmarks/osworld-verified/</link><guid isPermaLink="false">2026-06-26-osworld-2-launch</guid><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate><description>The XLANG team releases OSWorld 2.0, 108 long-horizon computer-use workflows that take skilled users a median of about 1.6 hours. The best agent, Claude Opus 4.8, completes only 20.6% of tasks at 500 steps, showing how much headroom the older benchmark no longer captures. Source: https://osworld-v2.xlang.ai/</description></item><item><title>SpreadsheetBench V1 leader reaches 83.11%</title><link>https://canagentswork.com/benchmarks/spreadsheetbench/</link><guid isPermaLink="false">2026-06-23-spreadsheetbench-passes-80</guid><pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate><description>Kingsoft Office's Qingqiu Agent scores 83.11% on the full 912-question set, a verified entry. The previous verified leader, Gemini in Google Sheets, scored 70.48% in March 2026. Source: https://spreadsheetbench.github.io</description></item><item><title>Artificial Analysis and IBM launch ITBench-AA</title><link>https://canagentswork.com/benchmarks/itbench-aa/</link><guid isPermaLink="false">2026-05-27-itbench-aa-launch</guid><pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate><description>59 Kubernetes incident-diagnosis tasks run in a fixed harness. Every frontier model scores below 50%. Claude Opus 4.7 leads at 47%, followed by GPT-5.5 at 46%. Source: https://artificialanalysis.ai/articles/itbench-aa-launch</description></item><item><title>METR reports a 17-hour time horizon and warns the suite is near its ceiling</title><link>https://canagentswork.com/benchmarks/metr-time-horizons/</link><guid isPermaLink="false">2026-05-08-metr-horizon-passes-16-hours</guid><pubDate>Fri, 08 May 2026 00:00:00 GMT</pubDate><description>METR adds Claude Mythos Preview (early) with a 50% horizon of 1,044.8 minutes (about 17.4 hours, CI 508.9 to 3,304.3) and posts a notice that measurements above 16 hours are unreliable with the current task suite. Claude Opus 4.6 had reached 718.8 minutes in February. Source: https://metr.org/time-horizons/</description></item><item><title>SREGym paper released with 90 live SRE problems</title><link>https://canagentswork.com/benchmarks/sregym/</link><guid isPermaLink="false">2026-05-08-sregym-launch</guid><pubDate>Fri, 08 May 2026 00:00:00 GMT</pubDate><description>UIUC and University of Toronto publish SREGym. Frontier agents reach 27% to 61% end-to-end success on the 90-problem suite. On failures unique to SREGym (hardware, metastable, concurrent), the best agent falls to 28%. Source: https://arxiv.org/abs/2605.07161</description></item><item><title>Terminal-Bench 2.1 fixes 28 tasks and replaces 2.0</title><link>https://canagentswork.com/benchmarks/terminal-bench-2/</link><guid isPermaLink="false">2026-05-06-terminal-bench-2-1-supersedes-2-0</guid><pubDate>Wed, 06 May 2026 00:00:00 GMT</pubDate><description>The maintainers release 2.1 after finding defects in 28 of the 89 tasks in 2.0: changed external dependencies, resource budgets too small for valid solutions, and instructions that did not match tests. Most agent-model pairs score higher on 2.1; Claude Code with Opus 4.6 gains 12.1 points. Terminal-Bench 3.0 (July 2026) and 4.0 (August 2026) followed. Source: https://www.tbench.ai/news/terminal-bench-2-1</description></item><item><title>MLE-bench pauses new leaderboard submissions</title><link>https://canagentswork.com/benchmarks/mle-bench/</link><guid isPermaLink="false">2026-04-24-mle-bench-submissions-paused</guid><pubDate>Fri, 24 Apr 2026 00:00:00 GMT</pubDate><description>OpenAI stops taking new MLE-bench leaderboard submissions while it builds a process to make sure submissions are fair and comparable. A v2 release with batched fixes is planned in the frontier-evals repo. Source: https://github.com/openai/mle-bench</description></item><item><title>Claude Mythos Preview reaches 100% on Cybench subset</title><link>https://canagentswork.com/benchmarks/cybench/</link><guid isPermaLink="false">2026-04-07-cybench-saturated</guid><pubDate>Tue, 07 Apr 2026 00:00:00 GMT</pubDate><description>Anthropic's system card reports 100% pass@1 on its 35-task Cybench subset with 10 trials per task, and says that, given the saturation of the benchmark, it is no longer sufficiently informative. Source: https://cdn.sanity.io/files/4zrzovbb/website/7624816413e9b4d2e3ba620c5a5e091b98b190a5.pdf</description></item><item><title>GPT-5.5 reaches 84.9% wins plus ties on GDPval</title><link>https://canagentswork.com/benchmarks/gdpval/</link><guid isPermaLink="false">2026-04-gdpval-gpt-5-5-reaches-85</guid><pubDate>Wed, 15 Apr 2026 00:00:00 GMT</pubDate><description>In the GPT-5.5 launch post OpenAI reports 84.9% wins plus ties against industry experts, with GPT-5.4 at 83.0% and Claude Opus 4.7 at 80.3%. The post does not say which grader produced the numbers. Source: https://openai.com/index/introducing-gpt-5-5/</description></item><item><title>GAIA leaderboard passes the 92% human score</title><link>https://canagentswork.com/benchmarks/gaia/</link><guid isPermaLink="false">2026-03-11-gaia-passes-human-score</guid><pubDate>Wed, 11 Mar 2026 00:00:00 GMT</pubDate><description>Alibaba Cloud's OPS-Agentic-Search posts 92.36% on the GAIA test set, the first entry above the 92% human figure from the paper. By September 2026 several self-submitted agents sit between 93% and 94%, so the benchmark no longer separates the best systems. Source: https://huggingface.co/datasets/gaia-benchmark/results_public</description></item><item><title>Spider 2.0-Snow leader reaches 96.70%</title><link>https://canagentswork.com/benchmarks/spider-2/</link><guid isPermaLink="false">2026-03-01-spider-2-snow-passes-96</guid><pubDate>Sun, 01 Mar 2026 00:00:00 GMT</pubDate><description>Genloop's Sentinel Agent v2 Pro tops the Spider 2.0-Snow table at 96.70%. Fifteen months earlier the maintainers' Spider-Agent with o1-preview scored 23.58%. The Snow variant now has little headroom. Source: https://spider2-sql.github.io</description></item><item><title>Martian launches Code Review Bench</title><link>https://canagentswork.com/benchmarks/martian-code-review-bench/</link><guid isPermaLink="false">2026-02-26-code-review-bench-launch</guid><pubDate>Thu, 26 Feb 2026 00:00:00 GMT</pubDate><description>Martian publishes an open benchmark for AI code review tools with an offline set of 50 PRs and 173 golden comments, plus an online set built from fresh GitHub PRs. The repo, judge prompts, and results are MIT licensed. Source: https://github.com/withmartian/code-review-benchmark</description></item><item><title>MLE-bench leader earns medals in 64.44% of competitions</title><link>https://canagentswork.com/benchmarks/mle-bench/</link><guid isPermaLink="false">2026-02-23-mle-bench-passes-60</guid><pubDate>Mon, 23 Feb 2026 00:00:00 GMT</pubDate><description>Baidu's Famou-Agent 2.0 with Gemini-3-Pro-Preview tops the MLE-bench table at 64.44% any-medal rate. Sixteen months earlier the best agent scored 17.12%. Source: https://github.com/openai/mle-bench</description></item><item><title>OpenAI stops reporting SWE-bench Verified</title><link>https://canagentswork.com/benchmarks/swe-bench-verified/</link><guid isPermaLink="false">2026-02-23-openai-retires-swe-bench-verified</guid><pubDate>Mon, 23 Feb 2026 00:00:00 GMT</pubDate><description>OpenAI says SWE-bench Verified no longer measures frontier coding progress. Its audit found flawed tests in 59.4% of 138 hard tasks, and probes showed frontier models from three labs reproducing gold patches from training data. OpenAI asks other developers to stop reporting it and points to SWE-Bench Pro instead. Source: https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/</description></item><item><title>Sierra fixes 50+ airline and retail tasks for τ³-bench</title><link>https://canagentswork.com/benchmarks/tau2-bench/</link><guid isPermaLink="false">2026-02-tau2-bench-task-fixes</guid><pubDate>Sun, 15 Feb 2026 00:00:00 GMT</pubDate><description>Sierra audited the airline and retail domains with the SABER (τ-Bench Verified) team at Amazon and fixed 27 airline and 26 retail tasks with wrong expected actions, ambiguous instructions, or impossible constraints. Airline pass^1 rose by 14 to 20 points for the three re-run models, so scores before and after are not comparable. Telecom was not changed. Source: https://taubench.com/blog/tau3-task-fixes.html</description></item><item><title>METR releases Time Horizon 1.1 with 228 tasks</title><link>https://canagentswork.com/benchmarks/metr-time-horizons/</link><guid isPermaLink="false">2026-01-29-metr-time-horizon-1-1</guid><pubDate>Thu, 29 Jan 2026 00:00:00 GMT</pubDate><description>METR grows the task suite from 170 to 228 tasks, doubles the number of 8-hour-plus tasks to 31, and moves from Vivaria to Inspect. Estimates shift: the post-2023 doubling time drops from 165 to 131 days, and Claude Opus 4.5 moves from 289 to 320 minutes. Source: https://metr.org/blog/2026-1-29-time-horizon-1-1/</description></item><item><title>Mercor releases APEX-Agents</title><link>https://canagentswork.com/benchmarks/apex-agents/</link><guid isPermaLink="false">2026-01-21-apex-agents-launch</guid><pubDate>Wed, 21 Jan 2026 00:00:00 GMT</pubDate><description>Mercor, with Box and Harvey, releases 480 long-horizon tasks from investment banking, consulting, and corporate law inside simulated project workspaces. The best agent, Gemini 3 Flash, passes 24.0% of tasks on one attempt. Source: https://www.mercor.com/blog/introducing-apex-agents/</description></item><item><title>HAL declares CORE-Bench solved</title><link>https://canagentswork.com/benchmarks/core-bench/</link><guid isPermaLink="false">2025-12-03-core-bench-solved</guid><pubDate>Wed, 03 Dec 2025 00:00:00 GMT</pubDate><description>Claude Opus 4.5 in a Claude Code scaffold scores 77.78% on CORE-Bench-Hard, nearly double its 42.22% with HAL's CORE-Agent. After HAL fixed grading errors in 8 tasks and removed 1 broken task, the score is 95.5%. HAL treats the benchmark as solved and plans to open a private test set. Source: https://hal.cs.princeton.edu/corebench_hard</description></item><item><title>Artificial Analysis launches GDPval-AA</title><link>https://canagentswork.com/benchmarks/gdpval-aa/</link><guid isPermaLink="false">2025-12-gdpval-aa-launch</guid><pubDate>Mon, 15 Dec 2025 00:00:00 GMT</pubDate><description>Artificial Analysis opens an independent Elo leaderboard that runs OpenAI's public GDPval tasks through its open-source Stirrup agent harness and grades deliverables with blind pairwise LLM judging. Claude Opus 4.5 leads at launch. Source: https://www.linkedin.com/posts/artificial-analysis_announcing-gdpval-aa-our-leaderboard-and-activity-7404608113624072193-Q82r</description></item><item><title>GPT-5.2 Thinking passes the 50% line on GDPval</title><link>https://canagentswork.com/benchmarks/gdpval/</link><guid isPermaLink="false">2025-12-gdpval-passes-expert-parity</guid><pubDate>Mon, 15 Dec 2025 00:00:00 GMT</pubDate><description>OpenAI reports that GPT-5.2 Thinking beats or ties industry professionals in 70.9% of GDPval comparisons, judged by expert humans. It is the first model OpenAI reports above the 50% parity line. Source: https://openai.com/index/introducing-gpt-5-2/</description></item><item><title>GPT-5.1-Codex-Max reaches 80% on SWE-Lancer IC Diamond</title><link>https://canagentswork.com/benchmarks/swe-lancer/</link><guid isPermaLink="false">2025-11-18-swe-lancer-codex-max-80</guid><pubDate>Tue, 18 Nov 2025 00:00:00 GMT</pubDate><description>OpenAI's GPT-5.1-Codex-Max system card reports about 80% pass@1 on IC SWE Diamond tasks, averaged over three runs, up from 55% for GPT-5 and 67% for GPT-5-Codex in the same chart. This is the last OpenAI system card to report SWE-Lancer. Source: https://deploymentsafety.openai.com/gpt-5-1-codex-max/swe-lancer</description></item><item><title>Andon Labs releases Vending-Bench 2</title><link>https://canagentswork.com/benchmarks/vending-bench-2/</link><guid isPermaLink="false">2025-11-18-vending-bench-2-launch</guid><pubDate>Tue, 18 Nov 2025 00:00:00 GMT</pubDate><description>The second Vending-Bench adds adversarial suppliers, negotiation, delivery delays, supplier bankruptcies, and refund demands, and scores the bank balance after one simulated year. Gemini 3 Pro led at launch with $5,478.16, ahead of Claude Sonnet 4.5 ($3,838.74) and Grok 4 ($1,999.46). Source: http://web.archive.org/web/20251118175816/https://andonlabs.com/evals/vending-bench-2</description></item><item><title>Terminal-Bench 2.0 and Harbor released</title><link>https://canagentswork.com/benchmarks/terminal-bench-2/</link><guid isPermaLink="false">2025-11-07-terminal-bench-2-launch</guid><pubDate>Fri, 07 Nov 2025 00:00:00 GMT</pubDate><description>The Terminal-Bench team releases 2.0, a harder and more carefully verified set of 89 terminal tasks, together with Harbor, a new package for running agents in cloud containers. The best verified row for a pre-launch model is Codex CLI with GPT-5 at 49.6%. Source: https://www.tbench.ai/news/announcement-2-0</description></item><item><title>CodeClash launches goal-oriented coding tournaments</title><link>https://canagentswork.com/benchmarks/codeclash/</link><guid isPermaLink="false">2025-11-02-codeclash-launch</guid><pubDate>Sun, 02 Nov 2025 00:00:00 GMT</pubDate><description>Stanford and Princeton researchers release CodeClash, where models evolve codebases over 15-round tournaments in six game arenas. Across 1,680 tournaments, Claude Sonnet 4.5 leads with the highest Elo, followed by GPT-5 and o3. Top models lose every round to expert human bots. Source: https://arxiv.org/abs/2511.00839</description></item><item><title>CVE-Bench v2.0 closes grading shortcuts</title><link>https://canagentswork.com/benchmarks/cve-bench/</link><guid isPermaLink="false">2025-10-30-cve-bench-v2-grading-fix</guid><pubDate>Thu, 30 Oct 2025 00:00:00 GMT</pubDate><description>The maintainers hardened the outbound-request check and the SQL-injection check after finding that agents could pass without real exploits. GPT-4o agent success rates fell by up to 32.5 points after the fixes. Source: https://open.substack.com/pub/ddkang/p/cve-bench-v20-making-evaluation-more</description></item><item><title>ImpossibleBench measures how often coding agents game their tests</title><link>https://canagentswork.com/benchmarks/impossiblebench/</link><guid isPermaLink="false">2025-10-23-impossiblebench-launch</guid><pubDate>Thu, 23 Oct 2025 00:00:00 GMT</pubDate><description>CMU and Anthropic researchers release tasks where tests contradict the specification, so any pass is a shortcut. GPT-5 passes 54.0% of the conflicting SWE-bench tasks by cheating. Stricter prompts cut cheating sharply. Source: https://arxiv.org/abs/2510.20270</description></item><item><title>GPT-5 reaches 95.8% telecom pass^1 on τ²-bench</title><link>https://canagentswork.com/benchmarks/tau2-bench/</link><guid isPermaLink="false">2025-10-02-tau2-bench-gpt-5-telecom-96</guid><pubDate>Thu, 02 Oct 2025 00:00:00 GMT</pubDate><description>Sierra's own run of GPT-5 (evaluated 2025-08-09, added to the leaderboard 2025-10-02) scored 95.8% pass^1 and 85.1% pass^4 on telecom, up from about 50% pass^1 four months earlier. Telecom pass^1 has stayed near 98% for the best models since then. Source: https://github.com/sierra-research/tau2-bench/blob/main/web/leaderboard/public/submissions/gpt-5_sierra_2025-08-09/submission.json</description></item><item><title>OpenAI releases GDPval</title><link>https://canagentswork.com/benchmarks/gdpval/</link><guid isPermaLink="false">2025-09-25-gdpval-launch</guid><pubDate>Thu, 25 Sep 2025 00:00:00 GMT</pubDate><description>OpenAI publishes GDPval, 1,320 real work tasks from 44 occupations in 9 US industries, with a public 220-task gold subset. Blinded experts rate Claude Opus 4.1 better than or equal to the professional's deliverable in 47.6% of comparisons. Source: https://openai.com/index/gdpval/</description></item><item><title>Scale AI releases SWE-Bench Pro</title><link>https://canagentswork.com/benchmarks/swe-bench-pro/</link><guid isPermaLink="false">2025-09-21-swe-bench-pro-launch</guid><pubDate>Sun, 21 Sep 2025 00:00:00 GMT</pubDate><description>Scale AI publishes SWE-Bench Pro with 1,865 long-horizon tasks from 41 repositories, including a 731-task public set from GPL repositories and a 276-task private set from startups. Under a 50-turn, $2 cap, the best public-set result is 23.3% (GPT-5, medium reasoning). Source: https://arxiv.org/abs/2509.16941</description></item><item><title>OSWorld becomes OSWorld-Verified</title><link>https://canagentswork.com/benchmarks/osworld-verified/</link><guid isPermaLink="false">2025-07-28-osworld-verified-launch</guid><pubDate>Mon, 28 Jul 2025 00:00:00 GMT</pubDate><description>The HKU XLANG team fixes about 300 reported problems in OSWorld tasks and checkers, moves evaluation to a parallel AWS setup, and re-runs all baselines. The report names CoACT-1 as the best agent at 60.76%, about 84% of the roughly 72% human figure. Source: https://xlang.ai/blog/osworld-verified</description></item><item><title>AIDev dataset of agent pull requests released</title><link>https://canagentswork.com/benchmarks/aidev-dataset/</link><guid isPermaLink="false">2025-07-20-aidev-dataset-release</guid><pubDate>Sun, 20 Jul 2025 00:00:00 GMT</pubDate><description>Queen's University researchers publish AIDev, 456,535 pull requests by five coding agents across 61,453 GitHub repositories. In popular repositories, agent PRs are merged less often than human PRs (Codex 65.3% versus 76.8% for humans). Source: https://arxiv.org/abs/2507.15003</description></item><item><title>Sierra releases τ²-bench with a dual-control telecom domain</title><link>https://canagentswork.com/benchmarks/tau2-bench/</link><guid isPermaLink="false">2025-06-09-tau2-bench-launch</guid><pubDate>Mon, 09 Jun 2025 00:00:00 GMT</pubDate><description>τ²-bench adds a telecom customer-service domain where both the agent and the simulated user hold tools. At launch the best telecom pass^1 was about 50% (o4-mini, gpt-4.1-mini, and Claude 3.7 Sonnet at 49%), and GPT-4.1 dropped from 74% on retail to 34% on telecom. Source: https://arxiv.org/abs/2506.07982</description></item><item><title>UC Berkeley releases CyberGym with 1,507 real vulnerabilities</title><link>https://canagentswork.com/benchmarks/cybergym/</link><guid isPermaLink="false">2025-06-03-cybergym-launch</guid><pubDate>Tue, 03 Jun 2025 00:00:00 GMT</pubDate><description>Agents must write proof-of-concept inputs that reproduce real OSS-Fuzz vulnerabilities in 188 projects. The best agent-model pairs at launch reproduce fewer than 20% of them. Source: https://arxiv.org/abs/2506.02548</description></item><item><title>PR Arena starts tracking agent pull requests on GitHub</title><link>https://canagentswork.com/benchmarks/pr-arena/</link><guid isPermaLink="false">2025-05-26-pr-arena-launch</guid><pubDate>Mon, 26 May 2025 00:00:00 GMT</pubDate><description>The PRarena repo records its first data point. At that time Codex had 51,548 ready PRs with 85.8% merged, and Copilot had 2,099 ready PRs with 75.94% merged. Source: https://raw.githubusercontent.com/aavetis/PRarena/main/data.csv</description></item><item><title>Nebius launches SWE-rebench with fresh GitHub issues</title><link>https://canagentswork.com/benchmarks/swe-rebench/</link><guid isPermaLink="false">2025-05-26-swe-rebench-launch</guid><pubDate>Mon, 26 May 2025 00:00:00 GMT</pubDate><description>Nebius publishes an automated pipeline that collects new issue and pull request pairs from Python repositories, plus a leaderboard that runs every model in the same ReAct scaffold five times. In the paper, GPT-4.1 leads the January 2025 window at 31.1%. Source: https://arxiv.org/abs/2505.20411</description></item><item><title>Salesforce releases CRMArena-Pro</title><link>https://canagentswork.com/benchmarks/crmarena-pro/</link><guid isPermaLink="false">2025-05-24-crmarena-pro-launch</guid><pubDate>Sat, 24 May 2025 00:00:00 GMT</pubDate><description>CRMArena-Pro expands CRMArena to 19 sales, service, and CPQ task types across B2B and B2C Salesforce orgs, adds multi-turn users and confidentiality checks. The best agent (gemini-2.5-pro with ReAct) completed 58.3% of single-turn B2C queries and about 35% in multi-turn. All models showed near-zero confidentiality awareness without special prompting. Source: https://arxiv.org/abs/2505.18878</description></item><item><title>Vals AI releases the Finance Agent Benchmark</title><link>https://canagentswork.com/benchmarks/vals-finance-agent/</link><guid isPermaLink="false">2025-05-20-vals-finance-agent-launch</guid><pubDate>Tue, 20 May 2025 00:00:00 GMT</pubDate><description>Vals AI publishes 537 expert-written financial research questions and an agent harness with EDGAR and web search tools. The best model, OpenAI o3, answered 46.8% correctly at an average cost of $3.79 per query. Source: https://arxiv.org/abs/2508.00828</description></item><item><title>OpenAI releases PaperBench for replicating ICML papers</title><link>https://canagentswork.com/benchmarks/paperbench/</link><guid isPermaLink="false">2025-04-02-paperbench-launch</guid><pubDate>Wed, 02 Apr 2025 00:00:00 GMT</pubDate><description>Agents must replicate 20 ICML 2024 papers from scratch. Claude 3.5 Sonnet with a basic scaffold scores 21.0%, and o1 with a 36-hour limit scores 26.0%. ML PhDs reached 41.4% on a 3-paper subset after 48 hours. Source: https://arxiv.org/abs/2504.01848</description></item><item><title>OpenAI releases BrowseComp</title><link>https://canagentswork.com/benchmarks/browsecomp/</link><guid isPermaLink="false">2025-04-browsecomp-launch</guid><pubDate>Tue, 15 Apr 2025 00:00:00 GMT</pubDate><description>OpenAI open-sources 1,266 hard-to-find, easy-to-verify browsing questions. Deep research answers 51.5%, o1 9.9%, and GPT-4o with browsing 1.9%. Human trainers solved 29.2% within two hours. Source: https://openai.com/index/browsecomp/</description></item><item><title>UIUC releases CVE-Bench with 40 critical web CVEs</title><link>https://canagentswork.com/benchmarks/cve-bench/</link><guid isPermaLink="false">2025-03-31-cve-bench-launch</guid><pubDate>Mon, 31 Mar 2025 00:00:00 GMT</pubDate><description>Agents must exploit real critical-severity vulnerabilities in sandboxed web applications. The best GPT-4o agent framework exploits 12.5% of CVEs with five attempts when given a vulnerability description. Source: https://github.com/uiuc-kang-lab/cve-bench</description></item><item><title>METR introduces the 50% task-completion time horizon</title><link>https://canagentswork.com/benchmarks/metr-time-horizons/</link><guid isPermaLink="false">2025-03-18-metr-time-horizon-paper</guid><pubDate>Tue, 18 Mar 2025 00:00:00 GMT</pubDate><description>METR's paper "Measuring AI Ability to Complete Long Software Tasks" defines the time horizon metric. Claude 3.7 Sonnet had a 50% horizon of around 50 minutes, and the frontier horizon had doubled about every seven months since 2019. Source: https://arxiv.org/abs/2503.14499</description></item><item><title>Ambig-SWE tests whether coding agents ask for clarification</title><link>https://canagentswork.com/benchmarks/ambig-swe/</link><guid isPermaLink="false">2025-02-18-ambig-swe-launch</guid><pubDate>Tue, 18 Feb 2025 00:00:00 GMT</pubDate><description>CMU researchers release an underspecified variant of SWE-Bench Verified with a simulated user. Models rarely ask questions unless prompted, and most cannot tell a vague issue from a complete one. When they do interact, resolve rates rise sharply. Source: https://arxiv.org/abs/2502.13069</description></item><item><title>OpenAI releases SWE-Lancer with $1 million of Upwork tasks</title><link>https://canagentswork.com/benchmarks/swe-lancer/</link><guid isPermaLink="false">2025-02-17-swe-lancer-launch</guid><pubDate>Mon, 17 Feb 2025 00:00:00 GMT</pubDate><description>OpenAI publishes 1,488 real freelance software tasks worth $1 million in actual payouts, with a public Diamond split worth $500,800. The best model, Claude 3.5 Sonnet, solves 26.2% of Diamond IC tasks and earns $208,050 across the Diamond set. Source: https://openai.com/index/swe-lancer/</description></item><item><title>IBM Research releases ITBench</title><link>https://canagentswork.com/benchmarks/itbench/</link><guid isPermaLink="false">2025-02-07-itbench-launch</guid><pubDate>Fri, 07 Feb 2025 00:00:00 GMT</pubDate><description>IBM Research and UIUC publish ITBench, a framework of live Kubernetes scenarios for SRE, compliance (CISO), and FinOps agents. The first paper reports that agents resolve 13.8% of SRE scenarios. Source: https://arxiv.org/abs/2502.05352</description></item><item><title>Stanford releases MedAgentBench</title><link>https://canagentswork.com/benchmarks/medagentbench/</link><guid isPermaLink="false">2025-01-24-medagentbench-launch</guid><pubDate>Fri, 24 Jan 2025 00:00:00 GMT</pubDate><description>MedAgentBench gives agents 300 physician-written tasks in a FHIR-compliant virtual EHR with 100 de-identified patients. Claude 3.5 Sonnet v2 led with a 69.67% success rate; models did much better on record lookups (up to 85%) than on tasks that change records. Source: https://arxiv.org/abs/2501.14654</description></item><item><title>AIOpsLab paper released with 48 cloud operations problems</title><link>https://canagentswork.com/benchmarks/aiopslab/</link><guid isPermaLink="false">2025-01-12-aiopslab-launch</guid><pubDate>Sun, 12 Jan 2025 00:00:00 GMT</pubDate><description>Microsoft Research and partners publish AIOpsLab, a framework that deploys microservices, injects faults, and evaluates agents on detection, localization, root cause analysis, and mitigation. The best agent in the paper reaches 59.32% accuracy. Source: https://arxiv.org/abs/2501.06706</description></item><item><title>Aider launches the polyglot leaderboard</title><link>https://canagentswork.com/benchmarks/aider-polyglot/</link><guid isPermaLink="false">2024-12-21-aider-polyglot-launch</guid><pubDate>Sat, 21 Dec 2024 00:00:00 GMT</pubDate><description>Paul Gauthier replaces aider's saturating Python-only benchmark with 225 hard Exercism exercises in six languages. o1 with high reasoning effort leads at 61.7%, ahead of Claude 3.5 Sonnet at 45.3%. Source: https://aider.chat/2024/12/21/polyglot.html</description></item><item><title>CMU releases TheAgentCompany</title><link>https://canagentswork.com/benchmarks/theagentcompany/</link><guid isPermaLink="false">2024-12-18-theagentcompany-launch</guid><pubDate>Wed, 18 Dec 2024 00:00:00 GMT</pubDate><description>CMU publishes a simulated software company with GitLab, Plane, ownCloud, RocketChat, and language-model coworkers, plus 175 work tasks. The best agent, OpenHands with Claude 3.5 Sonnet, completes 24.0% of tasks. Source: https://arxiv.org/abs/2412.14161</description></item><item><title>METR releases RE-Bench with 71 human expert baselines</title><link>https://canagentswork.com/benchmarks/re-bench/</link><guid isPermaLink="false">2024-11-22-re-bench-launch</guid><pubDate>Fri, 22 Nov 2024 00:00:00 GMT</pubDate><description>Seven ML research engineering environments with matched human and agent conditions. Agents score about 4 times the human average at a 2-hour budget, but humans narrowly pass the best agent at 8 hours and reach about twice the agent score at 32 hours. Source: https://metr.org/blog/2024-11-22-evaluating-r-d-capabilities-of-llms/</description></item><item><title>Spider 2.0 launches with 632 enterprise text-to-SQL problems</title><link>https://canagentswork.com/benchmarks/spider-2/</link><guid isPermaLink="false">2024-11-12-spider-2-launch</guid><pubDate>Tue, 12 Nov 2024 00:00:00 GMT</pubDate><description>HKU's XLANG Lab and partners release Spider 2.0. The paper reports that its code agent with o1-preview solves 21.3% of tasks, against 91.2% on Spider 1.0 and 73.0% on BIRD. Source: https://arxiv.org/abs/2411.07763</description></item><item><title>OpenAI releases MLE-bench with 75 Kaggle competitions</title><link>https://canagentswork.com/benchmarks/mle-bench/</link><guid isPermaLink="false">2024-10-09-mle-bench-launch</guid><pubDate>Wed, 09 Oct 2024 00:00:00 GMT</pubDate><description>The best setup at launch, o1-preview with the AIDE scaffold, reaches a medal in about 17% of competitions (16.9% in the paper, 17.12% in the repo table). Source: https://arxiv.org/abs/2410.07095</description></item><item><title>Princeton releases CORE-Bench for computational reproducibility</title><link>https://canagentswork.com/benchmarks/core-bench/</link><guid isPermaLink="false">2024-09-17-core-bench-launch</guid><pubDate>Tue, 17 Sep 2024 00:00:00 GMT</pubDate><description>270 tasks from 90 published papers at three difficulty levels. The best agent, CORE-Agent with GPT-4o, reaches 21.48% on the hardest level. Source: https://arxiv.org/abs/2409.11363</description></item><item><title>Stanford releases Cybench with 40 professional CTF tasks</title><link>https://canagentswork.com/benchmarks/cybench/</link><guid isPermaLink="false">2024-08-15-cybench-launch</guid><pubDate>Thu, 15 Aug 2024 00:00:00 GMT</pubDate><description>The best agent (Claude 3.5 Sonnet) solves 17.5% of tasks unguided; GPT-4o solves 12.5%. Agents solve tasks that took human teams up to 11 minutes; the hardest task took humans 24 hours 54 minutes. Source: https://arxiv.org/abs/2408.08926</description></item><item><title>OpenAI and the SWE-bench team release SWE-bench Verified</title><link>https://canagentswork.com/benchmarks/swe-bench-verified/</link><guid isPermaLink="false">2024-08-13-swe-bench-verified-launch</guid><pubDate>Tue, 13 Aug 2024 00:00:00 GMT</pubDate><description>93 developers screen 1,699 SWE-bench tasks. The 500 tasks that pass become SWE-bench Verified. GPT-4o resolves 33.2% with the best open-source scaffold, double its score on the original SWE-bench. Source: https://openai.com/index/introducing-swe-bench-verified/</description></item><item><title>SpreadsheetBench launches with 912 real spreadsheet questions</title><link>https://canagentswork.com/benchmarks/spreadsheetbench/</link><guid isPermaLink="false">2024-06-21-spreadsheetbench-launch</guid><pubDate>Fri, 21 Jun 2024 00:00:00 GMT</pubDate><description>Renmin University and partners release SpreadsheetBench. In the paper GPT-4o scores 18.35% (soft) and 15.02% (hard) overall, and Copilot in Excel about 20% on a subset, while Excel experts score 71.33% and 62.00% on a 50-question subset. Source: https://arxiv.org/abs/2406.14991</description></item><item><title>ServiceNow Research releases WorkArena and BrowserGym</title><link>https://canagentswork.com/benchmarks/workarena/</link><guid isPermaLink="false">2024-03-12-workarena-launch</guid><pubDate>Tue, 12 Mar 2024 00:00:00 GMT</pubDate><description>ServiceNow Research publishes WorkArena, 33 browser tasks on a live ServiceNow instance, together with the BrowserGym environment for web agents. WorkArena++ (682 compositional tasks) follows in July 2024. Source: https://arxiv.org/abs/2403.07718</description></item><item><title>Meta and Hugging Face release GAIA</title><link>https://canagentswork.com/benchmarks/gaia/</link><guid isPermaLink="false">2023-11-21-gaia-launch</guid><pubDate>Tue, 21 Nov 2023 00:00:00 GMT</pubDate><description>GAIA offers 466 questions that need browsing, file reading, and tool use. Humans score 92% while GPT-4 with plugins scores 15%. A leaderboard with private test answers opens on Hugging Face. Source: https://arxiv.org/abs/2311.12983</description></item></channel></rss>