Method
How we build the atlas, and the rules behind every color and label. If a rule changes, we say so on the timeline and version the change.
1. The territory: jobs
Jobs come from the SOC 2018 system through the O*NET 31.0 Database (997 civilian occupations; military occupations are out of scope because BLS OEWS does not cover them). We group occupations into 17 job families with ordered rules on SOC codes. Area on the atlas is US employment from the BLS OEWS May 2025 national estimates (830 detailed occupations). O*NET sub-occupations share the employment figure of their SOC group, so they change the pages, not the area.
2. Exploration: which benchmarks cover which jobs
Each benchmark maps to one or more job families: primary for its main family and partial for others. Some benchmarks also map to specific occupations. A family with no mapped benchmark shows as unexplored (hatched). That means nobody has measured the work in public in a way we can cite. It does not mean agents cannot do it.
3. Two grains: task and project
- Task level: bounded tasks with a clear spec, from minutes to hours of human time. Example: fix one GitHub issue.
- Project level: whole projects or deliverables judged by an acceptance standard, from hours to days. Example: a paid freelance project that a client must accept.
4. Benchmark status
Status compares the frontier result with a reference point that each benchmark declares. A ceiling is a hard maximum, such as 100% of tasks. A parity line is human-expert performance, such as 50% wins and ties against experts.
- Ceiling: open below 20%, emerging from 20%, strong from 60%, saturated from 90% of the ceiling.
- Parity: open below 50% of parity, emerging from 50%, strong at or above parity.
- No reference point (Elo scores, money balances, field signals, and lower-is-better behavior rates): unrated.
- Solved and retired come from the maintainers or a major evaluator, with a source and a date.
- Saturated
- The best result is at 90% or more of the ceiling.
- Strong
- The best result is at 60–90% of the ceiling, or at or above human parity.
- Emerging
- The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
- Open
- The best result is below 20% of the ceiling, or below half of human parity.
- Unrated
- No fixed reference point (for example Elo scores or field signals).
- Solved
- Its maintainers or a major evaluator declared it solved.
- Retired
- Maintainers or a major user stopped using it as a frontier measure.
- No results yet
- We have not recorded a result yet.
5. Readings
A reading is our written judgment of what the benchmarks mapped to a family show, at one grain. We write it, date it, and list its evidence. It is not a formula. Rules:
- Every reading above "Unexplored" cites at least one benchmark mapped to that family.
- A reading with evidence from fewer than two independent institutions carries the label thin evidence.
- We never claim that agents can do a whole job. The highest level is "Strong", and it always lists its limits.
How we pick a level, using only the benchmarks at the reading's grain:
- Unexplored: no benchmark at this grain maps to the family.
- Not yet: the benchmarks that map most directly to the family show frontier results in the open band, or the best agents lose to a human baseline.
- Partial: the evidence is mixed, or only adjacent benchmarks (partial mappings) show success, or success is shown only on narrow parts of the work.
- Strong: at least two primary-mapped benchmarks from independent institutions are strong, saturated, or solved, and no primary-mapped benchmark at this grain is open.
Benchmarks with no reference point (Elo scores, field signals, behavior rates) add context but do not set the level by themselves. A stale result still counts until a newer one replaces it. The reading says so when a result is old.
- Strong
- Agents succeed on most benchmark tasks. Limits remain in reliability, scope, or cost.
- Partial
- Agents succeed on some scoped tasks. Reliable or end-to-end work is not shown.
- Not yet
- The best agents fail most tasks on the best available benchmarks.
- Unexplored
- No benchmark covers this work yet.
6. Sources and provenance
- Primary sources first: the benchmark's paper, site, repo, or leaderboard, or its maintainers' posts.
- Lab-reported results come from a model or agent developer about its own system. We label them.
- Aggregators help us discover results. We do not cite them as the source of a number.
- Every result records its source URL, source kind, who reported it, and when we retrieved it. The build rejects a result without them.
- When sources disagree, we show the primary value and log the disagreement in public. See conflicts.
- We show the frontier and a few history points, and we link to the full leaderboard at the source.
7. What we do not do (yet)
- No composite score. One number would hide too much.
- No runs of our own. Every result comes from its source. We plan independent runs later, starting with DevOps.
- No logos without written permission. Names identify sources and do not imply endorsement.
8. Known limits
- Benchmarks measure tasks, not jobs. A job is a bundle of tasks, judgment, and relationships.
- Capability is not adoption. See usage research such as the Anthropic Economic Index for adoption.
- Coverage is biased toward work that is easy to benchmark, especially software.
- Labor data is for the United States. Most benchmarks are in English.
- Most results are single runs. Reliability over many runs is often unknown.
9. Corrections
If a number, summary, or mapping is wrong, tell us. We aim to correct errors within 48 hours of a report and record the change. Every entry shows the date we last checked it against its sources.