Benchmarks / GAIA

GAIA

Built by Meta and Hugging Face · Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, et al. · released 21 Nov 2023

466 questions for general AI assistants that need web browsing, file reading, and tool use to answer. Each question has one short, unambiguous answer. Questions are easy for people but were hard for models in 2023.

Frontier

93.7%

Test-set accuracy

Ops-Agentic-Search-2.0

22 Sep 2026 · Source: Unnamed submitter (GAIA leaderboard entry links to github.com/tosky001/OpenSearch-Agentic-Search) (third party)

Top test-set entry as of 2026-09-23. Level 1 98.92%, level 2 93.08%, level 3 85.71%. The organisation field is blank. Self-submitted answer file; not re-run by the maintainers. The next entries are CustomGPT.ai Research Lab v44 at 93.36% and Co-Sight Pro v1.0.1 at 93.02%.

Saturated

The best result is at 90% or more of the ceiling.

See the full leaderboard at Meta

GAIA: Test-set accuracy over time, 5 recorded results. 0%20%40%60%80%100%Jan 2024Jan 2025Jan 2026 GPT-4 with plugins: 15% (Nov 2023) h2oGPTe Agent v1.6.8: 65.1% (21 Dec 2024) Co-Sight_v2.1.0: 87% (13 Oct 2025) OPS-Agentic-Search: 92.4% (11 Mar 2026) Ops-Agentic-Search-2.0: 93.7% (22 Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/gaia/"><img src="https://canagentswork.com/og/benchmarks-gaia.png" width="600" height="315" alt="GAIA: the best result is 93.7% (Ops-Agentic-Search-2.0, 22 Sep 2026)." loading="lazy"></a>

Markdown:

[![GAIA: the best result is 93.7% (Ops-Agentic-Search-2.0, 22 Sep 2026).](https://canagentswork.com/og/benchmarks-gaia.png)](https://canagentswork.com/benchmarks/gaia/)

What it measures

Share of questions answered exactly right. Questions come in three difficulty levels and often attach a file (a spreadsheet, image, audio clip, or PDF). Answers are checked by quasi exact match against a hidden reference. 165 questions form a public validation set; 300 form the test set with private answers that powers the Hugging Face leaderboard.

Test-set accuracy: Share of the 300 private test questions answered correctly (average across levels 1 to 3). Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
browser, api-tools, documents
Grading
automated-tests
Tasks
466
Human reference
Human respondents scored 92% in the paper's study.
Contamination
The validation questions and answers are public and the maintainers ask people not to train on them. Test answers stay private, but the questions are public.
Reuse
Gated dataset on Hugging Face. Test answers are private. The maintainers ask that the public set not be reposted or used for training. (cite-only)

Limits to keep in mind

  • The leaderboard accepts self-submitted answer files from anyone. Top entries are often unnamed or commercial agents with little public detail, and the maintainers do not re-run them. Source
  • Top scores now exceed the 92% human figure, so the benchmark no longer separates the best agents. Source
  • Answers are short strings or numbers, so the benchmark does not test long deliverables or judgment calls. Source
  • The validation leaderboard was closed because it was no longer informative. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemTest-set accuracyDateSource
Ops-Agentic-Search-2.0 · frontier93.7%22 Sep 2026Unnamed submitter (GAIA leaderboard entry links to github.com/tosky001/OpenSearch-Agentic-Search) · third-party
OPS-Agentic-Search92.4%11 Mar 2026Alibaba Cloud (submitted to the GAIA leaderboard) · third-party
Co-Sight_v2.1.087%13 Oct 2025ZTE-AICloud (submitted to the GAIA leaderboard) · third-party
h2oGPTe Agent v1.6.865.1%21 Dec 2024h2o.ai (submitted to the GAIA leaderboard) · third-party
GPT-4 with plugins15%Nov 2023Meta · primary

Timeline

  • 11 Mar 2026 — GAIA leaderboard passes the 92% human score. Source
  • 21 Nov 2023 — Meta and Hugging Face release GAIA. Source

Where it sits in the atlas

Office and administrative supportData and analytics (partial)

Web research

Last checked 23 Sep 2026 against 3 primary sources, with a second independent check. See an error? Tell us.

How to cite

Credit the original work first: GAIA by Meta and Hugging Face (https://arxiv.org/abs/2311.12983).

Then, if you used this page:

Can Agents Work. "GAIA: frontier results and sources." https://canagentswork.com/benchmarks/gaia/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-gaia,
  title        = {{GAIA: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/gaia/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}