Benchmarks / GAIA
GAIA
Built by Meta and Hugging Face · Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, et al. · released 21 Nov 2023
466 questions for general AI assistants that need web browsing, file reading, and tool use to answer. Each question has one short, unambiguous answer. Questions are easy for people but were hard for models in 2023.
Frontier
93.7%
Test-set accuracy
Ops-Agentic-Search-2.0
22 Sep 2026 · Source: Unnamed submitter (GAIA leaderboard entry links to github.com/tosky001/OpenSearch-Agentic-Search) (third party)
Top test-set entry as of 2026-09-23. Level 1 98.92%, level 2 93.08%, level 3 85.71%. The organisation field is blank. Self-submitted answer file; not re-run by the maintainers. The next entries are CustomGPT.ai Research Lab v44 at 93.36% and Co-Sight Pro v1.0.1 at 93.02%.
The best result is at 90% or more of the ceiling.
What it measures
Share of questions answered exactly right. Questions come in three difficulty levels and often attach a file (a spreadsheet, image, audio clip, or PDF). Answers are checked by quasi exact match against a hidden reference. 165 questions form a public validation set; 300 form the test set with private answers that powers the Hugging Face leaderboard.
Test-set accuracy: Share of the 300 private test questions answered correctly (average across levels 1 to 3). Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- browser, api-tools, documents
- Grading
- automated-tests
- Tasks
- 466
- Human reference
- Human respondents scored 92% in the paper's study.
- Contamination
- The validation questions and answers are public and the maintainers ask people not to train on them. Test answers stay private, but the questions are public.
- Reuse
- Gated dataset on Hugging Face. Test answers are private. The maintainers ask that the public set not be reposted or used for training. (cite-only)
Limits to keep in mind
- The leaderboard accepts self-submitted answer files from anyone. Top entries are often unnamed or commercial agents with little public detail, and the maintainers do not re-run them. Source
- Top scores now exceed the 92% human figure, so the benchmark no longer separates the best agents. Source
- Answers are short strings or numbers, so the benchmark does not test long deliverables or judgment calls. Source
- The validation leaderboard was closed because it was no longer informative. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Test-set accuracy | Date | Source |
|---|---|---|---|
| Ops-Agentic-Search-2.0 · frontier | 93.7% | 22 Sep 2026 | Unnamed submitter (GAIA leaderboard entry links to github.com/tosky001/OpenSearch-Agentic-Search) · third-party |
| OPS-Agentic-Search | 92.4% | 11 Mar 2026 | Alibaba Cloud (submitted to the GAIA leaderboard) · third-party |
| Co-Sight_v2.1.0 | 87% | 13 Oct 2025 | ZTE-AICloud (submitted to the GAIA leaderboard) · third-party |
| h2oGPTe Agent v1.6.8 | 65.1% | 21 Dec 2024 | h2o.ai (submitted to the GAIA leaderboard) · third-party |
| GPT-4 with plugins | 15% | Nov 2023 | Meta · primary |
Timeline
Where it sits in the atlas
Office and administrative supportData and analytics (partial)
Go to the source
- Website huggingface.co
- Paper arxiv.org
- Full leaderboard huggingface.co
- Dataset huggingface.co
Last checked 23 Sep 2026 against 3 primary sources, with a second independent check. See an error? Tell us.
How to cite
Credit the original work first: GAIA by Meta and Hugging Face (https://arxiv.org/abs/2311.12983).
Then, if you used this page:
Can Agents Work. "GAIA: frontier results and sources." https://canagentswork.com/benchmarks/gaia/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-gaia,
title = {{GAIA: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/gaia/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}