AI benchmarks vs your own evals: how to read a model score / The benchmarks vendors quote
Agent benchmarks explained: GAIA, WebArena, tau-bench and METR time horizons
Agent benchmarks score whether an AI model can finish a task that takes several steps and tools. GAIA has 466 questions, WebArena has 812 website tasks, tau-bench has 165 customer service tasks and OSWorld has 369 desktop tasks. The METR time horizon reports the length of task, measured by how long it takes a skilled person, that a model completes half the time. At publication the best models scored between 12.24% and 61.2% on these tests, and the AI Index 2026 reports later best scores between 66.3% and 74.5%. Each score describes that benchmark's tasks on one date and one version. It leaves out repeat reliability, cost, and your own tools and rules, so use it to pick candidates and then test your own agent.
Published September 30, 2026. Editorial.
Key takeaways
- GAIA (2023) has 466 questions, WebArena (2023) has 812 website tasks, tau-bench (2024) has 165 customer service tasks and OSWorld (2024) has 369 desktop tasks. Each checks the final result of the task.
- At publication the best model scored 15% on GAIA against 92% for people, 14.41% on WebArena against 78.24%, and 12.24% on OSWorld against 72.36%, by each benchmark's own paper.
- On tau-bench the best model in the 2024 paper, gpt-4o, passed 61.2% of retail tasks on one try and fewer than 25% when the same task had to pass eight times in a row.
- The METR time horizon is the task length, in skilled human working time, at which a model succeeds half the time. METR's paper gives a doubling time of 207 days and its January 2026 update gives 196 days.
- Agent benchmarks get revised: OSWorld in July 2025, the METR task set in January 2026 and tau2-bench in July 2026. A score is comparable only with scores from the same version.
In November 2023 the authors of GAIA put a set of real-world questions to people and to GPT-4 equipped with plugins, which are add-on tools. People answered 92% correctly and the model answered 15% [1]. The questions, in the authors' words, were "conceptually simple for humans yet challenging for most advanced AIs" [1].
An agent is an AI model that works in steps: it calls a tool, such as a web browser or a database lookup, reads the result and decides what to do next. An agent benchmark is a fixed, public set of tasks that scores whether the whole job got done. Four names appear most often in announcements (GAIA, WebArena, tau-bench and OSWorld) along with one measure, the METR time horizon. This page explains each from its own paper, with the figures at publication and their dates. It belongs to our guide to AI benchmarks vs your own evals.
GAIA: questions that need tools and several steps
GAIA stands for General AI Assistants. Its paper, from authors at Meta, Hugging Face and other groups, appeared on arXiv on 21 November 2023 [1]. Stanford's AI Index 2026 dates the benchmark to May 2024 [9]; the arXiv date is the one on the paper itself.
The benchmark has 466 questions, each with a recorded answer. The authors published the questions and kept the answers to 300 of them private, to run a leaderboard, which is a public ranking table of scores [1]. A person needs between 6 minutes for the simplest questions and 17 minutes for the most complex ones [1].
At publication people scored 92% and GPT-4 with plugins scored 15% [1]. The AI Index reports that the best system reached 74.5% in September 2025, against the same 92% human figure [9].
WebArena: tasks on working copies of websites
WebArena, from Carnegie Mellon University, was published in July 2023. It gives an agent working copies of four kinds of website: online shopping, discussion forums, shared software development and content management [2]. The benchmark has 812 tasks, built from 241 templates [2].
Success is judged on the result. The paper describes its checks as "evaluating the functional correctness of task completions" [2]: whether the thing the task asked for now exists on the site, whichever route the agent took.
At publication the best GPT-4 agent completed 14.41% of tasks and people completed 78.24% [2]. The human figure comes from 170 tasks, one per sampled template, performed by five computer science graduate students. The AI Index reports 74.3% for the best agent in early 2026, "within 4 percentage points of the human baseline of 78.2%" [9].
A 2024 Princeton paper, "AI Agents That Matter", reports "pervasive shortcomings in the reproducibility of WebArena and HumanEval evaluations", meaning that published results were hard to repeat [10]. HumanEval is a coding test.
tau-bench: a customer service agent that has rules to follow
tau-bench was published in June 2024 by researchers at Sierra. It simulates a conversation between a customer, played by a language model, and an agent that has two things: tools for looking up and changing records, and a written policy it must follow [3]. There are two settings, a retail shop with 115 tasks and an airline with 50 [3].
A task is scored by the records. The benchmark "compares the database state at the end of a conversation with the annotated goal state" [3], so the agent passes when the orders and bookings end up exactly as they should.
On a single try, the best model in the paper, gpt-4o, passed 61.2% of retail tasks and 35.2% of airline tasks [3]. The paper then adds a measure called pass^k: the chance that all k repeated tries at the same task succeed. For gpt-4o in retail, pass^8 was below 25% [3]. An invented illustration shows why repeats matter: an agent that passes one try in 0.75 of cases passes three tries in a row in 0.75 x 0.75 x 0.75 = 0.421875 of cases, which is 42%. Our page on agent reliability across repeated runs covers the method.
The AI Index 2026 reports single-try scores for leading models between 62.9% and 70.2% [9]. The sources read for this page give no human score for tau-bench. The maintainers of the follow-up, tau2-bench, wrote in July 2026 that they had made over 75 task fixes and that results from before version 1.0.1 "are not comparable" with later ones [11].
OSWorld: tasks on a real computer desktop
OSWorld, published in April 2024, has 369 tasks that use real web and desktop programs, files on the computer, and work that uses several programs together [4]. The agent works inside a virtual machine, a computer simulated inside another computer, so that a mistake cannot harm a real one. Each task comes with its own starting state and its own checking program, and the paper limits every task to 15 steps and 30 minutes [4].
At publication people completed 72.36% of the tasks and the best model completed 12.24% [4].
The benchmark has changed since. On 28 July 2025 the project site announced OSWorld-Verified and asked teams to compare against the new results. It also says 8 tasks that depend on Google Drive may be left out, which gives 361 tasks (369 minus 8) [5]. The AI Index reports that the best score rose to 66.3%, "within 6 percentage points of human performance" [9], and the text does not say which version that is. A quoted OSWorld score needs its version and its task count beside it.
The METR time horizon: how long a task the agent can finish
METR measures something different from a percentage. It takes a set of tasks, records how long each one takes a skilled person, and then finds the task length at which an AI agent succeeds half the time. The paper's definition: "the time humans typically take to complete tasks that AI models can complete with 50% success rate" [6]. METR's page adds that the figure is "a measure of the difficulty of a task, rather than the time an AI spends to complete the task" [7].
The figures, each with its source:
- The first result. The March 2025 paper used 170 tasks. Its abstract gives Claude 3.7 Sonnet a 50% time horizon of 50 minutes, and METR's blog post of 19 March 2025 gives one hour for the same model. Both sources word the figure as approximate [6][12].
- The growth rate. The paper's abstract gives seven months as the approximate doubling time from 2019 onward. The paper's body gives 207 days, and METR's January 2026 update gives 196 days for the same trend [6][8].
- A stricter pass mark. At 80% success, the paper says horizons are 4 to 6 times shorter [6].
- The revision. On 29 January 2026 METR grew the set from 170 to 228 tasks and the number of tasks of 8 hours or more from 14 to 31. Estimates for individual models moved: two versions of GPT-4 fell by 35% and 57%, and GPT-5 and Opus 4.5 rose by 55% and 11%. Human times were measured for 5 of the 31 long tasks, and the rest are estimates [8].
- The limits METR states. Each task is run 6 times. METR says its tasks are much "cleaner" than real paid work. On 8 May 2026 the page added that "Measurements above 16 hrs are unreliable with our current task suite" [7].
METR's page gives an example of what a horizon means in practice. On tasks that take a skilled person between 90 minutes and 3 hours, a GPT-5 agent, listed there with a horizon of 2 hours and 17 minutes, succeeds every time on a third of the tasks, fails every time on another third, and gives mixed results on the rest. METR gives the horizon and the two shares as approximate figures [7].
The benchmarks side by side
| Benchmark | Tasks | How success is checked | Human score | Best model at publication |
|---|---|---|---|---|
| GAIA, November 2023 [1] | 466 questions | The answer is compared with a recorded answer | 92% | 15%, GPT-4 with plugins |
| WebArena, July 2023 [2] | 812 | The result on the website is correct | 78.24% | 14.41%, a GPT-4 agent |
| tau-bench, June 2024 [3] | 165 (115 retail, 50 airline) | The final records match the goal | None in the sources read | 61.2% retail and 35.2% airline, gpt-4o |
| OSWorld, April 2024 [4] | 369 | A checking program for each task | 72.36% | 12.24% |
| METR time horizon, March 2025 [6] | 170, then 228 | Success rate against human task time | Human time is the unit | 50 minutes at 50% success, Claude 3.7 Sonnet |
What these scores tell a buyer and what they leave out
The scores show that agents improved quickly. Between publication and the AI Index's latest figures, the best result went from 15% to 74.5% on GAIA, from 14.41% to 74.3% on WebArena and from 12.24% to 66.3% on OSWorld [9]. The same report summarises that agents still fail one attempt in three on structured benchmarks, a figure it gives as approximate [9].
Four things are missing from a single percentage.
Repeat reliability. Most scores are one try per task. The tau-bench result shows a model above 60% on one try and below 25% on eight tries in a row [3].
Cost. The Princeton paper states that "AI agent evaluations must be cost-controlled", because "simply calling the underlying model multiple times can increase accuracy" [10]. A score with no cost beside it may have been raised by retrying. Our page on cost, time and step limits for agents explains how to measure both.
The version. OSWorld, tau-bench and the METR measure were all revised, and each revision moved scores. Two scores from different versions cannot be ranked.
Your tools and your rules. No benchmark contains your systems, your policies or your customers. The Princeton authors make the wider point: "the benchmarking needs of model and downstream developers have been conflated" [10]. A downstream developer is a team that builds a product on top of a model, and its question is narrower than the model maker's. How to read a model vendor's benchmark claims lists the settings to ask about.
Which benchmark to read for your type of agent
No one benchmark is best for comparing agents. Read the one whose tasks most resemble your work:
- An agent that operates websites: WebArena.
- An agent that operates desktop programs and files: OSWorld.
- An agent that talks to customers, follows a policy and changes records: tau-bench.
- An assistant that answers questions by searching and using tools: GAIA.
- An agent that does long software tasks alone: the METR time horizon.
Then test your own agent on your own tasks. Our guide to AI agent evals is the method reference, how to evaluate an AI agent is the shorter business overview, and how to choose an AI model with your own evals covers choosing the model the agent runs on. For the question of how much to rely on an agent at all, read can you trust an AI agent.
How Reveneau applies this
Reveneau reads agent benchmarks to decide which models are worth trying, and decides with an eval suite written for the project. All of our code is written by AI, and every change must pass that suite before release. We write it from the client's specification before any code exists.
For an agent, the suite uses the client's own tools and rules. Each task states what the records must look like at the end, the same idea tau-bench uses, and each task is run several times so that a pass on one try is never taken as proof. We grade the judged checks with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader, which makes repeated runs practical.
Because AI does work that would otherwise need more engineers, a build takes a small team, and that saving goes into the client's price. Reveneau as a company takes responsibility for the whole project through production and after release. Our AI development service describes how an agent build is scoped.
Common questions
What is GAIA?
GAIA is a benchmark of 466 questions for general AI assistants, published on arXiv in November 2023 by authors from Meta, Hugging Face and other groups. The questions are simple for a person, who needs 6 to 17 minutes each, and need tools and several steps from an AI. At publication people scored 92% and GPT-4 with plugins scored 15%. The answers to 300 questions are kept private.
What is WebArena?
WebArena is a benchmark of 812 tasks carried out on working copies of four kinds of website: online shopping, discussion forums, shared software development and content management. Carnegie Mellon University published it in July 2023. A task passes when its result on the site is correct. The best GPT-4 agent in the paper completed 14.41% of tasks, against 78.24% for the people tested.
What is tau-bench?
tau-bench is a customer service benchmark published in June 2024 by researchers at Sierra. A language model plays the customer, and the agent has tools for changing records and a written policy to follow. It has 115 retail tasks and 50 airline tasks, and a task passes when the final database matches the goal. In the paper gpt-4o passed 61.2% of retail tasks on a single try.
What is OSWorld?
OSWorld is a benchmark of 369 tasks on a real computer desktop, published in April 2024. The agent uses web and desktop programs and files inside a virtual machine, and each task has its own checking program. At publication people completed 72.36% of tasks and the best model 12.24%. The project site announced a revised version, OSWorld-Verified, on 28 July 2025.
What is the METR time horizon?
The METR time horizon is the length of task, measured by how long it takes a skilled person, at which an AI agent succeeds a set share of the time, 50% in the headline figure. METR's March 2025 paper used 170 tasks and gave Claude 3.7 Sonnet a horizon of 50 minutes. The measure describes how hard a task the agent can finish, and METR says it is separate from how long the agent runs.
Which benchmark is best for comparing AI agents?
Each agent benchmark tests a different type of work, so none of them is best for comparing all AI agents. WebArena covers websites, OSWorld covers desktop programs, tau-bench covers customer service under a policy, GAIA covers questions that need tools, and the METR time horizon covers long software tasks. Read the one closest to your work, with its date and version, then test candidates on your own tasks.
What does a time horizon of two hours mean in practice?
A time horizon of two hours means the agent succeeds half the time on tasks that take a skilled person two hours. METR's page gives an example: a GPT-5 agent listed with a horizon of 2 hours and 17 minutes succeeds every time on a third of tasks that take a person 90 minutes to 3 hours, fails every time on another third, and gives mixed results on the rest. At 80% success, METR's paper says horizons are 4 to 6 times shorter.
Why do agent benchmark scores change when a benchmark is revised?
Agent benchmark scores change after a revision because the tasks or the grading changed. When METR grew its task set from 170 to 228 in January 2026, estimates for two GPT-4 versions fell by 35% and 57% while GPT-5 and Opus 4.5 rose by 55% and 11%. The tau2-bench maintainers say results from before version 1.0.1 are not comparable with later ones.
How close are AI agents to human scores on these benchmarks?
The AI Index 2026 reports agents approaching the human figures on several benchmarks. On GAIA the best system reached 74.5% in September 2025 against 92% for people. On WebArena the best agent reached 74.3% in early 2026 against a human figure of 78.2%. On OSWorld the best score rose to 66.3%, which the report places within 6 percentage points of human performance.
Do agent benchmark scores include the cost of running the agent?
Most agent benchmark scores leave cost out. The 2024 Princeton paper AI Agents That Matter states that agent evaluations must be cost-controlled, because calling the underlying model several times can raise accuracy. A score reached with many retries costs more per task than the same score reached on one try. Ask for the cost per task next to any accuracy figure.
Does a high agent benchmark score mean an agent will work in my business?
A high agent benchmark score shows the model can finish that benchmark's tasks on the date and version tested. Your business has its own tools, rules and customers, and none of them were in the test. Repeat reliability is also missing from a single score: in the tau-bench paper a model above 60% on one try passed fewer than 25% of retail tasks eight times in a row.
What should I do after reading an agent benchmark score?
After reading an agent benchmark score, write down the version, the task count and the date, then build a small test of your own. Choose real tasks from your business, state what the records must look like at the end of each one, and run each task several times. tau-bench scores agents the same way, by comparing the final database with a goal state.
References
- [1] Mialon and co-authors (Meta, Hugging Face and others), GAIA: a benchmark for General AI Assistants (arXiv, 21 November 2023): 466 questions with answers to 300 kept private, people at 92% against 15% for GPT-4 with plugins, and 6 to 17 minutes per question for a person.
- [2] Zhou and co-authors (Carnegie Mellon University), WebArena: A Realistic Web Environment for Building Autonomous Agents (arXiv, 25 July 2023): 812 tasks from 241 templates, four kinds of website, checks on functional correctness, the best GPT-4 agent at 14.41% against 78.24% for people, and how the human figure was measured.
- [3] Yao, Shinn, Razavi and Narasimhan (Sierra), tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (arXiv, 17 June 2024): the simulated user, 115 retail and 50 airline tasks, scoring by final database state, gpt-4o at 61.2% and 35.2%, and pass^8 below 25% in retail.
- [4] Xie and co-authors, OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments (arXiv, 11 April 2024): 369 tasks, the virtual machine, a starting state and checking program per task, the limit of 15 steps and 30 minutes, people at 72.36% and the best model at 12.24%.
- [5] OSWorld authors, OSWorld project site (update note dated 28 July 2025, read 30 September 2026): the move to OSWorld-Verified, the request to compare against new results, and the 361-task option without 8 Google Drive tasks.
- [6] Kwa, West, Becker and co-authors (METR), Measuring AI Ability to Complete Long Software Tasks (arXiv, 18 March 2025): the definition of the 50% time horizon, 170 tasks, Claude 3.7 Sonnet at 50 minutes, doubling every seven months (207 days in the body), and 80% horizons 4 to 6 times shorter.
- [7] METR, Task-Completion Time Horizons of Frontier AI Models (live page, last updated 8 May 2026): the plain definition, 6 runs per task, the GPT-5 example, tasks cleaner than real work, and the 16-hour limit.
- [8] METR, Time Horizon 1.1 (29 January 2026): 170 to 228 tasks, long tasks from 14 to 31, the 196-day doubling time, model estimates that moved by 11% to 57%, and human times measured for 5 of 31 long tasks.
- [9] Stanford HAI, AI Index Report 2026, Chapter 2: Technical Performance: GAIA at 74.5% in September 2025 and dated to May 2024, WebArena at 74.3% in early 2026, OSWorld at 66.3%, tau-bench single-try scores between 62.9% and 70.2%, and agents failing one attempt in three.
- [10] Kapoor, Stroebl, Siegel, Nadgir and Narayanan (Princeton University), AI Agents That Matter (arXiv, 1 July 2024): cost-controlled evaluation, repeated calls raising accuracy, reproducibility problems in WebArena evaluations, and the different needs of model developers and product developers.
- [11] Sierra Research, tau2-bench repository (README notice dated July 2026): over 75 task fixes, and results before version 1.0.1 not comparable with later ones. The maintainers' own account.
- [12] METR, Measuring AI Ability to Complete Long Tasks (blog post, 19 March 2025): Claude 3.7 Sonnet with a time horizon of one hour.
Related reading
Can you trust an AI agent with real work yet?
An agent that answers a question and an agent that takes an action are not the same risk. Here is how we decide where an agent is ready to act, and where it is not.
What founders get wrong about AI agents
An impressive agent demo and a reliable agent are two different things. Most of the work, and most of the risk, is in the final step before production, which nobody shows in the demo.
How autonomous are AI coding agents, really?
Engineers at the leading AI labs now say a model writes one hundred percent of their code. Read the quotes closely and a person is still involved in every one of them. Here is what the 2026 evidence supports, and what it does not.
More in The benchmarks vendors quote
SWE-bench explained: what a coding benchmark score means
SWE-bench is a public test of AI coding models. It gives a model a real bug report from an open software project and counts the task as solved when the project's tests pass after the model's change. The original set has 2,294 tasks from 12 Python projects. The score is the percentage of tasks solved, and it changes with the version of the test, the program wrapped around the model and the date. OpenAI stopped reporting the Verified version on 23 February 2026 after its own audit. A high score shows a model fixes public Python bugs that have tests. Your own requirements were outside the test.
MMLU, GPQA and ARC-AGI explained: knowledge and reasoning benchmarks
MMLU, GPQA, ARC-AGI and Humanity's Last Exam are public tests that model makers quote to show knowledge and reasoning. MMLU has 15,908 multiple-choice questions in 57 subjects, and the authors of a newer test write that models now score over 90% on it. GPQA has 448 graduate-level science questions on which experts scored 65%. ARC-AGI uses puzzles that ask the solver to work out a new rule. Humanity's Last Exam has 2,500 questions written by experts. Each score describes performance on that test's own questions. For a product that answers customers from your own documents, these scores say little, because none of them tests your documents, your search step or your customers' questions.