The benchmarks vendors quote

MMLU, GPQA and ARC-AGI explained: knowledge and reasoning benchmarks

MMLU, GPQA, ARC-AGI and Humanity's Last Exam are public tests that model makers quote to show knowledge and reasoning. MMLU has 15,908 multiple-choice questions in 57 subjects, and the authors of a newer test write that models now score over 90% on it. GPQA has 448 graduate-level science questions on which experts scored 65%. ARC-AGI uses puzzles that ask the solver to work out a new rule. Humanity's Last Exam has 2,500 questions written by experts. Each score describes performance on that test's own questions. For a product that answers customers from your own documents, these scores say little, because none of them tests your documents, your search step or your customers' questions.

Published September 30, 2026. Editorial.

Key takeaways

  • MMLU, published in 2020, has 15,908 four-option questions in 57 subjects. A 2024 re-check of 5,700 of them estimated that 6.49% contain errors, and the Humanity's Last Exam paper says models now score over 90%.
  • GPQA has 448 multiple-choice science questions. Experts scored 65% and skilled people from outside the subject scored 34% with web access, by its 2023 paper. On the 198-question Diamond set one question is 0.505 percentage points.
  • ARC-AGI tests working out a new rule from puzzles. On ARC-AGI-2, launched 24 March 2025, public reasoning systems scored single digits at launch, and the AI Index reports a leader at 84.6% by February 2026.
  • Humanity's Last Exam has 2,500 questions by its paper and site, and 2,700 by the AI Index 2026. The AI Index reports that accuracy went from under 10% to 38.3% in a single year.
  • A knowledge benchmark asks what a model remembers. A product that answers from your documents has to find the right passage and stay within it, so test it with your own customers' questions.

When the MMLU test was published in September 2020, the largest GPT-3 model scored less than 20 percentage points above random guessing on average [1]. The authors of a newer test, first published in January 2025, write that large language models (LLMs) "now achieve over 90% accuracy on popular benchmarks like MMLU" [10]. Between those two papers the test went from hard to one that no longer separates strong models, and new tests were written to replace it.

Four of these tests appear again and again in model announcements: MMLU, GPQA, ARC-AGI and Humanity's Last Exam. This page gives what each one contains, how it is scored, what people score on it and what is known about its limits, from each benchmark's own paper or site. It belongs to our guide to AI benchmarks vs your own evals.

MMLU: 57 subjects, four answer options per question

A benchmark is a fixed, public set of questions that many AI models are scored on. MMLU stands for Measuring Massive Multitask Language Understanding, the title of the paper that introduced it on 7 September 2020 [1].

The paper describes 57 subjects, "including elementary mathematics, US history, computer science, law, and more". The authors collected 15,908 questions in total. The test set, the part used for the published score, has 14,079 of them, and each subject has at least 100 test questions [1].

Every question is multiple choice with four options, so a model that guesses scores 25%. The score is the share of questions answered correctly. The human figure in the paper comes from workers on Amazon Mechanical Turk, a website where people are paid to do small online tasks, with no special training, who scored 34.5% [1]. The paper also says one subject draws on freely available practice questions for a professional exam, which means the questions came from public material.

What is known about the limits of MMLU

Wrong questions. A June 2024 paper from the University of Edinburgh and others, titled "Are We Done with MMLU?", re-checked 5,700 questions by hand, 100 from each of the 57 subjects. The authors estimate that 6.49% of MMLU questions contain errors, and they found errors in 57% of the Virology questions they analysed [2]. The 6.49% is an estimate from that sample.

Design quality. BetterBench, a Stanford study from November 2024, scored 24 benchmarks against 46 good practices chosen by its authors. MMLU received the lowest score in the assessment, a weighted average of 5.5, while GPQA received 11.0. The paper gives 15 as the highest possible score per category [3]. That score rates how a benchmark was designed and documented.

Every strong model scores near the top. Stanford's AI Index 2026 reports that on MMLU-Pro, a harder version from 2024, the leading 15 models all score above 87% as of early 2026, and the gap from first to fifteenth is "just over 4 percentage points" [12]. Researchers call this saturation: scores so high that the test no longer separates models. Our page on benchmark saturation and Goodhart's law explains it, and the page on benchmark contamination covers the evidence that models have seen MMLU questions in the text they learned from.

GPQA: 448 graduate-level science questions

GPQA was published on 20 November 2023. Its paper describes "448 multiple-choice questions written by domain experts in biology, physics, and chemistry", and states that 25% accuracy is random chance [4].

The paper reports three human and model figures:

  • Experts who hold or are studying for a doctorate in the subject scored 65%. The figure rises to 74% when clear mistakes that the experts identified afterwards are discounted.
  • Skilled people from outside the subject scored 34%, although they spent over 30 minutes on average with unrestricted access to the web. The paper's title calls the benchmark "Google-Proof".
  • The strongest GPT-4 setup the authors tested scored 39% [4].

GPQA comes in three sizes: an extended set of 546 questions, the main set of 448 and a set called Diamond with 198 [4]. Check which one an announcement names. On the 198-question Diamond set, one question is worth 0.505 percentage points (1 divided by 198), so a lead of one point between two models is two questions.

The paper states a limit of its own human figures: the accuracies on the main and Diamond sets "are skewed by selection effects", because those sets were chosen using the same people's answers [4].

ARC-AGI: puzzles that ask for a rule the solver has not seen before

ARC-AGI starts from an argument. In a November 2019 paper, François Chollet argued that skill at one task is a poor measure of intelligence, because skill depends on prior knowledge and experience, and that unlimited training data can raise a system's skill to any level while hiding how well it handles new problems. The paper defines intelligence as "skill-acquisition efficiency", meaning how efficiently a system learns a new skill, and presents a benchmark built on that idea, the Abstraction and Reasoning Corpus [5].

ARC-AGI-1. The ARC Prize Foundation, which runs the benchmark, describes "800 puzzle-like tasks, designed as grid-based reasoning problems": 400 for training and 400 for public evaluation. Another 100 semi-private and 100 private tasks are kept back. By the foundation's account the benchmark "remained unsolved by AI systems" from 2019 until late 2024, when a system called o3-preview scored 75% with a low computing budget and 87% with a higher one [6].

ARC-AGI-2. The second version launched on 24 March 2025 with 1,000 training tasks and three evaluation sets of 120 tasks each: public, semi-private and private [7][8]. The foundation tested the tasks on over 400 members of the public, and says every task was solved by at least 2 people in under 2 attempts [7]. AI systems also get two attempts per task [8]. At launch the foundation wrote that "Pure LLMs score 0% on ARC-AGI-2" and that public reasoning systems reached single-digit percentages [8]. The AI Index reports that by February 2026 the leader on the ARC leaderboard, the public ranking table, was Gemini 3 Deep Think at 84.6% [12]. From ARC-AGI-2 onward the foundation reports cost next to every score [7].

ARC-AGI-3. Released on 25 March 2026, the third version replaces fixed puzzles with "hundreds of interactive environments and thousands of game-style levels", and gives the AI "no instructions, no rules, and no stated goals" [9]. The launch post gives no exact count of environments and no model score.

Humanity's Last Exam: 2,500 expert-written questions

The Humanity's Last Exam paper, from the Center for AI Safety and Scale AI, was first published on 24 January 2025. It describes "2,500 questions across dozens of subjects, including mathematics, humanities, and the natural sciences", in multiple-choice and short-answer form so that a program can grade them [10]. The project site says the set was finalised at 2,500 questions on 3 April 2025 and that contributors came from over 500 institutions across 50 countries [11]. The AI Index 2026 gives the size as 2,700 questions [12]. The paper and the site both say 2,500.

The authors keep a private test set, also called a held-out set: questions that are never published, used "to assess potential model overfitting" [10]. Overfitting here means a model tuned to the public questions.

Scores rose fast. The AI Index reports that "in a single year, accuracy went from under 10% to 38.3%" [12]. The project site says a high score "would not alone suggest autonomous research capabilities", and it lists a rolling version, HLE-Rolling, released on 8 October 2025 [11]. The sources read for this page give no human score for the exam.

The benchmarks side by side

Benchmark Year Size Format and what it tests Known limit
MMLU 2020 15,908 questions in 57 subjects Multiple choice, four options; school and professional knowledge An estimated 6.49% of questions contain errors [2]
GPQA 2023 448 questions (Diamond: 198) Multiple choice; graduate science One Diamond question is 0.505 points
ARC-AGI-1 2019 800 public tasks, 200 kept back Grid puzzles; working out a new rule 87% reached in late 2024 [6]
ARC-AGI-2 2025 1,000 training tasks, 360 for evaluation Grid puzzles in the same format Leader at 84.6% by February 2026 [12]
Humanity's Last Exam 2025 2,500 questions (AI Index: 2,700) Multiple choice and short answer; expert knowledge From under 10% to 38.3% in one year [12]

Which of these benchmarks measures reasoning

ARC-AGI is the one designed to test reasoning apart from stored knowledge. Its tasks are grid puzzles instead of exam questions, the foundation says it chose tasks that people find easy and AI finds hard [8], and it removed tasks that a program could solve by trying every possibility [7]. MMLU, GPQA and Humanity's Last Exam mix the two: a model can answer by recalling a fact, by working the answer out, or by both.

Our position is that no score on this page equals "reasoning ability". A benchmark measures performance on its own tasks, and Chollet's point applies to every one of them: enough exposure to similar tasks raises the score without raising the ability to handle a new problem [5].

Are multiple-choice benchmarks useful

Multiple-choice tests are cheap to grade, because a program compares one letter with the answer key, and the result is exact. They show large gaps well. The move from under 20 points above guessing in 2020 to over 90% is a real change in what models can do [1][10].

They are weak in three ways. Guessing earns 25%. Answer keys contain errors, as the MMLU re-check found [2]. And a customer never offers four options: a product receives an open question and has to write the answer. Use these scores to rule out a weak model. For choosing between leading models they say little, and the AI Index notes that narrow gaps move competition "toward cost, reliability, and domain-specific performance" [12].

What a knowledge score says about a product that answers from your documents

It says little. Take an assistant that answers customers from a company's own policy documents. Three things decide whether it works, and a knowledge benchmark tests none of them.

Where the answer comes from. MMLU asks what a model remembers. Your assistant has to answer from the documents you supply and say so when they hold no answer.

The search step. Most such products first look up the relevant passages and then ask the model to write from them, a design called retrieval-augmented generation, or RAG. If the search returns the wrong passage, the answer is wrong whatever the model's exam score. RAG evaluation explains how to test that step.

The costly failure. On an exam a wrong answer is a wrong letter. In a product it is a confident statement the model made up, sent to a customer.

OpenAI's own documentation separates "Industry benchmarks for comparing models in isolation, like MMLU" from "Specific tests you implement to measure your LLM application's performance" [13]. The second kind is what you need: real questions from your customers, with the correct answers taken from your documents. How to choose an AI model with your own evals gives the steps, and our post on evaluating an AI product in production covers what to check once the product is live.

Where Reveneau fits

At Reveneau a public benchmark score is a reason to try a model, and the decision to use it comes from the client's own eval suite. All of our code is written by AI, and every change must pass a large eval suite, written from the client's specification before the code, before it is released. For a product that answers from a client's documents, that suite holds the client's own questions and the answers their documents support, so the test matches the work.

We grade the judged checks in the suite with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. A faster run means a new model can be tried against the whole suite soon after it appears. Because AI does work that would otherwise need more engineers, a build takes a small team, and that saving goes into the client's price. Reveneau as a company takes responsibility for the whole project through production and after release. To talk about an eval suite for your product, contact us.

Common questions

What is MMLU?

MMLU is a multiple-choice test of general knowledge for AI models, published in September 2020. The name stands for Measuring Massive Multitask Language Understanding. Its paper describes 15,908 questions across 57 subjects, from elementary mathematics to law, with four answer options each, so guessing scores 25%. Untrained human workers scored 34.5% in the paper. Leading models now score over 90%, by the Humanity's Last Exam paper.

What is GPQA?

GPQA is a set of 448 multiple-choice questions in biology, physics and chemistry, written by experts and published in November 2023. By its paper, experts with or studying for a doctorate in the subject scored 65%, and skilled people from outside the subject scored 34% after spending over 30 minutes on average with web access. The strongest GPT-4 setup the authors tested scored 39% at publication.

What is ARC-AGI?

ARC-AGI is a series of puzzle benchmarks built on François Chollet's 2019 argument that intelligence is the efficiency of learning a new skill. Each task is a grid puzzle whose rule the solver has to work out. ARC-AGI-1 has 800 public tasks, ARC-AGI-2 launched on 24 March 2025, and ARC-AGI-3, released on 25 March 2026, uses interactive game-style environments with no instructions.

What is Humanity's Last Exam?

Humanity's Last Exam is a benchmark of expert-written questions from the Center for AI Safety and Scale AI, first published in January 2025. Its paper and project site give 2,500 questions in multiple-choice and short-answer form, while the AI Index 2026 gives 2,700. The authors built it because models passed 90% on MMLU. The AI Index reports accuracy rising from under 10% to 38.3% in a single year.

Are multiple-choice benchmarks useful for judging an AI model?

Multiple-choice benchmarks are useful for ruling out a weak model and for showing large changes over time. A program can grade them exactly and cheaply. Their limits are that guessing earns 25% on a four-option test, that answer keys contain mistakes (an estimated 6.49% of MMLU questions, by a 2024 re-check), and that a real product receives open questions with no options to pick from.

Which benchmark measures reasoning instead of memory?

ARC-AGI is the benchmark designed to measure reasoning apart from stored knowledge. Its tasks are grid puzzles instead of exam questions, and the ARC Prize Foundation says it removed tasks that a program could solve by trying every possibility. MMLU, GPQA and Humanity's Last Exam mix recall and reasoning in one score. No single score equals reasoning ability, because each benchmark measures performance on its own tasks.

What do people score on MMLU, GPQA and ARC-AGI?

Human scores differ by benchmark and by who was tested. On MMLU, untrained Mechanical Turk workers scored 34.5%. On GPQA, experts scored 65% and skilled people from outside the subject scored 34%. On ARC-AGI-2, the ARC Prize Foundation says every task was solved by at least 2 people in under 2 attempts, in a study with over 400 members of the public.

What is GPQA Diamond?

GPQA Diamond is the smallest of the three GPQA sets, with 198 questions. The paper also defines a main set of 448 and an extended set of 546. On Diamond one question is worth 0.505 percentage points, so a lead of one point between two models equals two questions. When an announcement quotes a GPQA score, check which of the three sets was used.

Is MMLU still useful for comparing models in 2026?

MMLU separates leading models poorly in 2026. The Humanity's Last Exam paper says models score over 90% on it, and the AI Index 2026 reports that on the harder MMLU-Pro the top 15 models all score above 87%, within just over 4 percentage points of each other. A 2024 re-check also estimated that 6.49% of MMLU questions contain errors. MMLU still shows whether a model is far behind.

How do ARC-AGI-1, ARC-AGI-2 and ARC-AGI-3 differ?

The three ARC-AGI versions differ in size and form. ARC-AGI-1 has 800 public grid puzzles, and a system scored 87% on it in late 2024, by the ARC Prize Foundation's account. ARC-AGI-2 keeps the puzzle format with 1,000 training tasks and 360 evaluation tasks, and reports cost with every score. ARC-AGI-3 replaces fixed puzzles with hundreds of interactive environments that give the AI no instructions.

Does a high MMLU or GPQA score mean a model will answer my customers correctly?

A high MMLU or GPQA score shows what a model remembers and can work out on exam questions. Your product has a different job: answer from your documents, find the right passage first, and say when the documents hold no answer. No knowledge benchmark tests those steps. OpenAI's documentation separates industry benchmarks such as MMLU from the tests a team writes for its own application.

What should I do after reading a model's knowledge benchmark scores?

After reading a model's knowledge benchmark scores, use them to make a short list and then test that list on your own material. Collect real questions from your customers, write down the correct answer your documents support for each, and run every candidate model on the same set. Compare quality, cost and response time on that set, because the top 15 models are within just over 4 points of each other on MMLU-Pro, by the AI Index 2026.

References