Skip to content

AI Benchmarks Explained: MMLU, HumanEval, and What They Measure

AI benchmarks compared: MMLU spans 57 subjects and 14,000 questions; HumanEval measures code generation via pass@k; MMLU-Pro and GPQA raise the difficulty ceiling.

An Anthropic table compares AI models across the MMLU and HumanEval benchmarks.
Credit: Anthropic

An AI benchmark is a standardized test that measures model performance on a fixed task set. AI benchmarks like MMLU and HumanEval translate model behavior into a single number that vendors quote in launch posts and that researchers cite on leaderboards. The number is useful, but it carries hidden assumptions about evaluation conditions, scoring rules, and the integrity of the test set itself. A reader who treats an AI benchmark score as a direct measure of intelligence is reading half the story. The other half lives in the methodology: which questions were asked, how answers were extracted, whether the test set leaked into training data, and which version of the benchmark produced the figure. The most-cited AI benchmarks each measure a specific slice of capability, and knowing what each one actually counts is the prerequisite for reading a leaderboard with any skepticism. This article walks the anatomy of MMLU, HumanEval, GPQA, MMLU-Pro, GSM8K, and SWE-bench, then closes on contamination and how to read a score without being misled.

What an AI Benchmark Actually Measures

An AI benchmark is a fixed test set with a scoring protocol that converts model outputs into a number so that different systems can be compared on the same scale. The original MMLU paper frames the goal directly: the authors propose a new test to measure a text model's multitask accuracy, and to attain high accuracy on it a model must possess extensive world knowledge and problem-solving ability (arXiv 2009.03300, Measuring Massive Multitask Language Understanding). That framing captures the standard pattern across the field: a frozen test set, a deterministic scoring rule, and a single accuracy figure that maps to a leaderboard row. The mechanics matter because every choice in the protocol shapes what the score means. Holding the test set constant is what makes scores across vendors comparable. Holding the scoring rule constant is what makes scores across releases comparable. Holding the prompt format constant is what makes few-shot and zero-shot numbers comparable to each other. When any of those slip, the figure becomes a number about that one run rather than a measurement of the model. The same protocol logic underpins broader evaluation suites such as HELM from Stanford and the Open LLM Leaderboard on Hugging Face, which aggregate results from harnesses originally built at EleutherAI. For the broader picture of where evaluation sits inside model development, see how machine learning models are trained and evaluated.

  • Test set: the frozen pool of questions or problems used during evaluation; identical across all models being compared.
  • Scoring protocol: the deterministic rule that turns model outputs into a number, such as exact-match accuracy or unit-test pass rate.
  • Evaluation harness: the code that runs the test set against a model and records results, often the same across vendors for a given benchmark.
  • Multitask accuracy: the fraction of questions answered correctly across a heterogeneous mix of subjects or tasks.
  • Reporting convention: the choices about few-shot examples, chain-of-thought prompting, and answer extraction that vendors disclose alongside the headline score.

MMLU: 57 Subjects, 14,000 Questions, One Score

Hugging Face Open LLM Leaderboard showing model scores across MMLU, ARC, HellaSwag, and TruthfulQA benchmarks
Open LLM Leaderboard · Credit: Hugging Face

The MMLU AI benchmark tests multitask language understanding across 57 subjects, from elementary mathematics and US history to computer science and law, using 14,000 multiple-choice questions that require both world knowledge and problem-solving ability. OpenAI's own description of the dataset for GPT-4 puts the figures in the same shape: the MMLU benchmark is a suite of 14,000 multiple-choice problems spanning 57 subjects (OpenAI, GPT-4 research). Each question presents a stem and four answer choices, and the model's job is to pick the correct letter. Scoring is exact-match accuracy over the full test set, which makes MMLU one of the simplest benchmarks to administer and one of the easiest to compare across vendors. The 57-subject mix is what gives the score its breadth. A model that scores well on MMLU has to handle quantitative reasoning, factual recall, and domain-specific terminology in proportions that no single-subject test would expose. OpenAI reports that GPT-4o mini scores 82% on MMLU, which it describes as a textual intelligence and reasoning benchmark (OpenAI, GPT-4o mini: Advancing Cost-Efficient Intelligence). The figure illustrates how far accuracy on the benchmark has climbed since the test was introduced. Vendors of frontier systems such as Claude and Gemini cite MMLU alongside companion suites like BIG-bench, ARC, HellaSwag, Winogrande, TruthfulQA, DROP, and AGIEval when positioning a release. As scores approached the ceiling on the original MMLU, the research community responded with MMLU-Pro, which expands the answer set from four to ten options and removes trivial questions to restore signal at the top. For why interpretability matters when reading any single score, see how explainable AI differs from black-box models.

  1. Subject mix: 57 tasks covering humanities, social sciences, STEM, and professional domains such as law and medicine.
  2. Question format: single-correct multiple-choice with four answer options per question.
  3. Test set size: roughly 14,000 questions across all subjects, drawn from real exams and curated sources.
  4. Scoring rule: exact-match accuracy on the full test set, reported as a single percentage.
  5. Evaluation modes: zero-shot, few-shot, and chain-of-thought variants, each producing different headline numbers for the same model.

HumanEval: Measuring Code Generation with pass@k

HumanEval is the AI benchmark for code generation maintained by OpenAI, presenting the model with a function signature and docstring and requiring it to produce a correct implementation that passes a hidden unit-test suite. OpenAI's own repository describes it precisely as an evaluation harness for the HumanEval problem-solving dataset from the paper "Evaluating Large Language Models Trained on Code", with samples stored as JSON Lines records that carry a task_id and a completion field (GitHub, openai/human-eval README). The scoring metric is pass@k, the probability that at least one of k generated candidates passes every test case for a given problem. Pass@1 measures whether the first attempt is correct; pass@10 lets the model sample ten candidates and counts the problem solved if any one of them passes. That choice matters because real developer workflows often involve inspecting several completions rather than accepting the first. HumanEval reports its results as percentages, and OpenAI cites GPT-4o mini scoring 87.2% on HumanEval, which it describes as measuring coding performance (OpenAI, GPT-4o mini: Advancing Cost-Efficient Intelligence). Like MMLU, the benchmark has limits that vendor scores rarely surface. The original HumanEval set covers Python functions with self-contained unit tests, not full repositories, multi-file refactors, or production code review. Two newer evaluations, MBPP and SWE-bench, extend code generation testing toward the parts of software work that HumanEval skips. The continued popularity of pass@k comes from its directness: a unit test either passes or fails, the metric is well-defined across labs, and the dataset is small enough to run cheaply.

  1. Problem format: a Python function signature plus docstring, with hidden unit tests as the ground truth.
  2. Sampling: the model produces k completions per problem, often at a fixed temperature to control diversity.
  3. Pass@k metric: the unbiased estimator of the probability that at least one of k samples passes every test for a problem.
  4. Reporting: headline figures are usually pass@1 and pass@10, with pass@100 reserved for research papers.
  5. Limits: single-function scope, no multi-file context, no production code review tasks.

Beyond the Basics: GPQA, MMLU-Pro, GSM8K, and SWE-bench

As frontier models saturated MMLU and HumanEval, the research community designed harder AI benchmarks to restore signal at the top of leaderboards. GPQA targets graduate-level expertise in biology, chemistry, and physics with questions validated by domain experts. MMLU-Pro keeps the multitask spirit of MMLU but raises the bar: the MMLU-Pro paper documents the format change from four to ten answer options and the removal of trivial or noisy questions, with the goal of pushing models toward deeper reasoning rather than surface pattern matching (arXiv 2406.01574, MMLU-Pro). GSM8K shifts ground to grade-school mathematics word problems that require multi-step arithmetic reasoning and free-form numeric answers, which makes it a cleaner test of chain-of-thought behavior than a multiple-choice format permits. Beside GSM8K, the MATH dataset stretches the same reasoning into competition-level problems, while MMMU adds image-conditioned questions and ARC-AGI tests abstract pattern induction. SWE-bench moves further still, asking a model to resolve real GitHub issues in real Python repositories, with success measured by whether the generated patch makes the repository's existing tests pass. OpenAI's overview of SWE-bench Verified, a human-validated subset created with the authors of the original benchmark, frames it as a more reliable measurement of model abilities to solve real-world software issues (OpenAI, Introducing SWE-bench Verified). The harder benchmarks share a design pattern: tighter answer formats, expert curation, and problems that punish memorization. None of them replaces MMLU or HumanEval in vendor reporting yet, but together they reshape what a leaderboard near the top can credibly say about a model. For a wider view of where benchmark limitations sit inside model risk, see AI risks including benchmark limitations, bias, and hallucination.

BenchmarkWhat it testsFormatScoring
MMLUMultitask world knowledge across 57 subjectsMultiple-choice, 4 optionsExact-match accuracy
MMLU-ProHarder multitask reasoningMultiple-choice, 10 optionsExact-match accuracy
HumanEvalPython function generationSignature plus docstring; hidden unit testspass@k
GPQAGraduate-level science questionsMultiple-choice, expert-validatedExact-match accuracy
GSM8KGrade-school math word problemsFree-form numeric answerFinal-answer accuracy
SWE-benchReal GitHub issue resolutionRepository patch generationTest-suite pass rate

Data Contamination: Why Benchmark Scores Can Mislead

Card showing Data Contamination and Benchmark Misleading: Test-set leakage and Near-duplicates in training data

An AI benchmark loses validity when examples from its test set appear in a model's training data, a problem called contamination that inflates reported scores without reflecting genuine generalization. The risk is structural rather than accidental: large language models are trained on web-scale text, and popular benchmarks like MMLU and HumanEval are widely discussed on GitHub, Hugging Face, and academic mirrors, so test-set leakage is the default unless training filters explicitly exclude it. Even when the test items themselves are scrubbed, near-duplicates can survive, and a model that has seen a paraphrase of an MMLU question is no longer being measured on generalization. A separate quality issue compounds the problem. The MMLU-Redux paper re-annotates a subset of MMLU and reports a different angle on the test set: the authors create a subset of 5,700 manually re-annotated questions across all 57 MMLU subjects, and they estimate that 6.49% of MMLU questions contain errors (arXiv 2406.04127, MMLU-Redux). The combination of leakage and annotation noise means a reported MMLU figure carries an irreducible band of uncertainty even before any methodology choices are layered on top. Live human-preference rankings on Chatbot Arena and LMArena, along with LLM-judge benchmarks such as MT-Bench and AlpacaEval, offer an orthogonal signal because their prompts are not a fixed published set, which limits the same leakage pathway. As headline accuracy on MMLU climbed toward the ceiling, contamination concerns intensified, which is part of why re-annotation efforts like MMLU-Redux exist. The canonical MMLU dataset is published on Hugging Face (Hugging Face, MMLU dataset).

  • Test-set leakage: exact or paraphrased benchmark items present in the training corpus, inflating measured accuracy.
  • Annotation noise: errors in the gold-standard answers themselves, which cap how high a faithful model can score.
  • Decontamination filters: training-time string-matching or hash filters that try to remove known benchmark items.
  • Held-out variants: private or refreshed test sets such as SWE-bench Verified that reduce leakage risk for new releases.
  • Reporting hygiene: vendors that disclose their decontamination steps make their leaderboard claims auditable; those that do not, do not.

How to Read a Benchmark Score Without Being Misled

A published AI benchmark score is an artifact of evaluation conditions as much as model capability, so reading it critically requires understanding what the number includes and what it omits. Three questions decide whether a score is meaningful in context. First, which benchmark version was used: original MMLU and MMLU-Pro produce different numbers on the same model because the harder benchmark expands answer choices from four to ten and removes trivial items, per the MMLU-Pro paper (arXiv 2406.01574, MMLU-Pro). Second, which evaluation conditions were applied: few-shot example counts, chain-of-thought prompting, answer-extraction rules, and decoding temperature all move the score, sometimes by several percentage points, even on identical weights. Third, what was the test set's integrity at the time of the run: a model trained after an open benchmark's release should be assumed to have seen some of it, and the size of that effect is rarely disclosed. The MMLU-Redux re-annotation also matters here, because a 6.49% question-error rate sets a soft ceiling on how faithfully any score reflects ability rather than agreement with imperfect keys (arXiv 2406.04127, MMLU-Redux). Reading defensively means treating the headline number as one input among several, alongside task-specific evaluations and qualitative checks. For the wider governance frame around how AI evaluation standards are set and audited, see algorithmic accountability and AI system evaluation standards.

  1. Identify the benchmark version: distinguish original MMLU from MMLU-Pro, HumanEval from MBPP, SWE-bench from SWE-bench Verified.
  2. Read the methodology footnote: few-shot count, chain-of-thought, answer-extraction rule, and decoding parameters all belong in any honest report.
  3. Check for contamination disclosure: a vendor that reports decontamination steps is auditable; silence on the topic is itself information.
  4. Cross-reference task-specific evaluations: use code-quality, retrieval, or reasoning subtests alongside the headline number.
  5. Treat margin gaps with skepticism: differences inside 1 to 2 percentage points often fall inside reporting noise on these benchmarks.

Putting MMLU and HumanEval in Context

MMLU and HumanEval remain the most-cited AI benchmarks in vendor model cards precisely because they are well-understood, reproducible, and broad enough to surface meaningful differences between models of different sizes. Both are also old enough that frontier-scale systems sit near the ceiling, which is why newer evaluations from GPQA to SWE-bench have entered the standard reporting bundle. A reader looking at a launch post that quotes only MMLU and HumanEval is seeing a partial picture: the original MMLU paper itself frames its purpose as measuring multitask accuracy and notes that high accuracy requires both world knowledge and problem-solving ability (arXiv 2009.03300, MMLU), and even a strong score there leaves agentic behavior, long-horizon coding, and graduate-level science untested. The practical takeaway is to read every AI benchmark score as a slice. MMLU measures breadth of knowledge, HumanEval measures one-shot Python function synthesis, GPQA measures expert discrimination, and SWE-bench measures whether a patch passes existing tests in a real repository. No one of them captures a model. Taken together with a reading of the evaluation conditions, they sketch a credible profile. For an adjacent application where accuracy benchmarks are subject to similar scrutiny in a different domain, see accuracy benchmarks applied to facial recognition systems.

  • MMLU: broad knowledge plus reasoning, multiple-choice, single percentage on the leaderboard.
  • HumanEval: single-function Python generation, pass@k metric, sensitive to k and temperature.
  • GPQA and MMLU-Pro: harder discriminators for frontier models where the older benchmarks have saturated.
  • SWE-bench and SWE-bench Verified: real repository patches that test more than function-level synthesis.
  • Reading the bundle: a model's profile emerges from several benchmarks plus disclosed evaluation conditions, never one number alone.

References

Frequently Asked Questions

What does MMLU actually test in an AI model?

MMLU tests a model's multitask text accuracy across 57 subjects with 14,000 multiple-choice questions. A model that scores 80% on MMLU answers four out of five questions correctly across subjects as varied as elementary mathematics, US history, computer science, and law, but that score does not map directly to performance on open-ended tasks or real-world reasoning chains.

What is pass@k in the HumanEval benchmark?

Pass@k is the probability that at least one correct solution appears among k generated candidates for a given coding problem. Rather than requiring the model to get it right on the first try, pass@k lets researchers estimate how reliably a model can solve a coding task given multiple attempts, which better reflects real developer workflows where several completions can be inspected.

Why do AI benchmark scores sometimes look inflated or inconsistent?

Benchmark scores can be inflated when test set examples appear in a model's training data, a problem called contamination. Scores also vary based on evaluation conditions such as whether the model is given few-shot examples, how answers are extracted, and which version of the benchmark is used, so direct comparisons between vendor-reported numbers are often unreliable without examining methodology.

What is MMLU-Pro and how does it differ from the original MMLU?

MMLU-Pro extends the original MMLU AI benchmark by expanding answer choices from four to ten options and removing trivial questions. The larger choice set makes random guessing far less rewarding and pushes models to demonstrate deeper reasoning rather than surface-level pattern matching. It was designed because frontier models were approaching ceiling performance on standard MMLU, reducing its ability to distinguish among the strongest systems.

What is GPQA and why is it considered harder than MMLU?

GPQA is a graduate-level AI benchmark of questions validated by domain experts in biology, chemistry, and physics. Unlike MMLU, which spans a broad set of undergraduate-level topics, GPQA targets narrow expert knowledge so that even a well-informed non-expert with internet access struggles to answer correctly, making it a meaningful signal for frontier-model discrimination.

Share this guide

Julian Beaumont

Julian Beaumont covers artificial intelligence and large language models for techshooked, following the path from research paper to deployed feature. His standard is anti-hype: ask what a model actually does, what trained it, how it fails, and whether a benchmark measures what the announcement claims.