Skip to content

Agentic AI Benchmarks Explained: SWE-bench, GAIA, and Measuring AI Agents

Agentic AI benchmarks like SWE-bench, GAIA, and WebArena test multi-step task completion, distinct from static knowledge tests and chat-preference rankings.

SWE-bench Verified leaderboard showing percent-resolved scores by AI agent
SWE-bench Verified leaderboard · Credit: SWE-bench

Agentic AI benchmarks are evaluation suites that measure whether an AI agent can complete a multi-step task inside a real or simulated environment, rather than produce a correct answer to a single question. A static knowledge benchmark like MMLU presents one question and checks one answer against a fixed key. An agentic benchmark instead hands the system a goal, a set of tools, and an environment, then checks whether that environment ends up in the correct state after a sequence of actions. That is a different question than the one Chatbot Arena asks, since Chatbot Arena ranks models by which single chat response a human voter prefers.

Five benchmarks anchor this task completion shift: SWE-bench and SWE-bench Verified, which score real GitHub issue resolution; GAIA, which scores general-assistant reasoning and tool use; WebArena and OSWorld, which drop an agent into a live web or desktop environment; and tau-bench, which tests multi-turn tool use against a business's own policy rules.

SWE-bench and SWE-bench Verified: Resolving Real GitHub Issues

SWE-bench evaluates large language models on real-world software issues collected from GitHub, asking a model to generate a patch that resolves a described problem inside an actual codebase. The task format and its Docker-based evaluation harness are documented at www.swebench.com/SWE-bench, which also tracks results from open coding-agent scaffolds such as SWE-agent, distinct from everyday coding assistants like Cursor or Aider that a developer drives interactively, since a SWE-bench agent must complete the whole task on its own.

  1. The harness gives the agent a snapshot of a real repository at the commit right before an issue was fixed, plus the GitHub issue text describing the problem.
  2. The agent must locate the relevant files across the codebase and write a patch, often touching more than one file to resolve the bug.
  3. The evaluation harness applies the patch inside a Docker container and runs the project's own hidden test suite against it.
  4. The instance counts as resolved only if the specific tests tied to that issue now pass, a strict, execution-based pass-or-fail check rather than a human judgment call.

SWE-bench Verified narrows that pool to a human-filtered subset of 500 instances, built in collaboration with OpenAI, per www.swebench.com/verified.html. Human annotators reviewed each of the 500 instances to confirm the problem description reads clearly, the test patches are correct, and the task is genuinely solvable, addressing quality concerns found in parts of the larger, unfiltered SWE-bench dataset. It is a curated subset built for more reliable scoring, not a replacement for the full original benchmark, and coding-agent leaderboards typically report both figures side by side.

GAIA: General Assistant Tasks Beyond Coding

Among agentic AI benchmarks, the GAIA benchmark stands out for testing general AI assistants on real-world questions that require reasoning, web browsing, and tool use rather than software patching. The GAIA benchmark, introduced by researchers including Grégoire Mialon, Yann LeCun, and Thomas Wolf, comprises 466 questions in total, with 300 answers held back for a private leaderboard and the remainder public, per arxiv.org/abs/2311.12983. The paper's headline finding is stark: human respondents answered correctly 92 percent of the time on GAIA's questions, while GPT-4 equipped with plugins reached only 15 percent, a gap that illustrates how conceptually simple questions for a person can still defeat a system that has to plan and use tools reliably across several steps.

The GAIA benchmark organizes its questions into three difficulty levels, per arxiv.org/abs/2311.12983: Level 1 tasks are solvable by a strong standalone model without much tool use, Level 2 tasks require multiple steps of reasoning or a couple of tool calls chained together, and Level 3 tasks require a longer chain of tool orchestration and reasoning, closer to what a genuinely autonomous assistant would face on an open-ended real-world request. That level structure is what makes the GAIA benchmark useful as a task completion measure rather than a single-shot trivia test: it isolates how far an agent's planning and tool use can carry it before accuracy collapses.

BenchmarkTask typeEnvironmentPrimary metric
SWE-bench / SWE-bench VerifiedResolve a real GitHub issue with a code patchSnapshotted code repository run inside DockerPercentage of instances resolved
GAIAAnswer a real-world question requiring reasoning and tool useOpen web and available toolsExact-match accuracy against a reference answer
WebArenaComplete a task on a live websiteSelf-hosted, realistic web environmentEnd-to-end task success rate
OSWorldComplete a task on a real computerFull Ubuntu, Windows, or macOS desktopExecution-based success rate via custom evaluation functions
tau-benchResolve a customer request while following business rulesSimulated multi-turn conversation with API toolsPass at k across repeated trials

WebArena and OSWorld: Agents Acting Inside Real Environments

WebArena and OSWorld push agent evaluation further by dropping the agent into a live, stateful environment for a multi-step task instead of a static text prompt. WebArena, introduced by researchers including Shuyan Zhou and Frank F. Xu, is a realistic web environment built around fully functional websites in four domains: e-commerce, social forum discussions, collaborative software development, and content management, per arxiv.org/abs/2307.13854. An agent follows a natural-language instruction across that multi-step task, and its success rate is measured by whether the website's end state reflects the completed task, not by matching a string of expected text. The paper reported its best GPT-4-based agent at a 14.41 percent success rate against a 78.24 percent human baseline.

OSWorld extends the same idea to a full desktop: a real computer environment spanning Ubuntu, Windows, and macOS, covering 369 tasks that touch desktop software, file operations, and workflows spanning multiple applications, scored through 134 custom execution-based evaluation functions, per osworld-v1.xlang.ai. The benchmark comes from researchers at The University of Hong Kong, Salesforce Research, Carnegie Mellon University, and the University of Waterloo, whose reported baseline puts human success above 72 percent against roughly 12 percent for the best model tested at publication, a gap wider than WebArena's.

Environment-based evaluation differs from a single-turn benchmark in ways that go beyond raw difficulty.

  • Reset requirements. The environment must return to a known starting state before every trial, since a partially completed website or desktop carries state into the next attempt and would otherwise contaminate the result.
  • Non-reasoning failure modes. A task can fail for reasons unrelated to the agent's reasoning, such as a page layout changing or an unexpected dialog box appearing mid-task.
  • Automated state checking. Grading requires checking the environment's actual final state, such as a database row or a saved file, rather than a simple string match.
  • Heavier compute per trial. Running these benchmarks at scale costs more than a static question set, since each trial needs a live browser instance or a full virtual machine.

tau-bench and the Pass at K Reliability Problem

Roundup of agentic AI benchmarks SWE-bench, GAIA, WebArena, OSWorld, and tau-bench

Among the agentic AI benchmarks covered here, tau-bench tests an agent's ability to hold a multi-turn conversation with a simulated user while following a company's own policy rules and calling the right API tools. The benchmark evaluates tool-agent-user interaction in real-world domains such as retail, where a language agent equipped with API tools and a written policy document must serve a simulated customer end to end, per arxiv.org/abs/2406.12045. The paper's central contribution is the pass at k metric, which reruns the identical task k times and checks whether the agent succeeds consistently rather than just once. Its reported results found that even a strong agent such as GPT-4o resolved fewer than 50 percent of retail tasks on a single attempt, with the pass at 8 score falling below 25 percent, a sharp reminder that a single successful run says little about reliability.

  • Hidden luck. A single successful run can conceal a fortunate sequence of tool calls that would not necessarily repeat on the next attempt.
  • Policy compliance at scale. An agent that follows a business's rules correctly nine times out of ten is not production-safe if the tenth failure violates a rule the customer or the business relies on.
  • Reproducibility as a metric. Once an agent's actions can branch differently across identical-looking runs, reproducibility itself has to be measured rather than assumed.
  • Brittleness at higher k. A low pass at k score at higher values of k reveals fragility that a single pass at 1 number conceals entirely.

The pass at k family, reported as pass at 1, pass at 3, or pass at 8 depending on the paper, has become shared vocabulary for task completion reporting across SWE-bench, GAIA, and tau-bench, even though each benchmark applies it to a different underlying task. A model's pass at 1 score on SWE-bench and its pass at 8 score on tau-bench are not directly comparable numbers, since one measures single-attempt code-patch resolution and the other measures repeated-trial policy compliance, but both use the same statistical idea to express how much a result can be trusted.

Why Agent Evaluation Is Uniquely Hard

Agentic AI benchmarks are harder to build and to trust than static benchmarks for four structural reasons that recur across SWE-bench, GAIA, WebArena, OSWorld, and tau-bench.

  • Long horizons. An agent's task completion can require dozens of individual actions, and an early wrong step, such as opening the wrong file or misreading a policy clause, compounds into a failure that only surfaces several steps later, making it hard to isolate where the reasoning actually broke down.
  • Stateful environments. Unlike a fixed question-answer pair, the environment itself changes as the agent acts, so the evaluation harness must snapshot, reset, and isolate each trial to keep results comparable, which is exactly what SWE-bench's Docker containers and OSWorld's virtual machines are built to do.
  • Partial credit. A task that is mostly but not fully completed, such as an agent that resolves a GitHub issue's core bug but breaks an unrelated test, or one that books the correct flight but the wrong seat class, forces benchmark designers to choose between strict all-or-nothing scoring and a more forgiving partial-credit rubric. Most of the benchmarks covered here default to strict, execution-based pass or fail rather than a partial score.
  • Reproducibility. An agent that calls a live API, browses the open web, or starts from a slightly randomized environment state can behave differently across identical-looking runs, which is precisely why pass at k metrics exist and why WebArena and OSWorld self-host their environments rather than pointing agents at the live internet.

These four factors are why agent benchmark scores move more slowly and carry wider caveats than a single MMLU accuracy number. A static benchmark score either holds or it does not; an agent benchmark score has to account for the environment it ran in, how many trials it ran, and how strictly partial success was counted. That is also why labs such as Anthropic and startups such as Cognition Labs publish their own coding-agent scaffolds against SWE-bench rather than a single model's raw score, since the scaffold built around a model is what actually completes the task.

Further reading

References

Frequently Asked Questions

What is the difference between an agentic AI benchmark and a benchmark like MMLU?

MMLU and similar static benchmarks score a single answer to a single question against a fixed answer key. An agentic AI benchmark like SWE-bench or GAIA gives the system a multi-step goal, a set of tools, and an environment, then scores whether the environment ends up in the correct state after a sequence of actions, not whether one isolated response was correct.

Is Chatbot Arena an agentic AI benchmark?

No. Chatbot Arena ranks models by collecting human votes on which of two chat responses people prefer in a single conversational turn. It measures preference, not task completion, and does not place an agent inside an environment or check whether a goal was actually achieved.

What does SWE-bench Verified measure that the original SWE-bench does not guarantee?

SWE-bench Verified is a 500-instance subset that OpenAI and the SWE-bench team human-reviewed for clarity and solvability. This filtering addresses quality issues found in parts of the larger, unfiltered SWE-bench dataset, so a resolved-rate score on the Verified subset is harder to game with an unclear or unsolvable task.

Why do agent benchmarks report pass at k instead of a single accuracy score?

Pass at k reruns the same task k times and checks whether an agent succeeds consistently, not just once. A single successful attempt can hide a fragile sequence of tool calls that would not repeat, so reporting pass at higher k values, such as pass at 8, reveals reliability problems that a one-shot pass at 1 score conceals.

Share this guide

Julian Beaumont

Julian Beaumont covers artificial intelligence and large language models for techshooked, following the path from research paper to deployed feature. His standard is anti-hype: ask what a model actually does, what trained it, how it fails, and whether a benchmark measures what the announcement claims.