Skip to content

NVIDIA Blackwell Leads First Agentic AI Infrastructure Benchmark with 20x Efficiency Gain

NVIDIA's Blackwell Ultra GB300 NVL72 tops AgentPerf, the first benchmark built for agentic AI workloads, running 20x more agents per megawatt than Hopper.

NVIDIA DGX Blackwell server rack cluster in dramatic dark lighting
Credit: NVIDIA

NVIDIA's Blackwell Ultra GPU platform leads the first industry benchmark designed specifically for agentic AI workloads, running 20 times more concurrent agents per megawatt than NVIDIA's own previous-generation Hopper hardware, according to results published June 12 by Artificial Analysis.

The benchmark, called AgentPerf, comes from Artificial Analysis, an independent AI infrastructure research group. It is the first benchmark built to measure agentic workloads rather than single-prompt inference. NVIDIA says a conventional AI benchmark measures one large language model (LLM) call, a single request-and-response sprint. An agent chains dozens to hundreds of LLM calls together, passing growing context between steps while invoking external tools such as code execution, database search, and web browsing at each handoff.

NVIDIA video thumbnail shows the Grace Blackwell DGX Spark supercomputer
NVIDIA DGX Spark | A Grace Blackwell AI Supercomputer on your desk. Video: NVIDIA via YouTube.

AgentPerf was designed around real coding-agent task trajectories drawn from public code repositories across 12 or more programming languages. It measures how many concurrent agentic tasks a platform sustains while meeting defined responsiveness and output-token-rate thresholds, the number that matters for an enterprise pricing a GPU deployment against a productivity target. The benchmark's first test model is DeepSeek V4 Pro, a large mixture-of-experts model representative of the frontier class powering current commercial agents.

On that workload, NVIDIA's GB300 NVL72 system, a rack-scale unit connecting 72 Blackwell Ultra GPUs, delivered the top result. The performance gap against the H200 NVL72 (Hopper) widens as concurrent agent sessions scale, a pattern NVIDIA attributes to overlapping communication and compute through CUDA kernels and to TensorRT LLM's separation of input processing from token generation. Three inference providers (Baseten, DeepInfra, and Together AI) already run production agentic workloads on Blackwell, the company said. Together AI serves Cursor, an agentic coding assistant; DeepInfra runs Pam.ai, an AI workforce platform for car dealerships.

The practical implication for infrastructure buyers is that the performance gap between agentic and conversational workloads is now measurable, not theoretical. A buyer selecting GPU infrastructure to run 500 concurrent coding agents faces a different procurement decision than one optimizing for chatbot throughput, and AgentPerf gives that buyer a published, methodology-disclosed number. These dynamics are reshaping how AI accelerators are evaluated and purchased for inference at scale. NVIDIA noted that the Vera Rubin architecture, its next-generation successor to Blackwell, is now in full production. That is the first indication that the AgentPerf results represent a floor, not a ceiling, for the current hardware cycle.

Share this story

Julian Beaumont

Julian Beaumont covers artificial intelligence and large language models for techshooked, following the path from research paper to deployed feature. His standard is anti-hype: ask what a model actually does, what trained it, how it fails, and whether a benchmark measures what the announcement claims.