CoreWeave completed a full training run of DeepSeek-V3 671B in 2.02 minutes during the MLPerf Training v6.0 benchmark round, setting the fastest result on record for that workload across all cloud submissions. The run used 8,192 NVIDIA GB300 NVL72 GPUs connected with NVIDIA Spectrum-X Ethernet networking, the largest cluster of that hardware type ever submitted to MLPerf, per MLCommons, the independent body that administers the benchmark.
CoreWeave published the benchmark detail, noting three GB300 NVL72 configurations it submitted: 8,192 GPUs at 2.02 minutes, 4,096 GPUs at 3.09 minutes, and 2,048 GPUs at 5.54 minutes, all in the MXFP8 precision format. The company says it was the only participant in the v6.0 round to scale a GB300 NVL72 cluster beyond 2,048 GPUs on the DeepSeek-V3 workload, doubling to 4,096 and then to 8,192 while maintaining strong scaling efficiency across the increase.
MLPerf benchmarks carry weight in AI cloud vendor comparisons in part because the rules require submissions to run on production infrastructure, not purpose-built benchmark clusters. CoreWeave emphasizes this: the same hardware configuration its customers use daily ran the MLPerf jobs, making the result a snapshot of operational throughput rather than a curated one-off.
CoreWeave also submitted results for two dense models on NVIDIA Blackwell hardware. A run on a 4,096-GPU cluster finished the 405B parameter variant of Meta's Llama model family in 9.77 minutes, which CoreWeave says is 2.8x faster than its own v5.0 result on the same test. A 64-GPU cluster finished the 8B variant in 16.54 minutes, which CoreWeave says was 9.7% faster than the next submitter on an identical setup. The company attributes both gains to software-layer changes: CUDA Graphs for reduced CPU scheduling overhead, topology-aware workload placement in its SUNK orchestrator, and full-stack networking optimizations. The GB300 NVL72 configuration also reached near-parity with larger GB200 NVL72 deployments using 20% fewer GPUs, per CoreWeave's submission writeup.
The result lands in an ongoing competition among AI cloud providers to demonstrate training efficiency to foundation-model teams evaluating where to run large jobs. MLPerf figures carry weight because they are independently validated and directly comparable across vendors. CoreWeave, which is building its infrastructure primarily on NVIDIA Blackwell hardware, has a structural interest in benchmarking well on DeepSeek-V3: the open-weight model is one of the most heavily used frontier-scale training workloads in active deployment, and a top result on it is a credible signal to prospective customers evaluating frontier training capacity.













