AI datacenters/Benchmarks
AI datacenters
Benchmarks for AI datacenters
There is no TOP500 for AI datacenters. No one ranks them by independently measured results on a common workload. What exists is a set of narrower tools that each answer part of the question: audited training and inference benchmarks that are submitted by vendors and run on a system, not a campus; efficiency and reliability figures that operators publish for their own fleets or single runs; and two independent efforts that estimate facilities from permits and imagery or rate GPU clouds. The best numbers a reader can compare today are described below with what each leaves out.
Why a ranked list does not exist
The TOP500 works because every entrant runs one program on one kind of arithmetic. AI datacenters break each condition. The largest clusters are private and their owners do not submit anything. Labs run different models in different precisions, so a number measured on one workload says little about another. Vendors choose what to submit and when. And the headline figures operators do publish, peak FLOPS in low precision, describe what the hardware could do, not what a training job achieved.
The result is that the same cluster can be described by an accelerator count, a power figure, a peak-FLOPS figure and a benchmark time, and the four are measured by different people under different rules. This section keeps each with its source and basis instead of collapsing them into one score.
What is measured, by whom
| Effort | What it measures | Run by | Limits |
|---|---|---|---|
| MLPerf Training | Time to train a fixed model to a fixed quality on the submitted system. The latest round found was v6.0 (June 2026), with 95 systems from 24 organisations. | MLCommons, with a review committee that checks results and can audit suspected cherry-picking. | Submissions are voluntary and vendor-prepared; the committee picks the workloads; it measures a training job on a system, not a whole datacenter. |
| MLPerf Inference | Throughput and latency serving fixed models. v6.1 (September 2026) had a record 30 submitters and a 512-accelerator system. | MLCommons; random and nominated audits of closed-division results. | System scale only, vendor-submitted. |
| MLPerf Storage, MLPerf Power | How many accelerators a storage system can keep fed (with simulated accelerators), and energy used while running the benchmarks. | MLCommons. | Subsystem metrics. Power data is not attached to most results. |
| InferenceMAX / InferenceX | Nightly open inference runs across vendors' hardware: tokens per second per GPU and per user, cost per token, and tokens per provisioned megawatt. | SemiAnalysis, with vendor and cloud support. | Its power figures are design-power estimates, not measurements. Open code and logs, but vendor-relevant. |
| HPL-MxP and Green500 | Mixed-precision HPL, closer to AI arithmetic than FP64 HPL, and FLOPS per watt, both published with the TOP500. | The TOP500 project, from voluntary submissions. | The TOP500 June 2026 top ten holds one commercial machine (Microsoft Eagle); the biggest AI clusters do not submit. |
| Epoch AI data centers hub | Estimated compute and IT power for 93 AI datacenters (as of 2 October 2026), built from satellite imagery, permits and company statements. CC-BY. | Epoch AI, independent of operators. | Estimates with roughly 1.4x error bands on power; Epoch puts its coverage of global AI compute at 43%, with weak coverage of Microsoft, Amazon and China. It sorts sites by compute and power, which ranks estimates, not measurements. |
| ClusterMAX 3.0 | Tiered ratings of GPU cloud providers from hands-on tests (77 providers reviewed). | SemiAnalysis. | Rates clouds, not facilities. SemiAnalysis also sells consulting and subscriptions and the full report is paywalled. |
MLPerf Training records
Time to train a fixed model to a fixed quality on the submitted system, in rounds run by MLCommons. Each row is one published result, linked to its source, not a ranked table.
| Submitter | System | Workload | Round | Accelerators | Result | Published | Basis | Source |
|---|---|---|---|---|---|---|---|---|
| CoreWeave | CoreWeave_GB300_2048x4 (GB300 NVL72) | DeepSeek-V3 671B pretraining | v6.0 | 8,192 NVIDIA GB300 | 2.02 minutes | 2026-06-16 | Consortium audited | mlcommons.orgdeveloper.nvidia.com |
| Microsoft | Azure GB200 (128x GB200 NVL72) | Llama 3.1 405B pretraining | v6.0 | 8,192 NVIDIA GB200 | 7.07 minutes | 2026-06-16 | Consortium audited | github.comdeveloper.nvidia.com |
| CoreWeave | CoreWeave_GB300_1024x4 (1,024 nodes) | Llama 3.1 405B pretraining | v6.0 | 4,096 NVIDIA GB300 | 9.77 minutes | 2026-06-16 | Consortium audited | mlcommons.orgmlcommons.org |
| NVIDIA | Tyche-hsg (80x GB200 NVL72) | Llama 3.1 405B pretraining | v5.1 | 5,120 NVIDIA GB200 | 10 minutes | 2025-11-12 | Consortium audited | github.comdeveloper.nvidia.com |
| NVIDIA | Eos-dfw (1,024 nodes x 8 HGX H100) | Llama 3.1 405B pretraining | v5.0 | 8,192 NVIDIA H100 | 19.65 minutes | 2025-06-04 | Consortium audited | github.comgithub.com |
| CoreWeave | Carina (39x GB200 NVL72, 2,496 active GPUs) | Llama 3.1 405B pretraining | v5.0 | 2,496 NVIDIA GB200 | 27.33 minutes | 2025-06-04 | Consortium audited | github.comcoreweave.com |
| NVIDIA | Eos-dfw_n1452 (1,452 nodes x 8 H100) | GPT-3 175B pretraining | v4.0 | 11,616 NVIDIA H100 | 3.4 minutes | 2024-06-12 | Consortium audited | github.comdeveloper.nvidia.com |
MLPerf Inference records
Throughput and latency serving fixed models on the submitted system, in rounds run by MLCommons. Each row is one published result, linked to its source, not a ranked table.
| Submitter | System | Workload | Round | Accelerators | Result | Published | Basis | Source |
|---|---|---|---|---|---|---|---|---|
| Crusoe | Crusoe 64 nodes x 8 AMD Instinct MI355X, RoCE Ethernet | Largest system submitted (any benchmark) | v6.1 | 512 AMD Instinct MI355X | 512 accelerators | 2026-09-16 | Consortium audited | mlcommons.orgrocm.blogs.amd.com |
| Crusoe | Crusoe 64 nodes x 8 MI355X | gpt-oss-120b (offline) | v6.1 | 512 AMD Instinct MI355X | 5,749,440 tokens/s | 2026-09-16 | Vendor reported | rocm.blogs.amd.comstoragereview.com |
| Crusoe | Crusoe 64 nodes x 8 MI355X | DeepSeek-R1 (offline) | v6.1 | 512 AMD Instinct MI355X | 2,901,950 tokens/s | 2026-09-16 | Vendor reported | rocm.blogs.amd.comstoragereview.com |
| NVIDIA | 4x GB300 NVL72 (288 GPUs, TensorRT) | DeepSeek-R1 (offline) | v6.1 | 288 NVIDIA GB300 | 2,705,130 tokens/s | 2026-09-16 | Vendor reported | developer.nvidia.commlcommons.org |
| NVIDIA | Vera Rubin NVL72 (72x VR200-288GB), preview category | DeepSeek-R1 (offline, preview) | v6.1 | 72 NVIDIA Rubin VR200 | 1,183,327 tokens/s | 2026-09-16 | Vendor reported | developer.nvidia.commlcommons.org |
| NVIDIA | 4x GB300 NVL72 over Quantum-X800 InfiniBand (72 nodes) | Largest system submitted (any benchmark) | v6.0 | 288 NVIDIA GB300 | 288 accelerators | 2026-04-01 | Consortium audited | mlcommons.orgdeveloper.nvidia.com |
| NVIDIA | 4x GB300 NVL72 (288 Blackwell Ultra GPUs) | DeepSeek-R1 (offline) | v6.0 | 288 NVIDIA GB300 | 2,494,310 tokens/s | 2026-04-01 | Vendor reported | developer.nvidia.commlcommons.org |
| NVIDIA | 4x GB300 NVL72 (288 Blackwell Ultra GPUs) | DeepSeek-R1 (server) | v6.0 | 288 NVIDIA GB300 | 1,555,110 tokens/s | 2026-04-01 | Vendor reported | developer.nvidia.commlcommons.org |
| NVIDIA | GB300 NVL72 (72 Blackwell Ultra GPUs) | DeepSeek-R1 (offline, per GPU) | v5.1 | 72 NVIDIA GB300 | 5,842 tokens/s per GPU | 2025-09-09 | Vendor reported | developer.nvidia.commlcommons.org |
| NVIDIA | GB300 NVL72 (72 Blackwell Ultra GPUs) | DeepSeek-R1 (server, per GPU) | v5.1 | 72 NVIDIA GB300 | 2,907 tokens/s per GPU | 2025-09-09 | Vendor reported | developer.nvidia.commlcommons.org |
| CoreWeave | CoreWeave GB200 instance (2 Grace CPUs, 4 Blackwell GPUs) | Llama 3.1 405B | v5.0 | 4 NVIDIA GB200 | 800 tokens/s | 2025-04-02 | Vendor reported | coreweave.commlcommons.org |
| NVIDIA | DGX B200 (8x Blackwell) | Llama 2 70B (server) | v5.0 | 8 NVIDIA B200 | 98,443 tokens/s | 2025-04-02 | Vendor reported | developer.nvidia.commlcommons.org |
| NVIDIA | GB200 NVL72 (72 Blackwell GPUs) | Llama 2 70B (v4.1 run, unverified) | v5.0 | 72 NVIDIA GB200 | 869,203 tokens/s | 2025-04-02 | Vendor reported | developer.nvidia.commlcommons.org |
MLPerf Storage records
How many accelerators a storage system can keep fed on training data loading. Each row is one published result, linked to its source, not a ranked table.
| Submitter | System | Workload | Round | Accelerators | Result | Published | Basis | Source |
|---|---|---|---|---|---|---|---|---|
| Nebius | Nebius Object Storage (Enhanced class) | RetinaNet training (object storage) | v3.0 | 768 NVIDIA B200 (simulated) | 768 simulated accelerators | 2026-09-01 | Vendor reported | nebius.commlcommons.org |
| Pure Storage | Everpure FlashBlade//EXA, 30 data nodes | Checkpointing, 1.25T-parameter model (write) | v3.0 | 1,024 NVIDIA B200 (simulated) | 877.52 GiB/s | 2026-09-01 | Consortium audited | storagereview.commlcommons.org |
| Microsoft | Azure Managed Lustre, 4,096 TiB, 128 clients | Checkpointing (write) | v3.0 | - NVIDIA B200 (simulated) | 642.2 GiB/s | 2026-09-01 | Consortium audited | storagereview.commlcommons.org |
| Pure Storage | Everpure FlashBlade//EXA, 51 hosts | KV cache (read) | v3.0 | - | 1,623.4 GiB/s | 2026-09-01 | Consortium audited | storagereview.commlcommons.org |
MLPerf Power records
Energy used while running the performance benchmarks. Each row is one published result, linked to its source, not a ranked table.
| Submitter | System | Workload | Round | Accelerators | Result | Published | Basis | Source |
|---|---|---|---|---|---|---|---|---|
| MLCommons | All on-premises v3.0 checkpoint-write submissions | Checkpoint write, on-premises submissions (median GB/s per watt) | Storage v3.0 | - | 14 GB/s per watt | 2026-09-01 | Consortium audited | mlcommons.org |
| MLCommons | All on-premises v3.0 UNet3D-read submissions | UNet3D read, on-premises submissions (median GB/s per watt) | Storage v3.0 | - | 34 GB/s per watt | 2026-09-01 | Consortium audited | mlcommons.org |
InferenceMAX records
An open, continuously run inference benchmark across vendors' hardware. Each row is one published result, linked to its source, not a ranked table.
| Submitter | System | Workload | Round | Accelerators | Result | Published | Basis | Source |
|---|---|---|---|---|---|---|---|---|
| SemiAnalysis | AMD MI300X single node (MX4 weights) | gpt-oss 120B, 1K in / 8K out, 90 tok/s/user | v1 (2025-10) | - AMD Instinct MI300X | 750,000 tokens/s per provisioned MW | 2025-10-09 | Independent | inferencex.semianalysis.comnewsletter.semianalysis.com |
| SemiAnalysis | AMD MI355X single node (MX4 weights) | gpt-oss 120B, 1K in / 8K out, 90 tok/s/user | v1 (2025-10) | - AMD Instinct MI355X | 2,550,000 tokens/s per provisioned MW | 2025-10-09 | Independent | inferencex.semianalysis.comnewsletter.semianalysis.com |
| SemiAnalysis | NVIDIA HGX H100 single node | gpt-oss 120B (interactivity level not stated for this pair) | v1 (2025-10) | - NVIDIA H100 | 900,000 tokens/s per provisioned MW | 2025-10-09 | Independent | inferencex.semianalysis.comnewsletter.semianalysis.com |
| SemiAnalysis | NVIDIA HGX B200 single node | gpt-oss 120B, FP4 weights (interactivity level not stated for this pair) | v1 (2025-10) | - NVIDIA B200 | 2,800,000 tokens/s per provisioned MW | 2025-10-09 | Independent | inferencex.semianalysis.comnewsletter.semianalysis.com |
HPL-MxP records
The mixed-precision variant of the HPL benchmark that the TOP500 project publishes, closer to AI arithmetic than FP64 HPL. Each row is one published result, linked to its source, not a ranked table.
| Submitter | System | Workload | Round | Accelerators | Result | Published | Basis | Source |
|---|---|---|---|---|---|---|---|---|
| Lawrence Livermore National Laboratory | El Capitan | HPL-MxP mixed-precision (rank 1) | 2026-06 | - AMD Instinct MI300A | 16.7 EFlop/s | 2026-06 | Independent | top500.orgtop500.org |
| Argonne National Laboratory | Aurora | HPL-MxP mixed-precision (rank 2) | 2026-06 | - Intel Data Center GPU Max | 11.6 EFlop/s | 2026-06 | Independent | top500.orgtop500.org |
| Oak Ridge National Laboratory | Frontier | HPL-MxP mixed-precision (rank 3) | 2026-06 | - AMD Instinct MI250X | 11.4 EFlop/s | 2026-06 | Independent | top500.orgtop500.org |
| National Supercomputing Centre in Shenzhen | LineShine | HPL-MxP mixed-precision (rank 4) | 2026-06 | - LX2 CPU (no accelerator) | 7.92 EFlop/s | 2026-06 | Independent | top500.orgtop500.org |
| Lawrence Livermore National Laboratory | El Capitan | HPL-MxP mixed-precision (rank 1) | 2025-11 | - AMD Instinct MI300A | 16.68 EFlop/s | 2025-11 | Independent | hpl-mxp.orgtop500.org |
| Argonne National Laboratory | Aurora | HPL-MxP mixed-precision (rank 2) | 2025-11 | - Intel Data Center GPU Max | 11.64 EFlop/s | 2025-11 | Independent | hpl-mxp.orgtop500.org |
| Oak Ridge National Laboratory | Frontier | HPL-MxP mixed-precision (rank 3) | 2025-11 | - AMD Instinct MI250X | 11.39 EFlop/s | 2025-11 | Independent | hpl-mxp.orgtop500.org |
What operators publish about a running cluster
The richest primary source on a large training cluster is Meta's Llama 3 paper. It trained a 405-billion-parameter model on up to 16,384 H100 GPUs and reports a BF16 model FLOPs utilisation of 38 to 43 percent depending on configuration, effective training time above 90 percent, and 466 job interruptions in a 54-day snapshot (419 unexpected, about 78 percent of those traced to confirmed or suspected hardware faults). Other groups report the same kind of figure for single runs: PaLM (46.2 percent on 6,144 TPU v4 chips) and ByteDance's MegaScale (55.2 percent on 12,288 GPUs). Google defines a broader ML productivity goodput that folds in scheduling and runtime losses.
These are useful and are not a series: each is one run on one cluster, and utilisation depends on the model, precision and sequence length, so it cannot be compared across labs. Nobody publishes utilisation, goodput or failure rates as a standing, comparable figure for every cluster.
What operators publish about the building
Facility figures are mostly fleet averages. Google reports a 2025 fleet PUE of 1.09, Microsoft a PUE of 1.17 and WUE of 0.27 L/kWh for FY25, and Meta a PUE of 1.08 and WUE of 0.19 L/kWh for 2024, with its water figure based on withdrawal rather than consumption. Boundaries differ, so the three are not comparable, and per-site figures, accelerator counts and rack densities are mostly not disclosed.
That may change in Europe. A September 2026 Commission act sets up a labelling scheme for datacenter efficiency, with labels from August 2027, voluntary below 500 kW of IT demand. We have not confirmed its publication in the Official Journal.
What to compare today
- Accelerator count and IT power. Operator claims or Epoch estimates. Check which one a number is, and whether it is built, contracted or planned.
- Time to train. MLPerf, per workload, for a submitted system only.
- Tokens per second per GPU and per megawatt. InferenceX, with design-power assumptions.
- PUE and WUE. Operator averages with different definitions.
- Peak low-precision FLOPS. Sourceable from vendor sheets but not measured, and sparse and dense figures are easily confused.