Compute·Atlas

Supercomputers/Analysis/The shadow list

Analysis 04

The shadow list

Two clusters that appear on no public ranking anywhere hold 49,152 accelerators between them. Bracketed against measured systems in this dataset using the same silicon, they come to 1.69 EFlop/s to 2.29 EFlop/s of HPL-equivalent compute, against 7.27 EFlop/s for every ranked system here combined. That is 23 to 31% of the visible world, from two machines belonging to one company, disclosed in a single engineering blog post.

This is the page we refused to publish three times, and here is what changed. An estimate is only worth publishing if every input is traceable. The estimator below uses no vendor peak figures and no outside numbers: each bound comes from dividing two sourced fields on one of our own system rows. Where that is not possible, this page publishes nothing for that system and says so by name. Every disclosed cluster here could be anchored.

Phase 1 seed dataset, compiled by hand. These rows were built from public operator, laboratory and vendor sources. A mechanical second-reader pass has since fetched every cited source: 32 of 52 systems have a readable citation that names them, and 4 are genuinely weakly sourced. Every claim carries its source and a confidence tier. Treat anything below verified as a lead, not a citation.

Read the shape, not the spikes. This is the installed base of the 52 systems in this dataset: a curated seed chosen for supply-chain diversity, not a census of the market. At that sample size one machine entering or leaving service moves a share line by tens of points, so year-to-year jumps are composition effects rather than market events. The multi-year trends are the part worth quoting; the individual steps usually have one system's name on them, and the table view under the chart will tell you which.

Ranked systems
45
measured performance
Ranked total
7.27 EFlop/s
Unranked systems
7
5 disclose no hardware at all
Shadow estimate
1.69 EFlop/s to 2.29 EFlop/s
estimator anchor-bracket-v1
As a share of ranked
23 to 31%

The measured basis

Every estimate on this page rests on this table and nothing else. These are the systems in the dataset that publish both an Rmax and a single accelerator count, so dividing one by the other gives a measured HPL rate per accelerator. No vendor peak figure appears anywhere in the calculation.

  • NVIDIA
  • AMD
  • Intel
Measured HPL rate per accelerator Dot plot of measured HPL performance per accelerator against the year the system entered service, derived only from systems publishing both an Rmax and an accelerator count. 013253850201020132016201920222025 Tesla M2050 NVIDIA 0.3579799107142857 TFlop/s per accelerator, first light 2010 Tesla K20X NVIDIA 0.9412457191780822 TFlop/s per accelerator, first light 2012 Tesla V100 NVIDIA 5.476851851851852 TFlop/s per accelerator, first light 2018 Tesla V100 NVIDIA 5.374710648148148 TFlop/s per accelerator, first light 2018 Instinct MI250X AMD 35.953443877551024 TFlop/s per accelerator, first light 2022 Instinct MI250X AMD 31.875419744795163 TFlop/s per accelerator, first light 2022 A100 SXM4 64GB NVIDIA 17.447916666666668 TFlop/s per accelerator, first light 2022 Data Center GPU Max 1550 Intel 15.876004016064257 TFlop/s per accelerator, first light 2023 GH200 Grace Hopper Superchip NVIDIA 40.44828869047619 TFlop/s per accelerator, first light 2024 GH200 Grace Hopper SuperchipInstinct MI250XInstinct MI250X
Measured HPL throughput per accelerator, by the year the system entered service. Where the same part appears in two systems the gap between them is real measurement spread, and it becomes the range in the estimate below.
Table view: Measured HPL rate per accelerator
AcceleratorMeasured inYearCountSystem RmaxTFlop/s each
NVIDIA Tesla M2050Tianhe-1A20107,1682.57 PFlop/s0.36
NVIDIA Tesla K20XTitan201218,68817.59 PFlop/s0.94
NVIDIA Tesla V100Sierra201817,28094.64 PFlop/s5.48
NVIDIA Tesla V100Summit201827,648148.6 PFlop/s5.37
AMD Instinct MI250XFrontier202237,6321.35 EFlop/s35.95
AMD Instinct MI250XLUMI202211,912379.7 PFlop/s31.88
NVIDIA A100 SXM4 64GBLeonardo202213,824241.2 PFlop/s17.45
Intel Data Center GPU Max 1550Aurora202363,7441.01 EFlop/s15.88
NVIDIA GH200 Grace Hopper SuperchipAlps202410,752434.9 PFlop/s40.45
How this was calculated, in full

Step 1. For every system publishing both an Rmax and exactly one accelerator count, divide the first by the second. Systems with more than one accelerator supplier are excluded, because the split of their Rmax between parts is not observable.

Step 2. For an unranked system with a disclosed accelerator count, bracket it with the lowest and highest measured rate for that same part. Where a part appears in several measured systems, the spread between them is the range. Frontier and LUMI both run MI250X and differ by about 12%; Summit and Sierra both run V100 and differ by about 2%. That spread is real and it is what a range should represent.

Step 3. Where a part has only one measured observation, a point estimate would be false precision, so it is widened by 15% either way, taken from the largest observed spread in step 2.

Step 4. Where no measured rate exists for the part and no stated proxy applies, publish nothing for that system.

The one substitution, stated rather than hidden. This dataset contains no measured HPL rate for the H100, because no system here publishes both an H100 count and an Rmax. It does contain a measured rate for the GH200 superchip, from Alps. GH200 is a Hopper GPU packaged with a Grace CPU. HPL runs on the GPU, so the measured GH200 rate is close to a measured H100 rate. That substitution is the single weakest link in this page, and if you disagree with it, the number to change is in the CSV.

What this estimator refuses to do. It does not multiply a vendor peak figure by an assumed efficiency. That method needs a peak number, an efficiency assumption, a utilisation assumption and a hardware mix assumption, and four uncertain inputs multiplied together produce a number with no meaningful confidence interval that nonetheless gets quoted as fact.

The estimate

SystemOperatorAcceleratorCount HPL-equivalentAnchored on
Meta GenAI cluster (RoCE) Meta Platforms NVIDIA H100 SXM5 24,576 844.9 PFlop/s to 1.14 EFlop/s Alps
Meta GenAI cluster (InfiniBand) Meta Platforms NVIDIA H100 SXM5 24,576 844.9 PFlop/s to 1.14 EFlop/s Alps

The read

The headline number is large and it is also the least interesting thing on this page. Two machines, one operator, one blog post, and the result lands somewhere around a quarter of every ranked system in this dataset combined. Meta disclosed these because it wanted credit for open hardware, not because anyone required it. There is no reason to think they are unusual, and every reason to think the ones nobody blogged about are larger.

The more useful finding is the shape of the disclosure. Of the 7 unranked systems here, 5 disclose no hardware at all. Not a performance figure, not an accelerator count, nothing that could be turned into a number by any method. That is the normal case. The Meta clusters are in this analysis precisely because they are the exception, which means any shadow estimate anyone publishes, including this one, is built on the small minority of operators who chose to say something.

Two ways this number misleads, in opposite directions. It overstates, because these clusters would never actually achieve it: HPL needs tuning, a tightly coupled fabric and a reason to run it, and a training cluster has none of the three. It understates far more severely, because FP64 is not what the hardware was bought for. The same H100s deliver an order of magnitude more arithmetic in BF16 and more again in FP8, which is the precision that training actually uses. Read as "how much of a ranked machine is this hardware equivalent to", the figure is fair. Read as "how much compute is out there", it is far too low.

That is the deeper problem with the whole exercise, and it is not solvable by better estimation. The benchmark that defines the visible world measures a precision the invisible world does not use. Even a perfect shadow list would be denominated in the wrong unit. Why FP64 and FP8 diverged is worth reading alongside this page.

The measured-rate chart is worth a second look on its own terms, independent of the shadow argument. It is a clean generational curve built from nothing but sourced fields on system rows, and it shows something the marketing numbers do not: a factor of roughly forty in real delivered HPL throughput per accelerator between 2010 and 2024, and a visible spread between two systems running identical silicon.