Supercomputers/Analysis/The shadow list
Analysis 04
The shadow list
Two clusters that appear on no public ranking anywhere hold 49,152 accelerators between them. Bracketed against measured systems in this dataset using the same silicon, they come to 1.69 EFlop/s to 2.29 EFlop/s of HPL-equivalent compute, against 7.27 EFlop/s for every ranked system here combined. That is 23 to 31% of the visible world, from two machines belonging to one company, disclosed in a single engineering blog post.
This is the page we refused to publish three times, and here is what changed. An estimate is only worth publishing if every input is traceable. The estimator below uses no vendor peak figures and no outside numbers: each bound comes from dividing two sourced fields on one of our own system rows. Where that is not possible, this page publishes nothing for that system and says so by name. Every disclosed cluster here could be anchored.
Phase 1 seed dataset, compiled by hand. These rows were built from public operator, laboratory and vendor sources. A mechanical second-reader pass has since fetched every cited source: 32 of 52 systems have a readable citation that names them, and 4 are genuinely weakly sourced. Every claim carries its source and a confidence tier. Treat anything below verified as a lead, not a citation.
Read the shape, not the spikes. This is the installed base of the 52 systems in this dataset: a curated seed chosen for supply-chain diversity, not a census of the market. At that sample size one machine entering or leaving service moves a share line by tens of points, so year-to-year jumps are composition effects rather than market events. The multi-year trends are the part worth quoting; the individual steps usually have one system's name on them, and the table view under the chart will tell you which.
- Ranked systems
- 45
- Ranked total
- 7.27 EFlop/s
- Unranked systems
- 7
- Shadow estimate
- 1.69 EFlop/s to 2.29 EFlop/s
- As a share of ranked
- 23 to 31%
The measured basis
Every estimate on this page rests on this table and nothing else. These are the systems in the dataset that publish both an Rmax and a single accelerator count, so dividing one by the other gives a measured HPL rate per accelerator. No vendor peak figure appears anywhere in the calculation.
- NVIDIA
- AMD
- Intel
Table view: Measured HPL rate per accelerator
| Accelerator | Measured in | Year | Count | System Rmax | TFlop/s each |
|---|---|---|---|---|---|
| NVIDIA Tesla M2050 | Tianhe-1A | 2010 | 7,168 | 2.57 PFlop/s | 0.36 |
| NVIDIA Tesla K20X | Titan | 2012 | 18,688 | 17.59 PFlop/s | 0.94 |
| NVIDIA Tesla V100 | Sierra | 2018 | 17,280 | 94.64 PFlop/s | 5.48 |
| NVIDIA Tesla V100 | Summit | 2018 | 27,648 | 148.6 PFlop/s | 5.37 |
| AMD Instinct MI250X | Frontier | 2022 | 37,632 | 1.35 EFlop/s | 35.95 |
| AMD Instinct MI250X | LUMI | 2022 | 11,912 | 379.7 PFlop/s | 31.88 |
| NVIDIA A100 SXM4 64GB | Leonardo | 2022 | 13,824 | 241.2 PFlop/s | 17.45 |
| Intel Data Center GPU Max 1550 | Aurora | 2023 | 63,744 | 1.01 EFlop/s | 15.88 |
| NVIDIA GH200 Grace Hopper Superchip | Alps | 2024 | 10,752 | 434.9 PFlop/s | 40.45 |
How this was calculated, in full
Step 1. For every system publishing both an Rmax and exactly one accelerator count, divide the first by the second. Systems with more than one accelerator supplier are excluded, because the split of their Rmax between parts is not observable.
Step 2. For an unranked system with a disclosed accelerator count, bracket it with the lowest and highest measured rate for that same part. Where a part appears in several measured systems, the spread between them is the range. Frontier and LUMI both run MI250X and differ by about 12%; Summit and Sierra both run V100 and differ by about 2%. That spread is real and it is what a range should represent.
Step 3. Where a part has only one measured observation, a point estimate would be false precision, so it is widened by 15% either way, taken from the largest observed spread in step 2.
Step 4. Where no measured rate exists for the part and no stated proxy applies, publish nothing for that system.
The one substitution, stated rather than hidden. This dataset contains no measured HPL rate for the H100, because no system here publishes both an H100 count and an Rmax. It does contain a measured rate for the GH200 superchip, from Alps. GH200 is a Hopper GPU packaged with a Grace CPU. HPL runs on the GPU, so the measured GH200 rate is close to a measured H100 rate. That substitution is the single weakest link in this page, and if you disagree with it, the number to change is in the CSV.
What this estimator refuses to do. It does not multiply a vendor peak figure by an assumed efficiency. That method needs a peak number, an efficiency assumption, a utilisation assumption and a hardware mix assumption, and four uncertain inputs multiplied together produce a number with no meaningful confidence interval that nonetheless gets quoted as fact.
The estimate
| System | Operator | Accelerator | Count | HPL-equivalent | Anchored on |
|---|---|---|---|---|---|
| Meta GenAI cluster (RoCE) | Meta Platforms | NVIDIA H100 SXM5 | 24,576 | 844.9 PFlop/s to 1.14 EFlop/s | Alps |
| Meta GenAI cluster (InfiniBand) | Meta Platforms | NVIDIA H100 SXM5 | 24,576 | 844.9 PFlop/s to 1.14 EFlop/s | Alps |
The read
The headline number is large and it is also the least interesting thing on this page. Two machines, one operator, one blog post, and the result lands somewhere around a quarter of every ranked system in this dataset combined. Meta disclosed these because it wanted credit for open hardware, not because anyone required it. There is no reason to think they are unusual, and every reason to think the ones nobody blogged about are larger.
The more useful finding is the shape of the disclosure. Of the 7 unranked systems here, 5 disclose no hardware at all. Not a performance figure, not an accelerator count, nothing that could be turned into a number by any method. That is the normal case. The Meta clusters are in this analysis precisely because they are the exception, which means any shadow estimate anyone publishes, including this one, is built on the small minority of operators who chose to say something.
Two ways this number misleads, in opposite directions. It overstates, because these clusters would never actually achieve it: HPL needs tuning, a tightly coupled fabric and a reason to run it, and a training cluster has none of the three. It understates far more severely, because FP64 is not what the hardware was bought for. The same H100s deliver an order of magnitude more arithmetic in BF16 and more again in FP8, which is the precision that training actually uses. Read as "how much of a ranked machine is this hardware equivalent to", the figure is fair. Read as "how much compute is out there", it is far too low.
That is the deeper problem with the whole exercise, and it is not solvable by better estimation. The benchmark that defines the visible world measures a precision the invisible world does not use. Even a perfect shadow list would be denominated in the wrong unit. Why FP64 and FP8 diverged is worth reading alongside this page.
The measured-rate chart is worth a second look on its own terms, independent of the shadow argument. It is a clean generational curve built from nothing but sourced fields on system rows, and it shows something the marketing numbers do not: a factor of roughly forty in real delivered HPL throughput per accelerator between 2010 and 2024, and a visible spread between two systems running identical silicon.