· Compute Atlas
Same hardware, a different machine
El Capitan and xAI's Colossus 1 both run liquid-cooled accelerators from the same two silicon generations. One reports a benchmark result to the world twice a year. The other has never run one. That gap is the whole story.
Stand two machines side by side today and you would struggle to tell which one is the “supercomputer.” Both are rows of dark racks in a warehouse-scale building. Both are packed with the newest accelerators from NVIDIA or AMD. Both are cooled by liquid, not fans, because air stopped being enough for either of them years ago. Both draw tens of megawatts, enough to power a small city.
One of them is El Capitan, a US national laboratory system built to certify a nuclear arsenal by simulation. The other is Colossus 1, xAI’s Memphis cluster built to train large language models. Their accelerators come from the same two silicon generations. Their cooling is liquid on both sides. And yet one of them appears on a public ranking, has its performance independently measured twice a year, and publishes its power draw to the kilowatt. The other has never run the benchmark that would let it appear on that ranking at all, and its own operator is the only source for almost every fact about it.
That gap, not the hardware, is what actually separates an “AI data center” from a “supercomputer” in 2026. Here is what is the same, what is different, and how each one is actually built.
What is genuinely the same
The chips. 74% of the accelerated systems in this dataset use NVIDIA silicon, and that is true whether the system runs physics or a language model. A frontier AI cluster and a frontier HPC system are very often bidding for the same wafer allocation at the same foundry in the same quarter. AMD supplies the rest of the running installed base in roughly similar proportion on both sides. There is no separate chip industry for “AI” and “HPC”: there is one leading-edge accelerator supply, and both buyers are at the front of the queue.
The cooling. 84% of running systems in this dataset with a recorded cooling design use direct liquid or immersion cooling, and that share is not split by workload. El Capitan uses HPE’s warm-water direct liquid cooling. Colossus 1 uses cold-plate direct liquid cooling from Supermicro. Different vendor, same physical principle: once a rack holds a few hundred kilowatts of accelerators, air cannot move enough heat out of the room, and every large buyer on either side of this comparison has already made the same engineering decision. See power and cooling for why the crossover point is fixed by physics rather than by taste.
The power problem. The single largest disclosed draw in this dataset is 42 megawatts, and the constraint that actually gates a next-generation build is increasingly the same one on both sides: not the chips, but grid interconnection and substation capacity. A GPU cluster and a physics cluster wait in the same interconnection queue.
The scale of the fabric. Both need thousands of accelerators to talk to each other fast enough that the network, not the arithmetic, becomes the bottleneck. Whether that fabric is called Slingshot, InfiniBand or Spectrum-X, the engineering problem, collapsing the number of hops between any two accelerators at tens of thousands of endpoints, is the same problem. See the interconnect.
What is actually different
What a number means. Scientific simulation needs FP64, full double-precision arithmetic, because a climate model or a weapons code compounds rounding error over millions of timesteps and a wrong answer is not an answer. Neural network training does not need that and does not want it: it runs happily in BF16 or FP8, formats that pack far more arithmetic into the same chip and the same watt. The same physical accelerator can differ by a factor of thirty in FLOP/s depending only on which format you ask it to run. This is the single most consequential difference between the two machine classes, and it is covered in full in AI clusters versus HPC systems.
How the machine talks to itself. HPC is dominated by frequent, small messages between a node and its immediate neighbours in a decomposed physical domain, so latency is what matters. Large-scale training is dominated by infrequent, enormous all-reduce operations that sum gradients across every worker after every batch, so bandwidth is what matters. That is the actual engineering reason rack-scale NVLink domains exist: keep the most intense traffic inside a tightly coupled cluster of accelerators, and send only the less intense traffic out over the general fabric.
What a failure costs. An HPC job that loses a rank is a lost run; the whole job typically has to restart from a checkpoint. A training job at the scale Colossus 1 or a comparable cluster runs is populated by tens of thousands of accelerators for weeks at a time, so hardware failure during the run is a near-certainty rather than a risk, and the software is built to drop a straggler and keep going. Different failure tolerance is not a preference; it is a consequence of the population and the duration.
Who is in the building. A national laboratory runs a queue of thousands of jobs from hundreds of researchers and optimises for fair throughput across all of them. A frontier training cluster is frequently built for, and occupied by, one job. The scheduling problem is closer to capacity planning than to queueing theory.
And the one that actually explains why you have not heard of most of these machines: disclosure. A national laboratory submits a benchmark result because public accountability for public money is part of the deal. A commercial operator gains nothing from disclosing a competitive asset and running the benchmark costs days of a cluster’s time that is worth a great deal of money. This dataset’s estimate is that unranked clusters, almost entirely AI training systems, add up to 62% to 71% of the combined measured performance of every ranked system it tracks, built without a single vendor-peak figure, only measured rates on the same chips elsewhere in the dataset. The ranked world you can look up is a national-laboratory world with a handful of voluntary commercial entries in it. It is not the frontier of what has actually been built.
The clearest single example
Tuolumne and El Capitan are worth naming together, because they remove hardware as a variable entirely. Same MI300A accelerator, same HPE Cray EX255a chassis, same Slingshot-11 fabric, built under the same CORAL-2 procurement, installed at the same laboratory. LLNL’s own material is explicit that Tuolumne runs open, unclassified science, while El Capitan runs stockpile stewardship. Identical machine, on paper. Different mission, different classification, different access. If two systems with the same bill of materials can differ that much, hardware alone was never going to be a reliable way to draw this line. See the showcase for what each mission on this site’s dataset is actually for.
How each one is actually built
Strip away the workload and the two are built from the same layers, in the same order, for the same physical reasons:
- The node. A server board holding one or more CPUs and, in almost every modern leadership system on either side, accelerators attached over a fast local link. See inside a node for how that pairing actually works and why HBM memory sits where it does.
- The rack. Nodes packed as densely as the cooling design allows. A GB200 NVL72 rack and an MI300A rack are solving the same density problem with different chips.
- The fabric. Thousands of nodes wired so that the traffic pattern the workload actually produces (nearest-neighbour for a physics grid, all-reduce for a training run) does not bottleneck. This is the layer where AI and HPC system design diverges hardest, and it is also where a fabric acquisition (NVIDIA buying Mellanox, HPE buying Cray) can reshape an entire market’s concentration in a single year; see supply-chain concentration.
- Cooling and power. Liquid to the chip, a substation sized for the load, a PUE the site actually measures. This layer does not care what the machine is for.
- The software stack. Where the two genuinely part ways again: a scheduler tuned for thousands of competing jobs on one side, a stack tuned to keep one enormous job alive on the other.
Why this distinction is worth keeping
Calling every large accelerated cluster a “supercomputer” erases a real difference: precision, communication pattern, failure model and who gets to know it exists all diverge in a way that matters if you are trying to measure this industry rather than just photograph it. But calling them fundamentally different kinds of machines erases something just as real: they are drawing from the same chip supply, hitting the same cooling physics, and increasingly built by the same integrators, on the same order of megawatts, for the same reason, because a decomposed physical problem and a decomposed matrix multiplication both turn out to need thousands of accelerators talking to each other quickly.
The honest description is the one this dataset is built to hold: same hardware category, same physical constraints, two different machines defined mostly by what runs on them and who is allowed to know about it.