COMPUTE ATLAS Supercomputer supply-chain graph
1153 systems 482 sites Sourced data

Blog/2026-10-02

· Compute Atlas

Frontier: the first exaflop and its price

Frontier cleared 1.102 exaflops on HPL in May 2022 at 21.1 MW, meeting the 20 MW-per-exaflop goal as a ratio, not as a power budget. The route took a missed schedule and a hunt for bad hardware.


Frontier is the Oak Ridge National Laboratory machine that became the first system to pass one exaflop on the High-Performance Linpack (HPL) benchmark, in May 2022. It is an HPE Cray EX system built from AMD processors and AMD Instinct MI250X accelerators on HPE’s Slingshot-11 network. It held first place on the TOP500 from June 2022 to June 2024 and was remeasured at 1.353 exaflops in November 2024 and ranked second behind El Capitan in November 2025. The dataset holds it at 1,353 PFlop/s Rmax, 2,055.72 PFlop/s Rpeak and 24,607 kW.

The most useful thing to know about Frontier is that the headline achievement was real and the road to it was harder than the announcement suggested. The 20 megawatt target was met as an efficiency ratio, per exaflop, while the whole machine drew more than that. The first run that crossed the line finished at about 5 a.m. on May 27, 2022, after months of missed schedules. And the reliability work, which the lab itself has written up in unusual detail, continued for at least two years after acceptance.

How it came to exist

Frontier was the first system announced out of the second CORAL procurement (Collaboration of Oak Ridge, Argonne and Livermore). Cray filed an 8-K on May 7, 2019 describing two agreements with Oak Ridge under the Department of Energy’s CORAL-2 program: a non-recurring engineering and centers-of-excellence contract of about $102 million and a system delivery contract of about $548 million. The filing set out a phased path: an early access system in 2020, test and development systems in the fourth quarter of 2021, final Frontier phases in 2022, and a go/no-go decision point in the fourth quarter of 2020.

Press coverage described the award as “more than $600 million”. The Next Platform’s 2019 account split it into $500 million for the system and $100 million for engineering, including a $50 million storage system. The 8-K figures sum to roughly $650 million. We do not know which breakdown reflects final spending, and the lab does not publish a final cost. The DOE paid. HPE acquired Cray in 2019, per the dataset’s company rows, which is why the machine is an HPE product with Cray branding.

The 2019 plan promised more than 1.5 exaflops of peak performance and delivery in 2021. Delivery slipped. ORNL’s own account of the build says the target was July 2021, then August, then September, and that the first of 74 cabinets arrived on September 24, 2021, with the last on October 18. In spring 2021 the team found about 150 parts it needed that could not be sourced. A November 2021 target for finishing was also missed. By February 2022 the team told itself it would not make the TOP500 deadline on its current trajectory. A May 26 run reached 939.8 petaflops, and the run that broke one exaflop completed at about 5 a.m. on May 27. The result was published on May 30 in Hamburg.

Acceptance came later than the benchmark. OLCF’s June 2022 user Q&A lists four phases of acceptance testing: correctness tests, contractual application benchmarks, other applications agreed with DOE, and a sustained stability test. The Q&A also says the storage tier would likely be accepted after the rest of the system. ORNL’s account says Frontier opened for full user operations in April 2023.

The test-bed ladder

Frontier was built behind a sequence of smaller systems, and the dataset holds three of them.

  • Spock was the earliest: 36 nodes, each with one 64-core EPYC 7662 and four MI100 GPUs, on Slingshot-10. It was an early-access system for the Frontier programming environment. OLCF’s archived guide says it was decommissioned on March 15, 2023.
  • Crusher used the final node design: 192 nodes in two cabinets, each with a Trento EPYC and four MI250X (768 GPUs in total). ORNL described it in March 2022 as a 1.5-cabinet version of Frontier. Early application work on it included a roughly 15-fold speedup for the Cholla code against a 2019 Summit baseline. It was decommissioned on April 12, 2024.
  • Frontier TDS is one 128-node rack, and it is still on the TOP500. Al Geist’s August 2022 slides say it arrived early so staff could prepare codes while Frontier was delivered and stabilized, and that HPL was run on it first. On May 2, 2022 it reached 19.2 PFlop/s at an average of 309 kW, which put it first on the June 2022 Green500 at 62.68 GFlop/s per watt.

Cray’s 2019 filing had scheduled test and development systems for the fourth quarter of 2021; Geist says the TDS arrived early. The ladder moved software and tuning work ahead of the main machine, but it did not prevent the schedule slip.

What is actually inside

A Frontier node has one 64-core AMD “Optimized 3rd Gen EPYC” (Trento) CPU with 512 GB of DDR4 and four MI250X accelerators. Each MI250X holds two compute dies, which the software sees as separate GPUs, each with 110 compute units, 64 GB of HBM2e at 1.6 TB/s and 23.9 TFLOPS of FP64 according to the OLCF user guide. The dies on one package are linked at 200 GB/s over Infinity Fabric, and the CPU talks to the GPUs over Infinity Fabric too, not PCIe. Each GPU is tied to one of four Slingshot NICs, giving 100 GB/s of injection bandwidth per node.

The network is dragonfly. IEEE Spectrum describes custom 64-port switches and a maximum of three hops between nodes. Cabinets are HPE Cray EX with 128 nodes each, and Geist’s slides describe everything as water cooled, including memory modules and NICs. The same slides give warm-water cooling at 32 C and a data center PUE of 1.03.

Storage is the Orion Lustre file system on HPE ClusterStor E1000, with a flash tier and a disk tier, plus about 37 PB of node-local NVMe. The software environment is Cray OS on SUSE Linux with Slurm, Cray MPICH with GPU awareness, and the Cray, AMD and GCC compilers.

Node and cabinet counts differ between sources, covered below. Core counts are not a clean way to count the machine: TOP500 has listed 8,730,112 cores (June 2022), 8,699,904 (May 2023) and 9,066,176 (November 2024), and these include accelerator cores.

How it performs and what it does

HPL measures dense FP64 linear algebra. Frontier’s progression is 1.102 exaflops (May 2022), 1.194 (May 2023), and 1.353 (November 2024), which is 65.8 percent of the 2,055.72 PFlop/s Rpeak in the dataset. Geist’s slides record the 1.102 run as using 9,248 nodes at 21.1 MW on average. The same slides record a mixed-precision HPL-AI run of 6.86 exaflops on May 16, 2022 on 9,248 nodes at about 25 MW, and note that at launch the facility had to absorb a 20 MW surge in under five seconds. The November 2024 list reports 11.4 exaflops on HPL-MxP (called HPL-AI in 2022), and 14.05 PFlop/s on HPCG. These are different tests: HPL-MxP uses lower-precision arithmetic to solve the same type of problem, and HPCG is a memory-bound test that stresses data movement rather than arithmetic.

On real work, ORNL names ExaSMR, EXAALT, Pele, WDMApp, WarpX, ExaSky, EQSIM, E3SM and CANDLE among the early user codes. The WarpX team won the 2022 ACM Gordon Bell Prize. In acceptance-era testing, the Parthenon-Hydro mini-application held about 92 percent weak-scaling efficiency on 9,216 nodes, a code-level figure and not a system benchmark.

The 20 MW goal against the 21 MW reality

Frontier’s reliability and power stories are linked, and the 20 megawatt goal is easy to misstate. The target traces to a 2008 DARPA report; Geist’s slides say studies then predicted 150 to 500 MW for an exaflop and vendors were given the goal of 20 MW. It is a ratio: 20 MW per exaflop. DOE’s November 2022 summary says Frontier operates within it.

By that ratio the machine passes. 21.1 MW for 1.102 exaflops is about 19.1 MW per exaflop. ORNL’s digital-twin paper reports 22.7 MW for 1.194 exaflops and 22.8 MW for 1.206 exaflops, and the dataset’s 24,607 kW for 1.353 exaflops is about 18.2. But the machine itself is not a 20 MW machine. Geist’s slide lists 29 MW as the maximum power, and the same paper models 28.2 MW at full utilization of 9,472 nodes. Geist’s slides also give 14.5 and about 15 MW per exaflop, which is consistent with dividing the 29 MW maximum by the 2 EF peak (our inference, the slide does not say). So there are three different “MW per exaflop” numbers for one machine, depending on whether you divide by measured HPL, peak, and which power you use. HPL, which keeps every GPU busy, draws less than the facility is built to carry.

Reliability and software problems that were public

ORNL has been unusually candid. In the final months before the benchmark, the team traced a stall to message processing: the higher a sender’s number, the longer a recipient took to handle it. They also saw a sawtooth in power draw, a spike and drop every four minutes that appeared only when all nodes ran together. According to ORNL, a debug option in a software library, enabled by default, forced sequential address lookups, and turning it off cleared the problem. ORNL’s account quotes Justin Whitt: “Maybe there really were fundamental limits to the technology that we just didn’t anticipate.”

After acceptance the work shifted to defective hardware. A paper by ORNL and HPE staff states that the MI250X is rated at 560 W, twice the V100 used in Summit, with more memory and a smaller process, and concludes that the new hardware fails more often than Summit’s components did. Its three strategies ran between September 2022 and April 2024. A LAMMPS study of 244 jobs tried capping GPU power at 500 W, raising HBM controller voltage by 100 mV and lowering memory frequency. In the last two experiments (50 one-hour jobs each at 4,096 nodes, 500 W cap, +100 mV), 17 of 50 jobs failed at the default 1,800 MHz memory clock and 11 of 50 at 1,200 MHz. The study found 19 cases of defective hardware. A HACC job-completion study from June 2023 to January 2024 recorded 20 power faults in its first phase of 57 jobs at 4,000 to 9,072 nodes; after a part replacement campaign (the paper does not name the part), power faults fell by 75 percent. Over a six-month window the weekly node screen ran 1.9 million node tests, and a constant handful of “bad actor” nodes kept reappearing.

Geist’s slides set a target mean time between application failures of six hours, and note the 2009 worry that failures “may happen faster than you can checkpoint a job”. We found no published measured MTBF figure for Frontier in any source we could fetch.

The supply chain

Almost everything sits with two companies. AMD supplies the CPU and the accelerator, and the MI250X is made on TSMC’s 6 nm process, per IEEE Spectrum, so TSMC sits one step beneath. HPE, which owns Cray, supplies the cabinets, the Slingshot-11 network, the operating system and the ClusterStor-based Orion storage. The dataset records the same split in its component rows: 37,632 MI250X and 9,408 Trento CPUs from AMD, and integrator, interconnect, cooling, storage and OS from HPE. Geist’s slides say the same EX235a node design is used in LUMI and Adastra.

What does not add up or is not public

  • Node count. ORNL’s 2022 pages and the HPL-era papers say 9,408 nodes in 74 cabinets. OLCF’s current user guide says 77 cabinets and 9,856 nodes. A Wikipedia entry gives 9,472 CPUs (74 cabinets of 128), and ORNL’s digital-twin paper models 9,472 nodes. No fetched page explains the differences; 9,856 is 77 cabinets of 128, and 9,472 is 74 of 128, while 9,408 is not a multiple of 128.
  • Price. About $650 million by the 8-K sums, “more than $600 million” in press, and no final figure.
  • Storage. The 2019 plan was more than 1 EB at up to 10 TB/s. Delivered figures vary by page: 679 PB usable (user guide), 700 PB disk plus 11 PB flash (OLCF page), 716 PB center-wide (Geist’s slides).
  • Power. Average HPL power, maximum power and modeled peak power are three different numbers, as above.
  • Failure rates. The lab publishes failure causes and counts, not a headline MTBF, and it does not name the part behind the power faults.

What this dataset says

This section describes the dataset as published on 2 October 2026. The dataset changes as rows are added and corrected.

The dataset holds Frontier as reported confidence, status upgraded, with 9,066,176 cores and 24,607 kW, which reflects the post-November 2024 remeasurement. Its component rows list 37,632 AMD MI250X, 9,408 AMD Trento CPUs and HPE as integrator, interconnect, cooling, storage and OS vendor, so in the supply concentration and accelerator share analyses it is a two-supplier machine. Its list appearances are thin, only the first-place June 2022 TOP500 entry and a Green500 rank of 8 in November 2023, even though it held first or second place on every list we read from 2022 to 2025; any list-history view built on these rows will undercount it. Frontier TDS (TOP500 rank 29 in June 2022, Green500 rank 1) shows how a one-rack system can beat the full machine on efficiency. Crusher and Spock carry no performance figures in the dataset. Read the 24.6 MW row on power and cooling against the 29 MW maximum, and place the 2019 to 2022 timeline on build pipeline.

Sources

deep-diveexascalehpcpower-and-coolingaccelerators