· Compute Atlas
El Capitan: building the NNSA exascale system
El Capitan put CPU and GPU on one package, cost about $600 million, and runs classified weapons work. Its public numbers disagree on nodes, power and peak, and the lab does not reconcile them.
El Capitan is the National Nuclear Security Administration’s first exascale computer, installed at Lawrence Livermore National Laboratory (LLNL) and built by HPE around AMD’s MI300A. It was verified at 1.742 exaflops on the High Performance Linpack benchmark in November 2024, took first place on the TOP500, and held it until June 2026, when LineShine displaced it. The remeasured HPL result is 1.809 exaflops.
The thing most worth knowing is not the ranking. It is that El Capitan is the first machine at this scale built on an APU: a single package carrying 24 Zen 4 CPU cores, GPU chiplets and one pool of HBM3 memory. The two earlier US exascale machines, Frontier and Aurora, pair discrete CPUs and GPUs. The second thing worth knowing is that El Capitan is a classified system. Most of what the public knows about it comes from benchmarks and from a few unclassified jobs run before it moved behind the fence.
How it came to exist
The contract was announced on August 13, 2019: $600 million with Cray Inc. for a machine of more than 1.5 exaflops, based on the Shasta architecture and Slingshot interconnect, with delivery anticipated in late 2022 (DOE). It sat under CORAL-2, the second joint procurement by the Oak Ridge, Argonne and Livermore labs. Cray was acquired by HPE in September 2019, a few weeks after that announcement (see our post on how the fabric layer consolidated).
The silicon was settled seven months later. On March 4, 2020, The Next Platform reported AMD as sole supplier of CPUs and GPUs: Genoa Epyc processors with Zen 4 cores and Radeon Instinct accelerators linked over Infinity Fabric, a target above 2 exaflops peak, a contract limit under 40 megawatts, and a budget the publication put at $500 million. Intel, NVIDIA and IBM received no part of the award (The Next Platform, 2020). The discrete-GPU design in that report is not the machine that was built. The MI300A APU replaced it, and The Next Platform dates the public details to June 2022. LLNL’s own account says the HPE Rabbit near-node storage design influenced its choice of Cray’s CORAL-2 proposal (HPC@LLNL).
The schedule slipped about two years against the 2019 contract. Installation began in May 2023 and the machine was deployed in 2024 (ASC program). LLNL verified the HPL result on November 18, 2024, at SC24 (LLNL), and dedicated the machine on January 9, 2025 (LLNL). The laboratory has not published an explanation for the slip. Its facility work finished on time: the Exascale Computing Facility Modernization project completed in 2022 under a $100 million budget (LLNL, Road to El Capitan). LLNL shut down Sequoia on January 31, 2020 specifically to make room for it (ASC). Its direct predecessor Sierra was rated 125 petaflops peak; LLNL puts the gain at about 22 times.
What is inside
LLNL’s hardware page lists 11,520 nodes: 11,424 batch, 64 debug and 32 login. Each compute node has four MI300A APUs, 512 GiB of HBM3, and four Slingshot 200 Gb/s interfaces, for 100 GB/s of injection bandwidth. That is 46,080 APUs, 1,105,920 CPU cores and about 5.9 million GiB of memory. The same page gives 2,889.2 petaflops peak and 36.0 MW peak power across 90 compute cabinets, and lists 720 Rabbit modules, one per 16 nodes (LLNL hardware overview).
The MI300A has 24 Zen 4 cores in three chiplets, 228 CDNA 3 compute units in six chiplets, and 128 GB of HBM3 shared by both. The Next Platform puts per-APU FP64 at 122.6 teraflops using matrix cores. The CPU and GPU chiplets are on a 5 nm process and the we/O dies on 6 nm, both at TSMC (The Next Platform, 2024). What the APU removes is the host-to-device copy. LLNL’s chief technology officer put it this way: “APUs eliminate the need for data transfers between memory systems.” (S&TR). For a laboratory with decades of CPU-era weapons codes, that matters more than the peak figure. LLNL’s staff report that codes using its RAJA portability layer ran on El Capitan with little porting effort.
The fabric is HPE Slingshot in a dragonfly topology, with 64-port switches at 25.6 Tb/s bidirectional. Storage has two tiers: a shared Lustre file system and the Rabbits. A Rabbit is a near-node storage chassis with its own processor that can be configured per job, either as a dedicated Lustre file system or as local block storage. The Flux scheduler allocates it alongside compute, HPE software running in Kubernetes containers configures it, and data drains to the large tier when the job ends (LLNL, Road to El Capitan 4). The Lustre capacity and bandwidth are not in any source we could fetch.
The software stack is LLNL’s own. The operating system is TOSS 4, the Tri-Lab Operating System Stack, based on Red Hat Enterprise Linux, with Flux, Spack, RAJA and MFEM. Application codes named by the lab include ARES and MARBL (S&TR).
Cooling is water-based without chillers, enabled by the facility upgrade. LLNL’s S&TR gives 35 MW of draw against 85 MW of facility power capacity. A separate interview with LLNL staff gives about 85 MW for computing plus roughly 15 MW for cooling (Exascale Computing Project).
How it performs and what it does
The benchmark results, each measuring something different:
- HPL, FP64: 1.742 exaflops on the November 2024 and June 2025 lists, against 2.746 exaflops peak, with 29,581 kW reported. TOP500 lists 11,039,616 cores (TOP500, Nov 2024). Remeasured at 1.809 exaflops in November 2025 against 2.821 exaflops peak, 29,685 kW and 11,340,000 cores (TOP500, June 2026).
- HPCG: 17.41 petaflops, first place. This measures memory-bound sparse linear algebra, a better proxy for weapons-style codes than HPL.
- HPL-MxP: 16.7 exaflops, first place. Mixed precision, the AI-style number, not comparable with HPL (LLNL benchmark summary).
- Green500: 58.89 gigaflops per watt on the June 2025 list (TOP500 Green500), and 60.94 at rank 23 in November 2025 (LLNL).
The core counts look like arithmetic on APU counts. 11,039,616 divided by 252 (24 CPU cores plus 228 compute units) is 43,808, the APU count The Next Platform gives for the HPL run, about 98% of the machine. 11,340,000 divided by 252 is exactly 45,000. TOP500 does not state that count, so it is our inference that the November 2025 run used about 45,000 APUs.
The public workload record is thin because the machine is classified. LLNL’s stated mission is simulating the nuclear stockpile without explosive testing, and it names AI-driven workflows such as ICECap, which couples multiphysics simulation with National Ignition Facility data. Two unclassified jobs were run first. A rocket-plume CFD run, 33 engines at Mach 10, used 11,136 nodes and over 44,500 APUs, reached 100 trillion grid points, and was run before the machine moved to classified operations (LLNL). A tsunami digital twin won the 2025 Gordon Bell Prize, using 43,520 GPUs offline and 512 online, with 55.5 trillion unknowns and a 0.2 second inference time (UT Austin). The 10-billion-fold speedup is the authors’ claim over earlier methods.
The classified and unclassified split
Security shaped the whole family, not just the main machine. LLNL’s hardware page tags El Capitan, Tuolumne and rzAdams with different network labels (SCF, CZ and RZ) without expanding them. Tuolumne is the open-science companion: 1,152 nodes, 4,608 APUs, 288.9 petaflops peak, 9 cabinets and 3.6 MW, about a tenth of El Capitan. LLNL says it runs unclassified work through its Multiprogrammatic and Institutional Computing program, while El Capitan runs the stockpile mission. Tuolumne ranked 12th on the November 2025 TOP500 at 208.1 petaflops. rzAdams is a one-cabinet, 128-node version carrying the RZ label. El Dorado at Sandia serves the same readiness purpose. Earlier steppingstones, Tioga and rzVernal, carried the MI250X generation so codes could be ported before the APUs existed.
The sequencing is the unusual part. El Capitan was first run in an early-access, open-science mode while it was tested, and LLNL expected full deployment as a classified system in March 2025 (Pleasanton Weekly). Once it crosses the fence, public performance data stops. The benchmark entries are close to the only independent measurements.
The supply chain
HPE is integrator, fabric, cabinet and storage-software supplier: the Cray EX255a, Slingshot and the Rabbit. AMD supplies the entire compute layer in one package, and TSMC fabricates it, per The Next Platform. Red Hat appears as a procurement partner in LLNL’s dedication article, and TOSS is built on its enterprise Linux. The same MI300A and node design appear at HPC7, an Italian energy company’s seismic imaging machine, and LineShine is the system that ended El Capitan’s run at the top.
What does not add up or is not public
- Node count. LLNL’s hardware page says 11,520 nodes and 46,080 APUs. The Next Platform and an LLNL article on the rocket run say 11,136 nodes and 44,544 APUs. The page counts login and debug nodes, but the numbers still do not subtract cleanly, since 11,520 minus 11,136 is 384 and the page lists 11,424 batch nodes.
- Peak. Three figures circulate: 2.746 exaflops (TOP500, 2024), 2.79 (LLNL announcement) and 2.889 (hardware page). They scale with APU count at about 62.7 teraflops each.
- Power. 29.6 MW during HPL, 35 MW (S&TR), 36.0 MW peak (hardware page), contract limit under 40 MW. These measure different things and no source reconciles them.
- Cabinets. 87 compute racks (The Next Platform) against 90 cabinets (LLNL).
- Cost. $600 million is the 2019 DOE contract value. The Next Platform reported a $500 million budget in 2020. We did not find a final cost on any LLNL page we fetched.
- Not disclosed: Lustre capacity, Rabbit capacity, why the schedule slipped, and anything about classified workloads.
What this dataset says
This section describes the dataset as published on 2 October 2026. The dataset changes as rows are added and corrected.
The El Capitan row records 1,809 petaflops Rmax, 2,821.1 petaflops Rpeak, 11,340,000 cores and 29,685 kW, the current TOP500 figures. Researching this piece updated them from the November 2024 values of 1,742 petaflops and 29,581 kW. That works out to 60.9 gigaflops per watt, close to TOP500’s own figure (see efficiency). List appearances hold three rows: TOP500 rank 1 in 2024-11, rank 2 in 2026-06, and Green500 rank 26 in 2025-06. There is no capacity claim row, and the 2019 announcement to 2024-11 first operation sets a five-year gap for announcement to first light. Seven component rows cover the MI300A, integrated CPU, HPE EX255a, Slingshot-11, warm-water cooling, the Tri-Lab Operating System Stack and ClusterStor E1000 storage; the MI300A appears in eleven systems in the dataset (accelerator share, interconnect share). El Capitan is one of 21 systems tagged classified among the 765 then in the dataset, which is the category sector mix and the shadow list treat as least observable. The storage row names ClusterStor E1000, which no LLNL page we could fetch confirms (LLNL says only Lustre and Rabbit storage), so see data trust.
Sources
- DOE, NNSA $600 million contract announcement (Aug 2019)
- The Next Platform, AMD selection (Mar 2020)
- The Next Platform, El Capitan architecture (Nov 2024)
- LLNL HPC, El Capitan hardware overview
- LLNL HPC, El Capitan platform page
- LLNL, El Capitan verified fastest (Nov 2024)
- LLNL, dedication (Jan 2025)
- LLNL, three-benchmark article (June 2025)
- LLNL, November 2025 TOP500
- LLNL, Gordon Bell finalist rocket simulation
- LLNL S&TR, Introducing El Capitan
- LLNL, Road to El Capitan 1
- LLNL, Road to El Capitan 4 (storage)
- HPC@LLNL, Rabbit announcement page
- ASC program, El Capitan page
- ASC program, Sequoia decommissioning
- Exascale Computing Project, siting El Capitan
- TOP500, November 2024 list
- TOP500, June 2026 list
- TOP500, El Capitan system page
- TOP500, Green500 June 2025
- TOP500, June 2026 news (LineShine)
- TOP500, June 2025 news
- UT Austin, Gordon Bell tsunami digital twin
- Pleasanton Weekly, inside El Capitan (Jan 2025)