COMPUTE ATLAS Supercomputer supply-chain graph
1153 systems 482 sites Sourced data

Blog/2026-10-02

· Compute Atlas

Fugaku: the CPU-only No. 1 and its successor

Fugaku led four benchmark rankings with no accelerator and has posted the same 442 PFLOPS result since 2020. Its successor, FugakuNEXT, adds NVIDIA GPUs.


Fugaku, at the RIKEN Center for Computational Science in Kobe, is 158,976 single-socket Arm nodes with no GPU, no separate network card and no DDR memory. The TOP500 lists it at 7,630,848 cores, 442.01 PFLOPS of measured FP64 (HPL), 537.21 PFLOPS peak and 29.9 MW. It took first place on the June 2020 TOP500 list. RIKEN’s announcement of that result says it was the first machine to lead the TOP500, HPCG and Graph500 at the same time. It also led HPL-AI, a mixed-precision benchmark, which Fujitsu’s release calls the first measurement of one exaflop.

The one thing worth knowing is that the machine has not changed. The TOP500 entry shows the same 442.01 PFLOPS on every list from November 2020 to June 2026. Across its 13 list appearances it ranked No. 1 on the first four (the June 2020 entry was a partial 415.53 PFLOPS system), No. 2 on the next three, then 4, 4, 6, 7, 7 and 9, so it was never out of the top 10. Nothing was added to it; other machines passed it. That makes it a clean case study in the design choice it embodies, and RIKEN’s successor design now documents why the choice is being reversed.

How it came to exist

Fugaku was the “post-K” machine, the successor to the K computer. The Japanese government launched the FLAGSHIP 2020 project in Japanese fiscal year 2014. In summer 2014 it selected nine priority application areas (drug discovery, personalised medicine, earthquake and tsunami, weather and climate, energy, clean energy systems, materials, manufacturing, and fundamental physics). RIKEN was the responsible institution and Fujitsu was selected in September 2015 to produce the basic design. These dates come from RIKEN’s own annual report chapter.

Jack Dongarra’s 2020 report on the system dates its origins to 2006 planning for a K follow-on, and says processor co-design with Fujitsu began in 2011 with a goal of a 100x application speedup over K. Fujitsu revealed the A64FX processor in August 2018. The K computer was shut down in August 2019, because, in The Next Platform’s words, “running K while building Post-K became impossible” for lack of floor space and infrastructure.

Delivery was pipelined. Fujitsu says shipments began on 2 December 2019. Dongarra’s report says the first rack shipped on 3 December and all racks were on the floor by 13 May 2020. The TOP500 result was announced on 22 June 2020 at ISC. RIKEN described the machine then as used experimentally for COVID-19 research and scheduled for full operation in April 2021. The November 2020 list carried the full-system 442.01 PFLOPS result.

Who paid: the project was backed by Japan’s education and science ministry (MEXT), per Fujitsu’s 2019 shipping announcement. Neither RIKEN nor Fujitsu pages we could fetch state a total cost. Wikipedia cites a Nikkei report of September 2018 putting it at about 130 billion yen (roughly US$1 billion). That is the only cost figure we found, and it is a press estimate, not a RIKEN number.

What is inside

The processor. The A64FX is Fujitsu’s own Armv8.2-A core design with 512-bit SVE vector units, 48 compute cores plus assistant cores for the operating system and MPI, in four core-memory groups. It carries 32 GiB of HBM2 on the package at 1,024 GB/s, and has no DDR interface. Fujitsu is the designer; per Dongarra the chip uses TSMC’s 7 nm process with CoWoS packaging, and Broadcom SerDes. At 2.0 GHz, FP64 peak per chip is about 3 TFLOPS; the full machine’s normal-mode peak is 488 PFLOPS FP64, 977 PFLOPS FP32, 1.95 EFLOPS FP16 and 3.90 EOPS INT8, rising to 537 PFLOPS FP64 in boost mode at 2.2 GHz (RIKEN).

The network. The Tofu-D controller is on the processor die, so there is no separate NIC. Each node has 10 ports, each two lanes at 28 Gbps, in a six-dimensional mesh/torus. RIKEN quotes 0.49 to 0.54 microseconds node-to-node latency. In a Fujitsu slide in Matsuoka’s 2021 briefing to the US DOE advisory committee, the on-die logic is about 6% of the die area and 8 to 9 W per node including SerDes and optical cables, against 25 to 30 W for a 100GbE or InfiniBand adapter. The machine has roughly 200,000 Tofu cables, 97,632 optical and 119,232 electrical.

The system. There are 396 full racks of 384 nodes and 36 half racks of 192 nodes. Total memory is 4.85 PiB and aggregate memory bandwidth is 163 PB/s. Storage has three layers: a node-local and job-shared tier (Dongarra gives 15.9 PB of NVMe), the Lustre-based FEFS shared file system, and commercial cloud storage. The software is Red Hat Enterprise Linux 8 with RIKEN’s McKernel lightweight kernel running beside it. Dongarra describes a closed-coupled chilled configuration with a custom water-cooling unit, a water-cooled processor and memory unit, and a 1,920 square metre footprint.

Power. TOP500 records 28,334.5 kW for the June 2020 run and 29,899.23 kW for the full-system run, which works out to 14.79 GFLOPS per watt. TOP500 also lists an optimised run of 404.69 PFLOPS at 26,248.36 kW, or about 15.4 GFLOPS per watt. Before the machine existed, Fujitsu’s prototype took first place on the November 2019 Green500 at 16.876 GFLOPS per watt.

Performance and what it did

On HPL, Fugaku’s efficiency is 442.01 over 537.21, or 82.3%, a figure Matsuoka’s briefing also gives. HPCG, which stresses memory bandwidth, was 13.4 PFLOPS on 138,240 nodes in June 2020; TOP500’s system page lists 16,004.5 TFLOPS without naming the list. Graph500 was 70,980 GTEPS on 92,160 nodes in June 2020 and 102,955 GTEPS on all 158,976 nodes in July 2021, its third consecutive win, against 31,302 GTEPS for K. HPL-AI was 1.421 EFLOPS at launch and 2.00 EFLOPS by June 2021, which is 93.2% of FP16 peak. In Matsuoka’s comparison, Summit reached 1.15 EFLOPS on the same test at 33.2% of its FP16 peak. Matsuoka’s slide attributes the gap to A64FX using vector units for FP16 where the GPUs use matrix engines.

The metric the project cared about is application speedup over K. Matsuoka reported an average of about 70x across nine target applications, from 23x for the genome code to 131x for the molecular dynamics code. Fujitsu’s 2019 target was up to 100x K’s application performance at about three times its power.

Known uses include:

  • COVID-19. Fugaku was opened to COVID work a year before general operation. Makoto Tsubokura’s group simulated droplet and aerosol spread in trains, offices, classrooms and restaurants, and masks and ventilation. A RIKEN-led team won the ACM Gordon Bell Special Prize for COVID-19 research, announced at SC21 in December 2021. Matsuoka’s slides record Prime Minister Suga citing the simulations on 22 November when urging mask-wearing.
  • Large language models. Fugaku-LLM, released 10 May 2024, is a 13-billion-parameter model trained on 380 billion tokens using 13,824 nodes. The team chose CPUs because of “a global shortage of GPUs” and to support domestic semiconductors.
  • Open research programmes. The FY2020 to FY2022 Fugaku promotion programme listed projects from cosmology and plasma physics to turbine design and whole-brain simulation.

The supply chain

Fujitsu is the processor designer, the interconnect designer and the integrator, which is why the system’s bill of materials is so short. Fujitsu supplies the CPU, the fabric, the FEFS storage software and the water-cooled racks. RIKEN co-designed it and supplies the McKernel kernel. Red Hat supplies the base operating system. Per Dongarra the silicon is made by TSMC, with Broadcom SerDes. The dataset’s own note calls it the shortest bill of materials of any modern leadership system, but it is not free of outside dependence: the Arm instruction set and the leading-edge foundry are external. Who made the HBM2 stacks is not recorded in the dataset or in any source we read. See supply concentration.

What does not add up or is not public

  • Transistor count. TOP500’s coverage of Fujitsu’s 2018 disclosure says 8,786 million transistors. Dongarra’s report says 87.86 billion. The first is almost certainly right and the second a typo, but we are reporting a conflict, not resolving it.
  • First shipment date. 2 December 2019 (Fujitsu) against 3 December (Dongarra).
  • Network bandwidth. RIKEN’s specification page gives 6.35 GB/s per connection. The dataset’s Tofu-D part record says 6.8 GB/s per link per direction.
  • Total cost. Not published by RIKEN in anything we could fetch. The 130 billion yen figure is a newspaper report relayed by Wikipedia.
  • Start of production use. “First operational” can mean first rack, benchmark day, early access, or general service. RIKEN pointed to April 2021 for full operation. Dongarra says early users had access in the first quarter of 2020. The machine’s story depends on which you pick.
  • Facility power. TOP500 reports system power at the benchmark. We found no published PUE or facility draw for the Kobe site.

Why it stayed in the top 10, and what comes next

We could not find a published explanation of why Fugaku stayed in the top 10, so the following is our reading of the numbers, not RIKEN’s. Its position slid slowly rather than collapsing because HPL rewards sustained FP64 and Fugaku sustains 82% of peak. Newer machines passed it only as accelerated systems reached exascale FP64 scale. The ranking fell from 2 to 4 in November 2023, then to 6, 7 and 9 as more of those arrived.

The successor is the reversal. RIKEN announced FugakuNEXT with Fujitsu and NVIDIA on 22 August 2025. The CPU is the FUJITSU-MONAKA-X Arm chip, paired with NVIDIA GPUs. RIKEN states more than 600 EFLOPS of FP8 (sparse) AI performance, up to a hundredfold application speedup over Fugaku, the same roughly 40 MW power constraint, and operation around 2030. Its release does not state a budget. Separately, NVIDIA’s November 2025 release describes two RIKEN Blackwell systems (1,600 and 540 GPUs, GB200 NVL4, Quantum-X800 InfiniBand) due in spring 2026.

A technical report, version 1.1, dated 29 May 2026, gives the reasoning. MEXT launched the development project in January 2025, and RIKEN worked with Fujitsu from June 2025 and NVIDIA from August 2025 through February 2026. The CPU part keeps binary compatibility with Fugaku, to preserve its software. The report justifies GPUs because they have “a track record of use in large-scale supercomputers”, so users can port early and the ecosystem is open. Two findings in the report explain the turn:

  • RIKEN analysed 3,129,958 Fugaku jobs of an hour or longer from FY2024: only 734 exceeded 30% of FP peak and 93 exceeded 50%. It describes a strong memory-bandwidth-bound tendency. The report’s measurement table lists 1,024 GB/s for the A64FX and 4,000 GB/s in the H100 column.
  • The AI target changed the job. The stated goals are 50 EFLOPS or more of effective AI performance and Zetta-scale peak. The A64FX was designed to run FP16 on its vector units, and the report treats matrix engines (SME2 on the CPU, tensor hardware on the GPU) as the route to that number.

The interconnect also changes. Tofu-D is not carried over: the report says the network is not yet decided, and both proposals it describes are fat-tree designs, one using NVIDIA Spectrum-X. The accelerator count is also unsettled. One survey question to application developers assumed 13,600 GPUs in 3,400 nodes of 4, with pods of 4 to 576 GPUs, but that is a hypothetical in a questionnaire, not a design decision. Several items, including scale-up and scale-out network choices, are scheduled to be decided in 2027.

So the no-accelerator path ended for a mix of reasons that RIKEN’s own report separates: bandwidth-bound applications that left FP64 peak idle, AI workloads that need matrix throughput, and the practical value of a mainstream GPU software ecosystem. Fugaku’s CPU-only result was real. The report does not say it was a mistake. It says the target moved.

What this dataset says

This section describes the dataset as published on 2 October 2026. The dataset changes as rows are added and corrected.

The fugaku row matches TOP500 on the headline fields: 442.01 PFLOPS Rmax, 537.212 PFLOPS Rpeak, 7,630,848 cores and 29,899 kW, with the flops basis marked as HPL-measured. It has seven component rows, with Fujitsu on five of them (CPU, integrator, interconnect, storage and cooling) and Red Hat and RIKEN on the operating system. It has no accelerator row at all. Of the 497 operational systems in the dataset, 352 have one, and of the 311 ranked ones, 211. See accelerator share. At 29,899 kW it is the third-largest disclosed power draw among operational systems with a recorded figure, behind only LineShine and Aurora; see power and cooling.

The A64FX appears in eight systems’ component rows, six of them with Tofu-D, including Wisteria/BDEC-01, Flow and the JMA and CWA weather systems. The dataset’s list history for Fugaku is thin: one TOP500 row (June 2020, rank 1) and one Green500 row (November 2023, rank 54), against 13 TOP500 appearances on TOP500’s own page. The K computer row records a November 2011 appearance at rank 1 (it was also first in June 2011, at 8.162 PFLOPS) and decommissioning in August 2019, about eight years of service; compare service life. Its Rmax of 10.51 PFLOPS is the November 2011 figure.

For FugakuNEXT the dataset carries one capacity claim, 40 MW, scope unspecified, labelled full build-out and a design constraint, and no Rmax or Rpeak. The dataset’s verification log flags the RIKEN project page as unreachable, which matches my own fetch: that URL returned a 404. Treat the Fugaku row’s source link as stale; see data trust.

Sources

deep-divehpchistoryacceleratorspower-and-coolinginterconnect