Supercomputers/Systems/Helma
Germany · operational · ai training
Helma
Also known as NHR@FAU Helma
Operated by FAU Erlangen-Nuremberg at RRZE Erlangen .
Phase 1 seed dataset, compiled by hand. These rows were built from public operator, laboratory and vendor sources. A mechanical second-reader pass has since fetched every cited source: 1092 of 1153 systems have a readable citation that names them, and 40 are genuinely weakly sourced. Every claim carries its source and a confidence tier. Treat anything below verified as a lead, not a citation.
Measured performance
- Rmax
- 32.22 PFlop/s
- Rpeak
- 51.03 PFlop/s
- Rmax ÷ Rpeak
- 63.1%
- Cores
- 125,952
- Power
- 651.13 kW
- Per watt
- 49.5 GF/W
Figures are as last publicly reported for the configuration described below, not a live measurement. Where a system was upgraded in place, the post-upgrade configuration is shown and the earlier one appears in the timeline.
What this machine is made of
One row per supplier relationship. “Supplier at build” is the company that shipped the part at the time; where that company has since been acquired, the parent it rolls up to today is shown beside it. That distinction is what makes ticker-level aggregation possible across a thirty-year dataset.
| Role | Part | Supplier at build | Quantity | Confidence | Source |
|---|---|---|---|---|---|
| Integrator | - MEGWARE (system integrator) | MEGWARE | - | Verified | top500.orghpc.fau.de |
| Accelerator | NVIDIA H200 SXM Accelerators CUDA (Hopper GH100) · 16,896C · 1.98 GHz · TSMC 4N · 700 W 96 nodes with 4 NVIDIA H200 141GB GPUs each | NVIDIA | 384 reported | Reported | doc.nhr.fau.detop500.org |
| Accelerator | NVIDIA H100 SXM5 94GB Accelerators CUDA (Hopper GH100) · 16,896C · 1.98 GHz · TSMC 4N · 700 W 96 nodes with 4 NVIDIA H100 94GB GPUs each | NVIDIA | 384 reported | Reported | doc.nhr.fau.de |
| CPU | AMD EPYC 9004 (Genoa) CPUs x86-64 · 96C · 2.4 GHz · TSMC N5 · 360 W AMD EPYC 9554 dual-socket nodes | AMD | - | Verified | top500.orgdoc.nhr.fau.de |
| Interconnect | NVIDIA InfiniBand NDR (Quantum-2) Interconnect Fat tree or dragonfly+ · 400 Gb/s per port · Under 0.6 microseconds Infiniband NDR200 (4x NDR200 internal, 2x NDR400 external per GPU node) | NVIDIA | - | Verified | top500.orgdoc.nhr.fau.de |
The read
The GPU cluster of NHR@FAU at Friedrich-Alexander-Universitaet Erlangen-Nuernberg, integrated by MEGWARE: dual AMD EPYC 9554 nodes with four NVIDIA H100 or H200 GPUs each (96 nodes of each type, 768 GPUs in total per the operator's documentation), InfiniBand NDR200, alongside a CPU partition of dual EPYC 9965 nodes. The H200 half is reserved for a Bavarian foundation-model effort; funders were the State of Bavaria, the NHR programme, FAU and partner universities. TOP500 lists 125,952 cores, Rmax 32.22 PFlop/s, Rpeak 51.03 PFlop/s and 651 kW under AlmaLinux 9.5. It first entered at 79th in November 2024 and reached its best rank, 51st, on the June 2025 list; FAU calls it the fastest supercomputer at German universities. Workload is classed ai_training because the operator names AI/ML and the foundation-model project as its purpose alongside atomistic simulation.
Timeline
Public list appearances
Pointers only: rank and edition, linking to the canonical entry. We do not reproduce list tables. See sourcing policy.
Site & facility
Change history
- 2026-10-02 Added Helma, the NHR@FAU H100/H200 cluster, 51st on the June 2025 TOP500
Source check
We fetched this system's own citations and recorded whether each page actually mentions it. This is a corroboration signal, not a fact check, and it is published so you can see how well the sourcing holds up rather than take it on trust.
| Cited source | Result | Found on the page |
|---|---|---|
| hpc.fau.de | names system, part and count | NVIDIA H200 SXM, NVIDIA H100 SXM5 94GB · counts: 384, 384 |
| top500.org | names system and part | AMD EPYC 9004 (Genoa), NVIDIA H200 SXM, NVIDIA H100 SXM5 94GB, NVIDIA InfiniBand NDR (Quantum-2) |
Checked 2026-10-09 by pnpm verify. Re-run it and the table changes with the web.
Further reading
Operating a supercomputer
Acceptance testing, why "installed" and "in production" are six months apart, and the machine lifecycle.
Inside a node
Latency-optimised versus throughput-optimised processors, NUMA, and the end of the host-to-device copy.
The interconnect
Topologies, why latency and tail behaviour matter more than bandwidth, and how the fabric market consolidated.