Supercomputers/Systems/IBM Blue Vela
USA · operational · ai training
IBM Blue Vela
Operated by IBM .
Phase 1 seed dataset, compiled by hand. These rows were built from public operator, laboratory and vendor sources. A mechanical second-reader pass has since fetched every cited source: 1092 of 1153 systems have a readable citation that names them, and 40 are genuinely weakly sourced. Every claim carries its source and a confidence tier. Treat anything below verified as a lead, not a citation.
Measured performance
- Rmax
- -
- Rpeak
- -
- Rmax ÷ Rpeak
- -
- Cores
- -
- Power
- -
- Per watt
- - GF/W
Figures are as last publicly reported for the configuration described below, not a live measurement. Where a system was upgraded in place, the post-upgrade configuration is shown and the earlier one appears in the timeline.
What this machine is made of
One row per supplier relationship. “Supplier at build” is the company that shipped the part at the time; where that company has since been acquired, the parent it rolls up to today is shown beside it. That distinction is what makes ticker-level aggregation possible across a thirty-year dataset.
| Role | Part | Supplier at build | Quantity | Confidence | Source |
|---|---|---|---|---|---|
| Integrator | - Dell PowerEdge XE9680 nodes on the NVIDIA H100 SuperPOD reference architecture | Dell Technologies | - | Reported | blocksandfiles.comarxiv.org |
| Accelerator | NVIDIA H100 SXM5 Accelerators CUDA (Hopper GH100) · 16,896C · 1.98 GHz · TSMC 4N · 700 W 5,000 NVIDIA H100 80GB GPUs, eight per node | NVIDIA | 5,000 reported | Reported | blocksandfiles.comarxiv.org |
| CPU | - Dual 48-core 4th Gen Intel Xeon Scalable per node, 2TB RAM | Intel | - | Reported | arxiv.org |
| Interconnect | NVIDIA InfiniBand NDR (Quantum-2) Interconnect Fat tree or dragonfly+ · 400 Gb/s per port · Under 0.6 microseconds Separate compute and storage InfiniBand fabrics, non-blocking rail-optimised fat tree, ConnectX-7 adapters | NVIDIA | - | Reported | arxiv.orgblocksandfiles.com |
| Storage | - IBM Storage Scale System 6000, two chassis, almost 3 PB raw NVMe | IBM | - | Reported | blocksandfiles.comarxiv.org |
The read
IBM Research's on-premises H100 training cluster, built with Dell and NVIDIA on the H100 SuperPOD reference architecture: Dell PowerEdge XE9680 nodes with eight 80GB H100 GPUs and dual 48-core fourth-generation Xeon CPUs, in 32-node scalable units, four to a 128-node compute pod, on non-blocking rail-optimised InfiniBand with separate compute and storage fabrics. Storage is IBM Storage Scale System 6000, almost 3 PB of raw NVMe. IBM's own paper (January 2025 revision) says Blue Vela began coming online in April 2024, that it lifted IBM Research's GPU capacity 104 percent over 2023 and that it was to reach a cumulative 214 percent by the end of 2024. The 5,000-GPU figure comes from Blocks and Files summarising IBM material, not from a page IBM publishes with the count, so this stays reported. IBM says only that the hosting site runs on 100 percent renewable energy; no location or power figure is given, so USA records only the operator's home country. It has never submitted an HPL result.
Timeline
- 2024-04 First operational source
Change history
- 2026-10-02 IBM Blue Vela added (5,000 H100 on-premises training cluster)
Source check
We fetched this system's own citations and recorded whether each page actually mentions it. This is a corroboration signal, not a fact check, and it is published so you can see how well the sourcing holds up rather than take it on trust.
| Cited source | Result | Found on the page |
|---|---|---|
| arxiv.org | does not name this system | NVIDIA H100 SXM5, NVIDIA InfiniBand NDR (Quantum-2) |
| blocksandfiles.com | names system, part and count | NVIDIA H100 SXM5, NVIDIA InfiniBand NDR (Quantum-2) · counts: 5000 |
Checked 2026-10-09 by pnpm verify. Re-run it and the table changes with the web.
Further reading
Inside a node
Why accelerators exist, what the memory hierarchy costs, and how the CPU and accelerator merged onto one package.
The interconnect
Topologies, why latency and tail behaviour matter more than bandwidth, and how the fabric market consolidated.
Storage and I/O
Parallel file systems, why checkpointing dominates the write load, and the metadata failure mode nobody expects.