Supercomputers/Systems/Meta GenAI cluster (RoCE)
USA · operational · ai training
Meta GenAI cluster (RoCE)
Operated by Meta Platforms at Meta (undisclosed US site) .
Phase 1 seed dataset, compiled by hand. These rows were built from public operator, laboratory and vendor sources. A mechanical second-reader pass has since fetched every cited source: 32 of 52 systems have a readable citation that names them, and 4 are genuinely weakly sourced. Every claim carries its source and a confidence tier. Treat anything below verified as a lead, not a citation.
Measured performance
- Rmax
- -
- Rpeak
- -
- Rmax ÷ Rpeak
- -
- Cores
- -
- Power
- -
- Per watt
- - GF/W
Figures are as last publicly reported for the configuration described below, not a live measurement. Where a system was upgraded in place, the post-upgrade configuration is shown and the earlier one appears in the timeline.
What this machine is made of
One row per supplier relationship. “Supplier at build” is the company that shipped the part at the time; where that company has since been acquired, the parent it rolls up to today is shown beside it. That distinction is what makes ticker-level aggregation possible across a thirty-year dataset.
| Role | Part | Supplier at build | Quantity | Confidence | Source |
|---|---|---|---|---|---|
| Integrator | Meta Grand Teton Racks Grand Teton on OpenRack | Meta Platforms | - | Verified | engineering.fb.comopencompute.org |
| Accelerator | NVIDIA H100 SXM5 Accelerators 24,576-GPU cluster, NVIDIA H100 | NVIDIA | 24,576 reported | Reported | engineering.fb.com |
| Interconnect | Arista 7800R4 AI Networking Arista 7800 with Wedge400 and Minipack2 OCP rack switches, RoCE, 400 Gbps | Arista Networks | - | Reported | engineering.fb.com |
| Storage | Meta Tectonic Storage Tectonic distributed storage, FUSE API, flash | Meta Platforms | - | Reported | engineering.fb.com |
The read
Meta published the hardware and the fabric and no performance figure, which is the normal shape of a commercial disclosure and the reason this system has no rank. We record the accelerator count because the operator stated it, and we record no FLOPS at all because deriving one would require an HPL rate for this part that nobody has measured. See the shadow list for what that costs. What makes this cluster and its InfiniBand twin unusually valuable is that they are a controlled experiment: same operator, same 24,576 H100s, same 400 Gbps endpoints, same year, and deliberately different fabrics. Meta built one on RoCE over Ethernet with Arista 7800 switches specifically to compare it against the other. That is the InfiniBand-versus-Ethernet question the whole industry is arguing about, run once, properly, by someone with the budget to do it.
Site & facility
Source check
We fetched this system's own citations and recorded whether each page actually mentions it. This is a corroboration signal, not a fact check, and it is published so you can see how well the sourcing holds up rather than take it on trust.
| Cited source | Result | Found on the page |
|---|---|---|
| engineering.fb.com | names system, part and count | NVIDIA H100 SXM5, Arista 7800R4 AI, Meta Grand Teton, Meta Tectonic · counts: 24576 |
| opencompute.org | could not be fetched | - |
Checked 2026-09-08 by pnpm verify. Re-run it and the table changes with the web.
Further reading
Inside a node
Why accelerators exist, what the memory hierarchy costs, and how the CPU and accelerator merged onto one package.
The interconnect
Topologies, why latency and tail behaviour matter more than bandwidth, and how the fabric market consolidated.
Operating a supercomputer
Acceptance testing, why "installed" and "in production" are six months apart, and the machine lifecycle.