Supercomputers/Systems/xAI Colossus
USA · operational · ai training
xAI Colossus
Deep dive: xAI Colossus: fast to build, hard to count
Operated by xAI at xAI Colossus (Memphis) .
Phase 1 seed dataset, compiled by hand. These rows were built from public operator, laboratory and vendor sources. A mechanical second-reader pass has since fetched every cited source: 1092 of 1153 systems have a readable citation that names them, and 40 are genuinely weakly sourced. Every claim carries its source and a confidence tier. Treat anything below verified as a lead, not a citation.
Measured performance
- Rmax
- -
- Rpeak
- -
- Rmax ÷ Rpeak
- -
- Cores
- -
- Power
- -
- Per watt
- - GF/W
Figures are as last publicly reported for the configuration described below, not a live measurement. Where a system was upgraded in place, the post-upgrade configuration is shown and the earlier one appears in the timeline.
What this machine is made of
One row per supplier relationship. “Supplier at build” is the company that shipped the part at the time; where that company has since been acquired, the parent it rolls up to today is shown beside it. That distinction is what makes ticker-level aggregation possible across a thirty-year dataset.
| Role | Part | Supplier at build | Quantity | Confidence | Source |
|---|---|---|---|---|---|
| Integrator | - Supermicro liquid-cooled AI SuperCluster | Supermicro | - | Reported | supermicro.com |
| Accelerator | NVIDIA H100 SXM5 Accelerators CUDA (Hopper GH100) · 16,896C · 1.98 GHz · TSMC 4N · 700 W 100,000 NVIDIA Hopper GPUs | NVIDIA | 100,000 reported | Reported | nvidianews.nvidia.comsupermicro.com |
| Interconnect | NVIDIA Spectrum-4 Networking Ethernet switch ASIC · 51.2 Tb/s aggregate, 64 ports at 800G NVIDIA Spectrum-X Ethernet networking platform | NVIDIA | - | Verified | nvidianews.nvidia.comsupermicro.com |
| Cooling | Direct-to-chip cold plate Cooling Cold plate on the package, air for the remainder · Removes roughly 50-80 percent of node heat to liquid direct-to-chip (DLC) liquid-cooled AI systems | Supermicro | - | Reported | supermicro.com |
AI datacenter metrics
Figures beyond an accelerator count, each with who stated it. What can and cannot be compared across AI datacenters.
| Metric | Value | As of | Basis | Note | Source |
|---|---|---|---|---|---|
| Accelerators | 200,000 accelerators | 2024-10-28 | Design target | NVIDIA release: xAI "in the process of doubling" Colossus to 200,000 Hopper GPUs combined total; target, not a confirmed installed count | nvidianews.nvidia.com |
| Accelerators | 100,000 accelerators | 2024-10 | Regulatory filing | SpaceX S-1: first Colossus cluster of about 100,000 H100; NVIDIA release of 28 Oct 2024 independently says 100,000 Hopper GPUs; NVIDIA said expansion to 200,000 planned | sec.govnvidianews.nvidia.com |
| Stated investment | 12.9 USD bn | 2026-06-15 | Third-party estimate | Epoch AI modelled build cost, compute plus construction, 2025 USD, scaled from IT power; not a disclosed spend. | epoch.ai |
| Facility power | 425 MW | 2026-06-15 | Third-party estimate | Epoch AI modelled gross facility power (IT plus cooling and overhead) at this timeline date. | epoch.ai |
| Facility power | 300 MW | 2025-12 | Design target | MLGW: second 150 MW increment requested for 300 MW total, projected by Dec 2025 but pending TVA board approval as of 5 May 2025 | mlgw.com |
| Facility power | 150 MW | 2025-05 | Regulatory filing | MLGW utility update: 150 MW grid service energized at Paul Lowry Rd (8 MW plus 142 MW substations); better sourced than the press figure in capacity_claims | mlgw.com |
| IT power | 340 MW | 2026-06-15 | Third-party estimate | Epoch AI modelled IT power from satellite imagery and permits at this timeline date; not operator-stated. SpaceX S-1 states about 130 MW for the first cluster. | epoch.ai |
| IT power | 130 MW | 2026-05 | Regulatory filing | SpaceX S-1 (May 2026): about 130 MW compute power for the first cluster; nameplate GPU draw, excludes cooling and distribution losses; no date given | sec.gov |
How it is powered
| Source | Provider | MW | Status | As of | Basis | Note | Source link |
|---|---|---|---|---|---|---|---|
| Grid | Memphis Light, Gas and Water | 150 | operating | 2025-05 | Regulatory filing | Paul Lowry site draws 150 MW from the TVA/MLGW grid (8 MW existing plus 142 MW new substation); TVA board approved it. xAI must curtail when grid demand is high. | mlgw.comrenewableenergyworld.com |
| Grid | Memphis Light, Gas and Water | 150 | contracted | 2026-02 | Regulatory filing | Second 150 MW step (300 MW total). TVA board approved an extra 150 MW of firm power in Feb 2026; WSMV omits the site, MLGW lists the second step at Paul Lowry. | wsmv.commlgw.com |
| On-site gas | xAI | 247 | operating | 2025-07 | Regulatory filing | Shelby County permit for 15 Solar SMT-130 turbines, about 247 MW (MeasuredAI). Cleanview counts 12 turbines, about 198 MW. | measuredai.substack.comcleanview.co |
| Battery storage | Tesla | - | operating | 2026-05 | Operator stated | S-1 says Megapacks add redundancy and cover grid curtailment at COLOSSUS; no MW stated. MeasuredAI counts 240+ Megapacks, Cleanview 119 (about 120 MW). | sec.govmeasuredai.substack.com |
How AI campuses are powered: capacities are what the operator, utility or filing states, not what has been delivered.
The read
NVIDIA's own release and Supermicro's own page independently say the Memphis cluster holds 100,000 Hopper GPUs on Spectrum-X Ethernet, which is the headline fact here. Neither names the Hopper part; we map it to the H100 because that is the part xAI is widely reported to have bought, but the sources say only Hopper, and later reporting describes a mixed H100, H200 and GB200 fleet of roughly twice that size, so the count is a snapshot, not a current total. The July 2024 start date comes from trade press, not an operator statement. Power is deliberately blank: reports give 130 MW and 150 MW for the first phase and it is unclear whether either is facility or IT load. No FLOPS figure is published by anyone, so this system carries no rank and no Rmax.
Timeline
- 2024-07 First operational source
Site & facility
Where sources disagree
We record conflicts instead of picking a winner quietly. All open questions.
- First-phase power: 130 MW against 150 MW. It is unclear whether either is facility or IT load, so power is left blank.
Change history
- 2026-10-04 AI datacenter metrics: xAI Colossus
- 2026-09-19 xAI Colossus added, verified on NVIDIA and Supermicro
Source check
We fetched this system's own citations and recorded whether each page actually mentions it. This is a corroboration signal, not a fact check, and it is published so you can see how well the sourcing holds up rather than take it on trust.
| Cited source | Result | Found on the page |
|---|---|---|
| nvidianews.nvidia.com | names system, part and count | NVIDIA Spectrum-4 · counts: 100000 |
| supermicro.com | names system, part and count | Direct-to-chip cold plate · counts: 100000 |
Checked 2026-10-09 by pnpm verify. Re-run it and the table changes with the web.
Further reading
Inside a node
Why accelerators exist, what the memory hierarchy costs, and how the CPU and accelerator merged onto one package.
The interconnect
Topologies, why latency and tail behaviour matter more than bandwidth, and how the fabric market consolidated.
Power and cooling
Why air ran out at 20 kW a rack, how warm-water cooling removes the chillers, and what PUE conceals.