COMPUTE ATLAS Supercomputer supply-chain graph
1153 systems 482 sites Sourced data

Blog/2026-10-02

· Compute Atlas

xAI Colossus: fast to build, hard to count

Colossus 1 went from site start to 100,000 GPUs in a claimed 122 days. Its GPU counts, its power sources and Colossus 2's size rest mostly on executive statements and court records.


xAI Colossus 1 is a liquid-cooled GPU cluster in a former Electrolux appliance factory in South Memphis. NVIDIA and Supermicro both say it held 100,000 Hopper-generation GPUs on Spectrum-X Ethernet by late 2024, and NVIDIA says it took 122 days from the start of work on the facility to completion. That is the fastest build of a cluster that size on the public record. It is also almost entirely a record made of company statements.

The more useful thing to know is how thin the independent layer is. No one has published a benchmark result for either Colossus site. The GPU count for Colossus 2 has been reported anywhere from about 110,000 to 550,000, and the larger figure comes from the CEO’s own posts. The power picture is better documented than the compute, but only because regulators, a utility and a lawsuit forced it into the open.

This piece separates what xAI and its executives said from what a utility, a county health department, a federal court docket or an outside analyst reported.

How Colossus 1 came to exist

The timeline, with who said each part:

  • June 5, 2024. The Greater Memphis Chamber announced the project at a press conference, according to local public radio coverage (WKNO). The chamber’s chief executive described it as a combination of the two largest supercomputers, multiplied by four. The location was initially withheld.
  • July 2024. MLGW, the local utility, moved the site’s service from the 8 MW of an adjacent substation toward a 50 MW commitment, according to a utility-focused analyst newsletter. WKNO reported 50 MW already approved and a request for 100 MW more by September.
  • September 2, 2024. Elon Musk announced the cluster had come online over the Labor Day weekend, having been completed in 122 days (reported by WKNO).
  • October 28, 2024. NVIDIA’s release restated the build as 122 days from the start of the facility to completion, and 19 days from first rack installation to the start of training. It put the count at 100,000 Hopper GPUs, with a plan to double to 200,000.
  • November 2024. The Tennessee Valley Authority board approved MLGW’s request to supply 150 MW to the site.

A caution on the 122 days. Counting back from September 2 puts the start in early May 2024, which is a month before the public announcement. That is our arithmetic, not a stated date, and “facility start” is not defined in any source we could fetch. The 19-day figure is NVIDIA’s and measures first rack to training, not the whole installation. Both are vendor or executive claims. The dataset’s own first-operational date for the system is July 2024, which comes from trade press, not from an operator statement.

xAI paid for the build. The utility says transmission upgrades and a new substation were built at xAI’s expense. The site is the former Electrolux plant, which MLGW says sits in an industrial park with existing utility infrastructure and was built with state grant money and a local tax abatement.

What is inside

The only detailed public description of the hardware is a ServeTheHome walkthrough of the Supermicro build, and it says some specifications were left vague on purpose. What it does state:

  • Servers. Supermicro 4U Universal GPU systems, each with eight GPUs on an HGX tray, plus liquid cooling blocks for the GPUs, the PCIe switches and two CPUs.
  • Racks. Eight servers per rack, so 64 GPUs per rack, with a coolant distribution unit in each rack. Eight racks form a 512-GPU group. At 100,000 GPUs that works out to roughly 1,560 racks, again our arithmetic.
  • Cooling. Direct-to-chip liquid cooling with quick-disconnect trays and a 1U manifold in each rack. Supermicro’s own page confirms the direct-to-chip design and the 100,000 count.
  • Fabric. NVIDIA’s release names Spectrum-X Ethernet with SN5600 switches at 800 Gb/s ports and BlueField-3 SuperNICs. NVIDIA claims 95% data throughput under its congestion control, against 60% for standard Ethernet, and no packet loss from flow collisions. Those are NVIDIA’s measurements of its own product.

Neither company names the GPU model. The dataset maps “Hopper” to the H100, which is also what ServeTheHome says.

Storage, CPU model, software stack and the exact GPU mix over time are not disclosed in anything we could fetch. The fleet has also changed. A May 2026 report of Anthropic’s compute deal puts the site at more than 220,000 NVIDIA GPUs and more than 300 MW, which is double the 2024 figure and implies the cluster grew well beyond what the vendors described.

Power: grid, turbines and a permit dispute

This is the best-documented part of the story, mostly from sources that are not xAI.

The grid. MLGW’s own public update of May 5, 2025 says the Paul Lowry Road site was receiving 150 MW: 8 MW from an existing substation and 142 MW from a new one. xAI requested a second 150 MW for a 300 MW total, with a second substation under construction at xAI’s expense. TVA, MLGW and xAI have a signed agreement requiring xAI to cut grid consumption when demand is high. TVA’s approval in November 2024 was conditioned on energy storage, recycled water and a demand-response program, and its chief executive said TVA could not refuse a load, only decide when and under what conditions it was served. TVA approved the second 150 MW on February 12, 2026.

The turbines. MLGW states it has no authority over xAI’s behind-the-meter generation. Gas turbines began appearing on site in June 2024. By September 2024 WKNO counted at least 18 portable units and reported no air permits. The Shelby County Health Department told the station it regulates generators that stay in one place more than 364 days, and that mobile classification moved jurisdiction to the EPA. The Southern Environmental Law Center (SELC) then documented 35 turbines from aerial imagery on April 9, 2025, while the mayor had said 15 were operating. On July 2, 2025 the county approved a permit for 15 Solar SMT-130 turbines totalling 247 MW. SELC says up to 35 had run without permits.

xAI told reporters little. Inside Climate News reported that xAI claimed an operational waiver covering 364 days without a permit, and that over 2,000 public comments, mostly opposed, were filed on the permit. SELC sent a notice of intent to sue on behalf of the NAACP in June 2025. Allegations of unpermitted operation are the plaintiffs’ position, not a finding.

Colossus 2 and the conflicting counts

Colossus 2 is a separate site at 5420 Tulane Road in the Whitehaven area of Memphis, with turbine generation across the state line in Southaven, Mississippi. By MLGW’s May 2025 update xAI had not submitted a final grid request for it, was not consuming significant electricity there, and had floated requests ranging from 260 MW to 1.1 GW. MLGW also noted the large substation and a TVA combined-cycle plant nearby. That is consistent with the design choice SemiAnalysis later described: power from on-site generation, not the grid.

The numbers, in order of when they appeared and who gave them:

DateFigureWhose figure
Aug-Sep 2025200 MW of cooling capacity by August 22, enough for about 110,000 GB200 NVL72 GPUs, seven 35 MW turbines runningSemiAnalysis, an outside analyst, from site analysis
Jan 2026About 2 GW of combined training compute after a third buildingMusk, quoted by Mississippi’s governor’s announcement
Jan 2026550,000 GB200 and GB300 GPUs for Colossus 2, about $18 billionIntrol, repeating Musk’s announcements
Sep 25, 2026110,000 GB200 plus 440,000 GB300, total 550,000, with 220,000 more GB300 “next week” and a 1.21 million target by year-endMusk, on X (reported by Investing.com)

Three things explain the spread. First, 110,000 is an analyst’s estimate of what 200 MW of cooling could support, not a count of installed chips. Second, 550,000 appears in January 2026 as a plan and in September 2026 as an installed figure, so the same number changed meaning. Third, nobody outside xAI has counted. Introl’s own piece is internally loose: its headline says 555,000 GPUs for the planned 2 GW site while its breakdown gives 550,000 for Colossus 2 alone and 230,000 for Colossus 1.

A sanity check using only the dataset’s power figure: 550,000 GPUs on roughly 1 GW would be about 1.8 kW all-in per GPU. Whether that 1 GW is capacity, planned build-out or measured draw is not stated anywhere we could verify.

The Colossus 2 turbines are now in federal court. The NAACP, represented by SELC and Earthjustice, sued X.AI Corp. and MZX Tech on April 14, 2026 in the Northern District of Mississippi, alleging 27 turbines operating without air permits. SELC lists a January 14, 2026 xAI application for a 41-turbine permit in Southaven and six turbines added in March 2026. SemiAnalysis described different turbine models and sizes a year earlier, so turbine figures differ by source and date. We could not find a ruling.

Yahoo Finance reports Musk saying training had already moved to Colossus 2 when Colossus 1 was rented to Anthropic in May 2026.

The supply chain

  • NVIDIA: the GPUs (Hopper at Colossus 1, GB200 and GB300 reported at Colossus 2), and the Spectrum-X switches and BlueField-3 SuperNICs.
  • Supermicro: the liquid-cooled servers, racks and coolant distribution, and the integration of the whole 100,000-GPU build.
  • xAI: operator and owner, and the party that built the substations and the on-site generation.
  • Turbines: Solar Turbines SMT-130 units appear in the Shelby County permit and the court coverage. Broader supply of turbines, rental fleets and batteries is reported by analysts but we could not verify it from primary sources.

What does not add up or is not public

  • No benchmark. We found no published HPL or MLPerf result for either site, and the dataset records none. Any “most powerful” claim, including Musk’s quoted line in NVIDIA’s release, is unmeasured.
  • Colossus 1 count. 100,000 (NVIDIA, October 2024), 200,000 as the planned doubling, and 220,000-plus (reported from the Anthropic deal) are three different snapshots, not conflicting numbers.
  • Colossus 2 count. 110,000 to 550,000, discussed above. No technical disclosure exists.
  • Power draw. Grid delivery is documented (150 MW at Colossus 1). Facility load, IT load and turbine output over time are not. A 247 MW turbine permit is a ceiling, not a measurement.
  • Start dates. July 2024 in the dataset, September 2, 2024 in Musk’s statement as reported, “122 days” with no defined start.
  • Anthropic deal date. Reports give May 6 and May 8, 2026.

What this dataset says

This section describes the dataset as published on 2 October 2026. The dataset changes as rows are added and corrected.

The dataset holds two rows. Colossus 1 is marked verified, on NVIDIA’s release and Supermicro’s page, with 100,000 H100-class GPUs recorded as reported, Spectrum-4 networking, Supermicro direct-to-chip cooling and Supermicro as integrator. It has no rank, no Rmax and no power value, with 130 MW and 150 MW noted as unresolved. Colossus 2 is marked reported, with first light in January 2026 and the GPU quantity deliberately blank and flagged as estimated because of the 110,000 to 550,000 spread. Neither system has any list appearance, which is the gap the shadow-list estimate exists to fill. The capacity table carries 150 MW for Colossus 1 and 1,000 MW for Colossus 2, both scoped “unspecified” at first phase. See data trust for how those confidence labels are assigned, power and cooling for where direct liquid cooling sits across the dataset, and supply concentration for how a single-vendor build like this one feeds the totals.

Sources

deep-diveai-clusterspower-and-coolinginterconnectsupply-chain