Chapter 06 · 12 min read
Power and cooling
A leadership-class system draws 20 to 40 megawatts, which is a small town, and the constraint on the next generation is now more often the substation and the heat rejection plant than the silicon.
Frontier draws about 24.6 MW. Aurora draws about 38.7 MW. Fugaku draws about 29.9 MW.
For scale: 30 MW is roughly the continuous electrical demand of a town of twenty to thirty thousand people. At industrial rates it is on the order of $15–25 million a year in electricity alone, which over a seven-year life is a substantial fraction of the capital cost of the machine.
Every watt that goes in comes out as heat, and it comes out in a room.
Why air stopped working
Air cooling works up to a point, and the point is set by heat capacity. Air is a poor thermal fluid: to remove a given amount of heat you must move a great deal of it, and the fan power required to move air scales roughly with the cube of velocity.
A traditional data centre hall handles 5 to 10 kW per rack comfortably, and 20 kW with effort. A rack of current accelerators is 40 to 130 kW. A Cray EX cabinet is rated to around 400 kW.
At that density, air is not merely inefficient: it is physically incapable. There is not enough air, and the fans required to move what there is would consume a significant fraction of the power they were installed to cool.
Water carries roughly 3,500 times more heat per unit volume than air. That ratio is why every large system built in the last decade is liquid cooled, and why the cooling architecture is now a first-class component of the machine rather than a facility afterthought.
How direct liquid cooling actually works
Cold plates on the hot parts. A metal block with internal channels is clamped to the processor package. Coolant flows through it, picks up the heat, and leaves. Cold plates go on the CPUs, the accelerators, the voltage regulators and increasingly the memory and the optical transceivers: anything dense enough to matter.
A secondary loop inside the racks. Coolant circulates between the cold plates and a coolant distribution unit, which is the piece of equipment that separates the clean, treated, precisely-controlled water touching the electronics from the facility water outside.
A primary loop to the outside. The CDU passes heat through a heat exchanger to the building’s water, which carries it to the heat rejection plant.
Heat rejection. Either evaporative cooling towers, which are efficient and consume water, or dry coolers, which use no water and are less effective in hot weather. This choice is now a siting decision with political consequences, because data centre water consumption has become a live public issue in several regions.
The term to know is warm-water cooling. If the coolant supplied to the racks is warm enough (30 to 45 °C rather than chilled), then the heat can often be rejected to ambient air without mechanical refrigeration at all, for most or all of the year. Eliminating the chillers eliminates the single largest non-IT power draw in the building. This is why nearly every system in this dataset that records a cooling architecture records warm-water DLC.
PUE, and what it conceals
Power Usage Effectiveness is total facility power divided by IT power. A PUE of 1.5 means that for every watt reaching a server, half a watt is spent on cooling, power conversion and lighting. A PUE of 1.0 is the unreachable ideal.
LUMI reports a PUE around 1.03. That is extraordinary, and it is a product of place as much as engineering: Kajaani is in Finland, the ambient air is cold for most of the year, the electricity is hydroelectric, and the waste heat is sold into the municipal district heating network, where it displaces heat that would otherwise have been generated by burning something.
Three cautions about the metric, because it is quoted far more often than it is understood:
It is a ratio, not an efficiency. A facility that makes its servers less efficient (running fans harder inside the chassis, where the power counts as IT load) improves its PUE. The metric rewards moving power draw across the boundary as readily as reducing it.
It says nothing about the carbon. A PUE of 1.05 on coal-fired power is worse for the climate than 1.4 on hydroelectric. PUE measures overhead, not impact.
It says nothing about water. A facility using evaporative cooling can achieve a superb PUE while consuming millions of litres a year. There is a separate metric, WUE, which is reported far less often.
This site records PUE where an operator publishes it and leaves the field blank otherwise: blank meaning we found no source, not zero. That distinction matters if you are aggregating across sites.
Heat reuse
If the coolant leaves the machine at 45 °C, that is a usable temperature for building heating. Several European sites do exactly this: LUMI feeds the Kajaani district heating network, and similar arrangements exist at other Nordic and German facilities.
It is genuinely good engineering and it has a hard limit: it requires a heat customer, physically adjacent, whose demand is seasonal in a way the machine’s output is not. It works in northern Europe where district heating networks are common and winters are long. It works much less well in Texas or Arizona.
The constraint that now binds
For most of this field’s history, the limit on building a bigger machine was silicon: what the vendor could deliver, when.
That is no longer reliably true. The current long poles are frequently:
Utility interconnection. Getting 50 or 100 MW delivered to a site requires a substation, transmission capacity and a utility willing to commit. Lead times on medium-voltage switchgear and large transformers have run to years, and in several regions the utility queue is now measured in multi-year increments.
Grid capacity. In some markets there is simply no available capacity at the required scale without new transmission, which is a decade-scale infrastructure problem rather than a procurement one.
Water. Where evaporative cooling is used, water rights and local opposition are real constraints, and increasingly a reason to accept the efficiency penalty of dry coolers.
Construction. The building, the electrical rooms, the mechanical plant. For a purpose-built facility this is measured in years.
This is precisely the layer that a system-centric ranking stops short of: it names the processor and the fabric and says nothing about the substation. It is why this site carries sites as first-class entities with utility, substation capacity, PUE and cooling architecture, even though those fields are sparsely populated: the gap in the public record is itself the finding.
What the numbers look like
For a current leadership-class system:
- IT load: 20–40 MW
- Facility total: IT load × PUE, so 21–48 MW
- Rack density: 40–400 kW
- Coolant supply temperature: 30–45 °C for warm-water designs
- Flow: hundreds to thousands of litres per minute per cabinet row
- Annual energy: 175–350 GWh
- Annual electricity cost: $10–35 million depending on tariff
The power and cooling analysis (installed megawatts by cooling architecture and by vendor) is one of the two remaining derived pages on this site, and it is blocked on exactly the coverage problem this chapter describes: cooling architecture is recorded for only a handful of systems, because operators publish it far less consistently than they publish core counts.