COMPUTE ATLAS Supercomputer supply-chain graph
1153 systems 482 sites Sourced data

Blog/2026-10-02

· Compute Atlas

LineShine: a CPU-only 2.2 exaflop machine

LineShine tops the June 2026 TOP500 at 2.198 EFlop/s with no accelerators. The headline is documented; the chip's designer, fab and memory suppliers are not.


LineShine took first place on the June 2026 TOP500 list with 2.198 EFlop/s on High Performance Linpack, against a theoretical peak of 2.736 EFlop/s, using 13,789,440 cores and drawing 42.22 MW. It sits at the National Supercomputing Centre in Shenzhen (NSCS), was built by the Shenzhen Cloud Computing Center, and is the first system on the list to pass two exaflops using CPUs alone. TOP500’s own description is of a “previously unlisted system” built around a custom Chinese processor, the LX2, and a proprietary interconnect, LingQi.

The one thing most worth knowing is not the ranking. It is the gap between what is established and what is claimed. The benchmark result, the core count and the headline architecture are documented by TOP500 and by a technical paper from the people who run the machine. The two claims that make LineShine politically interesting, that a Chinese organisation designed the chip and that the whole stack is domestically sourced, rest on the operator’s own statements and on trade-press inference. Nobody has published who designed the LX2, who fabricated it, or who made its memory.

How it came to exist

The public timeline is short and thin on money.

  • Installation year 2025. The TOP500 system page gives 2025 as the installation year. That is the only dated statement about when hardware went in, and it is a single source.
  • 17 April 2026. An arXiv paper on training machine-learning interatomic potentials reports full-machine runs on LineShine and thanks NSCC-SZ for technical support. Its authors are at the Chinese Academy of Sciences, the University of the Chinese Academy of Sciences and Sun Yat-Sen University, not the centre itself.
  • 24 April 2026. NSCS held an event on domestic computing applications in Shenzhen. The centre’s report of it is dated 27 April, and names its director as the system’s chief designer. Wikipedia’s announcement date of 27 April appears to be the page date, not the event date.
  • 22 to 25 May 2026. Yutong Lu presented the architecture at the HACI 2026 workshop in Shenzhen, according to The Next Platform. The talk was not released publicly; the slides circulated through researchers outside China, and that outlet’s deep dive is built on them.
  • June 2026. The TOP500 list placed it first, ending El Capitan’s run, and made it the fifth system over one exaflop.

Who paid, what it cost and what the procurement contract looked like are not in any source we could fetch. The dataset’s first-operational date, April 2026, is the date of that event, and its announcement date is 24 April 2026; both were added while researching this piece.

NSCS has form here. Nebulae, also operated from Shenzhen, was a Sugon system with Intel Xeons and NVIDIA Tesla GPUs that placed second on the June 2010 list. Sixteen years on, the same centre has a machine with no foreign processor and no GPU. The centre was added to the US Entity List in April 2021, alongside six other Chinese supercomputing entities, on the stated basis that its procurement of US-origin items was contrary to US national security and foreign policy interests, as reported by SiliconANGLE and a trade-law firm summarising the Bureau of Industry and Security notice. The Federal Register notice itself would not load for us.

What is actually inside

The architecture figures below come from the arXiv paper unless marked otherwise, and from TOP500 for the benchmark configuration.

Processor. Each LX2 is an Armv9 part with SVE2 vector units and SME matrix units. It has 304 cores at 1.55 GHz in two compute dies of 152 cores, and each die is split into four NUMA domains of 38 cores. The paper quotes up to 60.3 TFlop/s FP64 and 120.6 TFlop/s FP32 per processor. Working back from the clock, 60.3 TFlop/s over 304 cores at 1.55 GHz is 128 floating-point operations per core per cycle.

Memory. The paper gives each LX2 eight on-package HBM stacks totalling 32 GB at 4 TB/s, plus 128 GB of off-package DDR per die, so 256 GB per socket. The Next Platform instead reports 64 GB of HBM per socket at 8 TB/s. The paper is the better source, and the arithmetic backs it: Wikipedia, relaying an HPCwire article we could not fetch, gives 1.4515 PB of HBM in total, which is exactly 22,680 nodes times two sockets times 32 GB. The same relay gives 11.6121 PB of DDR, which is 256 GB per socket. Next Platform suggests the DDR may be LPDDR5X from ChangXin Memory Technologies. That is speculation, not disclosure.

Nodes. Two sockets per node. Per ServeTheHome (STH), eight nodes make a blade, 16 blades a frame and two frames a cabinet, so 256 nodes per cabinet, which puts the machine at roughly 89 to 90 cabinets.

Interconnect. LingQi is a dual-plane, multi-rail fat-tree giving 1.6 Tb/s per node (paper). Next Platform adds four layers, an optical cross-link layer in 32 network frames, over 3.5 Pb/s of bisection bandwidth and 1.07 microseconds across one hop; these are single-source figures. It also says the fabric may be an InfiniBand variant or a modified Ethernet, and that is its guess. Its parenthetical of two 400 Gb/s ports does not add to 1.6 Tb/s; STH’s 800 Gb/s on-chip NIC per CPU does. The TOP500 entry says only “LingQi”.

Software. TOP500 lists Kylin OS, the Lclang 1.0.0 compiler, a local OpenBLAS build and OpenMPI customised for LingQi. The paper’s training runs used PyTorch 2.10.0, a vendor BLAS called KML and OpenMPI 4.1.6rc4.

Power and cooling. TOP500 gives 42,220 kW and 52.07 GFlop/s per watt. The NSCS report cites full liquid cooling as a core design feature, and claims, as we read the Chinese, the largest centralised liquid-cooling deployment in the world, without figures. Next Platform and STH both give 690 W per LX2. Taking that figure at face value, 22,680 nodes times two sockets times 690 W is about 31.3 MW, or roughly three quarters of the 42.22 MW total, leaving about 11 MW for the network, storage, cooling pumps and everything else. That is our arithmetic on a reported figure, not a disclosed power budget.

Storage. The NSCS report mentions exascale-class storage but gives no capacity. The only figure we found, 650 PB, is Wikipedia’s, relayed from HPCwire.

How it performs and what it does

The TOP500 entry gives Rmax 2,198.40 PFlop/s and Rpeak 2,735.82 PFlop/s, an HPL efficiency of about 80.4%. That is high for a machine of this size. TOP500 puts it more than 20% ahead of the second system, El Capitan, at 1.809 EFlop/s.

Two details complicate the node count. TOP500’s cores equal 22,680 nodes times two sockets times 304 cores. Rpeak divided by 22,680 is 120.6 TFlop/s, twice the paper’s per-processor figure. But the paper calls 20,480 nodes “the full machine”, and its percentages of peak fit that count. Next Platform notes the 2,200-node difference and does not resolve it. Either the machine grew between April and June, or the paper’s figure is a partition. Nobody has said which.

Benchmarks beyond HPL pull in different directions. On HPCG, which stresses memory and the network, LineShine reached 22.00 PFlop/s, ahead of El Capitan’s 17.41. On HPL-MxP, which allows low precision and rewards tensor hardware, it reached 7.92 EFlop/s, fourth, a 3.6x speedup over its own HPL figure; El Capitan reached 16.7 EFlop/s, a 9.2x speedup. That is what CPU-only buys and costs: strong double-precision and memory behaviour, and a smaller multiplier when a workload can use low-precision matrix hardware. A Digitimes headline says its AI lead is unclear, but the article is paywalled and we could not read the reasoning.

On workloads, the evidence is the arXiv paper and the NSCS report. The paper trains a billion-parameter interatomic-potential model and reports 1.20 EFlop/s peak in single precision on LineShine, 24.4% of theoretical peak, with 90.3% weak-scaling efficiency, but sustained performance over a 1,000-step run was 1,033.3 PFlop/s. The same paper also ran on a second, GPU-accelerated Chinese exascale machine it calls CNIS, with 5,632 nodes and eight accelerators each, whose accelerator vendor it does not name. So LineShine is one Chinese design choice, not the only one. The NSCS report lists application claims, among them global 1 km Earth-system simulation, 100-million-atom first-principles calculations at 81% parallel scalability, and a seismic-imaging code at 1.88 times an NVIDIA A100. These are the centre’s own statements, not independently reproduced.

The supply chain

Here the record is mostly gaps. The dataset names three companies on this system, and every other supplier role is unknown.

  • Operator: NSCS.
  • Integrator: Shenzhen Cloud Computing Center, credited by TOP500 as the builder.
  • Operating system: Kylin OS, per the TOP500 entry. KylinSoft is in the dataset as KylinSoft, and the Kylin Linux part exists, but the dataset does not currently link LineShine to it.
  • Processor: The LingKun LX2 design. The Next Platform says it was designed by NSCS “presumably” with Huawei’s HiSilicon division, which is its inference; STH says it is thought to be Huawei’s. TOP500 and the NSCS report name no designer, and the paper credits NSCC-SZ with developing the system. Huawei is in the dataset, but we do not link it to this system, because no source states it.
  • Foundry: Unconfirmed. Next Platform thinks a SMIC 7 nm process at the N+3 refinement is highly likely. That is its inference.
  • HBM and DRAM: Unnamed. This is the part of any claim of full self-sufficiency that is hardest to check from outside.
  • Interconnect: LingQi, proprietary, with no named silicon vendor.

What does not add up or is not public

Sources conflict in several places:

  • Node count: 22,680 (TOP500 core count, Next Platform) versus 20,480 (the paper, also quoted by Next Platform).
  • HBM per socket: 32 GB (paper, STH, and the arithmetic above) versus 64 GB (Next Platform).
  • Network ports: 1.6 Tb/s per node, but the Next Platform’s port description does not add up to it.
  • China’s history on the list: Chips and Cheese calls this China’s first TOP500 submission in nine years, while hpc.rs counts 30 Chinese systems on the current list. We did not reconcile those, and the difference may be about which entities count as a submitter.

What the operator has not said: who funded it, what it cost, what the full node count is, who designed and fabricated the chip, the memory suppliers, the storage capacity, the cooling design, and the software stack in any detail. The centre’s own phrase for the stack is “全栈自主可控”, roughly fully self-controlled across the stack. That is an assertion by the operator. It is consistent with what TOP500 lists, but nothing outside the centre can verify that no foreign-origin component, whether a memory die, an EDA tool or a manufacturing tool, appears anywhere in the supply chain.

For the export-control question, the careful position is this. The centre has been on the Entity List since April 2021, and the machine contains no NVIDIA or AMD accelerator and no Intel or AMD CPU on the evidence available. It is therefore a domestic-substitution result in design terms. Whether it is one in fabrication terms depends on the foundry question, and that is inference, not record. The Next Platform’s framing is that China had no choice but to build its own.

What this dataset says

This section describes the dataset as published on 2 October 2026. The dataset changes as rows are added and corrected.

The LineShine row holds Rmax 2,198.4 PFlop/s, Rpeak 2,735.82 PFlop/s, 13,789,440 cores and 42,220 kW, classed as HPC, with TOP500 rank 1 in the June 2026 list as its only list appearance. All of that matches the TOP500 entry. The first-operational date is April 2026, the month the centre presented the machine, although TOP500 gives an installation year of 2025, and there are no capacity claims. Four component edges are recorded: the LX2 processor, the LingQi fabric, Kylin OS and the integrator, all tagged “reported”, while the system row itself is tagged “verified”. The verification table marks the sources “partial”.

Two consequences are worth stating. First, the domestic substitution index scores LineShine as fully domestic: Chinese CPU, interconnect and operating system, with no accelerator role recorded. Second, the index counts a system only from its first-operational date, which was blank until this piece was researched, so until 2 October 2026 LineShine did not move the China line; it now does. The index also measures supplier nationality, not fabrication. Fugaku is the nearest precedent for a no-accelerator design, and Sunway TaihuLight was the last Chinese number one. The row’s notes originally said “since 2016”; TOP500 says 2017, and the row now follows TOP500. See also accelerator share and interconnect share.

Sources

deep-divehpcexascalesupply-chainaccelerators