Supercomputers/Systems/ByteDance MegaScale Cluster
China · operational · ai training
ByteDance MegaScale Cluster
Also known as ByteDance 12, 288-GPU LLM training cluster
Operated by ByteDance .
Phase 1 seed dataset, compiled by hand. These rows were built from public operator, laboratory and vendor sources. A mechanical second-reader pass has since fetched every cited source: 1092 of 1153 systems have a readable citation that names them, and 40 are genuinely weakly sourced. Every claim carries its source and a confidence tier. Treat anything below verified as a lead, not a citation.
Measured performance
- Rmax
- -
- Rpeak
- -
- Rmax ÷ Rpeak
- -
- Cores
- -
- Power
- -
- Per watt
- - GF/W
Figures are as last publicly reported for the configuration described below, not a live measurement. Where a system was upgraded in place, the post-upgrade configuration is shown and the earlier one appears in the timeline.
What this machine is made of
One row per supplier relationship. “Supplier at build” is the company that shipped the part at the time; where that company has since been acquired, the parent it rolls up to today is shown beside it. That distinction is what makes ticker-level aggregation possible across a thirty-year dataset.
| Role | Part | Supplier at build | Quantity | Confidence | Source |
|---|---|---|---|---|---|
| Integrator | - ByteDance built and operates the cluster | ByteDance | - | Reported | arxiv.org |
| Accelerator | - More than 10,000 NVIDIA Ampere GPUs, 12,288 used in the 175B MFU experiments | NVIDIA | 12,288 reported | Reported | arxiv.org |
| Interconnect | - Broadcom Tomahawk 4 switches with 64 x 400 Gbps ports in a three-layer CLOS network | Broadcom | - | Reported | arxiv.org |
AI datacenter metrics
Figures beyond an accelerator count, each with who stated it. What can and cannot be compared across AI datacenters.
| Metric | Value | As of | Basis | Note | Source |
|---|---|---|---|---|---|
| Accelerators | 12,288 accelerators | 2023-09 | Operator stated | MegaScale paper: largest production LLM cluster exceeds 10,000 NVIDIA Ampere GPUs; 12,288 GPUs used for the 175B MFU run. | arxiv.org |
The read
ByteDance's production LLM training cluster described in the MegaScale paper (ByteDance and Peking University, arXiv 2402.15627). The paper says that as of September 2023 the largest AI cluster in ByteDance production for LLM training contained more than 10,000 NVIDIA Ampere GPUs, and reports a 55.2 percent MFU training a 175B model on 12,288 GPUs. The data-center network uses Broadcom Tomahawk 4 switches (64 x 400 Gbps ports) in a three-layer CLOS topology joining more than 10,000 GPUs. The paper does not name a site, GPU SKU or power. ByteDance said it was also building clusters of NVIDIA Hopper GPUs.
Timeline
No dated events recorded.
Change history
- 2026-10-09 Added ByteDance MegaScale Cluster
Source check
We fetched this system's own citations and recorded whether each page actually mentions it. This is a corroboration signal, not a fact check, and it is published so you can see how well the sourcing holds up rather than take it on trust.
Every source cited on this page was readable, and none of them mentions this system. They are cited at page granularity, normally because no stable deep link could be found. That is a real weakness in this row and it is why the system carries a confidence tier below verified. See how this check works.
| Cited source | Result | Found on the page |
|---|---|---|
| arxiv.org | does not name this system | - |
Checked 2026-10-09 by pnpm verify. Re-run it and the table changes with the web.
Further reading
Inside a node
Why accelerators exist, what the memory hierarchy costs, and how the CPU and accelerator merged onto one package.
The interconnect
Topologies, why latency and tail behaviour matter more than bandwidth, and how the fabric market consolidated.
Operating a supercomputer
Acceptance testing, why "installed" and "in production" are six months apart, and the machine lifecycle.