AI datacenters/Accelerators
AI datacenters
AI accelerators compared
An accelerator count says little unless you know which accelerator. This table puts 31 of the chips that fill AI datacenters on one footing: dense peak FLOPS by precision, memory and bandwidth, power, and the size of the scale-up domain they are sold in. Vendors headline sparse figures that are twice the dense ones and are rarely reached on real workloads, so every number here is dense, and a figure halved from a sparse headline is marked as derived. All of these are peaks from datasheets, not measured performance.
- NVIDIA
- AMD
- Other vendors
Table view: Dense peak TFLOPS per accelerator by release year
| Accelerator | Vendor | Released | Dense TFLOPS |
|---|---|---|---|
| AMD Instinct MI355X | AMD | 2026 | 5033.2 |
| Microsoft Maia 200 | Microsoft | 2026 | 5000 |
| AWS Trainium3 | Amazon | 2025 | 2517 |
| Google TPU7x (Ironwood) | Alphabet | 2025 | 4614 |
| NVIDIA B300 SXM | NVIDIA | 2025 | 4500 |
| NVIDIA GB300 NVL72 | NVIDIA | 2025 | 5000 |
| AMD Instinct MI325X | AMD | 2024 | 2614.9 |
| AWS Trainium2 | Amazon | 2024 | 1299 |
| Cerebras WSE-3 | Cerebras | 2024 | 62500 |
| Google TPU v6e (Trillium) | Alphabet | 2024 | 918 |
| Intel Gaudi 3 | Intel | 2024 | 1678 |
| Meta MTIA v2 | Meta Platforms | 2024 | 177 |
| NVIDIA B200 SXM | NVIDIA | 2024 | 4500 |
| NVIDIA GB200 NVL72 | NVIDIA | 2024 | 5000 |
| NVIDIA H200 SXM | NVIDIA | 2024 | 1979 |
| AMD Instinct MI300X | AMD | 2023 | 2614.9 |
| Google TPU v5e | Alphabet | 2023 | 197 |
| Google TPU v5p | Alphabet | 2023 | 459 |
| NVIDIA H100 SXM5 | NVIDIA | 2022 | 1979 |
| AMD Instinct MI250X | AMD | 2021 | 383 |
| Google TPU v4 | Alphabet | 2020 | 275 |
| NVIDIA A100 SXM4 40GB | NVIDIA | 2020 | 312 |
| NVIDIA A100 SXM4 80GB | NVIDIA | 2020 | 312 |
The accelerators
| Accelerator | Vendor | Released | BF16 | FP8 | FP4 | FP64 | Memory (GB) | Bandwidth (TB/s) | TDP (W) | FP8 per kW | Scale-up |
|---|---|---|---|---|---|---|---|---|---|---|---|
| AMD Instinct MI355X | AMD | 2026 | 2,517 | 5,033 | 10,066 | 78.6 | 288 | 8 | 1,400 | 3,595 | 8 |
| Google TPU 8i | Alphabet | 2026 | - | - | 10,100 | - | 288 | 8.6 | - | - | 1,024 |
| Google TPU 8t | Alphabet | 2026 | - | - | 12,600 | - | 216 | 6.5 | - | - | 9,600 |
| Microsoft Maia 200 | Microsoft | 2026 | - | 5,000 | 10,000 | - | 216 | 7 | 750 | 6,667 | - |
| AWS Trainium3 | Amazon | 2025 | 671 | 2,517 | 2,517 | - | 144 | 4.9 | - | - | 144 |
| Google TPU7x (Ironwood) | Alphabet | 2025 | 2,307 | 4,614 | - | - | 192 | 7.4 | - | - | 9,216 |
| Huawei CloudMatrix 384 | Huawei | 2025 | - | - | - | - | - | - | - | - | 384 |
| NVIDIA B300 SXM | NVIDIA | 2025 | 2,250 | 4,500 | 13,500 | 1.2 | 270 | 7.7 | 1,100 | 4,091 | 8 |
| NVIDIA GB300 NVL72 | NVIDIA | 2025 | 2,500 | 5,000 | 15,000 | 1.3 | 279 | 8 | 1,400 | 3,571 | 72 |
| AMD Instinct MI325X | AMD | 2024 | 1,307 | 2,615 | - | 163.4 | 256 | 6 | 1,000 | 2,615 | - |
| AWS Trainium2 | Amazon | 2024 | 667 | 1,299 | - | - | 96 | 2.9 | - | - | 64 |
| Cerebras WSE-3 | Cerebras | 2024 | 62,500 | - | - | - | 44 | 21,000 | 23,000 | - | - |
| Google TPU v6e (Trillium) | Alphabet | 2024 | 918 | - | - | - | 32 | 1.6 | - | - | 256 |
| Intel Gaudi 3 | Intel | 2024 | 1,678 | 1,678 | - | - | 128 | 3.7 | 900 | 1,864 | - |
| Meta MTIA v2 | Meta Platforms | 2024 | 177 | - | - | - | 128 | 0.2 | 90 | - | - |
| NVIDIA B200 SXM | NVIDIA | 2024 | 2,250 | 4,500 | 9,000 | 37 | 180 | 7.7 | 1,000 | 4,500 | 8 |
| NVIDIA GB200 NVL72 | NVIDIA | 2024 | 2,500 | 5,000 | 10,000 | 40 | 186 | 8 | 1,200 | 4,167 | 72 |
| NVIDIA H200 SXM | NVIDIA | 2024 | 990 | 1,979 | - | 67 | 141 | 4.8 | 700 | 2,827 | 8 |
| AMD Instinct MI300X | AMD | 2023 | 1,307 | 2,615 | - | 163.4 | 192 | 5.3 | 750 | 3,487 | 8 |
| Google TPU v5e | Alphabet | 2023 | 197 | - | - | - | 16 | - | - | - | 256 |
| Google TPU v5p | Alphabet | 2023 | 459 | - | - | - | 95 | 2.8 | - | - | 8,960 |
| Microsoft Maia 100 | Microsoft | 2023 | - | - | - | - | 64 | 1.8 | 700 | - | - |
| NVIDIA H100 SXM5 | NVIDIA | 2022 | 990 | 1,979 | - | 67 | 80 | 3.4 | 700 | 2,827 | 8 |
| AMD Instinct MI250X | AMD | 2021 | 383 | - | - | 95.7 | 128 | 3.2 | 500 | - | - |
| Google TPU v4 | Alphabet | 2020 | 275 | - | - | - | 32 | 1.2 | - | - | 4,096 |
| NVIDIA A100 SXM4 40GB | NVIDIA | 2020 | 312 | - | - | 19.5 | 40 | 1.6 | 400 | - | 8 |
| NVIDIA A100 SXM4 80GB | NVIDIA | 2020 | 312 | - | - | 19.5 | 80 | 2 | 400 | - | 8 |
| AMD Instinct MI430X | AMD | - | - | - | - | 288 | 432 | 23.3 | - | - | - |
| AMD Instinct MI455X | AMD | - | 5,000 | - | - | 5 | 432 | 23.3 | - | - | - |
| Huawei Ascend 910C | Huawei | - | 781 | - | - | - | - | 3.2 | - | - | - |
| NVIDIA Vera Rubin | NVIDIA | - | 4,000 | 17,500 | 35,000 | 33 | 288 | 19.2 | - | - | 72 |
TFLOPS are dense peaks per accelerator package. FP8 per kW divides the dense FP8 peak by the accelerator's own TDP and ignores the rest of the node, network and cooling. Scale-up is the number of accelerators in one high-bandwidth domain (72 for GB200 NVL72).
Derived peak compute of the biggest clusters
A cluster's accelerator count times each chip's dense peak is the one compute figure that can be set against another cluster of a different chip. It is a ceiling, not a benchmark result: real training reaches a fraction of it (Meta reported 38 to 43 percent on a 16,384-GPU H100 run), and it is only shown where every accelerator in the system has both a recorded count and a recorded peak. Counts for campuses that are not yet running are plans or estimates, so they are listed apart.
Running
| System | Status | Accelerators | Precision | Derived peak (EFLOPS) |
|---|---|---|---|---|
| AWS Madison Mega Site (Canton) | operational | 326,600 | FP8 | 424 |
| Meta Rosemount (Minnesota) | operational | 85,400 | FP8 | 384 |
| AWS Ridgeland Campus | operational | 261,200 | FP8 | 339 |
| xAI Colossus | operational | 100,000 | FP8 | 198 |
| Together AI / Hypertec Cloud GB200 Cluster | operational | 36,000 | FP8 | 180 |
| Tesla Cortex | operational | 66,000 | FP8 | 131 |
| Mistral Campus AI (Bruyeres-le-Chatel) | operational | 18,000 | FP8 | 90 |
| SINES Data Campus GB300 Deployment | operational | 12,600 | FP8 | 63 |
| Meta GenAI cluster (RoCE) | operational | 24,576 | FP8 | 49 |
| Meta GenAI cluster (InfiniBand) | operational | 24,576 | FP8 | 49 |
| Industrial AI Cloud | operational | 10,000 | FP8 | 45 |
| Google TPU7x (Ironwood) pod | operational | 9,216 | FP8 | 43 |
| Taiwan AI Factory | operational | 7,000 | FP8 | 35 |
| Firebird Armenian AI Factory (DC-1) | operational | 6,144 | FP8 | 28 |
| Nscale Verne Iceland Cluster | operational | 4,600 | FP8 | 23 |
| TensorWave Tucson MI325X Cluster | operational | 8,192 | FP8 | 21 |
| Cerebras Oklahoma City | operational | 300 | BF16 | 19 |
| Naver B200 4K Cluster | operational | 4,000 | FP8 | 18 |
| Nebius Modiin Cluster | operational | 4,000 | FP8 | 18 |
| IBM Blue Vela | operational | 5,000 | FP8 | 9.89 |
| SDS-AI | operational | 2,032 | FP8 | 9.14 |
| Google TPU v4 ML Hub (Oklahoma) | operational | 32,768 | BF16 | 9.01 |
| Shenzhen 10,000-Card Ascend Cluster | operational | 10,000 | BF16 | 7.81 |
| CoreWeave AI cluster for Inflection AI | operational | 3,584 | FP8 | 7.09 |
| Databricks DBRX Training Cluster | operational | 3,072 | FP8 | 6.08 |
| E2E Chennai B200 Cluster | operational | 1,024 | FP8 | 4.61 |
| Haein Cluster | operational | 1,000 | FP8 | 4.50 |
| Google TPU v5p pod | operational | 8,960 | BF16 | 4.11 |
| Israel-1 | operational | 2,048 | FP8 | 4.05 |
| Condor Galaxy 3 | operational | 64 | BF16 | 4.00 |
| HiPerGator AI | operational | 504 | FP8 | 2.27 |
| CZI GPU Cluster | operational | 1,024 | FP8 | 2.03 |
| Ubilink H100 Cluster | operational | 1,024 | FP8 | 2.03 |
| Pre-Eos | operational | 1,024 | FP8 | 2.03 |
| Nabuchodonosor | operational | 1,016 | FP8 | 2.01 |
| SAKURAONE | operational | 800 | FP8 | 1.58 |
| Google TPU v4 supercomputer | operational | 4,096 | BF16 | 1.13 |
| BioHive-2 | operational | 504 | FP8 | 1.00 |
| DAIS | operational | 264 | FP8 | 0.85 |
| AI-Farabium | operational | 400 | FP8 | 0.79 |
Planned or under construction
| System | Status | Accelerators (planned) | Precision | Derived peak (EFLOPS) |
|---|---|---|---|---|
| Microsoft Azure AI datacenter (Narvik, Norway) | under construction | 30,000 | FP8 | 525 |
| Solstice | planned | 100,000 | FP8 | 450 |
| Microsoft Azure AI supercomputer (Loughton, UK) | under construction | 23,040 | FP8 | 115 |
| Project Ceiba | under construction | 20,736 | FP8 | 93 |
| Equinox | planned | 10,000 | FP8 | 45 |
Sources
| Accelerator | Figure | Value | Basis | Source | Note |
|---|---|---|---|---|---|
| NVIDIA A100 SXM4 40GB | bf16 tflops | 312 TFLOPS | Vendor stated (dense) | nvidia.comepoch.ai | Dense BF16 tensor core; datasheet gives 312 dense and 624 with sparsity. Epoch AI also lists 312. |
| NVIDIA A100 SXM4 40GB | fp64 tflops | 19.5 TFLOPS | Vendor stated (dense) | nvidia.comimages.nvidia.com | FP64 Tensor Core peak; non-tensor FP64 is 9.7 TFLOPS. |
| NVIDIA A100 SXM4 40GB | memory gb | 40 GB | Vendor stated (dense) | nvidia.comaws.amazon.com | HBM2. AWS P4d lists 320 GB across eight A100 GPUs. |
| NVIDIA A100 SXM4 40GB | memory bw tbs | 1.555 TB/s | Vendor stated (dense) | nvidia.comepoch.ai | Stated as 1,555 GB/s. Epoch AI lists 1.56 TB/s. |
| NVIDIA A100 SXM4 40GB | tdp w | 400 W | Vendor stated (dense) | nvidia.comepoch.ai | SXM4 module maximum thermal design power. |
| NVIDIA A100 SXM4 40GB | scaleup domain accelerators | 8 accelerators | Vendor stated (dense) | nvidia.comnvidia.com | NVSwitch domain of the 8-GPU DGX A100 baseboard; HGX A100 is also offered with 4 or 16 GPUs. |
| NVIDIA A100 SXM4 80GB | bf16 tflops | 312 TFLOPS | Vendor stated (dense) | nvidia.comepoch.ai | Dense BF16 tensor core; datasheet gives 312 dense and 624 with sparsity. Epoch AI also lists 312. |
| NVIDIA A100 SXM4 80GB | fp64 tflops | 19.5 TFLOPS | Vendor stated (dense) | nvidia.comimages.nvidia.com | FP64 Tensor Core peak; non-tensor FP64 is 9.7 TFLOPS. |
| NVIDIA A100 SXM4 80GB | memory gb | 80 GB | Vendor stated (dense) | nvidia.comaws.amazon.com | HBM2e. AWS P4de lists 640 GB across eight A100 80GB GPUs. |
| NVIDIA A100 SXM4 80GB | memory bw tbs | 2.039 TB/s | Vendor stated (dense) | nvidia.comepoch.ai | Stated as 2,039 GB/s. Epoch AI lists the same. |
| NVIDIA A100 SXM4 80GB | tdp w | 400 W | Vendor stated (dense) | nvidia.comepoch.ai | 400 W standard; the HGX custom thermal solution SKU supports up to 500 W. |
| NVIDIA A100 SXM4 80GB | scaleup domain accelerators | 8 accelerators | Vendor stated (dense) | nvidia.comnvidia.com | NVSwitch domain of the 8-GPU DGX A100 baseboard; HGX A100 is also offered with 4 or 16 GPUs. |
| NVIDIA H100 SXM5 | bf16 tflops | 989.5 TFLOPS | Derived from vendor sparse figure | nvidia.comepoch.ai | Vendor lists a sparse figure; halved for dense. Page lists 1,979 with sparsity. Epoch AI lists 989.4. |
| NVIDIA H100 SXM5 | fp8 tflops | 1,979 TFLOPS | Derived from vendor sparse figure | nvidia.comepoch.ai | Vendor lists a sparse figure; halved for dense. Page lists 3,958 with sparsity. Epoch AI lists 1,979. |
| NVIDIA H100 SXM5 | fp64 tflops | 67 TFLOPS | Vendor stated (dense) | nvidia.comnvidia.com | FP64 Tensor Core peak; vector FP64 is 34 TFLOPS. |
| NVIDIA H100 SXM5 | memory gb | 80 GB | Vendor stated (dense) | nvidia.comaws.amazon.com | HBM3. AWS P5 lists eight H100 GPUs with 640 GB. |
| NVIDIA H100 SXM5 | memory bw tbs | 3.35 TB/s | Vendor stated (dense) | nvidia.comepoch.ai | HBM3 bandwidth as stated by NVIDIA. |
| NVIDIA H100 SXM5 | tdp w | 700 W | Vendor stated (dense) | nvidia.comepoch.ai | Stated as up to 700 W, configurable. |
| NVIDIA H100 SXM5 | scaleup domain accelerators | 8 accelerators | Vendor stated (dense) | nvidia.com | NVSwitch domain of the 8-GPU DGX H100 / HGX H100 baseboard (HGX also offered with 4 GPUs). |
| NVIDIA H200 SXM | bf16 tflops | 989.5 TFLOPS | Derived from vendor sparse figure | nvidia.comepoch.ai | Vendor lists a sparse figure; halved for dense. Page lists 1,979 with sparsity. Epoch AI lists 989.5. |
| NVIDIA H200 SXM | fp8 tflops | 1,979 TFLOPS | Derived from vendor sparse figure | nvidia.comepoch.ai | Vendor lists a sparse figure; halved for dense. Page lists 3,958 with sparsity. Epoch AI lists 1,979. |
| NVIDIA H200 SXM | fp64 tflops | 67 TFLOPS | Vendor stated (dense) | nvidia.comdam-cdn.nvd.orangelogic.com | FP64 Tensor Core peak; vector FP64 is 34 TFLOPS. |
| NVIDIA H200 SXM | memory gb | 141 GB | Vendor stated (dense) | nvidia.comaws.amazon.com | HBM3e. AWS P5e lists eight H200 GPUs with 1,128 GB. |
| NVIDIA H200 SXM | memory bw tbs | 4.8 TB/s | Vendor stated (dense) | nvidia.comepoch.ai | HBM3e bandwidth as stated by NVIDIA. |
| NVIDIA H200 SXM | tdp w | 700 W | Vendor stated (dense) | nvidia.comepoch.ai | Stated as up to 700 W, configurable (NVL PCIe variant up to 600 W). |
| NVIDIA H200 SXM | scaleup domain accelerators | 8 accelerators | Vendor stated (dense) | nvidia.comdam-cdn.nvd.orangelogic.com | NVSwitch domain of an 8-GPU HGX H200 baseboard (HGX also offered with 4 GPUs). |
| NVIDIA B200 SXM | bf16 tflops | 2,250 TFLOPS | Derived from vendor sparse figure | dam-cdn.nvd.orangelogic.comepoch.ai | Vendor lists a sparse figure; halved for dense. HGX B200 datasheet lists 4.5 PFLOPS sparse. Epoch AI lists 2,250. |
| NVIDIA B200 SXM | fp8 tflops | 4,500 TFLOPS | Derived from vendor sparse figure | dam-cdn.nvd.orangelogic.comepoch.ai | Vendor lists a sparse figure; halved for dense. HGX B200 datasheet lists 9 PFLOPS sparse. Epoch AI lists 4,500. |
| NVIDIA B200 SXM | fp4 tflops | 9,000 TFLOPS | Derived from vendor sparse figure | dam-cdn.nvd.orangelogic.comepoch.ai | Vendor lists a sparse figure; halved for dense. HGX B200 datasheet lists 18 PFLOPS sparse (72 dense per 8 GPUs). Epoch AI lists 9,000. |
| NVIDIA B200 SXM | fp64 tflops | 37 TFLOPS | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comnvidia.com | HGX B200 FP64 / FP64 Tensor Core; HGX page gives 296 TFLOPS per 8 GPUs. Epoch AI lists 31. |
| NVIDIA B200 SXM | memory gb | 180 GB | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | HGX B200 HBM3E per GPU; NVIDIA cites 192 GB as the Blackwell maximum capacity. |
| NVIDIA B200 SXM | memory bw tbs | 7.7 TB/s | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | HGX B200 per-GPU HBM3E bandwidth; the maximum Blackwell spec is quoted as 8 TB/s. |
| NVIDIA B200 SXM | tdp w | 1,000 W | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | HGX B200 configurable up to 1,000 W (GB200 variant up to 1,200 W). |
| NVIDIA B200 SXM | scaleup domain accelerators | 8 accelerators | Vendor stated (dense) | nvidia.comdam-cdn.nvd.orangelogic.com | NVLink domain of the 8-GPU DGX B200 / HGX B200 baseboard. |
| NVIDIA B300 SXM | bf16 tflops | 2,250 TFLOPS | Derived from vendor sparse figure | dam-cdn.nvd.orangelogic.comepoch.ai | Vendor lists a sparse figure; halved for dense. HGX B300 datasheet lists 4.5 PFLOPS sparse. Epoch AI lists 2,250. |
| NVIDIA B300 SXM | fp8 tflops | 4,500 TFLOPS | Derived from vendor sparse figure | dam-cdn.nvd.orangelogic.comepoch.ai | Vendor lists a sparse figure; halved for dense. HGX B300 datasheet lists 9 PFLOPS sparse. Epoch AI lists 4,500. |
| NVIDIA B300 SXM | fp4 tflops | 13,500 TFLOPS | Derived from vendor sparse figure | nvidia.comdam-cdn.nvd.orangelogic.com | HGX page: 108 PFLOPS dense across 8 GPUs; datasheet rounds to 14 PFLOPS per GPU. Epoch AI lists 14,000. |
| NVIDIA B300 SXM | fp64 tflops | 1.2 TFLOPS | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | FP64 / FP64 Tensor Core; Blackwell Ultra cuts FP64 sharply versus B200 (37 TFLOPS). |
| NVIDIA B300 SXM | memory gb | 270 GB | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | HGX B300 HBM3E per GPU; 288 GB is the Blackwell Ultra maximum. |
| NVIDIA B300 SXM | memory bw tbs | 7.7 TB/s | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | HGX B300 per-GPU HBM3E bandwidth. |
| NVIDIA B300 SXM | tdp w | 1,100 W | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | HGX B300 configurable up to 1,100 W (GB300 variant up to 1,400 W). |
| NVIDIA B300 SXM | scaleup domain accelerators | 8 accelerators | Vendor stated (dense) | nvidia.comdam-cdn.nvd.orangelogic.com | NVLink domain of the 8-GPU HGX B300 baseboard. |
| NVIDIA GB200 NVL72 | bf16 tflops | 2,500 TFLOPS | Derived from vendor sparse figure | dam-cdn.nvd.orangelogic.comepoch.ai | Per GPU (two-die package). Vendor lists a sparse figure; halved for dense. Datasheet lists 5 PFLOPS sparse. |
| NVIDIA GB200 NVL72 | fp8 tflops | 5,000 TFLOPS | Derived from vendor sparse figure | dam-cdn.nvd.orangelogic.comepoch.ai | Per GPU. Vendor lists a sparse figure; halved for dense. Datasheet lists 10 PFLOPS sparse; AWS P6e states 360 PFLOPS dense FP8 for 72 GPUs. |
| NVIDIA GB200 NVL72 | fp4 tflops | 10,000 TFLOPS | Derived from vendor sparse figure | dam-cdn.nvd.orangelogic.comepoch.ai | Per GPU. Vendor lists a sparse figure; halved for dense. Datasheet lists 20 PFLOPS sparse. |
| NVIDIA GB200 NVL72 | fp64 tflops | 40 TFLOPS | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comnvidia.com | Per GPU FP64 / FP64 Tensor Core; GB200 page lists 2,880 TFLOPS for the rack. Epoch AI lists 45. |
| NVIDIA GB200 NVL72 | memory gb | 186 GB | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | Per GPU HBM3E in NVL72 configuration (372 GB per Grace Blackwell superchip). |
| NVIDIA GB200 NVL72 | memory bw tbs | 8 TB/s | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | Per GPU HBM3E bandwidth (16 TB/s per superchip). |
| NVIDIA GB200 NVL72 | tdp w | 1,200 W | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | Per GPU, configurable up to 1,200 W. |
| NVIDIA GB200 NVL72 | scaleup domain accelerators | 72 accelerators | Vendor stated (dense) | nvidia.comaws.amazon.com | 72 Blackwell GPUs in one NVLink domain; AWS P6e-GB200 UltraServers also state 72 GPUs per NVLink domain. |
| NVIDIA GB300 NVL72 | bf16 tflops | 2,500 TFLOPS | Derived from vendor sparse figure | dam-cdn.nvd.orangelogic.comepoch.ai | Per GPU. Vendor lists a sparse figure; halved for dense. Datasheet lists 5 PFLOPS sparse. |
| NVIDIA GB300 NVL72 | fp8 tflops | 5,000 TFLOPS | Derived from vendor sparse figure | dam-cdn.nvd.orangelogic.comepoch.ai | Per GPU. Vendor lists a sparse figure; halved for dense. Datasheet lists 10 PFLOPS sparse. |
| NVIDIA GB300 NVL72 | fp4 tflops | 15,000 TFLOPS | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | Per GPU dense NVFP4; datasheet lists 20 sparse | 15 dense PFLOPS. NVIDIA blog also states 15 PFLOPS dense. |
| NVIDIA GB300 NVL72 | fp64 tflops | 1.3 TFLOPS | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | Per GPU FP64 / FP64 Tensor Core. GB300 page lists 100 TFLOPS for the rack. |
| NVIDIA GB300 NVL72 | memory gb | 279 GB | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comdeveloper.nvidia.com | Datasheet per-GPU HBM3E; the NVIDIA blog and Epoch AI cite 288 GB as the maximum capacity. |
| NVIDIA GB300 NVL72 | memory bw tbs | 8 TB/s | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | Per GPU HBM3E bandwidth. |
| NVIDIA GB300 NVL72 | tdp w | 1,400 W | Vendor stated (dense) | dam-cdn.nvd.orangelogic.comepoch.ai | Per GPU, configurable up to 1,400 W. |
| NVIDIA GB300 NVL72 | scaleup domain accelerators | 72 accelerators | Vendor stated (dense) | nvidia.comdam-cdn.nvd.orangelogic.com | 72 Blackwell Ultra GPUs and 36 Grace CPUs in one NVLink domain. |
| NVIDIA Vera Rubin | bf16 tflops | 4,000 TFLOPS | Vendor stated (dense) | nvidia.comdam-cdn.nvd.orangelogic.com | Per Rubin GPU, dense (page footnote). NVIDIA states performance is projected and subject to change. |
| NVIDIA Vera Rubin | fp8 tflops | 17,500 TFLOPS | Vendor stated (dense) | nvidia.comdam-cdn.nvd.orangelogic.com | Per Rubin GPU, FP8/FP6 training figure, stated dense. Projected. |
| NVIDIA Vera Rubin | fp4 tflops | 35,000 TFLOPS | Vendor stated (dense) | nvidia.comdeveloper.nvidia.com | Per GPU NVFP4 training figure, stated dense. The 50 PFLOPS NVFP4 inference figure is not dense and is not recorded. |
| NVIDIA Vera Rubin | fp64 tflops | 33 TFLOPS | Vendor stated (dense) | nvidia.comdeveloper.nvidia.com | Per GPU hardware FP64 (vector). NVIDIA also cites 200 TFLOPS FP64 matrix via Tensor Core emulation, not recorded. |
| NVIDIA Vera Rubin | memory gb | 288 GB | Vendor stated (dense) | nvidia.comdam-cdn.nvd.orangelogic.com | Per Rubin GPU HBM4. |
| NVIDIA Vera Rubin | memory bw tbs | 19.2 TB/s | Vendor stated (dense) | nvidia.comdam-cdn.nvd.orangelogic.com | NVL72 page value; an older HGX page lists 22 TB/s, datasheet superchip 38.5 TB/s for two GPUs. |
| NVIDIA Vera Rubin | scaleup domain accelerators | 72 accelerators | Vendor stated (dense) | nvidia.comdam-cdn.nvd.orangelogic.com | Vera Rubin NVL72: 72 Rubin GPUs and 36 Vera CPUs in one NVLink 6 domain. HGX Rubin NVL8 is 8. |
| AMD Instinct MI250X | bf16 tflops | 383 TFLOPS | Vendor stated (dense) | amd.comepoch.ai | Per MI250X module (two GCDs). Epoch AI lists 383. |
| AMD Instinct MI250X | fp64 tflops | 95.7 TFLOPS | Vendor stated (dense) | amd.comdocs.olcf.ornl.gov | Matrix-core FP64 per module; vector FP64 is 47.9. ORNL lists 47.9 matrix per GCD. |
| AMD Instinct MI250X | memory gb | 128 GB | Vendor stated (dense) | amd.comdocs.olcf.ornl.gov | HBM2e per module. ORNL lists 64 GB per GCD. |
| AMD Instinct MI250X | memory bw tbs | 3.2 TB/s | Vendor stated (dense) | amd.comdocs.olcf.ornl.gov | Per module. ORNL lists 1.6 TB/s per GCD. |
| AMD Instinct MI250X | tdp w | 500 W | Vendor stated (dense) | amd.comepoch.ai | 500 W typical, 560 W peak per AMD. |
| AMD Instinct MI300X | bf16 tflops | 1,307.4 TFLOPS | Vendor stated (dense) | amd.comepoch.ai | Dense; datasheet lists 1,307.4 dense and 2,614.9 with sparsity. |
| AMD Instinct MI300X | fp8 tflops | 2,614.9 TFLOPS | Vendor stated (dense) | amd.comamd.com | Dense; datasheet lists 2,614.9 dense and 5,229.8 with sparsity. |
| AMD Instinct MI300X | fp64 tflops | 163.4 TFLOPS | Vendor stated (dense) | amd.comamd.com | Matrix FP64; vector FP64 is 81.7 TFLOPS. |
| AMD Instinct MI300X | memory gb | 192 GB | Vendor stated (dense) | amd.comepoch.ai | HBM3. |
| AMD Instinct MI300X | memory bw tbs | 5.3 TB/s | Vendor stated (dense) | amd.comepoch.ai | Peak theoretical HBM3 bandwidth. |
| AMD Instinct MI300X | tdp w | 750 W | Vendor stated (dense) | amd.comepoch.ai | Maximum typical board power. |
| AMD Instinct MI300X | scaleup domain accelerators | 8 accelerators | Vendor stated (dense) | amd.comamd.com | Sold as an Instinct Platform of eight accelerators on Infinity Fabric. |
| AMD Instinct MI325X | bf16 tflops | 1,307.4 TFLOPS | Vendor stated (dense) | amd.comepoch.ai | Dense. AMD page rounds to 1.3 PFLOPS (sparse 2.61); exact value from Epoch AI, same as MI300X at 2,100 MHz. |
| AMD Instinct MI325X | fp8 tflops | 2,614.9 TFLOPS | Vendor stated (dense) | amd.comepoch.ai | Dense. AMD page rounds to 2.61 PFLOPS (sparse 5.22); exact value inferred from MI300X datasheet at the same 2,100 MHz. |
| AMD Instinct MI325X | fp64 tflops | 163.4 TFLOPS | Vendor stated (dense) | amd.comamd.com | Matrix FP64; vector FP64 is 81.7 TFLOPS. |
| AMD Instinct MI325X | memory gb | 256 GB | Vendor stated (dense) | amd.comepoch.ai | HBM3E. |
| AMD Instinct MI325X | memory bw tbs | 6 TB/s | Vendor stated (dense) | amd.comepoch.ai | Peak theoretical HBM3E bandwidth. |
| AMD Instinct MI325X | tdp w | 1,000 W | Vendor stated (dense) | amd.comepoch.ai | Typical board power, 1,000 W peak. |
| AMD Instinct MI355X | bf16 tflops | 2,516.6 TFLOPS | Vendor stated (dense) | amd.comepoch.ai | Dense matrix BF16; brochure lists 2.5166 PFLOPS dense and 5.0332 with sparsity. |
| AMD Instinct MI355X | fp8 tflops | 5,033.2 TFLOPS | Vendor stated (dense) | amd.comamd.com | Dense OCP-FP8; brochure lists 5.0332 dense and 10.0664 PFLOPS with sparsity. |
| AMD Instinct MI355X | fp4 tflops | 10,066.3 TFLOPS | Vendor stated (dense) | amd.comamd.com | MXFP4; the brochure lists no sparse variant for FP4 (10.0663 PFLOPS). |
| AMD Instinct MI355X | fp64 tflops | 78.6 TFLOPS | Vendor stated (dense) | amd.comamd.com | FP64 matrix and vector are both listed at 78.6 TFLOPS. |
| AMD Instinct MI355X | memory gb | 288 GB | Vendor stated (dense) | amd.comepoch.ai | HBM3E. |
| AMD Instinct MI355X | memory bw tbs | 8 TB/s | Vendor stated (dense) | amd.comepoch.ai | Peak theoretical HBM3E bandwidth. |
| AMD Instinct MI355X | tdp w | 1,400 W | Vendor stated (dense) | amd.comepoch.ai | Typical board power (TBP). |
| AMD Instinct MI355X | scaleup domain accelerators | 8 accelerators | Vendor stated (dense) | amd.comamd.com | Brochure describes an 8-GPU MI355X platform on Infinity Fabric. |
| AMD Instinct MI430X | fp64 tflops | 288 TFLOPS | Vendor stated (dense) | amd.comamd.com | AMD states up to 288 TFLOPS hardware-based peak theoretical FP64. |
| AMD Instinct MI430X | memory gb | 432 GB | Vendor stated (dense) | amd.comamd.com | Integrated HBM4, stated as up to 432 GB. |
| AMD Instinct MI430X | memory bw tbs | 23.3 TB/s | Vendor stated (dense) | amd.comamd.com | Peak theoretical bandwidth, stated as up to 23.3 TB/s. |
| AMD Instinct MI455X | bf16 tflops | 5,000 TFLOPS | Vendor stated (dense) | amd.comamd.com | Page lists 5 PFLOPS dense BF16 matrix (10.1 PFLOPS with sparsity); stated to one significant figure. |
| AMD Instinct MI455X | fp64 tflops | 5 TFLOPS | Vendor stated (dense) | amd.comamd.com | Page lists matrix and vector FP64 at 5 TFLOPS, far below MI430X (288). |
| AMD Instinct MI455X | memory gb | 432 GB | Vendor stated (dense) | amd.comamd.com | HBM4, 12 stacks. |
| AMD Instinct MI455X | memory bw tbs | 23.3 TB/s | Vendor stated (dense) | amd.comamd.com | Peak theoretical HBM4 bandwidth. |
| Google TPU v4 | bf16 tflops | 275 TFLOPS | Vendor stated (dense) | cloud.google.comepoch.ai | Peak per chip (bf16 or int8), two TensorCores. Epoch AI lists 275. |
| Google TPU v4 | memory gb | 32 GB | Vendor stated (dense) | cloud.google.comepoch.ai | Stated as 32 GiB HBM2 (about 34.4 GB). |
| Google TPU v4 | memory bw tbs | 1.2 TB/s | Vendor stated (dense) | cloud.google.comepoch.ai | Stated as 1,200 GBps. |
| Google TPU v4 | scaleup domain accelerators | 4,096 accelerators | Vendor stated (dense) | cloud.google.comarxiv.org | TPU v4 pod of 4,096 chips on a reconfigurable optical 3D torus. |
| Google TPU v5p | bf16 tflops | 459 TFLOPS | Vendor stated (dense) | cloud.google.comepoch.ai | Peak per chip. Google's FP8 figure is emulated on v5p, so FP8 is not recorded. |
| Google TPU v5p | memory gb | 95 GB | Vendor stated (dense) | cloud.google.comepoch.ai | Stated as 95 GiB HBM (about 102 GB). |
| Google TPU v5p | memory bw tbs | 2.765 TB/s | Vendor stated (dense) | cloud.google.comepoch.ai | Stated as 2,765 GBps. |
| Google TPU v5p | scaleup domain accelerators | 8,960 accelerators | Vendor stated (dense) | cloud.google.com | Pod of 8,960 chips on a 3D torus; the largest schedulable job is 6,144 chips. |
| Google TPU v5e | bf16 tflops | 197 TFLOPS | Vendor stated (dense) | cloud.google.comepoch.ai | Peak per chip (one TensorCore). Epoch AI lists 197. |
| Google TPU v5e | memory gb | 16 GB | Vendor stated (dense) | cloud.google.comepoch.ai | HBM per chip. |
| Google TPU v5e | scaleup domain accelerators | 256 accelerators | Vendor stated (dense) | cloud.google.com | Pod size 256 chips on a 2D torus. |
| Google TPU v6e (Trillium) | bf16 tflops | 918 TFLOPS | Vendor stated (dense) | cloud.google.comepoch.ai | Peak per chip (one TensorCore). Epoch AI lists 918. |
| Google TPU v6e (Trillium) | memory gb | 32 GB | Vendor stated (dense) | cloud.google.comepoch.ai | HBM per chip. |
| Google TPU v6e (Trillium) | memory bw tbs | 1.638 TB/s | Vendor stated (dense) | cloud.google.comepoch.ai | Stated as 1,638 GBps. |
| Google TPU v6e (Trillium) | scaleup domain accelerators | 256 accelerators | Vendor stated (dense) | cloud.google.com | Pod size 256 chips on a 2D torus. |
| Google TPU7x (Ironwood) | bf16 tflops | 2,307 TFLOPS | Vendor stated (dense) | cloud.google.comepoch.ai | Peak per chip. Epoch AI lists 2,307. |
| Google TPU7x (Ironwood) | fp8 tflops | 4,614 TFLOPS | Vendor stated (dense) | cloud.google.comepoch.ai | Peak per chip, native FP8 on Ironwood. Epoch AI lists 4,614. |
| Google TPU7x (Ironwood) | memory gb | 192 GB | Vendor stated (dense) | cloud.google.comepoch.ai | Stated as 192 GiB HBM (about 206 GB). |
| Google TPU7x (Ironwood) | memory bw tbs | 7.38 TB/s | Vendor stated (dense) | cloud.google.comepoch.ai | Docs state 7,380 GBps; Google's blog and Epoch AI state 7.37 TB/s. |
| Google TPU7x (Ironwood) | scaleup domain accelerators | 9,216 accelerators | Vendor stated (dense) | cloud.google.comblog.google | Pod of 9,216 chips on a 3D torus. |
| Google TPU 8t | fp4 tflops | 12,600 TFLOPS | Vendor stated (dense) | cloud.google.com | Google lists peak FP4 of 12.6 PFLOPS with no sparsity qualifier; TPUs have no structured-sparsity matmul. |
| Google TPU 8t | memory gb | 216 GB | Vendor stated (dense) | cloud.google.com | HBM capacity. |
| Google TPU 8t | memory bw tbs | 6.528 TB/s | Vendor stated (dense) | cloud.google.com | Stated as 6,528 GB/s. |
| Google TPU 8t | scaleup domain accelerators | 9,600 accelerators | Vendor stated (dense) | cloud.google.com | Single 3D-torus superpod of 9,600 chips. |
| Google TPU 8i | fp4 tflops | 10,100 TFLOPS | Vendor stated (dense) | cloud.google.comepoch.ai | Google lists peak FP4 of 10.1 PFLOPS with no sparsity qualifier. Epoch AI lists 10.1 PFLOPS. |
| Google TPU 8i | memory gb | 288 GB | Vendor stated (dense) | cloud.google.comepoch.ai | HBM capacity. |
| Google TPU 8i | memory bw tbs | 8.601 TB/s | Vendor stated (dense) | cloud.google.comepoch.ai | Stated as 8,601 GB/s. |
| Google TPU 8i | scaleup domain accelerators | 1,024 accelerators | Vendor stated (dense) | cloud.google.com | Boardfly pod of up to 1,024 active chips. |
| AWS Trainium2 | bf16 tflops | 667 TFLOPS | Vendor stated (dense) | awsdocs-neuron.readthedocs-hosted.comepoch.ai | Dense BF16/FP16/TF32; Neuron docs list 2,563 TFLOPS separately with 4x sparsity. |
| AWS Trainium2 | fp8 tflops | 1,299 TFLOPS | Vendor stated (dense) | awsdocs-neuron.readthedocs-hosted.comepoch.ai | Dense FP8; the sparse figure is listed separately. |
| AWS Trainium2 | memory gb | 96 GB | Vendor stated (dense) | awsdocs-neuron.readthedocs-hosted.comepoch.ai | Stated as 96 GiB HBM per chip. |
| AWS Trainium2 | memory bw tbs | 2.9 TB/s | Vendor stated (dense) | awsdocs-neuron.readthedocs-hosted.comepoch.ai | Stated as 2.9 TB/s. |
| AWS Trainium2 | scaleup domain accelerators | 64 accelerators | Vendor stated (dense) | aws.amazon.com | Trn2 UltraServer links 64 chips over NeuronLink; a single Trn2 instance has 16. |
| AWS Trainium3 | bf16 tflops | 671 TFLOPS | Vendor stated (dense) | awsdocs-neuron.readthedocs-hosted.comepoch.ai | Dense BF16/FP16/TF32; Neuron docs list a separate sparse figure. |
| AWS Trainium3 | fp8 tflops | 2,517 TFLOPS | Vendor stated (dense) | awsdocs-neuron.readthedocs-hosted.comepoch.ai | MXFP8 dense, listed as 2,517 TFLOPS. |
| AWS Trainium3 | fp4 tflops | 2,517 TFLOPS | Vendor stated (dense) | awsdocs-neuron.readthedocs-hosted.comepoch.ai | MXFP4 listed at the same 2,517 TFLOPS as MXFP8. |
| AWS Trainium3 | memory gb | 144 GB | Vendor stated (dense) | awsdocs-neuron.readthedocs-hosted.comaws.amazon.com | Stated as 144 GiB (docs) or 144 GB (instance page) of HBM3e. |
| AWS Trainium3 | memory bw tbs | 4.9 TB/s | Vendor stated (dense) | awsdocs-neuron.readthedocs-hosted.comaws.amazon.com | Stated as 4.9 TB/s. |
| AWS Trainium3 | scaleup domain accelerators | 144 accelerators | Vendor stated (dense) | aws.amazon.comaws.amazon.com | Trn3 UltraServer scales to 144 chips on a NeuronSwitch all-to-all fabric. |
| Microsoft Maia 100 | memory gb | 64 GB | Vendor stated (dense) | techcommunity.microsoft.com | HBM2E, four stacks. |
| Microsoft Maia 100 | memory bw tbs | 1.8 TB/s | Vendor stated (dense) | techcommunity.microsoft.com | Stated as 1.8 TB/s. |
| Microsoft Maia 100 | tdp w | 700 W | Vendor stated (dense) | techcommunity.microsoft.com | Designed for up to 700 W but provisioned at 500 W. |
| Microsoft Maia 200 | fp4 tflops | 10,000 TFLOPS | Vendor stated (dense) | blogs.microsoft.comepoch.ai | Microsoft states over 10 PFLOPS FP4; Epoch AI lists 10,145. Recorded as the stated lower bound. |
| Microsoft Maia 200 | fp8 tflops | 5,000 TFLOPS | Vendor stated (dense) | blogs.microsoft.comepoch.ai | Microsoft states over 5 PFLOPS FP8; Epoch AI lists 5,072. Recorded as the stated lower bound. |
| Microsoft Maia 200 | memory gb | 216 GB | Vendor stated (dense) | blogs.microsoft.comepoch.ai | HBM3e. |
| Microsoft Maia 200 | memory bw tbs | 7 TB/s | Vendor stated (dense) | blogs.microsoft.comepoch.ai | Stated as 7 TB/s. |
| Microsoft Maia 200 | tdp w | 750 W | Vendor stated (dense) | blogs.microsoft.comepoch.ai | Stated as 750 W SoC TDP. |
| Meta MTIA v2 | bf16 tflops | 177 TFLOPS | Vendor stated (dense) | ai.meta.comepoch.ai | Dense FP16/BF16; Meta lists 354 TFLOPS with sparsity. |
| Meta MTIA v2 | memory gb | 128 GB | Vendor stated (dense) | ai.meta.com | Off-chip LPDDR5 (not HBM); on-chip SRAM is 256 MB. |
| Meta MTIA v2 | memory bw tbs | 0.205 TB/s | Vendor stated (dense) | ai.meta.com | Off-chip LPDDR5 bandwidth of 204.8 GB/s; on-chip SRAM is 2.7 TB/s. |
| Meta MTIA v2 | tdp w | 90 W | Vendor stated (dense) | ai.meta.comepoch.ai | Stated as 90 W. |
| Huawei Ascend 910C | bf16 tflops | 781 TFLOPS | Third party | newsletter.semianalysis.comtrendforce.com | SemiAnalysis: CloudMatrix 384 gives 300 PFLOPS dense BF16 across 384 chips (300000/384). TrendForce reports 800 TFLOPS FP16. No Huawei datasheet. |
| Huawei Ascend 910C | memory bw tbs | 3.2 TB/s | Third party | trendforce.com | TrendForce reports 3.2 TB/s citing an industry report; no Huawei datasheet fetched. |
| Huawei CloudMatrix 384 | scaleup domain accelerators | 384 accelerators | Vendor stated (dense) | arxiv.orgnewsletter.semianalysis.com | 384 Ascend NPUs in one all-to-all Unified Bus supernode; SemiAnalysis independently states 384 Ascend 910C chips. |
| Intel Gaudi 3 | bf16 tflops | 1,678 TFLOPS | Vendor stated (dense) | cdrdv2-public.intel.comepoch.ai | BF16 MME peak, listed with no sparsity qualifier. |
| Intel Gaudi 3 | fp8 tflops | 1,678 TFLOPS | Vendor stated (dense) | cdrdv2-public.intel.comepoch.ai | FP8 MME peak, listed with no sparsity qualifier. |
| Intel Gaudi 3 | memory gb | 128 GB | Vendor stated (dense) | cdrdv2-public.intel.comepoch.ai | HBM2e, eight stacks. |
| Intel Gaudi 3 | memory bw tbs | 3.7 TB/s | Vendor stated (dense) | cdrdv2-public.intel.comepoch.ai | Peak HBM bandwidth. |
| Intel Gaudi 3 | tdp w | 900 W | Vendor stated (dense) | cdrdv2-public.intel.comepoch.ai | OAM module, up to 900 W with passive or liquid cooling. |
| Cerebras WSE-3 | bf16 tflops | 62,500 TFLOPS | Derived from vendor sparse figure | cdn.sanity.ioarxiv.org | Datasheet gives 125 PFLOPS FP16 marked as sparse; halved for dense. BF16 is not listed separately. |
| Cerebras WSE-3 | memory gb | 44 GB | Vendor stated (dense) | cdn.sanity.ioarxiv.org | On-chip SRAM only; the wafer has no HBM. An independent arXiv paper also states 44 GB. |
| Cerebras WSE-3 | memory bw tbs | 21,000 TB/s | Vendor stated (dense) | cdn.sanity.ioarxiv.org | Stated as 21 PB/s on-chip SRAM bandwidth; an independent arXiv paper also states 21 PB/s. |