COMPUTE ATLAS Supercomputer supply-chain graph
1153 systems 482 sites Sourced data

AI datacenters/Accelerators

AI datacenters

AI accelerators compared

An accelerator count says little unless you know which accelerator. This table puts 31 of the chips that fill AI datacenters on one footing: dense peak FLOPS by precision, memory and bandwidth, power, and the size of the scale-up domain they are sold in. Vendors headline sparse figures that are twice the dense ones and are rarely reached on real workloads, so every number here is dense, and a figure halved from a sparse headline is marked as derived. All of these are peaks from datasheets, not measured performance.

  • NVIDIA
  • AMD
  • Google
  • Other vendors
Dense peak TFLOPS per accelerator by release year Log-scale dot plot of the dense FP8 peak, or BF16 where FP8 is not offered, of each accelerator in the table by its release year, coloured by vendor. 10100100010000100000201620182020202220242026 AMD Instinct MI355X AMD 5.03e+3 dense TFLOPS (FP8, else BF16), 2026.5 Microsoft Maia 200 Other vendors 5.00e+3 dense TFLOPS (FP8, else BF16), 2026.5 AWS Trainium3 Other vendors 2.52e+3 dense TFLOPS (FP8, else BF16), 2025.5 Google TPU7x (Ironwood) Alphabet 4.61e+3 dense TFLOPS (FP8, else BF16), 2025.5 NVIDIA B300 SXM NVIDIA 4.50e+3 dense TFLOPS (FP8, else BF16), 2025.5 NVIDIA GB300 NVL72 NVIDIA 5.00e+3 dense TFLOPS (FP8, else BF16), 2025.5 AMD Instinct MI325X AMD 2.61e+3 dense TFLOPS (FP8, else BF16), 2024.5 AWS Trainium2 Other vendors 1.30e+3 dense TFLOPS (FP8, else BF16), 2024.5 Cerebras WSE-3 Other vendors 6.25e+4 dense TFLOPS (FP8, else BF16), 2024.5 Google TPU v6e (Trillium) Alphabet 918 dense TFLOPS (FP8, else BF16), 2024.5 Intel Gaudi 3 Other vendors 1.68e+3 dense TFLOPS (FP8, else BF16), 2024.5 Meta MTIA v2 Other vendors 177 dense TFLOPS (FP8, else BF16), 2024.5 NVIDIA B200 SXM NVIDIA 4.50e+3 dense TFLOPS (FP8, else BF16), 2024.5 NVIDIA GB200 NVL72 NVIDIA 5.00e+3 dense TFLOPS (FP8, else BF16), 2024.5 NVIDIA H200 SXM NVIDIA 1.98e+3 dense TFLOPS (FP8, else BF16), 2024.5 AMD Instinct MI300X AMD 2.61e+3 dense TFLOPS (FP8, else BF16), 2023.5 Google TPU v5e Alphabet 197 dense TFLOPS (FP8, else BF16), 2023.5 Google TPU v5p Alphabet 459 dense TFLOPS (FP8, else BF16), 2023.5 NVIDIA H100 SXM5 NVIDIA 1.98e+3 dense TFLOPS (FP8, else BF16), 2022.5 AMD Instinct MI250X AMD 383 dense TFLOPS (FP8, else BF16), 2021.5 Google TPU v4 Alphabet 275 dense TFLOPS (FP8, else BF16), 2020.5 NVIDIA A100 SXM4 40GB NVIDIA 312 dense TFLOPS (FP8, else BF16), 2020.5 NVIDIA A100 SXM4 80GB NVIDIA 312 dense TFLOPS (FP8, else BF16), 2020.5 Cerebras WSE-3AMD Instinct MI355XMicrosoft Maia 200
Each dot is one accelerator: its dense FP8 peak where it offers FP8, otherwise dense BF16, on a logarithmic axis. Mixing the two precisions understates the newest chips' gain, since FP8 and FP4 arrived as the headline formats only recently.
Table view: Dense peak TFLOPS per accelerator by release year
AcceleratorVendorReleasedDense TFLOPS
AMD Instinct MI355XAMD20265033.2
Microsoft Maia 200Microsoft20265000
AWS Trainium3Amazon20252517
Google TPU7x (Ironwood)Alphabet20254614
NVIDIA B300 SXMNVIDIA20254500
NVIDIA GB300 NVL72NVIDIA20255000
AMD Instinct MI325XAMD20242614.9
AWS Trainium2Amazon20241299
Cerebras WSE-3Cerebras202462500
Google TPU v6e (Trillium)Alphabet2024918
Intel Gaudi 3Intel20241678
Meta MTIA v2Meta Platforms2024177
NVIDIA B200 SXMNVIDIA20244500
NVIDIA GB200 NVL72NVIDIA20245000
NVIDIA H200 SXMNVIDIA20241979
AMD Instinct MI300XAMD20232614.9
Google TPU v5eAlphabet2023197
Google TPU v5pAlphabet2023459
NVIDIA H100 SXM5NVIDIA20221979
AMD Instinct MI250XAMD2021383
Google TPU v4Alphabet2020275
NVIDIA A100 SXM4 40GBNVIDIA2020312
NVIDIA A100 SXM4 80GBNVIDIA2020312

The accelerators

AcceleratorVendorReleased BF16FP8FP4FP64 Memory (GB)Bandwidth (TB/s)TDP (W)FP8 per kWScale-up
AMD Instinct MI355X AMD 2026 2,517 5,033 10,066 78.6 288 8 1,400 3,595 8
Google TPU 8i Alphabet 2026 - - 10,100 - 288 8.6 - - 1,024
Google TPU 8t Alphabet 2026 - - 12,600 - 216 6.5 - - 9,600
Microsoft Maia 200 Microsoft 2026 - 5,000 10,000 - 216 7 750 6,667 -
AWS Trainium3 Amazon 2025 671 2,517 2,517 - 144 4.9 - - 144
Google TPU7x (Ironwood) Alphabet 2025 2,307 4,614 - - 192 7.4 - - 9,216
Huawei CloudMatrix 384 Huawei 2025 - - - - - - - - 384
NVIDIA B300 SXM NVIDIA 2025 2,250 4,500 13,500 1.2 270 7.7 1,100 4,091 8
NVIDIA GB300 NVL72 NVIDIA 2025 2,500 5,000 15,000 1.3 279 8 1,400 3,571 72
AMD Instinct MI325X AMD 2024 1,307 2,615 - 163.4 256 6 1,000 2,615 -
AWS Trainium2 Amazon 2024 667 1,299 - - 96 2.9 - - 64
Cerebras WSE-3 Cerebras 2024 62,500 - - - 44 21,000 23,000 - -
Google TPU v6e (Trillium) Alphabet 2024 918 - - - 32 1.6 - - 256
Intel Gaudi 3 Intel 2024 1,678 1,678 - - 128 3.7 900 1,864 -
Meta MTIA v2 Meta Platforms 2024 177 - - - 128 0.2 90 - -
NVIDIA B200 SXM NVIDIA 2024 2,250 4,500 9,000 37 180 7.7 1,000 4,500 8
NVIDIA GB200 NVL72 NVIDIA 2024 2,500 5,000 10,000 40 186 8 1,200 4,167 72
NVIDIA H200 SXM NVIDIA 2024 990 1,979 - 67 141 4.8 700 2,827 8
AMD Instinct MI300X AMD 2023 1,307 2,615 - 163.4 192 5.3 750 3,487 8
Google TPU v5e Alphabet 2023 197 - - - 16 - - - 256
Google TPU v5p Alphabet 2023 459 - - - 95 2.8 - - 8,960
Microsoft Maia 100 Microsoft 2023 - - - - 64 1.8 700 - -
NVIDIA H100 SXM5 NVIDIA 2022 990 1,979 - 67 80 3.4 700 2,827 8
AMD Instinct MI250X AMD 2021 383 - - 95.7 128 3.2 500 - -
Google TPU v4 Alphabet 2020 275 - - - 32 1.2 - - 4,096
NVIDIA A100 SXM4 40GB NVIDIA 2020 312 - - 19.5 40 1.6 400 - 8
NVIDIA A100 SXM4 80GB NVIDIA 2020 312 - - 19.5 80 2 400 - 8
AMD Instinct MI430X AMD - - - - 288 432 23.3 - - -
AMD Instinct MI455X AMD - 5,000 - - 5 432 23.3 - - -
Huawei Ascend 910C Huawei - 781 - - - - 3.2 - - -
NVIDIA Vera Rubin NVIDIA - 4,000 17,500 35,000 33 288 19.2 - - 72

TFLOPS are dense peaks per accelerator package. FP8 per kW divides the dense FP8 peak by the accelerator's own TDP and ignores the rest of the node, network and cooling. Scale-up is the number of accelerators in one high-bandwidth domain (72 for GB200 NVL72).

Derived peak compute of the biggest clusters

A cluster's accelerator count times each chip's dense peak is the one compute figure that can be set against another cluster of a different chip. It is a ceiling, not a benchmark result: real training reaches a fraction of it (Meta reported 38 to 43 percent on a 16,384-GPU H100 run), and it is only shown where every accelerator in the system has both a recorded count and a recorded peak. Counts for campuses that are not yet running are plans or estimates, so they are listed apart.

Running

SystemStatusAcceleratorsPrecisionDerived peak (EFLOPS)
AWS Madison Mega Site (Canton) operational 326,600 FP8 424
Meta Rosemount (Minnesota) operational 85,400 FP8 384
AWS Ridgeland Campus operational 261,200 FP8 339
xAI Colossus operational 100,000 FP8 198
Together AI / Hypertec Cloud GB200 Cluster operational 36,000 FP8 180
Tesla Cortex operational 66,000 FP8 131
Mistral Campus AI (Bruyeres-le-Chatel) operational 18,000 FP8 90
SINES Data Campus GB300 Deployment operational 12,600 FP8 63
Meta GenAI cluster (RoCE) operational 24,576 FP8 49
Meta GenAI cluster (InfiniBand) operational 24,576 FP8 49
Industrial AI Cloud operational 10,000 FP8 45
Google TPU7x (Ironwood) pod operational 9,216 FP8 43
Taiwan AI Factory operational 7,000 FP8 35
Firebird Armenian AI Factory (DC-1) operational 6,144 FP8 28
Nscale Verne Iceland Cluster operational 4,600 FP8 23
TensorWave Tucson MI325X Cluster operational 8,192 FP8 21
Cerebras Oklahoma City operational 300 BF16 19
Naver B200 4K Cluster operational 4,000 FP8 18
Nebius Modiin Cluster operational 4,000 FP8 18
IBM Blue Vela operational 5,000 FP8 9.89
SDS-AI operational 2,032 FP8 9.14
Google TPU v4 ML Hub (Oklahoma) operational 32,768 BF16 9.01
Shenzhen 10,000-Card Ascend Cluster operational 10,000 BF16 7.81
CoreWeave AI cluster for Inflection AI operational 3,584 FP8 7.09
Databricks DBRX Training Cluster operational 3,072 FP8 6.08
E2E Chennai B200 Cluster operational 1,024 FP8 4.61
Haein Cluster operational 1,000 FP8 4.50
Google TPU v5p pod operational 8,960 BF16 4.11
Israel-1 operational 2,048 FP8 4.05
Condor Galaxy 3 operational 64 BF16 4.00
HiPerGator AI operational 504 FP8 2.27
CZI GPU Cluster operational 1,024 FP8 2.03
Ubilink H100 Cluster operational 1,024 FP8 2.03
Pre-Eos operational 1,024 FP8 2.03
Nabuchodonosor operational 1,016 FP8 2.01
SAKURAONE operational 800 FP8 1.58
Google TPU v4 supercomputer operational 4,096 BF16 1.13
BioHive-2 operational 504 FP8 1.00
DAIS operational 264 FP8 0.85
AI-Farabium operational 400 FP8 0.79

Planned or under construction

SystemStatusAccelerators (planned)PrecisionDerived peak (EFLOPS)
Microsoft Azure AI datacenter (Narvik, Norway) under construction 30,000 FP8 525
Solstice planned 100,000 FP8 450
Microsoft Azure AI supercomputer (Loughton, UK) under construction 23,040 FP8 115
Project Ceiba under construction 20,736 FP8 93
Equinox planned 10,000 FP8 45

Sources

AcceleratorFigureValueBasisSourceNote
NVIDIA A100 SXM4 40GB bf16 tflops 312 TFLOPS Vendor stated (dense) nvidia.comepoch.ai Dense BF16 tensor core; datasheet gives 312 dense and 624 with sparsity. Epoch AI also lists 312.
NVIDIA A100 SXM4 40GB fp64 tflops 19.5 TFLOPS Vendor stated (dense) nvidia.comimages.nvidia.com FP64 Tensor Core peak; non-tensor FP64 is 9.7 TFLOPS.
NVIDIA A100 SXM4 40GB memory gb 40 GB Vendor stated (dense) nvidia.comaws.amazon.com HBM2. AWS P4d lists 320 GB across eight A100 GPUs.
NVIDIA A100 SXM4 40GB memory bw tbs 1.555 TB/s Vendor stated (dense) nvidia.comepoch.ai Stated as 1,555 GB/s. Epoch AI lists 1.56 TB/s.
NVIDIA A100 SXM4 40GB tdp w 400 W Vendor stated (dense) nvidia.comepoch.ai SXM4 module maximum thermal design power.
NVIDIA A100 SXM4 40GB scaleup domain accelerators 8 accelerators Vendor stated (dense) nvidia.comnvidia.com NVSwitch domain of the 8-GPU DGX A100 baseboard; HGX A100 is also offered with 4 or 16 GPUs.
NVIDIA A100 SXM4 80GB bf16 tflops 312 TFLOPS Vendor stated (dense) nvidia.comepoch.ai Dense BF16 tensor core; datasheet gives 312 dense and 624 with sparsity. Epoch AI also lists 312.
NVIDIA A100 SXM4 80GB fp64 tflops 19.5 TFLOPS Vendor stated (dense) nvidia.comimages.nvidia.com FP64 Tensor Core peak; non-tensor FP64 is 9.7 TFLOPS.
NVIDIA A100 SXM4 80GB memory gb 80 GB Vendor stated (dense) nvidia.comaws.amazon.com HBM2e. AWS P4de lists 640 GB across eight A100 80GB GPUs.
NVIDIA A100 SXM4 80GB memory bw tbs 2.039 TB/s Vendor stated (dense) nvidia.comepoch.ai Stated as 2,039 GB/s. Epoch AI lists the same.
NVIDIA A100 SXM4 80GB tdp w 400 W Vendor stated (dense) nvidia.comepoch.ai 400 W standard; the HGX custom thermal solution SKU supports up to 500 W.
NVIDIA A100 SXM4 80GB scaleup domain accelerators 8 accelerators Vendor stated (dense) nvidia.comnvidia.com NVSwitch domain of the 8-GPU DGX A100 baseboard; HGX A100 is also offered with 4 or 16 GPUs.
NVIDIA H100 SXM5 bf16 tflops 989.5 TFLOPS Derived from vendor sparse figure nvidia.comepoch.ai Vendor lists a sparse figure; halved for dense. Page lists 1,979 with sparsity. Epoch AI lists 989.4.
NVIDIA H100 SXM5 fp8 tflops 1,979 TFLOPS Derived from vendor sparse figure nvidia.comepoch.ai Vendor lists a sparse figure; halved for dense. Page lists 3,958 with sparsity. Epoch AI lists 1,979.
NVIDIA H100 SXM5 fp64 tflops 67 TFLOPS Vendor stated (dense) nvidia.comnvidia.com FP64 Tensor Core peak; vector FP64 is 34 TFLOPS.
NVIDIA H100 SXM5 memory gb 80 GB Vendor stated (dense) nvidia.comaws.amazon.com HBM3. AWS P5 lists eight H100 GPUs with 640 GB.
NVIDIA H100 SXM5 memory bw tbs 3.35 TB/s Vendor stated (dense) nvidia.comepoch.ai HBM3 bandwidth as stated by NVIDIA.
NVIDIA H100 SXM5 tdp w 700 W Vendor stated (dense) nvidia.comepoch.ai Stated as up to 700 W, configurable.
NVIDIA H100 SXM5 scaleup domain accelerators 8 accelerators Vendor stated (dense) nvidia.com NVSwitch domain of the 8-GPU DGX H100 / HGX H100 baseboard (HGX also offered with 4 GPUs).
NVIDIA H200 SXM bf16 tflops 989.5 TFLOPS Derived from vendor sparse figure nvidia.comepoch.ai Vendor lists a sparse figure; halved for dense. Page lists 1,979 with sparsity. Epoch AI lists 989.5.
NVIDIA H200 SXM fp8 tflops 1,979 TFLOPS Derived from vendor sparse figure nvidia.comepoch.ai Vendor lists a sparse figure; halved for dense. Page lists 3,958 with sparsity. Epoch AI lists 1,979.
NVIDIA H200 SXM fp64 tflops 67 TFLOPS Vendor stated (dense) nvidia.comdam-cdn.nvd.orangelogic.com FP64 Tensor Core peak; vector FP64 is 34 TFLOPS.
NVIDIA H200 SXM memory gb 141 GB Vendor stated (dense) nvidia.comaws.amazon.com HBM3e. AWS P5e lists eight H200 GPUs with 1,128 GB.
NVIDIA H200 SXM memory bw tbs 4.8 TB/s Vendor stated (dense) nvidia.comepoch.ai HBM3e bandwidth as stated by NVIDIA.
NVIDIA H200 SXM tdp w 700 W Vendor stated (dense) nvidia.comepoch.ai Stated as up to 700 W, configurable (NVL PCIe variant up to 600 W).
NVIDIA H200 SXM scaleup domain accelerators 8 accelerators Vendor stated (dense) nvidia.comdam-cdn.nvd.orangelogic.com NVSwitch domain of an 8-GPU HGX H200 baseboard (HGX also offered with 4 GPUs).
NVIDIA B200 SXM bf16 tflops 2,250 TFLOPS Derived from vendor sparse figure dam-cdn.nvd.orangelogic.comepoch.ai Vendor lists a sparse figure; halved for dense. HGX B200 datasheet lists 4.5 PFLOPS sparse. Epoch AI lists 2,250.
NVIDIA B200 SXM fp8 tflops 4,500 TFLOPS Derived from vendor sparse figure dam-cdn.nvd.orangelogic.comepoch.ai Vendor lists a sparse figure; halved for dense. HGX B200 datasheet lists 9 PFLOPS sparse. Epoch AI lists 4,500.
NVIDIA B200 SXM fp4 tflops 9,000 TFLOPS Derived from vendor sparse figure dam-cdn.nvd.orangelogic.comepoch.ai Vendor lists a sparse figure; halved for dense. HGX B200 datasheet lists 18 PFLOPS sparse (72 dense per 8 GPUs). Epoch AI lists 9,000.
NVIDIA B200 SXM fp64 tflops 37 TFLOPS Vendor stated (dense) dam-cdn.nvd.orangelogic.comnvidia.com HGX B200 FP64 / FP64 Tensor Core; HGX page gives 296 TFLOPS per 8 GPUs. Epoch AI lists 31.
NVIDIA B200 SXM memory gb 180 GB Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai HGX B200 HBM3E per GPU; NVIDIA cites 192 GB as the Blackwell maximum capacity.
NVIDIA B200 SXM memory bw tbs 7.7 TB/s Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai HGX B200 per-GPU HBM3E bandwidth; the maximum Blackwell spec is quoted as 8 TB/s.
NVIDIA B200 SXM tdp w 1,000 W Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai HGX B200 configurable up to 1,000 W (GB200 variant up to 1,200 W).
NVIDIA B200 SXM scaleup domain accelerators 8 accelerators Vendor stated (dense) nvidia.comdam-cdn.nvd.orangelogic.com NVLink domain of the 8-GPU DGX B200 / HGX B200 baseboard.
NVIDIA B300 SXM bf16 tflops 2,250 TFLOPS Derived from vendor sparse figure dam-cdn.nvd.orangelogic.comepoch.ai Vendor lists a sparse figure; halved for dense. HGX B300 datasheet lists 4.5 PFLOPS sparse. Epoch AI lists 2,250.
NVIDIA B300 SXM fp8 tflops 4,500 TFLOPS Derived from vendor sparse figure dam-cdn.nvd.orangelogic.comepoch.ai Vendor lists a sparse figure; halved for dense. HGX B300 datasheet lists 9 PFLOPS sparse. Epoch AI lists 4,500.
NVIDIA B300 SXM fp4 tflops 13,500 TFLOPS Derived from vendor sparse figure nvidia.comdam-cdn.nvd.orangelogic.com HGX page: 108 PFLOPS dense across 8 GPUs; datasheet rounds to 14 PFLOPS per GPU. Epoch AI lists 14,000.
NVIDIA B300 SXM fp64 tflops 1.2 TFLOPS Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai FP64 / FP64 Tensor Core; Blackwell Ultra cuts FP64 sharply versus B200 (37 TFLOPS).
NVIDIA B300 SXM memory gb 270 GB Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai HGX B300 HBM3E per GPU; 288 GB is the Blackwell Ultra maximum.
NVIDIA B300 SXM memory bw tbs 7.7 TB/s Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai HGX B300 per-GPU HBM3E bandwidth.
NVIDIA B300 SXM tdp w 1,100 W Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai HGX B300 configurable up to 1,100 W (GB300 variant up to 1,400 W).
NVIDIA B300 SXM scaleup domain accelerators 8 accelerators Vendor stated (dense) nvidia.comdam-cdn.nvd.orangelogic.com NVLink domain of the 8-GPU HGX B300 baseboard.
NVIDIA GB200 NVL72 bf16 tflops 2,500 TFLOPS Derived from vendor sparse figure dam-cdn.nvd.orangelogic.comepoch.ai Per GPU (two-die package). Vendor lists a sparse figure; halved for dense. Datasheet lists 5 PFLOPS sparse.
NVIDIA GB200 NVL72 fp8 tflops 5,000 TFLOPS Derived from vendor sparse figure dam-cdn.nvd.orangelogic.comepoch.ai Per GPU. Vendor lists a sparse figure; halved for dense. Datasheet lists 10 PFLOPS sparse; AWS P6e states 360 PFLOPS dense FP8 for 72 GPUs.
NVIDIA GB200 NVL72 fp4 tflops 10,000 TFLOPS Derived from vendor sparse figure dam-cdn.nvd.orangelogic.comepoch.ai Per GPU. Vendor lists a sparse figure; halved for dense. Datasheet lists 20 PFLOPS sparse.
NVIDIA GB200 NVL72 fp64 tflops 40 TFLOPS Vendor stated (dense) dam-cdn.nvd.orangelogic.comnvidia.com Per GPU FP64 / FP64 Tensor Core; GB200 page lists 2,880 TFLOPS for the rack. Epoch AI lists 45.
NVIDIA GB200 NVL72 memory gb 186 GB Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai Per GPU HBM3E in NVL72 configuration (372 GB per Grace Blackwell superchip).
NVIDIA GB200 NVL72 memory bw tbs 8 TB/s Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai Per GPU HBM3E bandwidth (16 TB/s per superchip).
NVIDIA GB200 NVL72 tdp w 1,200 W Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai Per GPU, configurable up to 1,200 W.
NVIDIA GB200 NVL72 scaleup domain accelerators 72 accelerators Vendor stated (dense) nvidia.comaws.amazon.com 72 Blackwell GPUs in one NVLink domain; AWS P6e-GB200 UltraServers also state 72 GPUs per NVLink domain.
NVIDIA GB300 NVL72 bf16 tflops 2,500 TFLOPS Derived from vendor sparse figure dam-cdn.nvd.orangelogic.comepoch.ai Per GPU. Vendor lists a sparse figure; halved for dense. Datasheet lists 5 PFLOPS sparse.
NVIDIA GB300 NVL72 fp8 tflops 5,000 TFLOPS Derived from vendor sparse figure dam-cdn.nvd.orangelogic.comepoch.ai Per GPU. Vendor lists a sparse figure; halved for dense. Datasheet lists 10 PFLOPS sparse.
NVIDIA GB300 NVL72 fp4 tflops 15,000 TFLOPS Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai Per GPU dense NVFP4; datasheet lists 20 sparse | 15 dense PFLOPS. NVIDIA blog also states 15 PFLOPS dense.
NVIDIA GB300 NVL72 fp64 tflops 1.3 TFLOPS Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai Per GPU FP64 / FP64 Tensor Core. GB300 page lists 100 TFLOPS for the rack.
NVIDIA GB300 NVL72 memory gb 279 GB Vendor stated (dense) dam-cdn.nvd.orangelogic.comdeveloper.nvidia.com Datasheet per-GPU HBM3E; the NVIDIA blog and Epoch AI cite 288 GB as the maximum capacity.
NVIDIA GB300 NVL72 memory bw tbs 8 TB/s Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai Per GPU HBM3E bandwidth.
NVIDIA GB300 NVL72 tdp w 1,400 W Vendor stated (dense) dam-cdn.nvd.orangelogic.comepoch.ai Per GPU, configurable up to 1,400 W.
NVIDIA GB300 NVL72 scaleup domain accelerators 72 accelerators Vendor stated (dense) nvidia.comdam-cdn.nvd.orangelogic.com 72 Blackwell Ultra GPUs and 36 Grace CPUs in one NVLink domain.
NVIDIA Vera Rubin bf16 tflops 4,000 TFLOPS Vendor stated (dense) nvidia.comdam-cdn.nvd.orangelogic.com Per Rubin GPU, dense (page footnote). NVIDIA states performance is projected and subject to change.
NVIDIA Vera Rubin fp8 tflops 17,500 TFLOPS Vendor stated (dense) nvidia.comdam-cdn.nvd.orangelogic.com Per Rubin GPU, FP8/FP6 training figure, stated dense. Projected.
NVIDIA Vera Rubin fp4 tflops 35,000 TFLOPS Vendor stated (dense) nvidia.comdeveloper.nvidia.com Per GPU NVFP4 training figure, stated dense. The 50 PFLOPS NVFP4 inference figure is not dense and is not recorded.
NVIDIA Vera Rubin fp64 tflops 33 TFLOPS Vendor stated (dense) nvidia.comdeveloper.nvidia.com Per GPU hardware FP64 (vector). NVIDIA also cites 200 TFLOPS FP64 matrix via Tensor Core emulation, not recorded.
NVIDIA Vera Rubin memory gb 288 GB Vendor stated (dense) nvidia.comdam-cdn.nvd.orangelogic.com Per Rubin GPU HBM4.
NVIDIA Vera Rubin memory bw tbs 19.2 TB/s Vendor stated (dense) nvidia.comdam-cdn.nvd.orangelogic.com NVL72 page value; an older HGX page lists 22 TB/s, datasheet superchip 38.5 TB/s for two GPUs.
NVIDIA Vera Rubin scaleup domain accelerators 72 accelerators Vendor stated (dense) nvidia.comdam-cdn.nvd.orangelogic.com Vera Rubin NVL72: 72 Rubin GPUs and 36 Vera CPUs in one NVLink 6 domain. HGX Rubin NVL8 is 8.
AMD Instinct MI250X bf16 tflops 383 TFLOPS Vendor stated (dense) amd.comepoch.ai Per MI250X module (two GCDs). Epoch AI lists 383.
AMD Instinct MI250X fp64 tflops 95.7 TFLOPS Vendor stated (dense) amd.comdocs.olcf.ornl.gov Matrix-core FP64 per module; vector FP64 is 47.9. ORNL lists 47.9 matrix per GCD.
AMD Instinct MI250X memory gb 128 GB Vendor stated (dense) amd.comdocs.olcf.ornl.gov HBM2e per module. ORNL lists 64 GB per GCD.
AMD Instinct MI250X memory bw tbs 3.2 TB/s Vendor stated (dense) amd.comdocs.olcf.ornl.gov Per module. ORNL lists 1.6 TB/s per GCD.
AMD Instinct MI250X tdp w 500 W Vendor stated (dense) amd.comepoch.ai 500 W typical, 560 W peak per AMD.
AMD Instinct MI300X bf16 tflops 1,307.4 TFLOPS Vendor stated (dense) amd.comepoch.ai Dense; datasheet lists 1,307.4 dense and 2,614.9 with sparsity.
AMD Instinct MI300X fp8 tflops 2,614.9 TFLOPS Vendor stated (dense) amd.comamd.com Dense; datasheet lists 2,614.9 dense and 5,229.8 with sparsity.
AMD Instinct MI300X fp64 tflops 163.4 TFLOPS Vendor stated (dense) amd.comamd.com Matrix FP64; vector FP64 is 81.7 TFLOPS.
AMD Instinct MI300X memory gb 192 GB Vendor stated (dense) amd.comepoch.ai HBM3.
AMD Instinct MI300X memory bw tbs 5.3 TB/s Vendor stated (dense) amd.comepoch.ai Peak theoretical HBM3 bandwidth.
AMD Instinct MI300X tdp w 750 W Vendor stated (dense) amd.comepoch.ai Maximum typical board power.
AMD Instinct MI300X scaleup domain accelerators 8 accelerators Vendor stated (dense) amd.comamd.com Sold as an Instinct Platform of eight accelerators on Infinity Fabric.
AMD Instinct MI325X bf16 tflops 1,307.4 TFLOPS Vendor stated (dense) amd.comepoch.ai Dense. AMD page rounds to 1.3 PFLOPS (sparse 2.61); exact value from Epoch AI, same as MI300X at 2,100 MHz.
AMD Instinct MI325X fp8 tflops 2,614.9 TFLOPS Vendor stated (dense) amd.comepoch.ai Dense. AMD page rounds to 2.61 PFLOPS (sparse 5.22); exact value inferred from MI300X datasheet at the same 2,100 MHz.
AMD Instinct MI325X fp64 tflops 163.4 TFLOPS Vendor stated (dense) amd.comamd.com Matrix FP64; vector FP64 is 81.7 TFLOPS.
AMD Instinct MI325X memory gb 256 GB Vendor stated (dense) amd.comepoch.ai HBM3E.
AMD Instinct MI325X memory bw tbs 6 TB/s Vendor stated (dense) amd.comepoch.ai Peak theoretical HBM3E bandwidth.
AMD Instinct MI325X tdp w 1,000 W Vendor stated (dense) amd.comepoch.ai Typical board power, 1,000 W peak.
AMD Instinct MI355X bf16 tflops 2,516.6 TFLOPS Vendor stated (dense) amd.comepoch.ai Dense matrix BF16; brochure lists 2.5166 PFLOPS dense and 5.0332 with sparsity.
AMD Instinct MI355X fp8 tflops 5,033.2 TFLOPS Vendor stated (dense) amd.comamd.com Dense OCP-FP8; brochure lists 5.0332 dense and 10.0664 PFLOPS with sparsity.
AMD Instinct MI355X fp4 tflops 10,066.3 TFLOPS Vendor stated (dense) amd.comamd.com MXFP4; the brochure lists no sparse variant for FP4 (10.0663 PFLOPS).
AMD Instinct MI355X fp64 tflops 78.6 TFLOPS Vendor stated (dense) amd.comamd.com FP64 matrix and vector are both listed at 78.6 TFLOPS.
AMD Instinct MI355X memory gb 288 GB Vendor stated (dense) amd.comepoch.ai HBM3E.
AMD Instinct MI355X memory bw tbs 8 TB/s Vendor stated (dense) amd.comepoch.ai Peak theoretical HBM3E bandwidth.
AMD Instinct MI355X tdp w 1,400 W Vendor stated (dense) amd.comepoch.ai Typical board power (TBP).
AMD Instinct MI355X scaleup domain accelerators 8 accelerators Vendor stated (dense) amd.comamd.com Brochure describes an 8-GPU MI355X platform on Infinity Fabric.
AMD Instinct MI430X fp64 tflops 288 TFLOPS Vendor stated (dense) amd.comamd.com AMD states up to 288 TFLOPS hardware-based peak theoretical FP64.
AMD Instinct MI430X memory gb 432 GB Vendor stated (dense) amd.comamd.com Integrated HBM4, stated as up to 432 GB.
AMD Instinct MI430X memory bw tbs 23.3 TB/s Vendor stated (dense) amd.comamd.com Peak theoretical bandwidth, stated as up to 23.3 TB/s.
AMD Instinct MI455X bf16 tflops 5,000 TFLOPS Vendor stated (dense) amd.comamd.com Page lists 5 PFLOPS dense BF16 matrix (10.1 PFLOPS with sparsity); stated to one significant figure.
AMD Instinct MI455X fp64 tflops 5 TFLOPS Vendor stated (dense) amd.comamd.com Page lists matrix and vector FP64 at 5 TFLOPS, far below MI430X (288).
AMD Instinct MI455X memory gb 432 GB Vendor stated (dense) amd.comamd.com HBM4, 12 stacks.
AMD Instinct MI455X memory bw tbs 23.3 TB/s Vendor stated (dense) amd.comamd.com Peak theoretical HBM4 bandwidth.
Google TPU v4 bf16 tflops 275 TFLOPS Vendor stated (dense) cloud.google.comepoch.ai Peak per chip (bf16 or int8), two TensorCores. Epoch AI lists 275.
Google TPU v4 memory gb 32 GB Vendor stated (dense) cloud.google.comepoch.ai Stated as 32 GiB HBM2 (about 34.4 GB).
Google TPU v4 memory bw tbs 1.2 TB/s Vendor stated (dense) cloud.google.comepoch.ai Stated as 1,200 GBps.
Google TPU v4 scaleup domain accelerators 4,096 accelerators Vendor stated (dense) cloud.google.comarxiv.org TPU v4 pod of 4,096 chips on a reconfigurable optical 3D torus.
Google TPU v5p bf16 tflops 459 TFLOPS Vendor stated (dense) cloud.google.comepoch.ai Peak per chip. Google's FP8 figure is emulated on v5p, so FP8 is not recorded.
Google TPU v5p memory gb 95 GB Vendor stated (dense) cloud.google.comepoch.ai Stated as 95 GiB HBM (about 102 GB).
Google TPU v5p memory bw tbs 2.765 TB/s Vendor stated (dense) cloud.google.comepoch.ai Stated as 2,765 GBps.
Google TPU v5p scaleup domain accelerators 8,960 accelerators Vendor stated (dense) cloud.google.com Pod of 8,960 chips on a 3D torus; the largest schedulable job is 6,144 chips.
Google TPU v5e bf16 tflops 197 TFLOPS Vendor stated (dense) cloud.google.comepoch.ai Peak per chip (one TensorCore). Epoch AI lists 197.
Google TPU v5e memory gb 16 GB Vendor stated (dense) cloud.google.comepoch.ai HBM per chip.
Google TPU v5e scaleup domain accelerators 256 accelerators Vendor stated (dense) cloud.google.com Pod size 256 chips on a 2D torus.
Google TPU v6e (Trillium) bf16 tflops 918 TFLOPS Vendor stated (dense) cloud.google.comepoch.ai Peak per chip (one TensorCore). Epoch AI lists 918.
Google TPU v6e (Trillium) memory gb 32 GB Vendor stated (dense) cloud.google.comepoch.ai HBM per chip.
Google TPU v6e (Trillium) memory bw tbs 1.638 TB/s Vendor stated (dense) cloud.google.comepoch.ai Stated as 1,638 GBps.
Google TPU v6e (Trillium) scaleup domain accelerators 256 accelerators Vendor stated (dense) cloud.google.com Pod size 256 chips on a 2D torus.
Google TPU7x (Ironwood) bf16 tflops 2,307 TFLOPS Vendor stated (dense) cloud.google.comepoch.ai Peak per chip. Epoch AI lists 2,307.
Google TPU7x (Ironwood) fp8 tflops 4,614 TFLOPS Vendor stated (dense) cloud.google.comepoch.ai Peak per chip, native FP8 on Ironwood. Epoch AI lists 4,614.
Google TPU7x (Ironwood) memory gb 192 GB Vendor stated (dense) cloud.google.comepoch.ai Stated as 192 GiB HBM (about 206 GB).
Google TPU7x (Ironwood) memory bw tbs 7.38 TB/s Vendor stated (dense) cloud.google.comepoch.ai Docs state 7,380 GBps; Google's blog and Epoch AI state 7.37 TB/s.
Google TPU7x (Ironwood) scaleup domain accelerators 9,216 accelerators Vendor stated (dense) cloud.google.comblog.google Pod of 9,216 chips on a 3D torus.
Google TPU 8t fp4 tflops 12,600 TFLOPS Vendor stated (dense) cloud.google.com Google lists peak FP4 of 12.6 PFLOPS with no sparsity qualifier; TPUs have no structured-sparsity matmul.
Google TPU 8t memory gb 216 GB Vendor stated (dense) cloud.google.com HBM capacity.
Google TPU 8t memory bw tbs 6.528 TB/s Vendor stated (dense) cloud.google.com Stated as 6,528 GB/s.
Google TPU 8t scaleup domain accelerators 9,600 accelerators Vendor stated (dense) cloud.google.com Single 3D-torus superpod of 9,600 chips.
Google TPU 8i fp4 tflops 10,100 TFLOPS Vendor stated (dense) cloud.google.comepoch.ai Google lists peak FP4 of 10.1 PFLOPS with no sparsity qualifier. Epoch AI lists 10.1 PFLOPS.
Google TPU 8i memory gb 288 GB Vendor stated (dense) cloud.google.comepoch.ai HBM capacity.
Google TPU 8i memory bw tbs 8.601 TB/s Vendor stated (dense) cloud.google.comepoch.ai Stated as 8,601 GB/s.
Google TPU 8i scaleup domain accelerators 1,024 accelerators Vendor stated (dense) cloud.google.com Boardfly pod of up to 1,024 active chips.
AWS Trainium2 bf16 tflops 667 TFLOPS Vendor stated (dense) awsdocs-neuron.readthedocs-hosted.comepoch.ai Dense BF16/FP16/TF32; Neuron docs list 2,563 TFLOPS separately with 4x sparsity.
AWS Trainium2 fp8 tflops 1,299 TFLOPS Vendor stated (dense) awsdocs-neuron.readthedocs-hosted.comepoch.ai Dense FP8; the sparse figure is listed separately.
AWS Trainium2 memory gb 96 GB Vendor stated (dense) awsdocs-neuron.readthedocs-hosted.comepoch.ai Stated as 96 GiB HBM per chip.
AWS Trainium2 memory bw tbs 2.9 TB/s Vendor stated (dense) awsdocs-neuron.readthedocs-hosted.comepoch.ai Stated as 2.9 TB/s.
AWS Trainium2 scaleup domain accelerators 64 accelerators Vendor stated (dense) aws.amazon.com Trn2 UltraServer links 64 chips over NeuronLink; a single Trn2 instance has 16.
AWS Trainium3 bf16 tflops 671 TFLOPS Vendor stated (dense) awsdocs-neuron.readthedocs-hosted.comepoch.ai Dense BF16/FP16/TF32; Neuron docs list a separate sparse figure.
AWS Trainium3 fp8 tflops 2,517 TFLOPS Vendor stated (dense) awsdocs-neuron.readthedocs-hosted.comepoch.ai MXFP8 dense, listed as 2,517 TFLOPS.
AWS Trainium3 fp4 tflops 2,517 TFLOPS Vendor stated (dense) awsdocs-neuron.readthedocs-hosted.comepoch.ai MXFP4 listed at the same 2,517 TFLOPS as MXFP8.
AWS Trainium3 memory gb 144 GB Vendor stated (dense) awsdocs-neuron.readthedocs-hosted.comaws.amazon.com Stated as 144 GiB (docs) or 144 GB (instance page) of HBM3e.
AWS Trainium3 memory bw tbs 4.9 TB/s Vendor stated (dense) awsdocs-neuron.readthedocs-hosted.comaws.amazon.com Stated as 4.9 TB/s.
AWS Trainium3 scaleup domain accelerators 144 accelerators Vendor stated (dense) aws.amazon.comaws.amazon.com Trn3 UltraServer scales to 144 chips on a NeuronSwitch all-to-all fabric.
Microsoft Maia 100 memory gb 64 GB Vendor stated (dense) techcommunity.microsoft.com HBM2E, four stacks.
Microsoft Maia 100 memory bw tbs 1.8 TB/s Vendor stated (dense) techcommunity.microsoft.com Stated as 1.8 TB/s.
Microsoft Maia 100 tdp w 700 W Vendor stated (dense) techcommunity.microsoft.com Designed for up to 700 W but provisioned at 500 W.
Microsoft Maia 200 fp4 tflops 10,000 TFLOPS Vendor stated (dense) blogs.microsoft.comepoch.ai Microsoft states over 10 PFLOPS FP4; Epoch AI lists 10,145. Recorded as the stated lower bound.
Microsoft Maia 200 fp8 tflops 5,000 TFLOPS Vendor stated (dense) blogs.microsoft.comepoch.ai Microsoft states over 5 PFLOPS FP8; Epoch AI lists 5,072. Recorded as the stated lower bound.
Microsoft Maia 200 memory gb 216 GB Vendor stated (dense) blogs.microsoft.comepoch.ai HBM3e.
Microsoft Maia 200 memory bw tbs 7 TB/s Vendor stated (dense) blogs.microsoft.comepoch.ai Stated as 7 TB/s.
Microsoft Maia 200 tdp w 750 W Vendor stated (dense) blogs.microsoft.comepoch.ai Stated as 750 W SoC TDP.
Meta MTIA v2 bf16 tflops 177 TFLOPS Vendor stated (dense) ai.meta.comepoch.ai Dense FP16/BF16; Meta lists 354 TFLOPS with sparsity.
Meta MTIA v2 memory gb 128 GB Vendor stated (dense) ai.meta.com Off-chip LPDDR5 (not HBM); on-chip SRAM is 256 MB.
Meta MTIA v2 memory bw tbs 0.205 TB/s Vendor stated (dense) ai.meta.com Off-chip LPDDR5 bandwidth of 204.8 GB/s; on-chip SRAM is 2.7 TB/s.
Meta MTIA v2 tdp w 90 W Vendor stated (dense) ai.meta.comepoch.ai Stated as 90 W.
Huawei Ascend 910C bf16 tflops 781 TFLOPS Third party newsletter.semianalysis.comtrendforce.com SemiAnalysis: CloudMatrix 384 gives 300 PFLOPS dense BF16 across 384 chips (300000/384). TrendForce reports 800 TFLOPS FP16. No Huawei datasheet.
Huawei Ascend 910C memory bw tbs 3.2 TB/s Third party trendforce.com TrendForce reports 3.2 TB/s citing an industry report; no Huawei datasheet fetched.
Huawei CloudMatrix 384 scaleup domain accelerators 384 accelerators Vendor stated (dense) arxiv.orgnewsletter.semianalysis.com 384 Ascend NPUs in one all-to-all Unified Bus supernode; SemiAnalysis independently states 384 Ascend 910C chips.
Intel Gaudi 3 bf16 tflops 1,678 TFLOPS Vendor stated (dense) cdrdv2-public.intel.comepoch.ai BF16 MME peak, listed with no sparsity qualifier.
Intel Gaudi 3 fp8 tflops 1,678 TFLOPS Vendor stated (dense) cdrdv2-public.intel.comepoch.ai FP8 MME peak, listed with no sparsity qualifier.
Intel Gaudi 3 memory gb 128 GB Vendor stated (dense) cdrdv2-public.intel.comepoch.ai HBM2e, eight stacks.
Intel Gaudi 3 memory bw tbs 3.7 TB/s Vendor stated (dense) cdrdv2-public.intel.comepoch.ai Peak HBM bandwidth.
Intel Gaudi 3 tdp w 900 W Vendor stated (dense) cdrdv2-public.intel.comepoch.ai OAM module, up to 900 W with passive or liquid cooling.
Cerebras WSE-3 bf16 tflops 62,500 TFLOPS Derived from vendor sparse figure cdn.sanity.ioarxiv.org Datasheet gives 125 PFLOPS FP16 marked as sparse; halved for dense. BF16 is not listed separately.
Cerebras WSE-3 memory gb 44 GB Vendor stated (dense) cdn.sanity.ioarxiv.org On-chip SRAM only; the wafer has no HBM. An independent arXiv paper also states 44 GB.
Cerebras WSE-3 memory bw tbs 21,000 TB/s Vendor stated (dense) cdn.sanity.ioarxiv.org Stated as 21 PB/s on-chip SRAM bandwidth; an independent arXiv paper also states 21 PB/s.