COMPUTE ATLAS Supercomputer supply-chain graph
1153 systems 482 sites Sourced data

Blog/2026-10-07

· Compute Atlas

No TOP500 for quantum: what each metric is worth

Quantum computers have no shared ranking. Our table holds 75 benchmark rows for 35 machines, and all 75 are vendor-reported or peer-reviewed, none labelled as an independent measurement. Here is what each metric means and what it can't tell you.


The TOP500 works because everyone runs one program. Machines solve a dense system of linear equations with the HPL implementation of LINPACK, the list is re-ranked twice a year, and the page itself warns that the score does not describe overall system performance. Quantum computing has nothing like it. Our own quantum benchmarks table holds 75 rows covering 35 machines: 59 are vendor-reported and 16 come from peer-reviewed papers, and none is labelled as an independent measurement of someone else’s hardware. The metrics also differ in kind. Quantum volume tests a whole device on random circuits, algorithmic qubits tests it on a set of named algorithms, error per layered gate (EPLG) tests how well a long chain of qubits works together, and CLOPS measures speed, not accuracy. What we do not know is which of these, if any, predicts performance on a useful workload, because no machine has yet run one that a classical computer could not also run.

Why a single ranking never formed

HPL is a fixed kernel, so a petaflop is a petaflop on any vendor’s hardware. Quantum hardware differs at the physical level: superconducting chips with fixed neighbours, trapped ions that can interact with any partner, neutral atoms, photons and annealers. A test that favours one design penalises another, and the field’s main measurable property, how often a gate goes wrong, moves with every recalibration. Meanwhile each vendor has an incentive to report the figure its machine flatters. The table that results is a set of self-descriptions in different units, and the first job of a reader is to ask which question each number answers.

The site’s rows carry a basis label for that reason. A peer-reviewed row means a paper, usually written by the team that built the machine, passed review. It does not mean a third party reproduced the number. The only wider checks on quantum claims have come from classical-simulation groups, which we cover below.

Quantum volume: one number for the whole stack

IBM researchers defined quantum volume in 2018 as a single number from a concrete protocol, the largest random circuit of equal width and depth that the computer implements successfully. It reflects gate fidelity, connectivity, the set of calibrated gates and compiler quality together. A test is passed when the average heavy-output probability sits above 2/3 with two-sigma confidence, as Quantinuum describes the procedure. A QV of 2^n means an n-qubit, n-layer circuit passed.

The progression in our table shows how the metric has been used:

MachineQVDateSource
Fraunhofer IBM Quantum System One (27 qubits)322021Fraunhofer
IQM Garnet32 (2^5)2024IQM whitepaper
Quantinuum H1-11,048,576 (2^20)Apr 2024Quantinuum
AQT Lynx32,768 (2^15)May 2026AQT
Quantinuum H233,554,432 (2^25)Sep 2025QCR

Two things follow. First, the recent QV records in our table all come from trapped-ion vendors. AQT’s test used 15 qubits, 305 circuits and about 173 minutes, and the company calls the result the second-highest QV globally. Second, the metric is bounded by classical simulation, because, as we read the protocol, “heavy” outputs are defined from each circuit’s ideal output distribution, which has to be computed classically. The original paper described near-term machines of modest size, up to roughly 50 qubits. By the definition, Quantinuum’s 2^25 figure involves 25-qubit circuits on a 56-qubit machine. QV is therefore a good measure of how clean a machine’s gates and connectivity are on small circuits and a poor guide to what a large machine can do. Quantinuum itself notes the benchmark is sensitive to qubit number, fidelity and connectivity, which is another way of saying all-to-all trapped-ion designs have a built-in advantage on it. IBM, whose researchers invented QV, no longer appears in our table with a recent QV figure, and we did not find a stated reason in the pages we fetched.

Algorithmic qubits: the application-flavoured alternative

IonQ’s #AQ comes from a different idea. The company defines #AQ = N as the largest circuit you can complete with N qubits and N-squared two-qubit gates, built from algorithms in the benchmark suite of the Quantum Economic Development Consortium. A circuit passes when its fidelity, minus statistical error, exceeds 1/e, about 0.37. IonQ states that #AQ has no formal relationship to QV. The underlying QED-C paper describes volumetric benchmarks that track fidelity against circuit width and depth on well-known algorithms and small applications.

Our table lists #AQ 9 for Harmony, 25 for Aria and 36 for Forte on IonQ’s own pages, and 64 on the 100-qubit Tempo, reported by Quantum Computing Report. The pass marks differ (1/e here, a 2/3 heavy-output probability for QV) and so do the circuits, so #AQ 64 and QV 2^25 cannot be compared. In our table only IonQ reports #AQ, so the number does not support cross-vendor comparison.

EPLG and CLOPS: IBM’s quality-at-scale and speed numbers

Two further metrics come from IBM. Layer fidelity, proposed in 2023 by McKay and colleagues, runs simultaneous randomized benchmarking over a connected chain of qubits and converts the result into an error per layered gate that does not depend on chain length. They tested chains of 80 and 100 qubits on IBM’s 127-qubit Eagle and 133-qubit Heron processors and reported EPLG between 1.7 and 1.2 percent. The value of EPLG is that it penalises hardware whose best qubit pair is good and whose typical pair is not, which a best-gate fidelity hides.

IBM now publishes a column called “2Q error (layered)” next to “2Q error (best)” on its compute resources page. In our table, Heron appears at 0.2921 percent on ibm_boston and Nighthawk at 0.2367 percent on ibm_phoenix. When we re-read the page on 7 October 2026, it listed 2.301E-3 for ibm_boston and 2.855E-3 for ibm_phoenix, so the figures move with each calibration and any single value is a snapshot. The best single-coupler error on ibm_boston is roughly a fifth of its layered error, a reminder of the gap between best-case and typical numbers. Whether IBM’s layered column is exactly the EPLG of the 2023 paper is how our table treats it; the page itself gives no definition.

CLOPS, proposed in a 2021 IBM paper on quality, speed and scale, counts circuit layers per second and captures both the classical control stack and the quantum hardware. IBM’s page lists 340K for Heron r2 and r3 and 2.0M for Nighthawk r2. IQM’s whitepaper reports 2,600 “virtual CLOPS” on Garnet. Because CLOPS depends on the software and control pipeline, a trapped-ion machine whose gates take far longer will score much lower without being worse at accuracy. Speed matters for error mitigation, which needs many repeated runs, and tells you nothing about whether the answers are right.

Aggregators and the government referee

Two efforts try to replace vendor self-reporting. Metriq, maintained by the Unitary Foundation, runs a common benchmark suite called metriq-gym across cloud providers, taking inspiration from Geekbench and MLCommons. A March 2026 paper describes the platform, a composite Metriq Score, and results collected from more than ten quantum computers from multiple vendors. It is the closest thing to a TOP500 mechanism, since results come from executed benchmarks rather than press releases. We do not yet include Metriq scores in our table, and we could not confirm how many devices or runs the live site currently holds.

DARPA’s Quantum Benchmarking Initiative takes a different approach. Its goal is to determine whether an industrially useful quantum computer is possible by 2033, meaning computational value exceeds cost, and it applies third-party verification to each performer’s path. The stage structure runs from concept (Stage A) to R&D plan (Stage B) to government verification of whether the machine can be built and run as designed (Stage C). On 6 November 2025 DARPA named 11 companies for Stage B: Atom Computing, Diraq, IBM, IonQ, Nord Quantique, Photonic Inc., Quantinuum, Quantum Motion, QuEra, Silicon Quantum Computing and Xanadu. Google is not on that list. QBI is the only evaluation we found where an outside team will judge hardware against a cost-benefit standard instead of a gate fidelity, but it does not publish scores and Stage B is a planning stage, so for now it tells us who is being examined, not how they rank.

Advantage claims and the classical reply

The benchmarks above all measure the machine. A quantum advantage claim measures the machine against a classical computer, and that is where numbers have been revised most often.

Google’s 2019 Sycamore claim was 200 seconds for a task estimated at 10,000 years on the world’s fastest supercomputer. IBM replied within weeks that such circuits could be simulated in days on Summit using secondary storage. In 2021, Pan and Zhang generated one million correlated bitstrings from the 53-qubit, 20-cycle circuit on 60 GPUs with a linear cross-entropy fidelity of 0.739, higher than Google’s original score. Scott Aaronson later described the 2019 experiment as long superseded by newer ones.

IBM’s own 2023 Nature paper reported accurate expectation values on a 127-qubit processor beyond brute-force classical simulation. Tindall, Fishman, Stoudenmire and Sels then showed a belief-propagation tensor-network method that, by the authors’ account, surpassed the quantum processor’s output in accuracy and precision for that experiment. D-Wave’s 2025 Science paper claimed a result needing nearly one million years on Frontier by matrix-product-state methods; Flatiron researchers published a classical tensor-network approach, and D-Wave rebutted on four points in a May 2026 response reported by QCR, including that the classical work covered only small sectors and avoided its largest geometries. A 2026 review of the IBM, D-Wave and Google cases from a tensor-network standpoint does not declare either side the winner and says improved classical methods will raise the bar for later claims.

The latest cycle has institutional structure. IBM and partners launched the open Quantum Advantage Tracker in February 2026, with 30 submissions on IBM and Quantinuum hardware and an admission that new classical methods overtook quantum runtimes within weeks on the peaked-circuits example. On 30 July 2026 IBM and Algorithmiq claimed advantage for a simulation of heterogeneous matter that had stood on the tracker for eight months, while conceding that classical methods disagreed with each other and no exact solution exists. Google’s October 2025 Quantum Echoes claim of a 13,000-fold speedup is, so far, the case where later work supports the claim: a 2026 preprint argues tensor networks with belief propagation cannot simulate it, although the author list includes Google Quantum AI researchers. We cover that result in our Willow deep dive.

The pattern is that an advantage number is an estimate of the best known classical cost on the day of the claim. USTC’s Zuchongzhi 3.0 estimate of 6.4 billion years on Frontier and Quantinuum’s 100x XEB claim over Google 2019 belong in the same bucket, as vendor estimates against classical cost.

What each number is worth

Quantum volume is a reasonable indicator of small-circuit hardware quality: the gap between 2^25 and 2^5 reflects error rates and connectivity together, though QV does not separate the two. It says little about machines above about 50 qubits. #AQ is a single-vendor figure with a lenient pass threshold. EPLG and the layered error are the most informative public numbers about large superconducting chips, because they test many qubits at once, but IBM publishes the most and the numbers drift. CLOPS measures throughput. Advantage claims are the only measures that reach toward useful work, and they have a short half-life.

Gate fidelity, which fills 27 rows of our table, is the most common figure and the easiest to misread. Rigetti’s 99.5 percent for Ankaa-3 is a whole-device median, while IBM’s 99.9385 percent for Heron is a best single coupler. Quantinuum’s 99.921 percent for Helios is stated across all qubit pairs. They are different statistics and should not be ranked against each other.

What we could not confirm

We could not find a shared workload that a majority of vendors run, nor any independent audit of a vendor-reported QV, #AQ or EPLG figure. We did not confirm the current number of devices on Metriq. We could not verify IBM’s definition of its “2Q error (layered)” column. We found no QV figure for Quantinuum’s Helios. The research note on benchmarks that we were asked to reuse was not at the path provided, so every source here was fetched anew. Stage C of DARPA’s program has been reported in the press, but we did not fetch a primary source and have left it out.

What to watch

Watch whether Metriq or a QBI-style evaluator publishes scores for named machines, whether Google or IBM move from “advantage candidate” to a result no classical group has matched a year later, and whether the Quantum Advantage Tracker’s challenge cycle shortens. Until a standard kernel exists, the safest reading rule is to ask three questions of any quantum number: who measured it, on how many qubits, and against what classical baseline.

Sources

deep-divequantumbenchmarks