CPU vs GPU vs TPU vs QPU: Choosing a Quantum Backend
Last updated · 15 min read · ZKSF team
The short version
- The chip follows the METHOD, not the other way round. A statevector wants a GPU, a stabilizer circuit a CPU, a neural wavefunction a TPU
- GPUs do not unlock more qubits. A statevector needs 16 x 2^n bytes whoever holds it, so 141 GB buys two or three qubits over a good CPU
- Below about 30 qubits the classical answer is better, because it is the noise-free one
- A QPU answers a different question, which is what a physical device does to your circuit
Everything here is runnable on your own circuit. Try it in the console
Every quantum job faces the same three-way choice: classical CPU, classical GPU, or a physical quantum processor. The choice is frequently presented as a matter of preference, or as a progression in which the QPU is the destination and the others are waiting rooms. Neither framing survives contact with the numbers.
The three are answering different questions, and the figures below come from our own runs across all three tiers rather than from vendor literature.
The distinction that resolves most confusion
A CPU and a GPU are both classical processors executing the same arithmetic; a GPU does it with more parallel lanes. A QPU is not a faster classical processor. It is a physical apparatus whose observable behaviour is the result.
This has a consequence people find counterintuitive: at every size a classical machine can reach, the classical machine gives the *better* answer, because it gives the noise-free one. A QPU at 20 qubits does not compute anything a laptop cannot compute exactly and more quickly. What it provides is evidence about a physical device.
If the question is what a quantum algorithm outputs, simulate. If the question is what a particular piece of hardware does when asked to run that algorithm, use hardware. Conflating these two questions is the most expensive error available in this field.
CPU: the default, and more capable than qubit count suggests
CPU simulation is correct whenever the representation fits in memory, and it fits more often than raw qubit count implies, because the relevant variable is circuit structure rather than width.
- Exact statevector, to roughly 32 qubits: milliseconds to seconds, exact to floating-point precision
- Clifford circuits: thousands of qubits, exact, effectively free. Error-correction and stabilizer work lives here
- Structured circuits via tensor networks: 50 to 128 qubits in seconds when entanglement stays modest, which covers most QAOA, VQE and Trotterised dynamics workloads
- Shallow circuits via Pauli propagation: expectation values on hundreds of qubits
Our benchmark suite makes the range concrete. Every row was produced on a single consumer laptop (Intel i7-12700H, 32 GB RAM) with the GPU switched off, so it represents a floor rather than a ceiling:
Circuit Qubits Engine (auto) Wall time Accuracy
GHZ (Clifford) 5,000 clifford 0.56 s exact
QAOA MaxCut, p=3 100 mps.quimb.cpu 5.9 s converged (dev 0.0)
Layered ansatz 80 mps.quimb.cpu 4.3 s converged (dev 0.0)
Exact statevector 26 exact.cpu 2.7 s exactThe router selected each engine from the circuit's structure, and every approximate result carried a convergence check: the run was repeated at double the bond dimension and the top outcome probabilities did not move, indicating the compression captured the state. None of these circuits used a GPU, a cluster, or any quantum hardware.
GPU: throughput, not additional qubits
Against a well-provisioned CPU machine that is a gain of two or three qubits, which the exponential erases immediately.
Device memory Exact statevector ceiling
16 GB 29 qubits
24 GB 30 qubits
80 GB 32 qubits
141 GB 33 qubitsWhat a GPU provides is speed, roughly 10 to 50 times faster on the dense linear algebra behind statevector updates and tensor contractions. That matters when many circuits run in sequence rather than when one circuit runs large: parameter sweeps, QML training loops, batched noise studies, and above all variational optimisation, where a single 150-iteration SPSA run is 301 separate circuit evaluations. At roughly $3 per GPU-hour with per-second billing, a 20-minute sweep costs about a dollar.
Two production runs illustrate both halves of that statement. An 8-qubit GHZ at 1,000 shots returned the expected near-even split (502 all-zeros against 498 all-ones, textbook shot noise around the ideal 500/500) in about half a second.
A 32-qubit GHZ at 100 shots held a full 64 GiB statevector, larger than a 32 GB laptop can allocate, and completed exactly in 8.2 seconds. The GPU did not change either answer. It moved the memory ceiling out by a few qubits and returned the result faster, which is precisely its role.
Circuit Qubits Shots Wall time Result
GHZ (exact.gpu) 8 1,000 0.5 s 502/498, ideal 500/500
GHZ (exact.gpu) 32 100 8.2 s 52/48; full 64 GiB statevectorThe corollary is that a GPU is the wrong purchase for a qubit-count problem and the right one for an iteration-count problem. Most people who want more qubits need a different method, not a different processor.
Where the two tiers cross, measured
The claim above is a shape rather than a number, so here is the number. The same batch was run on both tiers at five widths: 60 circuits per point at 1,024 shots, which is one optimiser step of a small variational model. The figures are engine compute seconds, not round trip, so queueing and transport are excluded.
Qubits exact.cpu exact.gpu
2 2.888 5.372
8 3.143 5.442
14 3.421 5.396
20 6.076 5.691
24 59.914 6.124The GPU is close to flat across the whole range, because at these widths fixed overhead dominates and the statevector is small enough that the card is never the constraint. The CPU is faster below about 20 qubits for the same reason in reverse: there is no transfer to pay for. They cross at roughly 20 qubits, and by 24 the CPU is 9.8 times slower.
Both tiers billed the same, because every circuit in the batch lands on the $0.0001 per-circuit minimum and the batch is 60 circuits either way. Above the crossing the GPU is therefore about ten times faster for identical money, which is the part that decides where a training loop should run. Below it, the CPU is the cheaper habit and the faster one.
TPU: the one that does not belong here
Any comparison of processor types written for a general audience lists four: CPU, GPU, TPU, QPU. Three of them belong together and one does not, and it is worth saying which, because the grouping is a habit of vocabulary rather than a technical fact.
A Tensor Processing Unit is an application-specific chip Google designed for neural networks, available by the hour on their cloud and not sold.
Where a GPU is a general parallel processor that turned out to suit machine learning, a TPU is built for the one operation that dominates it: large matrix multiplication, executed on a systolic array that streams operands through a grid of multiply-accumulate units rather than round-tripping them to memory between steps. That single specialisation is where the efficiency comes from, and it is also the whole limitation.
The decisive detail for our purposes is precision. TPUs are optimised for reduced-precision formats, bfloat16 and int8, because a neural network's accuracy survives a truncated mantissa. The network was fitted to noisy data and its output is a probability, so a fractionally wrong weight changes nothing that matters.
Statevector simulation is the opposite case. It carries 2^n complex amplitudes whose relative phases are the entire content of the calculation, accumulating error over every gate, and the quantity we ultimately publish is a bound on how far the answer sits from the truth. A chip engineered to be approximately right is a poor instrument for work whose product is a measured error.
Built for Arithmetic Fit for statevector
CPU anything fp32/fp64 complex yes, the reference
GPU dense parallel linear algebra fp32/fp64 complex yes, 10-50x faster
TPU neural network matmul bfloat16/int8 real yes, to 29q in complex64
QPU nothing classical can do not applicable a different questionNone of this makes the TPU a lesser chip. It is extremely good at the thing it exists for, and a comparison that treats all four as interchangeable options for the same job misses what each one is for.
What changed. We now run a TPU engine, and the argument above is why
When this post was written it ended by saying this service offers CPU and GPU engines and no TPU engine. That is no longer true, and the reason is more interesting than a simple reversal. Nothing above turned out to be wrong. A TPU is still the wrong shape for a statevector, for exactly the reasons given.
What we found was a different method that wants precisely what a TPU is good at.
A neural network quantum state represents the wavefunction as a neural network and optimises it by variational Monte Carlo. The inner loop is not a complex statevector being contracted; it is dense matrix multiplication against batches of sampled spin configurations, which is the workload a TPU was built for.
It also reaches somewhere the other classical engines cannot: tensor networks assume entanglement follows an area law, and a state that breaks that assumption is where an MPS gives up and a neural ansatz keeps going.
That is the whole reason the tier exists. The runs behind it, on CPU and on the TPU with the same H2 molecule on both, are on neural network quantum states, and the argument for why the bound rather than the energy is the number to read is in neural network quantum states with a bound.
The precision point in this post was not academic, either. We hit it directly, as a hard refusal from the hardware: Element type C128 is not supported on TPU.
The engine narrows to complex64 there and records on the certificate that it did, because reduced precision moves the variance floor, and a variance sitting at the arithmetic floor must not be read as a converged ground state. The chip's reduced-precision bias is real; it is survivable for this method and would not be for a 30-qubit statevector.
A measured run, so the claim is checkable. An 8-spin transverse-field Ising ring on one TPU v5e chip: 60.4 seconds of variational Monte Carlo, three independent restarts agreeing to 1.1e-3, Gelman-Rubin R-hat of 1.0014. It returned a ceiling of -10.2499758 on the true ground-state energy, and the exact answer for that Hamiltonian is -10.2516623, so the ceiling holds.
And then we built the statevector engine anyway, 25 September 2026
exact.tpu does on a TPU the thing this post says a TPU is poor at, so it owes an explanation. The precision objection stands: it holds the state in complex64, where a CPU or GPU engine holds complex128, and against Aer at 18 qubits and 1,499 gates that is a difference of 7e-7 rather than the 2e-15 the same engine reaches in double precision. For sampling shot counts in the thousands that difference is far below shot noise; for anything that needs the sixteenth decimal it is not, and the certificate records which precision ran.
The real limit turned out to be memory shape rather than arithmetic. A TPU pads the last two dimensions of every array, so the obvious way of writing a gate, reshaping the state to put the target qubit last, padded roughly five hundredfold and asked for 287 GB at 26 qubits. Laying the state out as fixed-size rows instead, and reaching the low qubits by matrix multiplication rather than by reshaping, brought it back to 29 qubits on one v5e chip. It refuses 30.
So the ceiling is 29 qubits against 30 on a CPU and 32 on a GPU, which is the same conclusion the table above reaches: a TPU does not buy width. What it buys is a third kind of machine for the same circuit, and a batch that shares one machine. The published runs are on the applications pages, where the same QAOA instances that ran on CPU, GPU and five QPUs now carry a TPU row, including the satellite instance where it returned the smallest gap on the table and the logistics instance where it found the exact optimum.
The point estimate itself landed 0.00063 BELOW the exact value, which is 1.4 standard errors of ordinary Monte Carlo noise and is why the certified claim is the ceiling rather than the estimate.
So the four-way comparison resolves differently than it first appears. CPU, GPU and TPU are classical processors differing in how they parallelise arithmetic, and the QPU is not a processor in that sense at all.
But which classical chip suits you is decided by the METHOD, not by the chip: a statevector wants a GPU, a stabilizer circuit wants a CPU, and a neural wavefunction wants a TPU. Asking which chip is fastest, without naming the method, remains the wrong question.
The useful way to read the four is that the first three are classical processors differing in how they parallelise identical arithmetic, and the fourth is not a processor in that sense at all.
QPU: physics, not compute
Real hardware is not a faster simulator. At accessible sizes it is slower, noisier and more expensive: $0.30 per task plus per-shot fees, plus a queue measured in minutes to hours. What it uniquely provides is physical reality.
Hardware spend is justified in three cases, and it is worth being strict about them.
- Measuring algorithmic degradation under genuine device noise. A depolarizing model in a simulator is an approximation of a device's actual error process, which includes crosstalk, leakage, drift and correlated errors that no simple model captures. When the noise itself is the object of study, only the device will do
- Error-correction experiments requiring physical qubits by definition. A logical qubit's performance is a claim about hardware
- Circuits whose entanglement structure defeats every classical method, a condition that should be verified rather than assumed, since the verification is cheap and the assumption is not
The same GHZ family run on two real devices shows what that reality looks like. On a Rigetti Cepheus superconducting processor, a 3-qubit GHZ at 50 shots returned 45 shots in the two ideal GHZ states and 5 in bit-flipped states, the physical signature of device noise, after roughly 53 minutes in the scheduled queue. On an IonQ Forte-1 trapped-ion processor, a 2-qubit GHZ at 100 shots returned 98 shots in the ideal outcomes and 2 in error, after about 5 hours queued.
Neither run was faster or cheaper than the simulator, and neither was intended to be. What they provide is the measured noise behaviour of two qubit technologies, reported as raw counts against an exact reference. The technologies themselves are compared in Transmon or trapped ion?.
The decision procedure
Four questions settle the choice, in order.
- Is the question about an algorithm or about a device? About a device: QPU. About an algorithm: classical, and the remaining questions apply
- What structure does the circuit have? Clifford implies stabilizer simulation at any width. Bounded entanglement implies tensor networks. Shallow with an expectation-value target implies Pauli propagation. None of the above, under 32 qubits, implies exact statevector. None of the above, over 32 qubits, is the genuine frontier and is discussed in The 34-qubit wall
- One circuit or many? Many, particularly under an optimizer, is the GPU case. One large circuit is not
- Does the result carry an error bound? Statevector and stabilizer results are exact. Tensor-network and Pauli-propagation results are approximate and must arrive with a computed bound, or they are not usable as evidence
Stage Backend Typical size Purpose
Develop and debug CPU (exact) 10-25 qubits correctness
Scale structurally CPU/GPU (MPS) 60-100+ qubits convergence checks
Parameter sweeps GPU any (batched) throughput
Noise / hardware study QPU as needed physical realityOnce the tier is settled, the per shot cost calculator turns it into a number for your own workload.
One problem on all four tiers
The argument above is easier to follow with a single problem sent to every tier. The H2 molecule is two qubits, so it fits everywhere. CPU and GPU compute it exactly, a Google Cloud TPU solves it from its Hamiltonian using neural network quantum states, and three quantum processors run it on real hardware.
Run on our engines
The H2 molecule at its equilibrium bond length, whose exact electronic ground state is -1.857275 Ha. Two qubits, so it fits every device we offer. Submitted to each kind of compute we offer, on 16 September 2026. Every figure below is a real job on the service, priced as any customer would be priced.
| Device | Engine | Kind | Qubits | Result | Cost |
|---|---|---|---|---|---|
| exact.cpu | CPU | 2 | ZZ = -1.0000, the ideal value certificate | $0.0001 | |
![]() | exact.gpu | GPU | 2 | ZZ = -1.0000, the ideal value certificate | $0.0001 |
![]() | qpu.iqm.garnet | QPU | 2 | ZZ = -0.9326, superconducting, 4,096 shots certificate | $6.239 |
![]() | qpu.rigetti | QPU | 2 | ZZ = -0.5420, superconducting, 4,096 shots * certificate | $2.041 |
![]() | qpu.rigetti | QPU | 2 | ZZ = -0.5107, the same circuit re-run * certificate | $2.041 |
| qpu.aqt.ibex | QPU | 2 | ZZ = -0.9200, trapped ion, 100 shots certificate | $2.650 | |
| neural.cpu | CPU | 2 | -1.116981 Ha total, 0.0203 Ha above exact certificate | $0.0001 | |
![]() | neural.tpu | TPU | 2 | -1.116981 Ha total, 0.0203 Ha above exact certificate | $0.074 |
* The two Rigetti rows are one circuit run twice, an internal reproduction of the published benchmark notebook. A depolarizing noise model puts both versions at about -0.99, so the shortfall is not the circuit shape, but the identical program has not yet run on both devices. The steps are in the docs.
A note on the hardware certificates: they state Hellinger fidelity against the exact distribution. For an optimisation circuit that distribution is spread across many outcomes rather than concentrated on one, so the figure is low by construction and is not a measure of whether the device found a good answer. The result column above is.
The same problem is yours to run: every instance here is seeded, so it rebuilds exactly. Open the console and a cost estimate is free before anything executes.
Note what each tier is actually doing. CPU and GPU return the ideal value because two qubits is trivially exact; the TPU row is not simulating the circuit at all but searching for the ground state directly; and the three hardware rows carry device noise. The full write-up is on the chemistry benchmark.
What this implies for a research budget
A workflow following this progression treats simulation and hardware as complementary rather than competing. Simulators establish what a QPU run should produce, which is the only way to know whether the QPU run succeeded; QPU runs ground the simulation in physical behaviour the model omits.
For most research programmes in 2026 the resulting QPU share of total spend is under 10 percent. That figure is not a recommendation to avoid hardware.
It is a consequence of hardware being the right instrument for a narrow and well-defined set of questions, and the wrong one for everything else. It is also why hardware here is passed through at provider cost with no markup. A platform that profits from hardware routing has an incentive to route work there, and that incentive should not exist.
The complete run logs behind every number above, including the exact circuits and their costs, are published in the documentation.
Common questions
Is a QPU faster than a CPU?
Not in the sense the question implies, and for most work the CPU wins outright. A QPU is not a faster classical processor, it is a different kind of device.
For a circuit under about 30 qubits a CPU returns an exact, noise-free answer in seconds for $0.0001, while a hardware run costs a $0.30 task fee before a single shot and comes back with counts shaped by gate and readout error. The QPU is faster at nothing a simulator can already do; it is the only option when the question is about the device itself rather than about the algorithm.
What is the difference between a CPU, a GPU and a QPU?
A CPU and a GPU are both classical processors running the same arithmetic, and a GPU simply runs more of it in parallel. A QPU is different in kind. It holds physical qubits and returns measured samples rather than computed numbers.
In practice that means a CPU simulates exactly to about 30 qubits, a GPU to 32 on a 141 GB card, and a QPU offers 12 to 256 physical qubits depending on the machine, with noise on every one of them.
Does a GPU let you simulate more qubits than a CPU?
Barely, and this is the most common misconception about GPU quantum simulation. An exact statevector needs 16 x 2^n bytes no matter which processor holds it, so the ceiling is set by memory rather than by the chip.
A well-provisioned CPU machine reaches about 30 qubits, which is 16 GiB, and 32 qubits needs 64 GiB. That is two extra qubits for four times the memory, and the pattern does not improve. What the GPU actually buys is throughput: many circuits at once, which is why it is the right backend for parameter sweeps and optimizer loops rather than for one large circuit.
Run your own 100-qubit circuit, with an error bar.




