Quantum Machine Learning on Simulators and real Quantum hardware
Last updated · 12 min read · ZKSF team
The short version
- Quantum machine learning trains a circuit the way you train a network. The rotation angles are the weights: run the circuit, measure it, score it against a loss, adjust the angles, repeat
- Five methods measured on the platform, criteria fixed before each run. A quantum circuit Born machine (QCBM) reached a total variation distance of 0.0117 to its target, a QGAN 0.146 with visible mode collapse, a patch GAN 0.231 on an 8x8 image
- A QGAN is a GAN whose generator is a quantum circuit, trained against a classical discriminator that learns to tell its samples from real data
- Three were re-run on hardware, and the Born machine on three processors. It went from 0.068 noiseless to 0.102 on Rigetti and 0.104 on IQM Garnet, and dropped 3 of 100 shots outside its target on trapped-ion AQT against Rigetti's 102 of 1,000
- Training cost is counted in circuit evaluations. A 100-parameter model is tens of thousands of runs per training loop, which is why this work happens on simulators
- The gradients vanish as the model grows, provably. On a generic circuit the signal shrinks exponentially with qubit count, so a bigger model can be harder to train rather than better
- GPUs earn their keep above about 20 qubits. Below that a CPU is faster, because fixed overhead dominates a small circuit, and above it the CPU falls away sharply
- A reservoir readout changed the accuracy by 3.3 points at no extra cost. Same 200 circuits and same 512 shots, read as local observables instead of a full bitstring distribution
Everything here is runnable on your own circuit. Try it in the console
Quantum machine learning (QML) trains a parameterised quantum circuit in a manner loosely analogous to neural network training: a circuit with tunable rotation angles is defined, run, measured, scored against a loss function, and its angles adjusted to reduce that loss. The tunable circuit is variously called a variational quantum circuit, an ansatz, or a quantum neural network.
The left side follows a single optimization step. The parameter-shift rule computes each gradient exactly from two circuit runs, so a model with P tunable angles needs 2P+1 executions per step (201 runs at P=100), and thousands of steps to converge. On a queued, noisy QPU that is slow and costly, which is why the work runs on a simulator, where statevector evolution is the dense linear algebra a GPU accelerates: for circuits around 20 to 30 qubits, GPU training drops from hours to minutes. The right side shows the central obstacle. For many randomly initialized circuits the gradient magnitude falls exponentially as qubit count grows, a proven effect known as a barren plateau, not a symptom of poor tuning, and the landscape flattens until no usable slope remains. The listed structural responses (problem-informed circuit design, careful initialization, local rather than global cost functions, and layerwise training) are each studied by running many simulated circuits and measuring how the gradients actually behave.
Nearly all of this work happens on simulators. The reasons are structural, and understanding them is more useful than any particular architecture.
Five methods, measured
Each of these was run against the production API on exact.cpu at 512 shots, with central-difference gradients, and each had its success criterion written down before it ran. The circuit counts are the whole training loop, not one evaluation.
What was sampled. The reservoir is quantum reservoir computing: two moons, 200 samples on a 140/60 split, through a fixed 4-qubit entangling map with no trainable parameters at all: nothing in the circuit learns, only the classical readout does. The Born machine is a quantum circuit Born machine (QCBM), a quantum generative model that samples straight from the Born rule with no data fed in: 2 qubits and 6 trainable rotations over 60 Adam steps. The QGAN is a quantum generative adversarial network, a GAN whose generator is a quantum circuit rather than a neural network: a classical discriminator learns to separate its samples from the target while the circuit learns to fool it, and the two alternate. Ours is a 2-qubit generator of 4 rotations against a two-layer discriminator over 40 steps. The policy is quantum reinforcement learning, 2 qubits and 4 weights trained by REINFORCE over 60 decisions. The patch GAN is 4 sub-generators writing one 8x8 image, entangled within each patch and not across them.
method measure before after circuits
reservoir, naive test accuracy, 16 features - 0.783 200
reservoir, local obs test accuracy, 10 features - 0.817 200
QCBM TVD to target 0.459 0.0117 806
QGAN TVD to target 0.269 0.146 722
QGAN discriminator gap 0.0069 -0.0006 -
RL policy mean return, 60 decisions - 0.611 540
patch GAN TVD to an 8x8 target 0.828 0.231 3,224Two baselines belong beside that table. The raw two-moons coordinates, with no quantum step at all, give 0.850 on the same split, above both reservoir readouts. A random policy over the same 60 decisions returns 0.533, so the 0.611 is a margin inside the noise of that sample size rather than an established separation.
The two reservoir rows differ only in the arithmetic applied to identical counts: the same 200 circuits and the same 512 shots, read once as all 16 bitstring probabilities and once as the 10 local observables. The naive readout estimates each of 16 features from 512 shots, carrying about 0.022 of sampling noise each; the local observables use every shot for every feature, so their noise does not grow with qubit count.
The same circuits on a quantum processor
Training on hardware is not affordable at this scale: every gradient evaluation is its own task, so the Born machine's 806 circuits would be several hundred dollars in task fees alone. The converged circuit is a single task, so three of the five were trained on the simulator and then re-run once on a quantum processing unit (QPU), the chip itself rather than a model of it: a Rigetti superconducting device, at 1,000 shots.
method quantity noiseless Rigetti
QCBM TVD to target 0.068 0.102
QGAN TVD to target 0.270 0.286
RL policy P(action 1) at state +1.2 0.049 0.199
RL policy P(action 1) at state -1.2 0.708 0.650The Born machine lost 0.034 of total variation distance to device noise, and the QGAN 0.016. Both policy probes moved toward 0.5, by 0.15 and 0.06, which is readout error rather than a change of decision: the hardware reproduced the policy it was given. That policy chooses the opposite action to the reward convention it was trained under, and the same weights on a noiseless statevector choose the same way, which places the failure in the training rather than in the device.
The Born machine is the one to look at in full, because it is the only method here whose whole life cycle is visible: trained from a random start on a simulator, then submitted as a converged circuit to three processors across two modalities.
Run on our engines
A quantum circuit Born machine (QCBM) learning a two-qubit target distribution from a random start, then sampled. Submitted to each kind of compute we offer, on 18 September 2026. Every figure below is a real job on the service, priced as any customer would be priced.
| Device | Engine | Kind | Qubits | Result | Cost |
|---|---|---|---|---|---|
| exact.cpu | CPU | 2 | TVD 0.459 untrained, 0.0117 after 60 Adam steps | $0.0806 | |
![]() | exact.gpu | GPU | 2 | TVD 0.041, the same converged circuit at 1,000 shots certificate | $0.0001 |
![]() | qpu.rigetti | QPU | 2 | TVD 0.102, the converged circuit re-run at 1,000 shots | $0.7250 |
![]() | qpu.iqm.garnet | QPU | 2 | TVD 0.104, the same circuit and shot count on a second superconducting device certificate | $1.7500 |
| qpu.aqt.ibex | QPU | 2 | 3 of 100 shots outside the target, against 102 of 1,000 on Rigetti and 38 of 1,000 on Garnet certificate | $2.6500 | |
![]() | neural.tpu | TPU | — | a Born machine is a circuit, not a Hamiltonian, so it does not reach the neural tier | — |
A note on the hardware certificates: they state Hellinger fidelity against the exact distribution. For an optimisation circuit that distribution is spread across many outcomes rather than concentrated on one, so the figure is low by construction and is not a measure of whether the device found a good answer. The result column above is.
The same problem is yours to run: every instance here is seeded, so it rebuilds exactly. Open the console and a cost estimate is free before anything executes.
The cost column is the argument for the split. Training was 806 circuits on exact.cpu for eight cents; one hardware task re-running the converged circuit cost nine times that on Rigetti, twenty-two times on IQM Garnet and thirty-three times on AQT IBEX Q1. Training on any of them would have been the same multiple applied 806 times, which is the whole reason the loop stays on a simulator and only the answer goes to a device. The neural tier is absent by construction rather than by omission, because a Born machine is a circuit and the neural engines take a Hamiltonian.
What those engines do take is the other half of this field. Everything above points a quantum circuit at a machine-learning problem; a neural network can instead be used as the wavefunction itself, trained to find a Hamiltonian's ground state. That is AI applied to quantum rather than quantum applied to AI, and it runs on the same account and the same billing.
Run on our engines
A spin Hamiltonian solved on the TPU tier, and the H2 molecule solved on both neural engines for comparison. On 25 September the identical spin request ran on neural.gpu, an NVIDIA GPU, reaching energy -6.647038 with a ceiling of -6.630361. Submitted to each kind of compute we offer, on 16, 18 and 25 September 2026. Every figure below is a real job on the service, priced as any customer would be priced.
| Device | Engine | Kind | Qubits | Result | Cost |
|---|---|---|---|---|---|
![]() | neural.tpu | TPU | 6 | energy -6.918861, ceiling -6.886729, against an exact ground state of -7.048804 | $0.0395 |
![]() | neural.gpu | GPU | 6 | energy -6.647038, ceiling -6.630361, against an exact ground state of -7.048804 certificate | $0.0253 |
| neural.cpu | CPU | 2 | H2 at -1.116981 Ha, 0.0203 Ha above exact certificate | $0.0001 | |
![]() | neural.tpu | TPU | 2 | the same H2 molecule, -1.116981 Ha certificate | $0.0740 |
The same problem is yours to run: every instance here is seeded, so it rebuilds exactly. Open the console and a cost estimate is free before anything executes.
The 6-qubit row is a spin Hamiltonian on the TPU tier, reported with the certified ceiling beside the energy rather than as a bare number. The two-qubit rows are the H2 molecule solved on both neural engines, and they agree to seven significant figures, which indicates the method converged rather than the hardware differing. The write-up is on neural network quantum states.
Each method above has its own page carrying the runs, the classical baseline it was measured against and where it stops, together with three more that are not covered here: quantum kernels, reservoir computing and neural network tomography. They are collected under AI quantum computing use cases.
Why training is expensive in circuit evaluations
Each optimisation step requires a loss value and a gradient. The standard method for obtaining a gradient on quantum hardware is the parameter-shift rule, which evaluates the circuit twice per parameter: once at theta + pi/2 and once at theta - pi/2, with the difference giving the exact derivative. This is not finite differences; it is exact, which is the rule's appeal.
The cost is that it scales with parameter count. A model with 100 parameters requires 200 circuit executions per gradient, and thousands of steps to converge puts the total in the hundreds of thousands. On a queued, per-shot-billed device this is prohibitive on both time and money.
Simultaneous perturbation stochastic approximation, or SPSA, is the standard alternative. It perturbs every parameter at once along a random direction and costs two evaluations per step regardless of parameter count. The gradient estimate is noisy, so more steps are needed, but the trade is overwhelmingly favourable: 150 iterations of SPSA is 301 evaluations against 30,001 for parameter shift on a 100-parameter model.
Method Evaluations per step 100-param model, 150 steps
Parameter shift 2 x parameters 30,001
SPSA 2 301The choice of optimizer is therefore a hundred-fold cost decision, and it is routinely made on convergence grounds without reference to the invoice. On simulators the difference is minutes against hours; on hardware it is the difference between a feasible experiment and an infeasible one.
Where the data goes in
A question specific to QML, with no classical analogue, is how classical data enters the circuit. The encoding choice determines what function class the model can represent, and it matters more than the ansatz.
Angle encoding writes each feature into a rotation angle, using one qubit per feature. Amplitude encoding writes a 2^n-dimensional vector into the amplitudes of n qubits, which is exponentially compact but requires a state-preparation circuit whose depth generally cancels the saving.
Basis encoding writes bits directly. Repeated or data re-uploading encodings interleave data and trainable layers, and the resulting model is provably a truncated Fourier series in the input, with the number of repetitions setting the accessible frequency spectrum.
That last result is worth knowing because it makes the model's expressivity explicit rather than mysterious. A single-layer angle encoding gives access to a small number of frequencies, and no amount of training will fit a function outside that span.
The barren plateau problem
Any accurate account of QML must name its central difficulty. For randomly initialised parameterised circuits drawn from sufficiently expressive families, the variance of the loss gradient decays exponentially with qubit count. The landscape becomes flat to within the precision that finite shot counts can resolve, and gradient descent has no slope to follow.
The known responses are structural.
- Problem-informed ansatze, whose structure reflects the problem's symmetries, rather than generic hardware-efficient layers
- Local cost functions. Global observables produce plateaus at shallow depth; local ones provably avoid them for depths growing logarithmically in qubit count
- Careful initialisation, including identity-block strategies that begin the circuit near the identity where gradients are well behaved
- Layerwise training, growing depth incrementally rather than optimising a full-depth circuit from random initialisation
Each of these is studied by running many simulated circuits and measuring how gradients behave as structure changes, which places the simulator at the centre of the research rather than at its periphery. Measuring an exponentially small gradient requires an exact simulator; on hardware the shot noise would swamp the quantity being measured.
Where GPUs are effective
Statevector simulation of a QML circuit is dense linear algebra, which is precisely the GPU workload. Training loops that sweep parameter settings or evaluate batches of data points parallelise naturally.
The crossing point is measurable rather than a matter of opinion, and it sits at roughly 20 qubits: below it a CPU is faster because fixed overhead dominates a small circuit, above it the CPU falls away sharply while the GPU stays flat. The measured times at five widths, on identical batches, are in CPU vs GPU vs TPU vs QPU. Note that the GPU does not raise the qubit ceiling meaningfully. It raises throughput, and throughput is the binding constraint in a training loop.
The quantum kernel below is the small end of that curve, measured. Its 1,111 two-qubit circuits for seed 1 returned 0.675 on exact.cpu, on exact.gpu (an NVIDIA GPU) and on exact.tpu (a Google TPU) alike. The GPU billed $0.1111, the same as the CPU; the TPU billed $1.0336, because each 200-circuit batch is a machine of its own. At two qubits the hardware changes the bill, not the answer.
The unsettled question
QML remains openly unresolved on its central question: whether variational models deliver practical advantage over classical methods at useful scale. Several early advantage claims were subsequently matched by classical kernel methods or by dequantised algorithms, and the barren plateau results place real constraints on naive scaling.
The honest position is that the field has a well-defined set of open problems and no demonstrated advantage on a practically relevant dataset. This is a reason to work on it carefully rather than a reason to dismiss it, and it makes the tooling question sharper. Separating genuine signal from artifact requires exact, noise-free, reproducible evaluation, which is what a simulator provides and hardware currently does not.
A small one, run rather than cited
Two quantum methods on the same data, the same split and the same three seeds, against four classical baselines. The dataset is two moons, which is not linearly separable, drawn by numpy's seeded generator with noise 0.12: twenty-two training points and forty test points, with the two features scaled into [0, π] on the training range, on exact.cpu.
The first is a variational classifier, two qubits and four trainable parameters, at 1,024 shots over twenty steps of plain gradient descent with central-difference gradients, 4,000 circuit evaluations a seed. The second trains nothing at all. A quantum kernel embeds each point with a fixed depth-2 ZZ feature map and measures the overlap of every pair, and a classical support vector machine does the learning on the resulting matrix. That is 1,111 circuits a seed at 2,048 shots, and it is the method most of the early advantage claims were made with.
Test accuracy:
Seed VQC (4 params) Kernel SVM Logistic k-NN MLP (25 params) RBF kernel
1 0.925 0.675 0.900 0.950 0.875 0.950
2 0.825 0.750 0.775 0.875 0.900 0.875
3 0.700 0.750 0.900 0.950 1.000 0.950VQC is the variational quantum classifier described above; MLP is a multilayer perceptron, a small classical neural network; k-NN is k-nearest neighbours, which labels a point by the labels of the points closest to it and trains nothing at all. Logistic regression, k-NN and the MLP are written out in numpy; the RBF kernel is an untuned support vector machine. The variational classifier took 270 seconds a seed against 0.07 seconds for the MLP, and k-NN, which has no parameters and does no training at all, matched or beat it on every seed. The kernel trailed every classical baseline on every seed.
The kernel's measured accuracy was checked against the same kernel computed exactly, with no shots at all, and the two agree on all three seeds: the largest error in any single measured entry was 0.037, too small to move a prediction. The overlap is also exactly one for a point against itself, to fourteen decimal places, which is what a correctly built kernel has to give.
What reproduces. The seed fixes the data, the split, the classifier's starting angles and the baselines' initial weights. The service samples shots without a seed, so the 1,024-shot figures re-run to within shot noise rather than to the digit. Sent with an observable instead, exact.cpu returns each expectation computed from the statevector, and the same training then gives 0.925, 0.850 and 0.700 on every run. The scripts are variational_classifier.py, quantum_kernel.py and kernel_control.py in examples/ai.
The hyperparameter that decides the answer
The kernel result turns on how far the input data is scaled before it is written into rotation angles, a constant that carries no information about the problem at all. Same feature map, same data, exact arithmetic, first seed:
Feature map Features scaled into Kernel SVM accuracy
depth 1 [0, π] 0.700
depth 1 [0, π/2] 0.875
depth 1 [0, 1] 0.875
depth 1 [0, 1/2] 0.750
depth 2 [0, π] 0.675
depth 2 [0, π/2] 0.925
depth 2 [0, 1] 0.825
depth 2 [0, 1/2] 0.750The best of those is 0.925, and it is not a fair number: it was found by trying eight configurations and keeping the one that scored highest on the test set. The radial basis function kernel reaches 0.950 on the same split with no tuning whatsoever. A comparison that tunes one side and not the other is the most common way these results are overstated, in both directions.
What this is and is not. Four parameters on twenty-two points is a small model on a small problem, and two moons is low-dimensional classical geometry, which is where a quantum kernel is least expected to help: the separations proved in the literature are on datasets constructed for the purpose. So this is not evidence that these methods never work. It is one configuration of each, stated in full so it can be argued with or repeated, and it says nothing about barren plateaus, nothing about scaling, and nothing about what a wider circuit would do.
An exact simulator with GPU acceleration and a documented accuracy statement is the appropriate primary instrument for this work. It is not a fallback for when hardware access is unavailable.
Run your own 100-qubit circuit, with an error bar.




