ZKSF QCaaS Docs
Quantum Computing as a Service (QCaaS), documented for institutional and individual users of the platform.
On this page (70 sections)
- 4.1How to reproduce these
- 4.2Molecular ground state
- 4.3Portfolio selection
- 4.4Satellite tasking
- 4.5Scheduling and allocation
- 4.6Traffic route assignment
- 4.7Routing and job-shop encoding cost
- 4.8Cryptographic resource estimate
- 4.9QCBM, a generative model
- 4.10QGAN
- 4.11Quantum image generation
- 4.12Quantum reinforcement learning
- 4.13NNQS ground states
- 4.14QML classifiers and kernels
- 4.15Quantum reservoir computing
- 4.16NNQS tomography
On this page
Console guide
Working in the browser, with no install and no code: signing up, running a job, history and export, billing and your profile.
1.1Create your account
Four steps, once, at app.zksf.org.
- Sign up with an email address and a password
- Confirm the emailed code. The account stays unusable until you enter it. If the code never arrives, press Resend code on the same screen. Signing in, or signing up again with the same address, sends a fresh code as well
- Add credit under Billing. There is no subscription and no free tier yet, so a job submitted with $0.05 or less in the account is refused with a 402, and a TPU job with less than its $0.25 task fee. Billing and credit covers top-ups, refunds and what a balance buys
- Only for code: create an API key in Profile. The console needs none. The SDK and the REST API do: every example in the Python SDK part passes it to
qsim_sdk.Client, and the REST API takes it as a bearer token
1.2Run a job
The console at app.zksf.org shares one job history with the SDK, so a run started in either appears in both. It runs a single job, or a batch of gate circuits pasted one after another and split into batches of 200 for you. It also sweeps a pasted photonic or Pulser program over its free parameters, one line of values per run. Training loops, which submit many jobs per optimiser step, are still SDK only.
The panel works top to bottom, and the first control decides the rest of it.
- What are you running? Five kinds, and the editor and engine list follow the choice.
- Gate circuit: OpenQASM 2.0 or 3. A program with
iforwhileon measured bits needs 3; see dynamic circuits.- Getting it in. Build it in the Circuit builder beside the code box, paste it, type it, or press Upload circuit file for a
.qasmor.txt - From a framework. A circuit you already hold becomes text with
qasm2.dumps(circuit)in Qiskit, ortape.to_openqasm()in PennyLane. Either can also run here directly, without pasting anything - The builder. It and the code stay in step: drag a gate onto a wire, or tap it and then a spot; drag a placed gate to move it, an end to another wire to retarget it, or off the circuit to delete it; and use the arrow keys, Delete and Ctrl+Z from the keyboard
- Beyond the builder. A circuit using something it does not offer, such as a barrier, a custom gate or a classical condition, runs from the code box as written
- Getting it in. Build it in the Circuit builder beside the code box, paste it, type it, or press Upload circuit file for a
- Analog sequence: a Pulser sequence for neutral atoms. Pick a ready-made one, or choose Your own pulse sequence and paste the output of
seq.to_abstract_repr() - Photonic: a Perceval circuit and the photons entering it. The same menu offers Your own linear-optics program, pasted as
{"circuit": ..., "input": ...} - Optimisation problem: no program at all, just a Hamiltonian or a register for maximum independent set. The service runs the loop
- Gate circuit: OpenQASM 2.0 or 3. A program with
- Shots and Engine. The engine list shows only engines that run the kind you chose, so a photonic engine does not appear under a gate circuit. Leaving a gate circuit on
autolets the router pick; naming an engine overrides it. Two switches appear where they apply: Certified error bound onmps.quimb.cpu, and Apply zero-noise error mitigation onnoisy.cpuand hardware, which needs an observable - Estimate (free). Returns the engine, the predicted seconds and the exact price before anything is charged. Estimate before every hardware run
- Run. The preview on the right shows the circuit, pulse schedule, optics or register you are about to submit. Simulation returns in seconds; hardware queues, and the job keeps updating in place
- Read the result. Counts appear as bars, with the accuracy statement beneath them. Download certificate produces the public record; its URL resolves without an account, so it can go in a paper or a ticket
1.3Templates
Six ready-made gate circuits, two neutral-atom sequences and two photonic programs, in the picker above the editor. Choosing one loads the program, selects an engine that runs it, and fills in any observable it needs. The QASM is editable afterwards and the diagram redraws as you type.
The neutral atom and photonic templates are not circuits, so selecting one swaps the QASM editor and gate palette for the register and pulse schedule, or the optical circuit and input state.
- The engine picker still applies. An analog template runs on
analog.pulser.cpu,qpu.quera.aquilaorqpu.pasqal.fresnel; a photonic one onphotonic.slos.cpuorqpu.quandela.belenos - Run the same program on a simulator and then on hardware to compare them. The simulator returns the ideal distribution, the hardware returns the measured counts, and the certificate reports the fidelity between them
- To supply your own, pick Your own pulse sequence or Your own linear-optics program from the same menu and paste or upload it. The SDK takes them too
| Template | What it does | Engine | Walkthrough |
|---|---|---|---|
| GHZ / Bell state | Entangles three qubits so they always agree, the hello-world of entanglement. | clifford (exact) | Read |
| Grover search | Finds a marked item in an unsorted space; one iteration finds it with certainty. | clifford (exact) | Read |
| QAOA MaxCut | Splits a four-node ring to cut the most links; real combinatorial optimization. | exact.cpu | Read |
| VQE, H2 molecule | Finds hydrogen’s ground-state energy by tuning one angle; lands on -1.137 Ha. | pauli.cpu | Read |
| Bernstein-Vazirani | Reads a hidden bitstring in a single query where a classical computer needs one per bit. | clifford (exact) | Read |
| Quantum teleportation | Moves a qubit’s state across the circuit using entanglement and corrections alone. | exact.cpu | Read |
| Rydberg blockade, three atoms | Three atoms close enough to block each other settle into the only pattern that fits: excited, ground, excited. | analog.pulser.cpu | SDK |
| Antiferromagnetic chain, seven atoms | Sweeps a chain of seven atoms into the alternating order the blockade forces on it; lands on 1010101 about 91% of the time. | analog.pulser.cpu | Read |
| Hong-Ou-Mandel, simulator | Two identical photons meet on a beamsplitter and must leave together; the ideal coincidence term is exactly zero. | photonic.slos.cpu | SDK |
| Hong-Ou-Mandel, Belenos hardware | The same program on the real photonic processor, so you can read the ideal answer and the measured one side by side. | qpu.quandela.belenos | Read |
Every template runs on the same engines, billing and certificates as any other job, so the circuit can be edited or scaled up from there.
1.4History, export and cancelling
Every job on your account is listed in the console, including the ones submitted from code, and the cards at the top of the page count them.
- Your job history sits under the run panel, newest first, with what each run actually cost after any refund. Click Engine, Date, Cost or Status to sort the whole history, and click again to reverse it. Export CSV opens a dialog that downloads every job in a date range, for all engines, one tier or one engine; dates before your first job or after today cannot be chosen. Every job is reachable later by id from the SDK
- Jobs in queue, the card at the top, counts every job that has not finished. It turns amber when a job that is not waiting on quantum hardware has been queued for more than a day. Hardware queues can take days; other engines do not. Click it to list just those jobs
- CPU, GPU, TPU and QPU jobs run, the cards beside it, count every job that finished with a result on each tier. Click one to list just those jobs. Spend this month is what this month’s jobs cost after refunds
- Cancel appears on a quantum hardware job that is still waiting in its provider’s queue (not yet on Pasqal FRESNEL or error-mitigated runs). The charge is refunded in full once the provider confirms the job never ran. If it starts before the cancel reaches the device, it finishes, returns its result and is charged as normal, and the job says so
1.5Billing and credit
Pay as you go: you add credit, and each job is paid from it. There is no subscription and no free tier yet. Open Billing from the menu under your avatar.
- Available credit is your balance, shown to four decimal places because a simulation can cost a fraction of a cent
- Add credits: slide to, or tap, an amount from $20 to $1,000, then press Continue to payment. Payment is taken on Stripe’s checkout page and we never see your card. Back in the console the credit appears once the payment is confirmed, and a cancelled checkout charges nothing
- Spend by tier splits what your jobs cost across CPU, GPU, TPU and QPU. It covers this month, or the past year when this month has no charges yet, and says which
- Payment history lists every top-up with its date and amount
- What your balance buys is a calculator for what a given top-up buys on each engine
What you are charged. Price a run before submitting it with Estimate in the console, or estimate() for a gate circuit in code; both are free. For a quantum processor the estimate is the exact charge: it is the provider’s own task and shot price, charged when you submit. For a simulation on CPU, GPU or TPU it is an estimate, because the run is billed by the seconds it actually takes. The response says which, in exact, and the console labels it Exact price or Estimated cost. Nothing is taken up front for a simulation. It is charged for the seconds it runs, as it runs, at its tier’s rate:
- CPU: $0.0001 per circuit covers its first second, then $0.69 per CPU hour. A circuit under one second costs $0.0001
- GPU: $3 per GPU hour up to 30 qubits and $8 per GPU hour for 31 or 32 qubits, billed by the second, with the same $0.0001 minimum per circuit
- TPU: $0.25 per task for the first 480 seconds, then $1.85 per chip-hour, billed by the second
- QPU: the provider’s own task and shot price, charged when you submit
Every running job on your account stops when your balance reaches $0.05, and is charged for the seconds it ran. Your job history shows what each run cost. Some things cost nothing, and any charge goes back to your balance by itself:
- A job that fails because of a fault on our side, or that we refuse when you submit it, is not charged
- A hardware job you cancel while it is still queued is refunded in full once the provider confirms it never ran
A job whose own program stops with an error while it runs is charged for the seconds it ran, like any other run (ZKSF-4008). Any job that consumes compute is charged.
Jobs run until they complete. If a request includes max_runtime_seconds or max_seconds, the value is accepted but not applied: run time is difficult to predict in advance, and stopping a job part-way would charge for the time already used without returning a result. A job continues until it finishes or your balance reaches $0.05.
When your balance runs low. Available credit on the Billing screen is shown in three colours:
- Green: more than $5
- Amber: $5 or less. The Spend this month card on the dashboard also says Your balance is running low, please top up
- Red: $0.50 or less. The card says Your balance is nearly zero, please top up
Selecting the card opens Billing.
At $0.05 every running job stops, including jobs still in progress, and a new job is refused when you submit it. A stopped job keeps the charge for the seconds it ran. The console then shows a notice, You are out of balance. Please top up to continue., with a Top up button that opens Billing. Add credit and submit the job again.
For any issue with a payment or a top-up, email support@zksf.org.
1.6Profile and deleting an account
Open Profile from the menu under your avatar.
- Email is the address you signed up with, shown for reference
- Reset password emails a six-digit code. Enter it on the same screen with a new password of at least 10 characters, mixing upper and lower case and including a number, then press Set new password. It applies the next time you sign in. Signed out? Use Forgot password? on the sign-in screen instead
- API keys: Create API key makes a key for the SDK and the REST API, valid for 30 days. It is shown once, so copy it then. The list shows each key’s name, first characters and expiry, and Revoke stops a key working immediately. An account can hold 10 active keys. Create an API key shows where it goes
Deleting your account
Press Delete my account, type DELETE, then press Delete permanently. It is immediate and irreversible: your sign-in, email address, balance and API keys all go.
An account that still holds credit cannot be deleted. Email support@zksf.org first and we will settle the balance. Certificates already issued stay publicly verifiable, and they carry no personal data.
Your jobs, circuits and results are kept for 30 days after deletion in case you need them back: email support@zksf.org from the address that held the account. After 30 days they are removed for good and cannot be retrieved.
1.7Mobile app (Android)
A third client, signed in with the same account, so jobs and history are shared with the web console. It runs gate circuits: paste or upload OpenQASM 2 (.qasm or .txt), choose an engine, estimate, run, and export the same certificates. Zero-noise mitigation is available on hardware runs there too.
It runs one circuit at a time by hand. Sweeps, training loops and anything wired into a larger script need qsim-sdk on a desktop; it cannot run on a phone.
Engines and results
What runs a job and what comes back, however it was submitted: every engine and its limits, the router, the certificates, and why a job is refused.





OpenQASM 3 and dynamic circuits
The console’s code box, the SDK and the API take OpenQASM 2.0 and OpenQASM 3 in the same field. A program with no mid-circuit control flow is converted to 2.0 when it arrives, so it runs on every engine, the GPU tier and quantum hardware included.
OPENQASM 3.0;
include "stdgates.inc";
qubit[2] q;
bit[2] c;
h q[0];
c[0] = measure q[0];
if (c[0]) { x q[1]; } // runs only when q0 read 1
c[1] = measure q[1]; // so the result is 00 or 11, never 01- Dynamic circuits run on the Aer engines. A program with
if,while,fororswitchon measured bits runs as written onexact.cpu,noisy.cpuandmps.aer.cpu. Left on automatic, it goes toexact.cpu - Every other engine refuses it at submission, with a 400 that names the engines that take it, before anything is charged
- Counts only. An expectation value or error mitigation needs a circuit without mid-circuit measurement, so a dynamic program asking for either is refused
- From the SDK, a Qiskit circuit built with
if_testorwhile_loopis sent as OpenQASM 3 automatically (qsim-sdk 0.11.2), and any other circuit as 2.0
2.1The engines
What each engine is for and where it stops. The console’s engine list shows the ones that run the kind of program you chose, and code names them the same way, as engine=. Leave the choice to the router and it picks for you.
| Engine | Best for | Limits |
|---|---|---|
| CPUexact.cpu | Ground truth, any gate set | ≤ 29 qubits |
GPUexact.gpu | Exact results, GPU-accelerated | ≤ 32 qubits, up to 100,000 shots; for higher limits contact support@zksf.org |
TPUexact.tpu | Exact results on a Google TPU | ≤ 29 qubits, single precision, up to 100,000 shots; a machine is created per job, so a short circuit spends most of its time starting up and a batch shares one machine |
| CPUclifford | Stabilizer/error-correction circuits | Clifford gates only, 5,000+ qubits |
| CPUmps.quimb.cpu | QAOA, ansätze, structured circuits | entanglement-dependent, ~128 qubits; certified error bounds |
| CPUmps.aer.cpu | Independent tensor-network cross-check | Supports up to 63 qubits |
| CPUpauli.cpu | Expectation values of large circuits | needs an observable; 100 to 200 qubits, no bitstring counts |
| CPUnoisy.cpu | Preview hardware noise before you pay | device-noise model; optional error mitigation (ZNE) |
| CPUanalog.pulser.cpu | Neutral-atom analog dynamics | takes a Pulser sequence, not a circuit; ≤ 14 atoms, exact |
GPUanalog.pulser.gpu | The same dynamics, further | takes a Pulser sequence, not a circuit; ≤ 20 atoms, exact |
| CPUphotonic.slos.cpu | Linear-optics interference | takes a Perceval circuit and an input Fock state; ≤ 12 modes, exact |
GPUphotonic.gpu | Linear optics at hardware width | takes a Perceval circuit and an input Fock state; ≤ 24 modes and 12 photons, sampled exactly without enumerating every outcome |
QPUqpu.rigetti | Real superconducting hardware | Rigetti Cepheus 108 qubits, 10 to 50,000 shots; scheduled queue, physical noise |
| Real trapped-ion hardware | IonQ Forte Enterprise 1, 36 qubits, 100-5,000 shots; scheduled queue | |
QPUqpu.iqm.emerald | Real superconducting hardware | IQM Emerald, 54 qubits, up to 20,000 shots; scheduled queue |
QPUqpu.iqm.garnet | Real superconducting hardware | IQM Garnet, 20 qubits, up to 20,000 shots; scheduled queue |
| Real trapped-ion hardware | AQT IBEX Q1, 12 qubits, up to 2,000 shots; scheduled queue | |
| Real neutral-atom hardware | QuEra Aquila, 256 atoms, up to 1,000 shots; takes a Pulser sequence; runs Friday 04:00 UTC to Monday 04:00 UTC (pauses 12:00-13:00 and 00:00-01:00 daily), so a submission outside that is accepted and waits | |
| Real neutral-atom hardware | Pasqal FRESNEL, 100 atoms, at most 100 shots a job (for higher limits contact support@zksf.org); takes a Pulser sequence; billed as machine time at EUR 500/hour and a contractual 0.25 Hz, so about EUR 0.56 a shot and shot counts here are small by design | |
QPUqpu.quandela.belenos | Real photonic hardware | Quandela Belenos, 12 qubits: ≤ 24 modes and ≤ 12 photons, two modes per qubit; inputs only on connected modes |
| CPUneural.cpu | Ground states beyond the area law | Takes a Hamiltonian, not a circuit: use POST /solve. Up to 40 spins |
TPUneural.tpu | The same, on a Google TPU | Takes a Hamiltonian, not a circuit: use POST /solve. $0.25 per task for the first 480 s, then $1.85 per chip-hour |
GPUneural.gpu | The same, on a GPU | Takes a Hamiltonian, not a circuit: use POST /solve. Up to 40 spins, and no machine to create first, so a short search starts in seconds |
Analog neutral-atom sequences
Neutral-atom hardware has no gates. You send a register of atoms and a schedule of laser pulses, written with Pulser.
- In the console, choose Analog sequence, then a ready-made sequence or Your own pulse sequence. From code, see Pulser sequences
- Always name the engine. The router reads gate-circuit features these do not have
- Every sequence is validated against the target device’s own constraints before it is submitted
The same sequence runs on all three engines. Only the engine changes.
| Engine | Atoms | Shots | Billed | Before you submit |
|---|---|---|---|---|
| analog.pulser.cpu | 14 | any | $0.69 / hour | No queue. Exact, so only shot noise applies. Start here. |
| qpu.quera.aquila | 256 | ≤ 1,000 | $0.01 / shot | Runs Friday 04:00 to Monday 04:00 UTC, pauses 12:00-13:00 and 00:00-01:00 daily. A job outside that is accepted and waits. |
| qpu.pasqal.fresnel | 100 | ≤ 100 | ~EUR 0.56 / shot | Machine time, not shots. Keep counts small. Estimate first. |
FRESNEL bills EUR 500 an hour at a contractual 0.25 Hz repetition rate, which is where the per-shot figure comes from.
Aquila reports empty sites as well as ground and Rydberg states. Shots with an empty site are dropped by declared post-selection, and the retained fraction is reported with the result: a fidelity computed from those counts is a fidelity over the shots that were kept.
Local detuning
Pull chosen atoms' detuning down through a fixed spatial pattern. This is how maximum independent set is encoded on a neutral-atom machine.
- Build the sequence on a device that has a DMM.
WeightedAnalogDevicematches Aquila at 256 atoms - Site weights must lie between 0 and 1
- The waveform must stay at or below zero. The field lowers a detuning, never raises it
- Pulser enforces the same sign, so a legal Pulser sequence converts unchanged
Photonic linear optics
Photons enter chosen modes, interfere through beamsplitters and phase shifters, and you measure which modes they leave by. A program is two things, both required.
- A Perceval circuit
- An input Fock state. A gate circuit carries its own initial state; this does not
The input state is a list of photon counts, one per mode: [1, 0, 1] is one photon into mode 0 and one into mode 2. Both halves are hashed, so two runs differing only in the input cannot share a certificate. In the console, choose Photonic; from code, see Perceval programs.
| Engine | Limit | Use it for |
|---|---|---|
| photonic.slos.cpu | ≤ 12 modes | Exact. The reference a photonic certificate is measured against. |
| qpu.quandela.belenos | ≤ 24 modes, ≤ 12 photons | Real hardware. Returns its own declared noise figures with the counts. |
Two rules for Belenos, both enforced before you are charged.
- Photons enter only on connected modes: 0, 2, 4, 6, 8, 9, 12, 13, 16, 18, 20, 22. Anything else is refused at submission
- Modes 0 and 1 are not both available. A two-photon interference run therefore uses modes 0 and 2 and permutes them together
Quandela sells the same machine as MosaiQ 12, a 12-qubit processor. Dual-rail encoding spends two modes per qubit, so 24 modes carrying 12 photons is 12 qubits counted the other way.
In a Hong-Ou-Mandel run, two identical photons meeting on a beamsplitter must bunch, so both leave by the same mode and the coincidence term |1,1,0> is exactly zero. The coincidence fraction is therefore a direct fidelity measure, and it is what the photonic certificate reports.
On Belenos itself the same experiment bunched 95.4%, 97.0% and 97.7% across three runs. All three are written up in what a photonic quantum computer actually returns.
2.2The router, which is the default
The router is one of the 24 engines. Send a circuit without naming an engine and it picks the cheapest engine that can return a certifiable error bound, judging the circuit's structure rather than its size: whether every gate is Clifford, how much entanglement it can generate, whether an observable was supplied, and how wide it is. If no engine can return a meaningful error bar, the job is refused with the reason.
The choice matters for cost: a Clifford circuit at 500 qubits and an exact statevector at 28 differ in price by orders of magnitude.
In the console, leaving a gate circuit on auto uses the router, and choosing an engine overrides it. From code, leave engine out, or name one. Every result names the engine that produced it, and a named engine is always used exactly as given.
2.3Understand the error certificate (ZCC-v0.1)
Every approximate result carries a ZCC-v0.1 statement of how far it can be from the truth. By default it is a fast convergence check. For a measured, single-run bound, tick Certified error bound on mps.quimb.cpu in the console, or pass certified=True from code.
job["result"]["error_info"]
# {'method': 'MPS (quimb)', 'max_bond': 64,
# 'convergence_deviation': 0.0, 'converged': True}client.run(qc, engine="mps.quimb.cpu", certified=True)
# error_info = {'protocol': 'ZCC-v0.1',
# 'method': 'MPS, measured discarded-weight bound',
# 'truncation_weight': 3.2e-08, # exact discarded SVD weight
# 'error_bound': 2.5e-04, # max per-outcome prob error
# 'certified': True}convergence_deviation: how much the top outcome probabilities moved when we re-ran at double the bond dimension. 0.0 means the approximation captured the statecertified=True: builds the state with renormalization off and reads the accumulated discarded SVD weight from the norm deficit. Per-outcome probability error is bounded bysqrt(2 · truncation_weight), measured in a single run rather than extrapolated.- Each truncation is optimal by Eckart-Young, and the weight is read from the state rather than estimated
- The bound assumes truncation errors accumulate incoherently. Deep circuits do not always satisfy that, so it is an empirically supported bound, not a worst-case guarantee
- Never exceeded across 334 runs checked against exact simulation, 290 of them built to falsify it, up to 20 qubits, the largest size where an exact reference exists
- Pauli propagation reports the same protocol:
error_boundis the total discarded coefficient mass, which bounds the expectation-value error - Exact engines (statevector, stabilizer) report
truncation_error: 0.0; only shot noise applies
Any finished job can be exported as a downloadable PDF certificate with a public verification link, so a result you report can be independently checked by anyone. See the certification page for the full protocol, or read the paper for the derivations and the falsification testing behind it.
2.4Understand the hardware certificate (ZHF-v0.1)
QPU results carry a different statement, which is not how close an approximation is to the truth but how closely a real device’s measured counts track the exact ideal distribution. That is ZHF-v0.1.
client.run(qc, engine="qpu.ionq", shots=100)
# certificate = {'protocol': 'ZHF-v0.1',
# 'device': 'IonQ Forte-1',
# 'fidelity_mode': 'direct',
# 'fidelity': 0.9774}Those are the real figures from an archived run, certificate df1d4c698a954051. It ran on IonQ’s Forte-1, which has since been retired; qpu.ionq now runs on Forte Enterprise 1.
fidelity: the Hellinger fidelity between the measured hardware counts and the exact ideal distribution computed for the same circuit, in [0, 1]. 1.0 would mean the measured counts matched the ideal distribution exactlyfidelity_mode: "direct": the circuit was small enough (≤ 24 qubits) to also simulate exactly, so the fidelity is measured directly against that ideal reference, not asserted from a vendor specification- Above that qubit count, the certificate states that direct verification is not available instead of reporting a fidelity
- It applies to analog machines too, where the reference is an exact Schrödinger evolution rather than a statevector. A QuEra Aquila run on 11 September 2026 returned
fidelity: 0.6294, certificate d3abecb2c8d14c2f
Read a fidelity together with its shot count. That Aquila run kept 19 shots, and at 19 shots a flawless device would still score a median of only 0.875, with a 5th-to-95th percentile band of 0.762 to 0.951, purely from sampling.
The figure means something here because 0.6294 falls below that band, and none of 20,000 simulated perfect runs scored as low. At 100 shots the same perfect device scores 0.98 and at 1,000 shots 0.994, so a fidelity quoted without its shot count is not comparable to one that has it.
Like ZCC-v0.1, any finished QPU job can be exported as a downloadable PDF certificate with a public verification link. See the certification page for both protocols.
2.5Error mitigation (opt-in)
Any run carrying an observable can request zero-noise extrapolation. The circuit is evaluated at several amplified noise levels, then extrapolated back to the zero-noise limit.
- Works on
noisy.cpuand all five gate QPUs - Not available on the three non-gate engines. Belenos returns photon counts rather than an observable; the two neutral-atom machines run a continuous pulse schedule with no gates to fold
- Costs 3x a single run on hardware. It submits three folded circuits
- Requires an observable
- In the console, tick Apply zero-noise error mitigation, which appears for those engines, and enter the observable under it
client.run(qc, engine="qpu.rigetti", shots=200,
observable=[[1.0, "ZZ"]], mitigate=True)
# error_info: { protocol: "ZCC-Estimate-v0.2",
# technique: "zero-noise extrapolation",
# raw_expectation: 0.81, mitigated_expectation: 0.82,
# error_bound: 5.0e-02, estimated: true, certified: false }The result carries a ZCC-Estimate-v0.2 uncertainty, computed without reference to the true answer, so the statement carries over to hardware unchanged.
- The uncertainty is the larger of two figures: shot noise propagated through the extrapolation, and the disagreement between linear, quadratic and Richardson extrapolations of the same data
certifiedreadsfalseby design. An extrapolated value cannot carry a hard bound the way the exact, stabilizer and tensor-network paths can- The exported certificate states this explicitly. It is a well-founded statistical estimate, not a guarantee
2.6Reading a result as an image
A circuit can encode an image: one qubit per pixel, the pixel value written as a rotation, so the share of shots in which that qubit measures 1 is the pixel value. Pixels are not entangled with each other, so a pixel’s position in the image is the qubit’s position in the circuit.
The result returns as counts, like any other circuit, and counts alone do not indicate whether they represent an image. Declare it on submission:
job = client.run(circuit, shots=300, engine="qpu.rigetti",
params={"render": "image", "width": 32})In the console this is a checkbox beside the other circuit options. The job page then draws the image and offers it as a PNG.
- Write the value as
ry(value * pi)on qubity * width + x. A value of 0 measures 0 every time, 1 measures 1 every time, and a half measures either - One circuit holds as many pixels as the device has qubits, so a 108-qubit processor takes about a 10×10 image in a single job
- A larger image is several jobs, and the console renders one job at a time. A 32×32 is eleven circuits, shown as eleven strips. To join them, use
qsim_sdk.images.stitch, which takes the counts from each job in submission order and returns the whole picture - The conversion is arithmetic on counts the service has already certified. It introduces no error of its own, and the certificate covers the circuit, not the image
from qsim_sdk import images
counts = [client.job(i)["result"]["counts"] for i in job_ids]
pixels = images.stitch(counts, width=32)
open("out.pgm", "wb").write(images.to_pgm(pixels))2.7Solve a Hamiltonian for its ground state
Every path above takes a program you wrote. POST /solve takes a Hamiltonian and runs the variational loop for you. In the console it is the fourth kind under What are you running?, sharing one job history, one balance and one certificate format with everything else.
You name the ansatz rather than sending it, because OpenQASM 2 cannot express an unbound parameter.
POST /solve
{
"hamiltonian": [[-1.0523, "II"], [0.3979, "IZ"], [-0.3979, "ZI"],
[-0.0113, "ZZ"], [0.1809, "XX"]],
"qubits": 2, "ansatz": "real_amplitudes", "reps": 2,
"max_iterations": 60, "shots": 1024
}
# result
{ "energy": -1.857,
"ground_state": { "ceiling": -1.856, "basis": "..." },
"iterations": 60, "evaluations": 121 }- Read the ceiling, not the energy. The variational principle puts the true ground state at or below whatever the optimiser reports. Adding the simulation's own error bound gives a number the true answer cannot exceed. No bound from the engine means no ceiling is stated
- The run time is the price. A search is metered by the second like any CPU job. SPSA costs 2 evaluations per iteration whatever the parameter count, plus 1 final run, so 60 iterations is 121 evaluations and the time they take
- Every term names every qubit, padded with
I. Ragged Pauli strings are refused - A neural ansatz is available here too. See below
| Field | Accepts |
|---|---|
| qubits | 1 to 24 |
| ansatz | real_amplitudes, efficient_su2, two_local |
| reps | 1 to 6 |
| max_iterations | 1 to 200 |
A neural network as the ansatz
Everything above varies a circuit. The neural engines vary a neural network wavefunction instead, trained by variational Monte Carlo.
That matters where tensor networks stop. They assume entanglement follows an area law, and a state that breaks that assumption is where a neural ansatz reaches and an MPS does not.
They take a Hamiltonian, so they are reached through /solve, not /jobs. A circuit sent to one is refused with a message that points here.
POST /solve
{ "hamiltonian": [...], "qubits": 9, "engine": "neural.cpu" }Two devices, one method
neural.cpu runs on CPU. neural.tpu runs the identical code on a Google TPU, billed only while your job runs. Variational Monte Carlo is a training loop dominated by dense matrix multiplication against batches of sampled configurations.
Two differences apply when choosing between them. TPU arithmetic has no complex128, so the engine narrows to complex64 there and records that it did: a lower precision moves the variance floor, and a variance sitting at the arithmetic floor must not be read as a converged eigenstate. A TPU also has to start up before it computes, so a short job can spend more time starting than solving.
TPU time is $0.25 per task for the first 480 seconds, then $1.85 per chip-hour, billed by the second. CPU is billed at the ordinary CPU rate. The bound is the same on both.
Read the ceiling as the bound and the floor as a diagnostic. The result states which is which. The ceiling is unconditional, because the variational principle puts the true ground state at or below the measured energy, always.
The floor comes from Weinstein's inequality applied to the measured energy variance, and holds only while the trial state sits nearer the ground state than the first excited state, which a run cannot verify about itself. Independent restarts agreeing is evidence, not proof, and the restarts travel with the result so you can weigh it yourself.
A run whose Markov chains did not mix is refused rather than returned. When sampling is invalid the estimate is not an expectation value at all, so neither the ceiling nor the floor applies to it.
Maximum independent set
The other problem this takes. It goes to neutral atoms, not gate circuits. The blockade already forbids two atoms within a radius from both being excited, which is the independent-set constraint itself, so there is no encoding and no penalty terms.
POST /mis
{ "vertices": [[0,0], [6,0], [12,0], [18,0]], "shots": 200,
"engine": "qpu.quera.aquila" }
# result.mis
{ "best_set": [0, 2], "size": 2, "valid_fraction": 0.86,
"edges": [[0,1], [1,2], [2,3]], "blockade_um": 10.03 }- The graph is not an input. Vertices are atom positions in micrometres. An edge exists where two atoms fall inside the blockade radius, so the geometry is the problem. The family that maps directly is the unit-disk graphs
- The radius is not an input either. It follows from the drive as Rb = (C6 / Ω)1/6, which is 10.03 um at the default 12 rad/us amplitude. It comes back in
blockade_um. Lay your problem out against that number, not a round one - An arbitrary graph needs a layout step first. That is its own hard problem and this does not solve it for you
- Check
valid_fraction. Every shot is tested against the edges and invalid ones are excluded. A low fraction means the sweep was too fast for that graph, not that no solution exists - Optional
weights(non-negative) give the weighted problem, applied through the shifting field
2.8Error correction
Hold a logical qubit in a code for a number of rounds under circuit-level noise, and measure how often it comes back wrong. The answer is a logical error rate with a 95% confidence interval.
import qsim_sdk
client = qsim_sdk.Client(token="YOUR_TOKEN")
client.estimate_qec("surface", distance=5, shots=100_000)
job = client.qec("surface", distance=5, shots=100_000, physical_error=1e-3)
res = job["result"]
res["logical_error_rate"] # 0.0
res["confidence_interval_95"] # [0.0, 3.69e-05]
res["logical_errors"], res["shots"] # 0, 100000Read the interval, not the rate. A rate is a proportion measured from a finite number of shots, so at a hundred thousand shots with no failures it reads 0. The statement that can be made is that the rate is below the top of the interval, and that is the number the certificate carries.
Parameters: code is one of surface, surface_x, surface_unrotated or repetition. distance runs from 3 to 15, and must be odd for a rotated surface code. rounds defaults to the distance, which is the convention these rates are published at: fewer rounds flatters the code, because time-like errors have had less opportunity to accumulate. physical_error is the probability per operation, and measurement_error sets readout separately when that is the weak part of a device.
Decoding is minimum-weight perfect matching. Colour and qLDPC codes are refused rather than decoded that way, because matching is the wrong decoder for them and a wrong decoder returns a number indistinguishable from a right one. The run measures memory only: prepare, hold, read out. Logical gates and lattice surgery are not offered here.
Priced as ordinary CPU work, metered on runtime at $0.69 per CPU hour, billed by the second while it runs and not taken up front. It is metered rather than priced per circuit because it is the one job on this tier that routinely runs for minutes rather than for a moment. The cost driver is detectors times shots, so client.estimate_qec(...) is worth calling: the same endpoint spans a run that finishes in under a second and one that holds a worker for a quarter of an hour.
The second question is the one every roadmap is quoted in and nobody answers for a specific device: how many physical qubits is one logical qubit, at my error rate. That needs the suppression factor, which is how much each step of distance buys, and it is measurable from three runs.
job = client.qec_scaling(
physical_error=1e-3, # your device
target_logical_error_rate=1e-9, # what you need
)
res = job["result"]
res["suppression"]["lambda"] # 1.93, measured across distances 3, 5 and 7
res["distance"] # 25
res["physical_qubits"] # 1249
res["basis"] # "extrapolated from a suppression factor
# measured up to distance 5"Read basis. The suppression factor is measured, the distance that follows from it is arithmetic, and a target below anything those runs could observe is reached by extrapolating along the measured slope rather than by measuring it. Saying which is which is the difference between this and a number copied from a paper.
A distance that produced no failures is excluded from the fit rather than replaced by its confidence bound, because substituting a bound for a missing measurement turns a measurement into an assumption without saying so. At or above threshold, where the suppression factor comes out at or below 1, the answer is that no number of physical qubits reaches the target until the physical error rate comes down. That is a result, not a failure.
Omit target_logical_error_rate and it returns the suppression factor alone, which is the curve behind the threshold plot: run it at several error rates to find where your device crosses from distance helping to distance hurting.
2.9Why a job is refused
qsim_sdk.JobRejected: intractable classically at this structure: 60 qubits, depth 99, 826 entangling gates (13.8/qubit). Options: reduce entangling density below ~8/qubit for MPS, restrict to Clifford gates, or shrink to <= 29 qubits for exact simulation.
Deep, wide, highly entangled circuits are intractable for every classical method. A job that no engine can answer to a stated error bound is refused rather than run, and the rejection names the reason and the options for changing the circuit.
* This check covers classical simulability. It reads width, depth and entangling density to decide whether any classical engine can answer the circuit to a stated error bound. It does not model a real processor's coherence time. A circuit sent to a QPU can be accepted, priced and run even when it is long enough that the qubits decay before it finishes, and the result then comes back as noise. Read the certificate's fidelity figure to see whether that happened.
Error codes
A rejection like the one above names what to change in the circuit. The failures below arrive with a code and a message.
| Code | What happened | What to do |
|---|---|---|
| ZKSF-2005 | A cancel could not reach the hardware provider. Nothing was changed: the job and its charge are exactly as they were | Try the cancel again shortly |
| ZKSF-2011 | The photonic device is busy with another job | Resubmit shortly |
| ZKSF-2021 | Pasqal FRESNEL is in a maintenance window, which Pasqal schedules | Resubmit later |
| ZKSF-2030 | QuEra Aquila is outside its execution window | Resubmit between Friday 04:00 and Monday 04:00 UTC |
| ZKSF-4004 | The ansatz could not be built from what you sent | Check the ansatz name, qubit count and reps |
| ZKSF-4008 | Your program stopped with an error while it ran. It is charged for the seconds it ran | Check the program, or run a smaller part of it to find the step that fails |
| ZKSF-4009 | Your job could not be completed due to a technical issue with the circuit | Contact support@zksf.org for details |
| ZKSF-5001 | The TPU tier is not available right now | For a Hamiltonian, use neural.cpu, which runs the identical method and returns the same bound. For a circuit, use exact.gpu or exact.cpu, which return the same exact distribution |
| ZKSF-6001 | Your balance does not cover the job’s cost: it reached $0.05, so the job stopped, or it was refused at submit | Top up in Billing and submit the job again |
For these codes, and any other issue not listed here, email support@zksf.org.
2.10Worked use cases
Three problems that exist outside quantum computing, and what it takes to run each one here. Each says how the problem is expressed, which engine to start on, what comes back, and where the approach currently stops. The first two are reachable from the console, and every one of them from the SDK.
1. The ground-state energy of a molecule
Applications: reaction feasibility, catalyst screening and ligand binding.
Getting the input. Run the electronic structure in PySCF, map it to Pauli operators with OpenFermion or Qiskit Nature (Jordan-Wigner or Bravyi-Kitaev). You get a list of coefficients and Pauli strings, which is what the panel takes.
- Choose Optimisation problem, then Ground state of a Hamiltonian
- Paste the terms, one per line as
coefficient pauli. Every term must name every qubit, padding withI. The panel derives the qubit count from the strings and refuses ragged ones - Ansatz:
real_amplitudesfor a molecular Hamiltonian with real coefficients, which is most of them. Raise Repetitions if the energy plateaus above what you expect; each repetition adds parameters and cost - Start on
exact.cpu, which gives a noiseless answer and a real error bound. Move to hardware only once the ansatz and iteration count are settled - Read the ceiling, not just the energy. The variational principle puts the true ground state at or below the value found; adding the simulation’s own error bound gives a number the true answer cannot exceed
job = client.solve(
hamiltonian=[[-1.0523, "II"], [0.3979, "IZ"], [-0.3979, "ZI"],
[-0.0113, "ZZ"], [0.1809, "XX"]],
qubits=2, ansatz="real_amplitudes", max_iterations=200)
job["result"]["energy"] # -1.8482 Ha, exact is -1.8571
job["result"]["ground_state"]["ceiling"] # what it cannot exceedWhere it stops. 24 qubits, roughly a dozen spin orbitals. H2, LiH and BeH2 in a minimal basis, not an enzyme active site. A poor ansatz converges above the true ground state without indicating that it has, so the result reports the ceiling rather than stating that the minimum was found.
2. Scheduling around conflicts, and allocating scarce channels
Applications: transmitters that interfere at close range, delivery slots that must not overlap, the largest set of compatible tasks for one shift. All are the same problem, selecting as many items as possible with no two in conflict, which is NP-hard classically.
Getting the input. Lay the items out so distance encodes conflict. Two atoms conflict when they are within the blockade radius. For spatial problems the map is the input already: use real coordinates, scaled so the interference range lands at 10 um.
- Choose Optimisation problem, then Maximum independent set
- Enter one
x, yper line in micrometres. The preview draws the edges your layout actually creates, and the count updates as you type: check it matches the conflicts you meant before running anything - Run on
analog.pulser.cpufirst. No queue, exact to 14 atoms, so you can confirm the encoding on a small instance for a hundredth of a cent - Switch to
qpu.quera.aquilafor the real instance, up to 256 items. Aquila runs Friday 04:00 UTC to Monday 04:00 UTC, pausing 12:00-13:00 each day and 00:00-01:00 each night, so a submission outside that is accepted and waits - Read valid fraction alongside the answer. It is the share of shots that obeyed every conflict; a low number means the sweep was too fast for that instance and the run should be repeated slower, not that no solution exists
job = client.submit_mis(
vertices=[[0, 0], [6, 0], [12, 0], [18, 0], [9, 9], [0, 12]],
shots=200, engine="qpu.quera.aquila")
job["result"]["mis"]["best_set"] # the items to keep
job["result"]["mis"]["valid_fraction"] # how much to trust itWhere it stops. Conflicts that are already geometric map directly. An arbitrary conflict graph must be laid out into positions that reproduce it, which is its own hard problem. The family that maps without a layout step is the unit-disk graphs.
3. A quantum layer inside an existing model
The structure common to quantum machine-learning results, including generative and image-model demonstrations: a parameterised circuit as one layer of an otherwise ordinary network.
Getting the input. A Perceval circuit with free parameters becomes a torch nn.Module. Its output is the probability of each detection pattern, in a fixed order, which is the feature vector. The rest is an ordinary training loop.
- SDK only:
pip install "qsim-sdk[ml]". A training run is thousands of submissions, so it needs a script rather than the console. PyTorch covers the layer in full - Declare the parameters you want trained with
pcvl.P("name")and wrap the circuit inPhotonicLayer - Train against
photonic.slos.cpu. Every forward pass is one job and every backward pass is one job of two points per parameter, so the arithmetic is yours to check before you move - Only then set
engine="qpu.quandela.belenos", and do the sum first: 200 steps of a six-parameter model is 2,600 tasks
from qsim_sdk.ml import PhotonicLayer
layer = PhotonicLayer(client, circuit, [1, 1], shots=4000,
engine="photonic.slos.cpu")
opt = torch.optim.Adam(layer.parameters(), lr=0.05)
for step in range(200):
opt.zero_grad()
criterion(layer(), target).backward()
opt.step()Where it stops. Belenos is 12 dual-rail qubits, so the feature space is small. On today’s hardware a quantum feature map does not generally beat a classical baseline on real data, so train a classical model alongside it and compare the two.
2.11Worked example: real test runs on every tier
Everything above, shown on real jobs: thirteen runs across CPU, GPU, TPU and seven quantum processors, with the environment, the results, and what they cost. You can reproduce them from the console or the SDK.
Run A (CPU): the benchmark suite, on a laptop
Environment: an ordinary consumer laptop (Intel i7-12700H, 32 GB RAM), CPU only: no GPU, no cluster, no quantum hardware. Circuits were submitted with the engine field left blank, so the router chose the method each time.
| Circuit | Qubits | Depth | Engine picked | Wall time | Accuracy |
|---|---|---|---|---|---|
| GHZ (Clifford) | 5,000 | 5,000 | clifford | 0.56 s | exact |
| QAOA MaxCut, p=3 | 100 | 304 | mps.quimb.cpu | 5.9 s | converged (dev 0.0) |
| Layered ansatz | 80 | 9 | mps.quimb.cpu | 4.3 s | converged (dev 0.0) |
| Exact statevector | 26 | 9 | exact.cpu | 2.7 s | exact (ground truth) |
Reading the accuracy column: exact means the method makes no approximation at all, so only shot noise applies. Converged (dev 0.0) means the job was automatically re-run at double the bond dimension and the top outcome probabilities did not move; the compression captured the state. Every approximate job you run gets this same check, in its error_info.
Run B (GPU): 8-qubit GHZ
Environment: the GPU tier, submitted through the console exactly as any signed-in user would. The circuit is an 8-qubit GHZ state: one Hadamard, a chain of CNOTs, measure everything.
OPENQASM 2.0; include "qelib1.inc"; qreg q[8]; creg c[8]; h q[0]; cx q[0],q[1]; cx q[1],q[2]; cx q[2],q[3]; cx q[3],q[4]; cx q[4],q[5]; cx q[5],q[6]; cx q[6],q[7]; measure q -> c;
Steps: paste the circuit, set shots to 1,000, pick exact.gpu from the engine menu, run. The result:
counts: {"00000000": 502, "11111111": 498}
error_info: {"method": "exact statevector (GPU)",
"truncation_error": 0, "shot_noise_only": true,
"shots": 1000}
resources: {"device": "GPU", "wall_seconds": 0.54}A GHZ state should collapse to all-zeros or all-ones with equal probability, so the ideal split over 1,000 shots is 500/500; the measured 502/498 is within shot noise. Execution took half a second on the GPU. The first GPU job after a quiet period can take a little longer to start; that wait is not billed.
Run C (GPU): 32 qubits, the memory wall
The GPU tier’s exact statevector reaches 32 qubits, the point where a 32 GB laptop runs out of memory. A 32-qubit GHZ circuit (100 shots) holds a full 64 GiB statevector on the GPU, well past what a consumer machine can allocate.
counts: {"11111111111111111111111111111111": 54,
"00000000000000000000000000000000": 46}
error_info: {"method": "exact statevector (GPU)",
"truncation_error": 0, "shot_noise_only": true,
"shots": 100}
resources: {"device": "GPU", "qubits": 32,
"statevector": "64 GiB", "wall_seconds": 8.2}54/46 across the two GHZ states is the expected 50/50 split within shot noise, computed exactly (no approximation, truncation error 0) in 8.2 seconds, on GPU memory a laptop does not have.
Run D (QPU): Rigetti superconducting hardware
The same experiment on a real quantum processor: a 3-qubit GHZ circuit, 50 shots, engine qpu.rigetti, submitted through the console. Hardware runs are scheduled on the physical device, so this one returned after roughly 53 minutes in the queue; you can close the page and the result attaches to the job. The result:
counts: {"000": 23, "111": 22,
"110": 3, "010": 1, "100": 1}
error_info: {"method": "quantum hardware (Rigetti)",
"device": "Cepheus-1-108Q",
"shot_noise_and_device_noise": true,
"shots": 50,
"note": "raw hardware counts; no error mitigation applied"}45 of 50 shots landed in the two GHZ states; the remaining 5 show bit-flip errors from device noise, reported as measured. The same circuit on the simulators above gives the noiseless comparison. It can also be submitted with zero-noise error mitigation (error mitigation) to return a mitigated observable with a stated uncertainty, at three times the single-run cost.
Run E (QPU): IonQ Forte-1 trapped-ion hardware
The same family of experiment on a different kind of quantum computer: a 2-qubit GHZ (Bell) circuit, 100 shots, engine qpu.ionq, on IonQ’s Forte-1 trapped-ion processor. Trapped-ion queues are scheduled the same way; this run returned after about 5 hours. The result:
counts: {"00": 54, "11": 44, "01": 2}
error_info: {"method": "quantum hardware (IonQ)",
"device": "Forte-1",
"shot_noise_and_device_noise": true,
"shots": 100,
"note": "raw hardware counts; no error mitigation applied"}98 of 100 shots landed in the two GHZ states (00 and 11), with just 2 shots of device noise, so about 98% of shots fell in the ideal outcomes. Trapped-ion qubits have different noise characteristics from superconducting ones; running the identical circuit on both, alongside the exact simulator, measures that difference.
Run F (CPU): 192 qubits with Pauli propagation
Statevector simulation stops around 34 qubits because memory doubles with every qubit. Pauli propagation does not track the full state; it evolves the observable backward through the circuit, so its cost scales with circuit structure rather than qubit count. That reaches well past 100 qubits on an ordinary CPU.
Here is a 192-qubit GHZ circuit with an all-Z observable, engine pauli.cpu, which you can run yourself from the SDK in three lines:
import qsim_sdk
from qiskit import QuantumCircuit
n = 192
qc = QuantumCircuit(n)
qc.h(0)
for i in range(n - 1):
qc.cx(i, i + 1) # a 192-qubit GHZ state
client = qsim_sdk.Client(token="YOUR_TOKEN")
job = client.run(
qc,
engine="pauli.cpu",
observable=[[1.0, "Z" * n]], # measure <Z...Z> across all 192 qubits
)
print(job["result"]["expectation"]) # 1.0
print(job["result"]["error_info"]) # certified, error_bound 0expectation: 1.0
error_info: {"protocol": "ZCC-v0.1",
"method": "Pauli propagation (Heisenberg picture)",
"truncated_coefficient_mass": 0, "error_bound": 0,
"certified": true, "converged": true}The exact answer is known analytically: for a GHZ state the all-Z expectation value is exactly +1, and the engine returns 1.0 with a ZCC-v0.1 error bound of 0 (no Pauli terms were truncated). The run completed in under a minute and cost $0.0001, a hundredth of a cent, at the standard pauli.cpu rate. The result carries a public, independently verifiable certificate: api.zksf.org/certify/5b8b2c4309d44d41.
Run G (QPU): Quandela Belenos photonic hardware
Every hardware run above sends gates to a register of qubits. Belenos is addressed differently: single photons go through a mesh of beam splitters and phase shifters, and the detector counts how many leave by each output mode.
It is still a 12-qubit machine. Dual-rail encoding spends two modes per qubit, so 24 modes carrying 12 photons is 12 qubits counted the other way. What differs is the interface, not the qubit count.
The program is a circuit plus an input Fock state, so it has its own call, client.run_photonic(...), set out under Photonic linear optics above.
The experiment is Hong-Ou-Mandel interference, which is the photonic analogue of the GHZ test used above. Two identical photons meeting at a balanced beam splitter must leave together, so the ideal output is half |2,0> and half |0,2> with the coincidence term |1,1> at exactly zero. A coincidence indicates the two photons were not perfectly identical, so the coincidence fraction is a direct measure of photon quality.
counts: {"|1,0,0>": 4972, "|0,1,0>": 4742,
"|2,0,0>": 135, "|0,2,0>": 127,
"|0,0,1>": 15, "|1,1,0>": 8, "|1,0,1>": 1}
error_info: {"method": "photonic hardware (qpu:belenos)",
"device": "Belenos", "shots": 10000,
"device_noise_model": {"indistinguishability": 0.859,
"g2": 0.015,
"transmittance": 0.054},
"note": "raw hardware counts; no error mitigation applied"}
resources: {"modes": 3, "photons": 2,
"cost_eur": 0.31, "cost_usd": 0.3348}Almost every detection landed in a single-photon state. That is photon loss in the optics, not a failed run: the device reports its own transmittance as 0.054, so most detections arrive with one photon missing.
Only 271 detections out of 10,000, about 2.7%, registered both photons that were sent. Those are the only events in which interference can be observed at all.
Within those 271: 262 were bunched into |2,0,0> or |0,2,0>, and 9 were coincidences. The certificate reports a Hellinger fidelity of 0.9666 against the exact SLOS reference, conditioned on full photon detection, and states that conditioning: api.zksf.org/certify/063a6fcb82a9464a.
The conditioning is stated because it changes what the fidelity means.
- The ideal distribution gives zero probability to any outcome missing a photon
- Comparing raw counts against it therefore measures the transmittance of the apparatus, not interference, and returns a number near zero however good the photons are
- Conditioning on shots where every photon arrived is what the reported fidelity measures
- Belenos also reports the inputs to that number: indistinguishability 0.859, g² 0.015
The run took about 72 seconds end to end and cost €0.31, roughly $0.33 at 10,000 shots. Three further runs, and what changes between them, are in the photonic write-up.
Run H (QPU): IQM Garnet, a QAOA schedule on superconducting hardware
Runs D to G all send the same small GHZ circuit, which measures the device rather than solving anything. This one is an application: satellite observation tasking at 14 requests, seed 20260902, whose exact optimum is value 32.5128. The circuit is a QAOA ansatz over those 14 qubits, engine qpu.iqm.garnet, 500 shots, on IQM Garnet (superconducting, 20 qubits).
best sampled value: 26.8778 (exact optimum 32.5128, gap 5.64)
error_info: {"method": "quantum hardware (IQM Garnet (superconducting))",
"device": "Garnet", "shots": 500}
certificate: ZHF-v0.1, fidelity 0.09589, mode "direct"Read the two numbers separately. The certified fidelity is 0.096, which says the measured distribution over 214 outcomes is far from the ideal one. The answer is still usable, because QAOA does not need a faithful distribution: it needs one good sample, and every sample is checked against the problem classically before it is reported. A device can therefore return a usable optimisation result and a low distribution fidelity at the same time, and both are certified here rather than only the flattering one.
Certificate d758835eec8d4a64.
Run I (QPU): IQM Emerald, the same problem on a wider chip
Identical problem, identical ansatz, identical shot count, on IQM Emerald: same manufacturer, 54 qubits instead of 20. Changing one argument, engine, is the whole difference between this run and Run H, which is what makes the two comparable.
best sampled value: 26.9421 (exact optimum 32.5128, gap 5.57)
error_info: {"method": "quantum hardware (IQM Emerald (superconducting))",
"device": "Emerald", "shots": 500}
certificate: ZHF-v0.1, fidelity 0.097992, mode "direct"Emerald returns the better value of the two, 26.9421 against Garnet’s 26.8778, and the best of any quantum processor on this problem. The margin is small, and one run of 500 shots does not establish that one chip is better than the other at this width. What it does establish is that both reach the same neighbourhood and that neither closes the gap to the classical optimum. The classical engines on the same problem are on the space applications page.
Certificate a323a0c50cfe4520.
Run J (QPU): AQT IBEX Q1, an expectation value on trapped ions
A different measurement again: not counts and not an optimisation, but a single expectation value. The circuit prepares the H2 ground state at equilibrium on two qubits and measures ZZ, whose ideal value is exactly -1. Engine qpu.aqt.ibex, 100 shots, on AQT IBEX Q1 (trapped ion, 12 qubits).
ZZ = -0.9200 (ideal -1.0000)
error_info: {"method": "quantum hardware (AQT IBEX Q1 (trapped ion))",
"device": "Ibex-Q1", "shots": 100}
certificate: ZHF-v0.1, fidelity 0.948081, mode "direct"The same circuit on IQM Garnet returned ZZ = -0.9326 at 4,096 shots with a fidelity of 0.966196, certificate 485065a7ddce4945. Two modalities, one observable, both certified against the same exact reference. The shot counts differ by a factor of 40, which matters when comparing them: a device with a 2,000-shot ceiling and one with a 20,000-shot ceiling are not measured to the same precision, and the interval on an expectation value narrows as 1/√N.
Certificate 4924b8d2e8a743b8.
Run K (QPU): QuEra Aquila, an independent set solved by the blockade
Neutral-atom hardware does not run gates. The problem is the geometry: two atoms inside the blockade radius cannot both be excited, which is the independent-set constraint itself, so there is no encoding and no penalty terms. This run is nine atoms on a 3×3 grid at 6 um, submitted through POST /mis with engine qpu.quera.aquila and 1,000 shots.
At 6 um spacing the orthogonal neighbours are inside the 10.03 um blockade and so are the diagonals at 8.49 um, so the graph is the king’s graph on a 3×3 board: 9 vertices, 20 edges. The answer is checkable by hand. The four corners are 12 um apart, outside the blockade, and no independent set of five exists, so the maximum independent set is 4.
best_set: [0, 2, 6, 8] the four corners, size 4 valid_fraction: 0.9324 883 of 947 shots obeyed every edge optimal shots: 574 of 947 60.6% returned a set of size 4
Read valid_fraction first. Every shot is tested against the edges the register implements and invalid ones are excluded from the answer, so that fraction is the quality figure for the sweep: a low number means the adiabatic ramp was too fast for the instance, not that no solution exists. Here 93.2% of shots came back a valid independent set and most of those were optimal.
A ring is not a shape this machine can hold, which the same run established. Aquila addresses atoms in rows, so any two atoms share a y coordinate exactly or sit at least 4 um apart in y; a regular octagon puts two of them 3.75 um apart and is refused before submission. Scaling that octagon up until its rows clear 4 um pushes the sides past the blockade radius and the edges disappear. Lay a register out against those two rules rather than against a drawing.
A separate Aquila run carries a device fidelity rather than a solution: an analog Rydberg sequence on 4 atoms returned 0.6294 against an exact Schrödinger reference, certificate d3abecb2c8d14c2f, on 19 retained shots. At 19 shots that figure carries a wide interval, which is why the retained shot count is reported beside it.
Run L (QPU): Pasqal FRESNEL, the same modality on a second machine
FRESNEL takes the same Pulser sequence Aquila takes, so moving a sequence between them changes one argument. This run is a 4-atom sequence, engine qpu.pasqal.fresnel, 20 shots, on FRESNEL_CAN1, certified against an exact Schrödinger evolution computed with QuTiP.
top_outcomes: [["0000", 4], ["0010", 3], ["0001", 3],
["0110", 3], ["1001", 2], ["1000", 2]]
error_info: {"method": "neutral-atom hardware (FRESNEL_CAN1)",
"device": "FRESNEL_CAN1", "shots": 20}
certificate: ZHF-v0.1, fidelity 0.939827, mode "direct"Twenty shots is the sample, and it sets what the number can say. Sampling alone holds even a flawless device to about 0.88 at this width, so 0.9398 says FRESNEL is consistent with a high-fidelity device. It does not establish that FRESNEL is half again better than the 0.6294 above, because at this width and this shot count the two cannot be separated. Tightening it takes more shots, and the reason this run has 20 is the machine’s operating model rather than its speed: FRESNEL is driven as machine time at a contractual 0.25 Hz repetition rate, so one shot is four seconds and 20 shots is 80 seconds on the device.
Certificate bae45351a4a44538.
Run M (TPU): a neural network wavefunction on a Google TPU
The neural engines carry no circuit at all. They vary a neural network wavefunction by variational Monte Carlo and are reached through POST /solve. The problem here is the same H2 molecule as Run J, whose exact total ground-state energy is -1.137281 Ha, run once on neural.tpu and once on neural.cpu.
neural.tpu -1.116981 Ha total 0.0203 Ha above exact neural.cpu -1.116981 Ha total 0.0203 Ha above exact certified variational energies, where the two differ: neural.tpu -0.784595 ZCC-Estimate-v0.2 2ffa6319a68046e8 neural.cpu -0.784611 ZCC-Estimate-v0.2 22e9ae2aa0b24380
The same method on both, and the CPU is marginally lower. The two agree to five decimal places, and the CPU’s energy is the better of them by 1.5e-05 Ha. That direction is expected rather than surprising: a TPU has no complex128, so the engine narrows to complex64 there and records that it did, and a lower precision moves the variance floor the bound is read from.
This width is not what the tier is for. A TPU has to be provisioned before it computes anything, which at two qubits is most of the job, so a search this small finishes sooner on the CPU tier. The TPU earns its place where the network is large enough that dense matrix multiplication against batches of sampled configurations dominates the run. Either way the $0.25 task fee covers the first 480 seconds.
Both runs report a ceiling rather than only an energy: the variational principle puts the true ground state at or below the measured value, so the ceiling is the number the answer cannot exceed. How to read that bound, and the conditional floor beside it, is under Neural engines.
Summary: one circuit family, every tier
| Tier | Circuit | Qubits / modes | Result | Time |
|---|---|---|---|---|
| CPU | GHZ (Clifford) | 5,000 | exact | 0.56 s |
| CPU | QAOA MaxCut, p=3 | 100 | converged (dev 0.0) | 5.9 s |
| GPU | GHZ, 1,000 shots | 8 | 502/498 (ideal 500/500) | 0.5 s |
| GPU | GHZ, 100 shots | 32 | 54/46 (ideal 50/50); 64 GiB statevector | 8.2 s |
| CPU | GHZ (MPS), certified | 2,000 | converged; bound 2.1e-08 | 873 s |
| CPU | QAOA ring, p=1 (non-Clifford) | 1,000 | converged; bound 3.2e-07 | 312 s |
| CPU | Ising Trotter, p=4 (non-Clifford) | 1,000 | bound 2.5e-05 | 203 s |
| QPU | GHZ, 50 shots (Rigetti Cepheus) | 3 | 45/50 in GHZ states, 5 noise | ~53 min incl. queue |
| QPU | GHZ, 50 shots (IQM Garnet) | 3 | 48/50 in GHZ states, 2 noise | ~25 s incl. queue |
| QPU | GHZ, 50 shots (IQM Emerald) | 3 | 48/50 in GHZ states, 2 noise | ~60 min incl. queue |
| QPU | GHZ, 100 shots (IonQ Forte Enterprise 1) | 3 | 95/100 in GHZ states, 5 noise | ~5 min incl. queue |
| QPU | GHZ, 20 shots (AQT IBEX Q1) | 3 | 20/20 in GHZ states, no noise observed | ~80 min incl. queue |
| QPU | Hong-Ou-Mandel, 10,000 shots (Quandela Belenos) | 3 modes, 2 photons | 262 of 271 two-photon detections bunched, 9 coincidences | ~72 s |
| QPU | Rydberg sequence, 20 shots (QuEra Aquila) | 4 atoms | fidelity 0.6294 on 19 retained shots | ~2 h in window |
| QPU | Independent set, 1,000 shots (QuEra Aquila) | 9 atoms | found the optimum, size 4; 93.2% of shots valid | ~1 h in window |
| QPU | Rydberg sequence, 20 shots (Pasqal FRESNEL) | 4 atoms | fidelity 0.9398; 20 shots cannot separate it from perfect | 80 s on device |
| QPU | QAOA tasking, 500 shots (IQM Garnet / Emerald) | 14 | 26.8778 and 26.9421 against an exact optimum of 32.5128 | minutes |
| TPU | H2 ground state, neural wavefunction | 2 | -1.116981 Ha, 0.0203 Ha above exact; matches neural.cpu | provisioning-dominated at this width |
| CPU | GHZ, Pauli propagation, all-Z | 192 | <Z…Z> = 1.0, certified (bound 0) | < 1 min |
Python SDK
Working in code: install, a token, the same engines from a script or a notebook, sweeps and training loops, and Qiskit, PennyLane and PyTorch.
3.1Install
pip install qsim-sdk
Python 3.10 or newer. Two extras are optional, because most scripts need neither:
pip install "qsim-sdk[ml]" # PyTorch layers, qsim_sdk.ml pip install "qsim-sdk[multiframework]" # Cirq, PennyLane and pyQuil circuits
3.2Create an API key
Sign in at app.zksf.org, open Profile and press Create API key. A key is valid for 30 days and is shown only once, so copy it straight away; if it is lost, revoke it there and create another. Keep it out of your source: read it from an environment variable, as the examples in the SDK repository do.
import os import qsim_sdk client = qsim_sdk.Client(token=os.environ["ZKSF_TOKEN"])
From qsim-sdk 0.8.0, qsim_sdk.Client() reads ZKSF_TOKEN by itself when no token is passed, as the Qiskit and PennyLane integrations do.
3.3Run your first circuit
import qsim_sdk
from qiskit import QuantumCircuit
qc = QuantumCircuit(60) # 60 qubits
qc.h(0)
for i in range(59):
qc.cx(i, i + 1)
client = qsim_sdk.Client(token="YOUR_TOKEN")
job = client.run(qc, shots=1000)
print(job["result"]["counts"]) # the answer
print(job["result"]["error_info"]) # how much to trust itThe circuit is Clifford, so the router runs it on a stabilizer engine: exactly, in milliseconds. A 60-qubit statevector would need 16 million terabytes of memory.
3.4Estimate before you spend
est = client.estimate(qc)
# {'engine': 'mps.quimb.cpu', 'predicted_seconds': 5.1,
# 'predicted_cost_usd': 0.00014, 'reason': '60 qubits, low
# entangling density -> MPS attempt with convergence check'}Estimates are free and conservative. No job is charged before one is available.
About estimates
A simulation is billed by the seconds it actually runs, so the charge can differ. Simulation estimates vary with the width, depth and structure of each circuit and are a reference point only. Your final charge is set by your job's actual runtime, as described in our pricing. Estimates are often higher than the runtime charged, but can also be lower. QPU estimates stay close to the final charge, because each provider sets its price per shot
That prices one circuit. The minimum charge applies per circuit, so a 600-point sweep costs 600 minimums. Reading the single-circuit figure as the price of a sweep understates it by the number of points. Price a sweep with estimate_batch, which returns the total, the cost of each point, and how many requests the sweep will take.
est = client.estimate_batch(circuits, shots=1024, engine="exact.cpu")
# {'points': 12000, 'total_usd': 1.2, 'batches': 60, ...}
est["per_point_usd"] # the cost of each pointEngines scale to zero when nothing is queued, which is what keeps simulation at a fraction of a cent. The first job after an idle period waits while a worker starts; later jobs are faster. The SDK prints a line to stderr once a job has waited more than fifteen seconds, and Client(on_poll=...) drives your own progress output instead.
3.5Qiskit, PennyLane and PyTorch
Code you already have does not need rewriting. client.run takes a Qiskit circuit as it is, and with the multiframework extra it takes Cirq, PennyLane and pyQuil circuits too and converts them. Beyond that, two companion packages make the engines look native to Qiskit and to PennyLane, and qsim_sdk.ml turns a quantum program into a PyTorch layer.
Qiskit
qiskit-zksf registers the engines as Qiskit backends and as V2 primitives, so an existing Qiskit program runs here by naming a different backend rather than being rewritten. It covers gate circuits, on the simulators and on the five gate QPUs.
pip install qiskit-zksf
from qiskit_zksf import ZKSFProvider
backend = ZKSFProvider(token="YOUR_TOKEN").backend("zksf_auto")
job = backend.run(qc, shots=1000)
job.result().get_counts() # outcomes
job.error_info() # how far they may be offfrom qiskit_zksf import ZKSFProvider provider = ZKSFProvider() result = provider.sampler().run([circuit], shots=1000).result() result[0].data.meas.get_counts() # outcomes result[0].metadata["error_info"] # how far they may be off result = provider.estimator().run([(circuit, observable)]).result() result[0].data.evs # expectation value result[0].data.stds # the measured bound on it
zksf_autois the router. Every gate engine also has a backend of its own, such aszksf_mps,zksf_cliffordorzksf_iqm_garnet, andprovider.backends(hardware=True)lists the quantum processors- Qubit limits match the service’s, so a circuit too large for a backend fails at transpile time rather than after a round trip
provider.estimate(qc, shots=1000)is the free estimate, which Qiskit has no concept of- Every parameter binding in a primitive run is a separate billed job, so a large sweep is refused unless you raise
max_jobs_per_run
Without installing any of this, qasm2.dumps(circuit) or qasm3.dumps(circuit) produces text the console’s code box takes, which is all a one-off run needs. A circuit with mid-circuit if or while exports only as 3; see dynamic circuits and Run a job.
PennyLane
pennylane-zksf registers the engines as a PennyLane device, so an existing QNode runs here by naming a different device. A parameter sweep goes in one request, because PennyLane hands the device a batch of circuits whenever it broadcasts.
pip install pennylane-zksf
import pennylane as qml
dev = qml.device("zksf.simulator", wires=3, shots=1024)
@qml.qnode(dev)
def circuit(theta):
qml.RY(theta, wires=0)
qml.CNOT(wires=[0, 1])
return qml.expval(qml.PauliZ(0))
circuit(0.5)
dev.error_info() # the accuracy statement for that run- Pin an engine with
engine=, a quantum processor included:qml.device("zksf.simulator", wires=2, shots=1000, engine="qpu.iqm.garnet") - Measurements:
qml.counts,qml.sample,qml.probs, andqml.expvalof a Pauli observable, one per circuit. An unpinned device sends expectation values topauli.cpu - No gradients: each execution is a billed job. Compute gradients on a local device and use this one for the runs that need the reach or the bound
Without installing any of this, tape.to_openqasm() produces the OpenQASM 2 text the console’s code box takes, which is all a one-off run needs. The box takes OpenQASM 3 as well. See Run a job.
pennylane-zksf on PyPI · source
PyTorch
qsim_sdk.ml wraps a parameterised program as an nn.Module, so a quantum circuit becomes a differentiable layer in an ordinary PyTorch model. PhotonicLayer takes a Perceval circuit and SequenceLayer a Pulser sequence, and the output is the probability of each outcome in a fixed order.
- Forward pass: 1 job. Backward pass: 1 job carrying 2 points per parameter
- Gradients are central differences, not the parameter-shift rule. The shift rule is exact only where an output is a sinusoid of the parameter, true for a Pauli rotation and false for a beamsplitter angle or a pulse amplitude
- torch is not an SDK dependency.
pip install "qsim-sdk[ml]"
import perceval as pcvl
import torch
from qsim_sdk.ml import PhotonicLayer
circuit = pcvl.Circuit(2) // (0, pcvl.BS(theta=pcvl.P("theta")))
layer = PhotonicLayer(client, circuit, [1, 1], shots=4000,
engine="photonic.slos.cpu")
opt = torch.optim.SGD(layer.parameters(), lr=0.3)
for step in range(200):
opt.zero_grad()
loss = criterion(layer(), target) # layer() = probability per output state
loss.backward()
opt.step()Shot noise sets the floor on how small a gradient you can resolve, so an estimate from N shots carries an error of order 1/√N, so an optimiser that stalls at low shot counts is usually reading noise rather than a flat landscape. Raise shots, or raise the step size eps, before treating the model as untrainable.
The cost of training on hardware, and where the approach stops, are set out in a quantum layer inside an existing model.
3.6Pick an engine from code
Leave engine out and the router chooses. Name one and it is used exactly as given, including where another engine would suit the circuit better. Engine parameters pass straight through.
client.run(qc) # router chooses; the result names the engine client.run(qc, engine="exact.cpu") # used exactly as given client.run(qc, engine="mps.quimb.cpu", max_bond=128) # engine parameters pass through
Which call you use depends on what the engine takes. What each is for, and its limits, are in The engines.
client.run(circuit, engine=...)
exact.gpu
exact.tpu
qpu.rigetti
qpu.iqm.emerald
qpu.iqm.garnetclient.run_sequence(sequence, engine=...)
analog.pulser.gpuclient.run_photonic(circuit, input_state, engine=...)
photonic.gpu
qpu.quandela.belenosclient.solve(hamiltonian, qubits, engine=...)
neural.tpu
neural.gpu3.7Pulse sequences and photonic programs
Neutral-atom and photonic engines take programs that are not gate circuits, so they have calls of their own. What each program is, and the device rules it has to follow, are under Analog neutral-atom sequences and Photonic linear optics.
Pulser sequences
- Use
client.run_sequence(...), notrun() - No Pulser installed? Pass
seq.to_abstract_repr()as a JSON string
from pulser import Pulse, Register, Sequence
from pulser.devices import AnalogDevice
reg = Register.square(2, spacing=6.0).with_automatic_layout(AnalogDevice)
seq = Sequence(reg, AnalogDevice)
seq.declare_channel("ising", "rydberg_global")
seq.add(Pulse.ConstantPulse(1000, 6.0, 0.0, 0.0), "ising")
job = client.run_sequence(seq, shots=1000)
job["result"]["counts"] # bitstrings, 1 = RydbergThe same call reaches all three engines. Change one argument.
Local detuning adds a spatial pattern to the same sequence, within the rules under Analog neutral-atom sequences:
from pulser.register.weight_maps import DetuningMap seq.config_detuning_map(DetuningMap(coords, [1.0, 0.5, 0.0]), "dmm_0") seq.add_dmm_detuning(ConstantWaveform(600, -3.0), "dmm_0")
job = client.run_sequence(seq, shots=100, engine="qpu.quera.aquila") # Aquila reports empty sites as well as ground and Rydberg states. Shots with # an empty site are dropped by declared post-selection, and the retained # fraction is reported alongside the result: a fidelity computed from these # counts is a fidelity over the shots that were kept. job["result"]["counts"] job["result"]["error_info"]["post_selection"]["retention"]
Perceval programs
Use client.run_photonic(circuit, input_state, ...). The input state takes a plain list, so no Perceval objects are needed to say where the photons go.
import perceval as pcvl
from perceval.components import BS, PERM
# Hong-Ou-Mandel: two photons meet on a balanced beamsplitter
circuit = pcvl.Circuit(3) // (1, PERM([1, 0])) // (0, BS.H())
job = client.run_photonic(circuit, [1, 0, 1], shots=1000)
job["result"]["counts"] # {'|2,0,0>': 498, '|0,2,0>': 502}
job["result"]["error_info"]
# {'method': 'exact linear-optics simulation (SLOS)',
# 'truncation_error': 0.0, 'shot_noise_only': True,
# 'modes': 3, 'photons': 2}
# Real hardware, same call:
client.run_photonic(circuit, [1, 0, 1], shots=1000,
engine="qpu.quandela.belenos")From another language, post a photonic body to /jobs carrying {"circuit": ..., "input": ...} as a JSON string, both halves serialised by Perceval. There is no estimate() for photonic or analog programs yet: the cost model reads gate-circuit features they do not have.
3.8Verbatim execution
By default the provider compiles your circuit. It may substitute gates, requantise angles and route onto different physical qubits. Pass verbatim=True and the device runs exactly what you wrote.
client.run(qc, engine="qpu.rigetti", shots=1000,
verbatim=True)- Use it for benchmarking and device characterisation. A benchmark of a compiled circuit measures the compiler as much as the machine
- Supported on all five gate QPUs
- Requires the circuit to already be in that device's native gates on its own connectivity
- If it is not, the provider rejects the task and does not bill it
- Certificates are unaffected. It is still a circuit with a computable ideal distribution
3.9Parameter sweeps and training loops
A variational solver, a quantum kernel or a photonic generative model evaluates one parameterised program hundreds of times as an optimiser moves its parameters. Send the whole step as one job rather than one submission per evaluation.
Quantum machine learning on simulators covers which of these approaches currently work, which remain research, and the shot cost of a training loop.
| Program kind | Call | You send |
|---|---|---|
| Gate circuit | run_sweep | One bound circuit per point. OpenQASM 2 has no free parameters. |
| Pulser or Perceval | run_parametric_sweep | The program once, plus a list of bindings. Bound server-side. |
Each bound point hashes to its own value, so a certificate still identifies exactly which parameters produced it.
theta = Parameter("theta")
qc = QuantumCircuit(2)
qc.h(0); qc.rz(theta, 0); qc.cx(0, 1); qc.measure_all()
job = client.run_sweep(qc, [{theta: v} for v in values], shots=1024)
qsim_sdk.counts(job) # one per value, in the order you sent them
job["summary"] # {"total": N, "succeeded": ..., "rejected": ..., "errored": ...}- A failed point does not fail the batch. It returns
status: "error"with a reason, the rest return normally, andcounts()putsNonein its slot so the list still lines up with the values you sent - Only points that returned a result are charged
- One certificate per evaluation, not per batch. A bound is a property of one execution. Name the one you want with
POST /jobs/{job_id}/certificate?index=N - 200 points per batch. Use
run_batchfor a list of unrelated circuits rather than a sweep
# photonic: parameters declared with pcvl.P("name")
circuit = pcvl.Circuit(2) // (0, pcvl.BS(theta=pcvl.P("theta")))
job = client.run_parametric_sweep(
circuit,
[{"theta": t} for t in numpy.linspace(0, 3.14, 40)],
input_state=[1, 1],
engine="photonic.slos.cpu",
)
job["results"][7]["result"]["counts"] # the point for bindings[7]
# analog: variables declared with seq.declare_variable("amp")
job = client.run_parametric_sweep(
seq, [{"amp": a} for a in amplitudes], engine="analog.pulser.cpu")Sweeps and batches run on simulation engines only: run_sweep, run_parametric_sweep or run_batch naming a qpu.* engine is refused with a 400. Each hardware task is queued and billed on its own, so settle the sweep in simulation, then send the point you want to the machine.
In the console, tick Run several circuits as a batch on the circuit editor and paste the circuits one after another, each starting with its own OPENQASM line, 2.0 or 3. It prices them with one quote and splits them into batches of 200 for you. A parameter sweep over one program is SDK only.
To train a program as one layer of a PyTorch model instead, see PyTorch.
3.10SDK and HTTP API reference
| Call | What it does | Since |
|---|---|---|
| Client(base_url, token) | Construct once, reuse. | 0.1 |
| estimate(circuit, shots) | Free feasibility and cost forecast. | 0.1 |
| estimate_batch(circuits, shots) | The same for a whole sweep: total, per point, batches. | 0.9 |
| run(circuit, shots, engine, **params) | Submit and wait. Raises JobRejected / JobFailed. | 0.1 |
| submit(...) / job(job_id) | Non-blocking variant. | 0.1 |
| run_batch(circuits) | Many unrelated circuits in one request. | 0.2 |
| run_sweep(circuit, bindings) | One parameterised Qiskit circuit at many values. Binds locally. | 0.2 |
| run_sequence(sequence, shots) | A neutral-atom Pulser sequence. Above. | 0.4 |
| run_photonic(circuit, input, shots) | A linear-optics circuit and its photons. Above. | 0.5 |
| run_parametric_sweep(program, bindings) | One parameterised Pulser or Perceval program, bound server-side. | 0.6 |
| qsim_sdk.ml.PhotonicLayer / SequenceLayer | A program as a torch nn.Module. Needs the [ml] extra. Above. | 0.6 |
| solve(hamiltonian, qubits) | Ground state, loop run for you. Read the ceiling. | 0.7 |
| run_mis(vertices) | Maximum independent set on a neutral-atom register. | 0.7 |
| qsim_sdk.counts(job) / expectations(job) | Batch results in submission order, None where a point returned nothing. | 0.2 |
| estimate_solve(hamiltonian, qubits) | The price of a ground-state search before submitting it. Free. | 0.8 |
| pending() / jobs(limit, cursor) / iter_jobs() | Every unfinished job, and the job history newest first. | 0.8 |
| cancel(job_id) | Cancel a hardware job still waiting in its provider’s queue. Raises CancelRefused with the reason when it cannot be. | 0.8 |
| summary(since) / balance() | Jobs run per tier and net spend this month; available credit. | 0.8 |
| certificate(job_id, index) | Issue a finished job’s public certificate: verify URL and PDF. | 0.8 |
Every call has a submit_ twin that returns a job id instead of waiting.
The SDK is a thin wrapper over a plain HTTP API, so any language can call it directly. The full specification is published in machine-readable form, browsable interactively at api.zksf.org/docs or as raw OpenAPI 3.1 at openapi.json. Certificates are readable as JSON at /certify/{cert_id}/json, alongside the HTML verification page and the PDF, so a third party can check one programmatically rather than by eye.
Over HTTP, pending() is GET /jobs/pending, every unfinished job on your account across its whole history, which is what the console’s jobs-in-queue card counts. cancel() is POST /jobs/{job_id}/cancel, which asks the provider to cancel a hardware job that is still waiting in its queue: it answers 409 with the reason when the job cannot be cancelled, and the charge is refunded only once the provider confirms the job never ran. A job waiting at a provider reads queued, with queue_position when the provider publishes one, and moves to running when the device starts it.
zcc-verify checks a certificate independently, recomputing its declared bound from the measurement it reports, with no dependencies and no account: pip install zcc-verify.
Questions, bugs, or a circuit that is refused and should not be: support@zksf.org. Benchmark challenges welcome.
3.11Citing ZKSF
If a result produced here appears in published work, cite the archived client. The DOI below is the concept DOI, which always resolves to the most recent release and so does not go stale.
10.5281/zenodo.21836619
The full record, including per-version DOIs, is at doi.org/10.5281/zenodo.21836619. To pin an exact release for reproducibility, use that release’s own version DOI from the record page rather than the concept DOI. The repository also ships a CITATION.cff, so GitHub’s “Cite this repository” button produces BibTeX and APA directly.
Individual jobs are separately citable. Every certificate has a permanent public URL under api.zksf.org/certify/ that resolves without an account, so a reviewer can check the exact accuracy statement behind a specific number rather than taking the figure on trust.
Quantum Computing use cases
Every benchmark published under /applications, with the call that reproduces it and what to change to run your own instance instead.
4.1How to reproduce these
Every benchmark on the applications pages is a job on this service, priced as any customer is priced. Each section below gives the call that produced the published figure and what to change to run your own instance. Start every one of them on a simulator engine: the answer is exact, the cost is a hundredth of a cent, and only then does a hardware engine tell you something the simulator cannot.
Reproducing our own instances, exactly
This section has two modes. The snippets under each use case take yourcircuit or QUBO and run it on our engines. To land on the exact numbers an applications page publishes, take our instance from the list below and run it the same way. Both use the same calls; only the input differs.
Instance files ship in the SDK repository under examples/benchmarks/instances/.
- Space, satellite tasking.
satellite-tasking.ipynbbuilds it from seed 20260902: 14 requests, 14 qubits, exhaustive optimum 32.5128. QAOA p=2, COBYLA, three restarts, scored as the best feasible value in the returned shots - Finance, portfolio.
portfolio-optimisation.ipynb, seed 20260902: 12 assets, choose 4, optimum 0.199796, and the percentile column is that objective ranked against all 495 feasible portfolios - Traffic. traffic_instances.json, the 10-vehicle entry: 30 qubits, 228 couplings, exhaustive optimum congestion 11. One binary per vehicle-route pair, one route per vehicle as a penalty, and a coupling for each road segment two routes share. Past exact simulation at 30 qubits, so the angles come from the 24-qubit entry of the same file
- Logistics and manufacturing. sector_instances.py builds both and solves them exhaustively; the snippet below is the logistics half in full. Check against the QUBO minimum it prints
The same instances on the newest engines
Four engines were added on 25 September 2026, and the published tables gained a row for each from a production job on the instance already described above. Nothing changes but the engine argument, which is the point: the same call, the same instance, a different machine.
engine="exact.tpu"ran space (31.2164, gap 1.30), finance (0.199796, the provable optimum), logistics (-18.7897, the optimum), manufacturing (-12.8804, gap 2.81) and the 7-bit elliptic-curve key recovery. The charge on each is on the table it belongs to. A machine is created per job, so expect two to five minutes of wall time for a circuit that computes in seconds, and send a batch when you have oneengine="analog.pulser.gpu"ran the nine-atom maximum independent set, the identical Pulser sequence as the CPU row, 500 of 500 shots validengine="neural.gpu"ran the six-spin ring throughPOST /solve, energy -6.647038 with a ceiling of -6.630361. It takes a Hamiltonian, not a circuit, exactly as the other neural engines doengine="photonic.gpu"takes the same Perceval circuit and input Fock state asphotonic.slos.cpu, to 24 modes and 12 photons
Traffic is the one instance where the engine changes the problem. Thirty qubits is past what exact.tpu holds, so its row is a 29-qubit variant with the last vehicle's third route removed, whose exact optimum is congestion 12 rather than 11.
import itertools, numpy as np
# Logistics: 4 vehicles, 4 stops, one each. Cost is distance driven.
rng = np.random.default_rng(20260919)
depots, stops = rng.uniform(0, 10, (4, 2)), rng.uniform(0, 10, (4, 2))
dist = np.linalg.norm(depots[:, None, :] - stops[None, :, :], axis=2)
pen = float(dist.max()) * 4.0
Q = np.zeros((16, 16)); f = lambda a, b: a * 4 + b
for a in range(4):
for b in range(4): Q[f(a,b), f(a,b)] = float(dist[a,b]) - 2*pen
for a in range(4): # one stop per vehicle
for b1 in range(4):
for b2 in range(b1+1, 4): Q[f(a,b1), f(a,b2)] += pen; Q[f(a,b2), f(a,b1)] += pen
for b in range(4): # one vehicle per stop
for a1 in range(4):
for a2 in range(a1+1, 4): Q[f(a1,b), f(a2,b)] += pen; Q[f(a2,b), f(a1,b)] += pen
# -292.3588: the number the Rigetti row is scored against
print(min(float(np.array(x) @ Q @ np.array(x))
for x in itertools.product([0,1], repeat=16)))- Manufacturing is the same shape on seed 20260920: four jobs over four slots, cost is weighted lateness against each job's due slot, and two jobs sharing a machine in one slot is penalised. Its QUBO minimum is -351.2615
- Score every shot by the QUBO's own value, penalties included, and report the best. A 16-binary assignment problem has 24 feasible states in 65,536, so filtering on feasibility returns nothing from 500 shots and tells you nothing about the device
- Angles are re-optimised per run. The optimiser samples a simulator, so two runs of the same instance converge to different angles. Same problem, same depth, same shot count; not a byte-identical circuit
pip install qsim-sdk import qsim_sdk client = qsim_sdk.Client(token="YOUR_TOKEN") # Profile > Create API key client.estimate(circuit, shots=512) # free, before anything is charged client.balance() # what is left on the account
- What the console does too. Gate circuits, batches of gate circuits, ground states, sweeps of one coefficient, maximum independent set and state reconstruction all have a form at app.zksf.org, and both share one job history
- SDK only. Anything that submits more than one job per step. A training loop is one job per gradient evaluation, so the console cannot run one. It can send a list of gate circuits as a batch, but not a parameter sweep, and it cannot stitch an image across several jobs
- The scripts behind the AI cases. Each AI case below has its full script, raw results and saved weights in examples/ai on GitHub
- The extras.
pip install "qsim-sdk[ml]"adds the PyTorch layers used by five of the sections below;pip install "qsim-sdk[multiframework]"accepts Cirq, PennyLane and pyQuil circuits
4.2Molecular ground state
Reproduces chemicals and pharmaceuticals. A Hamiltonian in, an energy and a ceiling out: the service runs the variational loop, so no circuit is submitted. H2 at equilibrium is two qubits and its exact total is -1.137281 Ha.
job = client.solve(
hamiltonian=[[-1.0523, "II"], [0.3979, "IZ"], [-0.3979, "ZI"],
[-0.0113, "ZZ"], [0.1809, "XX"]],
qubits=2, ansatz="real_amplitudes", reps=2,
max_iterations=200, shots=1024, engine="exact.cpu")
job["result"]["energy"] # the optimiser's best
job["result"]["ground_state"]["ceiling"] # what the true answer cannot exceed- Your own molecule. Run the electronic structure in PySCF and map it to Pauli operators with OpenFermion or Qiskit Nature. Every term must name every qubit, padded with
I - Read the ceiling, not the energy. The variational principle puts the true ground state at or below what the optimiser reports; adding the run's own error bound gives the number it cannot exceed
- Cost follows the evaluation count. SPSA runs 2 evaluations an iteration plus 1 final run, so 60 iterations is 121 evaluations, and the job is metered by the second they take
4.3Portfolio selection
Reproduces finance and trading. A cardinality-constrained selection becomes a QUBO and then a Pauli Hamiltonian, so it takes the same call as a molecule. 24 qubits is 24 assets.
# Q is your QUBO matrix. z_i = 1 - 2*x_i maps {0,1} to Pauli Z.
terms = [[float(Q[i][i]) / 2.0, "I" * i + "Z" + "I" * (n - i - 1)]
for i in range(n)]
job = client.solve(hamiltonian=terms, qubits=n,
ansatz="real_amplitudes", reps=2,
max_iterations=60, engine="exact.cpu")- Your own universe. Build Q from your covariance matrix and expected returns, with the cardinality constraint as a penalty term. The notebook
portfolio-optimisation.ipynbbuilds it end to end on free packages - Check the encoding before you pay for it. Solve the same QUBO by exhaustive enumeration at small n and confirm the Hamiltonian reproduces that answer
4.4Satellite tasking
Reproduces space and satellites. This one is a gate circuit rather than a Hamiltonian: a QAOA ansatz over the tasking instance, submitted like any other circuit, which is why it reaches the gate QPUs.
for engine in ("exact.cpu", "exact.gpu", "qpu.iqm.garnet", "qpu.iqm.emerald"):
est = client.estimate(qaoa_circuit, shots=500, engine=engine)
print(engine, est["predicted_cost_usd"])
job = client.run(qaoa_circuit, shots=500, engine="exact.cpu")
qsim_sdk.counts(job)- Your own instance. Seed your own requests, windows and values. The published run is 14 requests at seed 20260902, whose exact optimum is 32.5128
- Score every sample classically. QAOA needs one good sample rather than a faithful distribution, so check each bitstring against the instance before reporting a value
4.5Scheduling and allocation
Reproduces scheduling and allocation. Positions in, an independent set out. The blockade enforces the constraint, so there is no QUBO and no penalty weight, and the service builds the sweep and decodes the answer.
job = client.run_mis(
vertices=[[0, 0], [6, 0], [12, 0], [0, 6], [6, 6],
[12, 6], [0, 12], [6, 12], [12, 12]],
shots=1000, engine="qpu.quera.aquila")
job["result"]["mis"]["best_set"] # the items to keep
job["result"]["mis"]["valid_fraction"] # share of shots obeying every edge
job["result"]["mis"]["blockade_um"] # the radius your drive produced- Start on the emulator.
analog.pulser.cpuis exact to 14 atoms and costs a hundredth of a cent, so confirm your layout produces the edges you meant first - Your own problem. Lay items out so distance encodes conflict, scaled so the interference range lands at the blockade radius the call returns rather than at a round number
- Two device rules. Atoms must be at least 4 um apart, and any two must share a y coordinate exactly or sit at least 4 um apart in y
4.6Traffic route assignment
Reproduces traffic and smart cities. Ten vehicles is 30 qubits as a QAOA circuit, and exhaustive search returns the exact optimum in 0.18 seconds at that size, which is the comparison the page makes.
client.estimate(your_qaoa_circuit, shots=1000) job = client.run(your_qaoa_circuit, shots=1000, engine="exact.cpu") # Score every returned assignment against the conflict list yourself, # then compare against the exhaustive optimum at the same size.
- Check whether your graph is unit-disk before assuming the neutral-atom route applies. A road conflict graph generally is not, which is what sends this problem to gate hardware as a QUBO instead
4.7Routing and job-shop encoding cost
Reproduces logistics and manufacturing. Both answer a question that needs no quantum run at all: how many qubits the problem costs to encode. The arithmetic follows from the formulation, so it is the same on any machine.
# Vehicle routing, one binary per (vehicle, stop, position): qubits = n_vehicles * n_stops * n_positions # Job-shop, one binary per (job, machine, time slot): qubits = n_jobs * n_machines * n_timeslots # 100 stops -> 10,000 qubits. A 15x15 shop -> 268,425.
- Do this first, every time. An encoding that needs more qubits than any device has is settled before a single job is submitted, and the notebook
encoding-cost.ipynbdoes the arithmetic for both
4.8Cryptographic resource estimate
Reproduces cryptography and digital assets. Small elliptic-curve keys are recovered exactly, then the same circuit is run against a device-noise model to see what a real machine would do to it.
client.estimate(ecdlp_circuit, shots=64)
# Exact, so a failure is the algorithm's and not the device's
client.run(ecdlp_circuit, engine="exact.cpu", shots=64)
# Then ask what a real device would do to it
client.run(ecdlp_circuit, engine="noisy.cpu",
noise="superconducting", shots=4096)- Report the noise floor with any recovery. Every shot is verified classically for free, so a run with no signal still finds the key with probability 1 - (1 - 1/n)^s for a group of order n and s usable shots. A recovery published without that figure cannot be told from luck
- Count gates, not qubits. Width grows about three qubits per bit of key; the two-qubit gate count grows by a factor of about five per bit, and that is what a device has to survive
- The 7-bit key on the GPU tier. The full script is examples/crypto:
python ecdlp_gpu.py. It is fixed end to end: the first cyclic curve of order 128, secret 67, the circuit flattened tocxandu(21 qubits, 163,192 two-qubit gates), 64 shots onexact.gpu. Our run recovered 67 from 24 usable shots in 66 seconds. Name the classical registers anything butxory: OpenQASM 2.0 has one namespace, and those are gate names
4.9QCBM, a generative model
Reproduces QCBM and quantum generative models. A quantum circuit Born machine has no input: the measurement probabilities are the distribution, and training moves them toward a target. Published run: TVD 0.459 to 0.0117 in 806 circuits.
import torch
from qiskit import QuantumCircuit
from qiskit.circuit import ParameterVector
from qsim_sdk.ml import CircuitLayer
torch.manual_seed(0) # fixes every torch draw; shots stay stochastic
w = ParameterVector("w", 6)
qc = QuantumCircuit(2)
qc.ry(w[0], 0); qc.ry(w[1], 1); qc.cx(0, 1)
qc.ry(w[2], 0); qc.ry(w[3], 1); qc.cx(0, 1)
qc.ry(w[4], 0); qc.ry(w[5], 1)
qc.measure_all()
outcomes = ["00", "01", "10", "11"]
layer = CircuitLayer(client, qc, shots=512, engine="exact.cpu", eps=0.25,
init=[0.3, 1.9, 0.7, 2.4, 1.1, 0.5], outcomes=outcomes)
opt = torch.optim.Adam(layer.parameters(), lr=0.25)
want = torch.tensor([0.5, 0.0, 0.0, 0.5])
for _ in range(60):
loss = (layer() - want).abs().sum() / 2 # total variation distance
opt.zero_grad(); loss.backward(); opt.step()- SDK only. This is 806 jobs across 60 steps, one per gradient evaluation, so it cannot run from the console. Needs the
[ml]extra - Your own target. Replace
wantwith any distribution over the same outcomes. Widen the circuit and the outcome list together - Then try it on hardware. Submit the converged circuit once with
engine="qpu.rigetti": one task instead of hundreds. The published figures at 1,000 shots are TVD 0.102 on Rigetti ($0.7250) and 0.104 onqpu.iqm.garnet($1.7500), from the same circuit - And on the GPU tier. The same submission with
engine="exact.gpu"returns TVD 0.041 for $0.0001. That price is the platform floor rather than the GPU rate, which is $3.00 an hour against the CPU tier's $0.69: a two-qubit run ends before either reaches a cent, so the two simulator tiers bill the same here and stop doing so once a circuit runs for real time - Buy a trapped ion by the shot. This target forbids two of its four outcomes, so shots landing on 01 or 10 measure fidelity directly rather than distance. On
qpu.aqt.ibexat 100 shots that was 3 of 100, against 102 of 1,000 on Rigetti and 38 of 1,000 on Garnet, for $2.6500. Every device charges the same $0.30 to accept a job and then $0.0235 a shot on AQT against $0.000425 on Rigetti, so the same 1,000-shot run is $23.80 there. Pick the shot count from the question: a fidelity question needs few shots, a distribution needs many
4.10QGAN
Reproduces QGAN and synthetic data. The generator is a circuit and the discriminator is ordinary PyTorch, so only the generator is submitted. Published run: TVD 0.269 to 0.146 over 40 alternating steps.
w = ParameterVector("w", 4)
qc = QuantumCircuit(2)
qc.ry(w[0], 0); qc.ry(w[1], 1); qc.cx(0, 1)
qc.ry(w[2], 0); qc.ry(w[3], 1)
qc.measure_all()
gen = CircuitLayer(client, qc, shots=512, engine="exact.cpu", eps=0.25,
init=[2.2, 0.4, 1.7, 2.9], outcomes=outcomes)
disc = torch.nn.Sequential(torch.nn.Linear(4, 8), torch.nn.ReLU(),
torch.nn.Linear(8, 1)) # logits, scored with BCEWithLogitsLoss
g_opt = torch.optim.Adam(gen.parameters(), lr=0.20)
d_opt = torch.optim.Adam(disc.parameters(), lr=0.05)
for step in range(40):
fake = gen()
# 1. discriminator: separate fake from target
# 2. generator: move fake toward fooling it
# Track BOTH the discriminator gap and the TVD.- SDK only. One job per generator evaluation, 722 in the published run
- Track two quantities. A shrinking discriminator gap alone proves nothing, because a collapsing discriminator produces the same signal. Report the distance to the target as well
- Seed your own runs to compare them. The published run was not seeded. The discriminator's weights are drawn at random, so
torch.manual_seedbefore the loop is what makes two of your runs start alike; the shots are sampled fresh on the device, so they land close rather than on the same decimal - Then try it on hardware. The hardware rows ran a second training of the same model; its converged circuit was run on each device to compare outputs. At 1,000 shots it gives TVD 0.286 on
qpu.rigetti($0.7250) and 0.240 onqpu.iqm.garnet($1.7500), either side of the 0.270 the noiseless simulator gives at the same shot count.exact.gpureturns 0.278 for $0.0001, which is the platform floor and not the GPU rate - The photonic version swaps
CircuitLayerforPhotonicLayer(client, circuit, input_state)onphotonic.slos.cpu, which is the exact reference a Belenos run is certified against
The same model in linear optics, and on the photonic processor
Three modes, two photons, three trainable beamsplitters, and the gate target carried onto Fock states. Photons enter modes 0 and 2 rather than 0 and 1 because Belenos has single-photon sources on alternating modes, so the same program runs on hardware unchanged. Seeded, so your five starting points are ours.
import perceval as pcvl, torch
from qsim_sdk.ml import PhotonicLayer
OUTCOMES = ["|2,0,0>", "|1,1,0>", "|1,0,1>", "|0,2,0>", "|0,1,1>", "|0,0,2>"]
TARGET = torch.tensor([0.4, 0.1, 0.0, 0.0, 0.1, 0.4])
circuit = (pcvl.Circuit(3)
// (0, pcvl.BS(theta=pcvl.P("t0")))
// (1, pcvl.BS(theta=pcvl.P("t1")))
// (0, pcvl.BS(theta=pcvl.P("t2"))))
for seed in range(5): # the seed IS the experiment
torch.manual_seed(seed)
init = (torch.rand(3) * 2.0).tolist()
layer = PhotonicLayer(client, circuit, [1, 0, 1], shots=1000,
engine="photonic.slos.cpu", eps=0.25,
init=init, outcomes=OUTCOMES)
opt = torch.optim.Adam(layer.parameters(), lr=0.20)
for _ in range(30):
loss = (layer() - TARGET).abs().sum() / 2
opt.zero_grad(); loss.backward(); opt.step()- Seeds 0 to 4 gave P(target) 0.875, 1.000, 0.989, 0.932 and 0.766. Each start draws its initial angles from
torch.manual_seed(0)to(4), so your five starting points are ours. Shots are sampled fresh on every run, so your final values land close to these, not exactly on them. The optimum is exactly reachable here, so the ceiling is 1.0 rather than something you have to discover - Then the device. Submit the converged angles with
engine="qpu.quandela.belenos". Use a large shot count: Belenos charges EUR 0.30 a job whatever the sample count, so 100,000 shots is $0.4584 against $0.3449 for 1,000 - Post-select on coincidences, or the number is about loss. Detection is heralded and most attempts lose a photon: our run recorded 2,686 two-photon events in 100,000 attempts, 2.7%. Keep only the states with two photons and renormalise before scoring. Read the raw counts instead and a perfectly trained circuit scores a TVD near 0.97, which is the transmission of the device rather than anything the model learned
- What we measured. 97.4% of coincidences on the four intended outcomes, so the trained support survives; the two bunched outcomes come back 0.538 against 0.130 where the simulator has them level, and a second identical run gave 0.516 and 0.140. The single-photon counts lean the same way, which points at the two paths having different transmission
4.11Quantum image generation
Reproduces quantum AI image generation. Two opposite methods: one trains patch generators, the other spends one qubit per pixel and trains nothing. The untrained one measures the device, since its error per pixel is readout error.
# Write pixel (x, y) as a rotation on qubit y*width + x.
qc = QuantumCircuit(n)
for y in range(height):
for x in range(width):
qc.ry(pixels[y][x] * math.pi, y * width + x)
qc.measure_all()
job = client.run(qc, shots=300, engine="qpu.rigetti",
params={"render": "image", "width": width})
from qsim_sdk import images
grid = images.from_counts(job["result"]["counts"], width=width)
open("out.pgm", "wb").write(images.to_pgm(grid))- The console renders one job. Tick the image box and the job page draws it and offers a PNG, which covers any image that fits in a single circuit: about 100 pixels on a 108-qubit device
- SDK only. A larger image is several jobs. A 32×32 is eleven circuits, and joining them is
images.stitch(counts_list, width=32), which takes the counts in submission order - The trained method is four
CircuitLayersub-generators, one per patch, entangled within a patch and independent across them, which is what keeps each circuit small enough to run
4.12Quantum reinforcement learning
Reproduces quantum RL and control policies. The state enters as input angles and the action comes out as a measurement, so the circuit replaces the policy network and nothing else in the loop changes.
x = ParameterVector("x", 2) # the observed state
w = ParameterVector("w", 4) # what trains
qc = QuantumCircuit(2, 2)
qc.ry(x[0], 0); qc.ry(x[1], 1)
qc.ry(w[0], 0); qc.ry(w[1], 1); qc.cx(0, 1)
qc.ry(w[2], 0); qc.ry(w[3], 1)
qc.measure([0, 1], [0, 1])
# inputs= names the data parameters, so forward(x) holds them fixed
# and trains only the weights.
layer = CircuitLayer(client, qc, inputs=[p.name for p in x],
shots=512, engine="exact.cpu", eps=0.25,
init=[0.5, 1.2, 0.9, 0.3], outcomes=outcomes)
opt = torch.optim.Adam(layer.parameters(), lr=0.20)
p_one = layer(batch_of_states)[:, 2] + layer(batch_of_states)[:, 3]
loss = (-(r - r.mean()) * torch.log(p_one.clamp(1e-6, 1))).mean()- SDK only. One job per decision, and
inputs=has no console equivalent - Compare against a random policy on the same task. The published run returns 0.611 against 0.533 random over 60 decisions, which is inside the spread of that sample size: separating them needs hundreds of episodes across several seeds
- Seed the episodes. States and actions are both drawn, so
torch.manual_seed(0)before the loop is what makes two of your runs comparable to each other. Report the seed with the number - Then try it on hardware. Submit the converged weights at the two probe states. Every engine puts them on opposite sides of 0.5 and in the same order: 0.158 and 0.624 on
exact.gpu, 0.199 and 0.650 onqpu.rigetti, 0.177 and 0.576 onqpu.iqm.garnet. Each probe is its own job, so the price is one run's: $0.0001, $0.7250 and $1.7500 respectively. The decision survives the move between devices and the margin is what changes
4.13NNQS ground states
Reproduces NNQS for materials and chemistry. A neural network represents the wavefunction and variational Monte Carlo trains it, which reaches states where a tensor network gives up. Takes a Hamiltonian, not a circuit.
job = client.solve(
hamiltonian=terms, qubits=n,
engine="neural.cpu", # or "neural.tpu"
params={"n_iter": 300, "n_samples": 4096,
"n_restarts": 3, "alpha": 2,
"ansatz": "rbm"}) # or "rbm_symm"
job["result"]["ground_state"]["ceiling"] # unconditional
job["result"]["ground_state"]["floor"] # conditional, a diagnostic
# Many Hamiltonians on one machine, paid for provisioning once:
client.run_solve_batch(problems=[{"qubits": n, "hamiltonian": t} for t in sweep],
engine="neural.tpu")- Read the ceiling as the bound and the floor as a diagnostic. The ceiling is unconditional; the floor holds only while the trial state sits nearer the ground state than the first excited state, which a run cannot verify about itself
rbm_symmis the largest accuracy lever and is refused unless your Hamiltonian is verifiably translation invariant. On a six-spin Ising ring it reached 0.0018 above exact against the plain network's 0.0067, on a sixth of the parameters- A sweep belongs in one job. On the TPU tier most of a short job is the machine being created and deleted, so sending a phase diagram point by point pays that each time. The console runs a sweep of one coefficient; an arbitrary list of Hamiltonians is
run_solve_batch
4.14QML classifiers and kernels
Reproduces QML and classification. Two methods on one dataset: a variational classifier that trains rotation angles, and a kernel that trains nothing in the circuit at all and hands a matrix to a classical SVM. Both read the same seeded data, built first.
import numpy as np
from qiskit import QuantumCircuit
def moons(n, noise, rng): # two interleaving half-circles
k = n // 2
t_out, t_in = np.pi * rng.random(k), np.pi * rng.random(n - k)
x = np.vstack([np.c_[np.cos(t_out), np.sin(t_out)],
np.c_[1 - np.cos(t_in), 0.5 - np.sin(t_in)]])
y = np.r_[np.zeros(k, dtype=int), np.ones(n - k, dtype=int)]
x = x + rng.normal(0, noise, x.shape)
order = rng.permutation(n)
return x[order], y[order]
seed = 1 # published: seeds 1, 2 and 3
rng = np.random.default_rng(seed) # the order of these draws is part of the data
xtr, ytr = moons(22, 0.12, rng)
xte, yte = moons(40, 0.12, rng)
lo, hi = xtr.min(0), xtr.max(0) # into [0, pi] on the training range
Xtr, Xte = (xtr - lo) / (hi - lo) * np.pi, (xte - lo) / (hi - lo) * np.pi
theta = rng.normal(0, 0.5, 4) # starting angles, drawn after the data
def circuit(x, th): # angles to six decimals, as published
qc = QuantumCircuit(2)
qc.ry(round(x[0], 6), 0); qc.ry(round(x[1], 6), 1)
qc.ry(round(th[0], 6), 0); qc.ry(round(th[1], 6), 1)
qc.cx(0, 1)
qc.ry(round(th[2], 6), 0); qc.ry(round(th[3], 6), 1)
return qc
def z0(circuits): # <Z> on qubit 0, computed, not sampled
job = client.run_batch(circuits, engine="exact.cpu", observable=[[1.0, "IZ"]])
return np.array([r["result"]["expectation"] for r in job["results"]])
eps, lr = 0.2, 0.35
for step in range(20): # plain gradient descent on squared error
shifted = [theta + s * eps * np.eye(4)[p] for p in range(4) for s in (1, -1)]
v = z0([circuit(x, t) for x in Xtr for t in [theta] + shifted]).reshape(22, 9)
dL = 2 * ((v[:, 0] + 1) / 2 - ytr) / 22
grad = [(dL * (v[:, 1 + 2 * p] - v[:, 2 + 2 * p]) / (2 * eps) / 2).sum() for p in range(4)]
theta = theta - lr * np.array(grad)
acc = ((z0([circuit(x, theta) for x in Xte]) > 0).astype(int) == yte).mean()
acc # seeds 1 to 3: 0.925, 0.850, 0.700, on every run- What the seed fixes.
default_rng(seed)draws the 22 training points, then the 40 test points, then the four starting angles, so the data, the split and the start are identical on every machine - Exact, so it reproduces to the digit. Sent with an observable,
exact.cpucomputes each<Z>from the statevector instead of sampling it, and nothing random is left. 198 circuits a step fit one batch; 4,000 circuits in all, $0.4000 a seed - The published 1,024-shot run sends
shots=1024with no observable and reads<Z>from the counts. It scored 0.925, 0.825 and 0.700. The service samples shots without a seed, so that configuration re-runs to within shot noise, one standard deviation of 0.031 on each<Z>, not to the digit - The full script is variational_classifier.py:
python variational_classifier.py 1for seed 1, exact;--shots 1024for the sampled configuration;--localfor the exact run computed on your machine, free. It runs logistic regression, k-NN and an MLP on the same split, written out in numpy
from sklearn.svm import SVC
# Xtr, Xte, ytr, yte from the classifier above: both methods see identical points.
def feature_map(a): # ZZ feature map, depth 2
qc = QuantumCircuit(2)
for _ in range(2):
qc.h(0); qc.h(1)
qc.rz(2 * a[0], 0); qc.rz(2 * a[1], 1)
qc.cx(0, 1); qc.rz(2 * (np.pi - a[0]) * (np.pi - a[1]), 1); qc.cx(0, 1)
return qc
def overlap(a, b): # U(a) then U(b) undone: P(00) is K(a, b)
qc = feature_map(a).compose(feature_map(b).inverse())
qc.measure_all()
return qc
# Every pair of points is its own overlap circuit: 231 train pairs
# and 880 test pairs, so 1,111 circuits. A batch takes at most 200,
# so this is 6 requests. estimate_batch splits for you; run_batch does not.
train = [(i, j) for i in range(22) for j in range(i + 1, 22)]
test = [(i, j) for i in range(40) for j in range(22)]
circuits = ([overlap(Xtr[i], Xtr[j]) for i, j in train] +
[overlap(Xte[i], Xtr[j]) for i, j in test])
est = client.estimate_batch(circuits, shots=2048, engine="exact.cpu")
est["total_usd"], est["per_point_usd"]
N = qsim_sdk.MAX_BATCH_CIRCUITS # 200
p00 = []
for i in range(0, len(circuits), N):
job = client.run_batch(circuits[i:i + N], shots=2048, engine="exact.cpu")
p00 += [c.get("00", 0) / sum(c.values()) for c in qsim_sdk.counts(job)]
Ktr = np.eye(22)
for (i, j), v in zip(train, p00):
Ktr[i, j] = Ktr[j, i] = v
Kte = np.zeros((40, 22))
for (i, j), v in zip(test, p00[len(train):]):
Kte[i, j] = v
SVC(kernel="precomputed").fit(Ktr, ytr).score(Kte, yte) # seeds 1 to 3: 0.675, 0.750, 0.750
SVC(kernel="rbf").fit(xtr, ytr).score(xte, yte) # RBF baseline: 0.950, 0.875, 0.950- From the console too. Tick Run several circuits as a batch and paste the circuits one after another: 1,111 go in as 6 batch jobs. Building the pairs and the kernel matrix still takes code
- Split the submission yourself.
estimate_batchchunks internally and prices any number of circuits.run_batchdoes not, and a list longer thanqsim_sdk.MAX_BATCH_CIRCUITSis refused with a 422 - The full script is quantum_kernel.py:
python quantum_kernel.py 1for seed 1 at 2,048 shots, 1,111 circuits a seed, $0.1111;--exactsends the observable|00><00|instead, so the kernel is computed rather than sampled and re-runs to the digit. kernel_control.py computes the same exact kernel on your machine, with no submissions, so it is free - Sampled and exact agree here. On all three seeds the 2,048-shot kernel scored the same as the exact one; the largest error in a single entry was 0.037
- On a GPU or a TPU.
python quantum_kernel.py 1 --engine exact.gpu, or--engine exact.tpu. Seed 1 returns 0.675 on both, the same asexact.cpu. The GPU costs $0.1111, the same as the CPU; the TPU costs $1.0336, because each 200-circuit batch runs on its own machine.--exactneedsexact.cpu - Tune both sides or neither. On seed 1 the same kernel scores from 0.675 to 0.925 across eight input scalings and depths, a constant that carries no information about the problem, so compare against a tuned classical baseline rather than a default one.
kernel_control.pyprints the sweep
4.15Quantum reservoir computing
Reproduces QRC, quantum reservoir computing. A fixed entangling circuit with no trainable parameters: nothing in the circuit learns, and the entire model is the classical readout on top of it.
import numpy as np
from qiskit import QuantumCircuit
from qiskit.circuit import ParameterVector
from sklearn.datasets import make_moons
from sklearn.linear_model import LogisticRegression
X, y = make_moons(n_samples=200, noise=0.2, random_state=0)
X = X * (np.pi / 2) # into an angle range
x = ParameterVector("x", 2) # the reservoir: nothing here trains
base = QuantumCircuit(4)
base.h(range(4))
base.rz(2.0 * x[0], 0); base.rz(2.0 * x[1], 1)
base.cx(0, 1); base.cx(1, 2); base.cx(2, 3)
base.rz(2.0 * x[0] * x[1], 3)
base.cx(2, 3); base.cx(1, 2); base.cx(0, 1)
base.ry(0.7, 0); base.ry(1.3, 1); base.ry(0.9, 2); base.ry(1.1, 3)
base.measure_all()
circuits = [base.assign_parameters({x[0]: a, x[1]: b}) for a, b in X]
# 200 circuits, one request: exactly the batch cap.
job = client.run_batch(circuits, shots=512, engine="exact.cpu")
all_counts = qsim_sdk.counts(job)
def z(bits, i): # qiskit is little-endian: bits[-1] is qubit 0
return 1 - 2 * int(bits[::-1][i])
def features(c): # <Z_i> and <Z_i Z_j>: every shot feeds every feature
n = sum(c.values())
p = {b: v / n for b, v in c.items()}
singles = [sum(q * z(b, i) for b, q in p.items()) for i in range(4)]
pairs = [sum(q * z(b, i) * z(b, j) for b, q in p.items())
for i in range(4) for j in range(i + 1, 4)]
return singles + pairs
F = [features(c) for c in all_counts]
clf = LogisticRegression(max_iter=2000).fit(F[:140], y[:140])
clf.score(F[140:], y[140:]) # published: 0.817- SDK only. One job per sample, batched
- The full script is
reservoir_localin ai_proofs.py:python ai_proofs.py reservoir_local, 200 circuits. The bitstring-distribution readout that scored 0.783 isreservoir - Read local observables, not the bitstring distribution. Identical counts gave 0.783 read one way and 0.817 the other. The outcome space is 2^n against a fixed shot budget, so a distribution readout stops working as width grows; local observables give n^2 features and each uses every shot
- Publish the classical baseline. The raw two-dimensional coordinates reach 0.850 on the same split
4.16NNQS tomography
Reproduces NNQS tomography. Measurement records in, a reconstructed state out. The records come from hardware: run a circuit on a QPU, measure it in several bases, and reconstruct on the neural tier.
job = client.run_tomography(
n_spins=2,
bases=["XX", "XY", "XZ", "YX", "YY", "YZ",
"ZX", "ZY", "ZZ", ...], # one per shot
outcomes=[[+1, -1], [-1, -1], ...], # one per shot, +1/-1
engine="neural.cpu")
job["result"]["agreement"]["agreed"] # observables reproduced
job["result"]["agreement"]["observables"] # each with its interval
job["result"]["certifiable"] # always False, and says why- Measure a complete set of bases. An under-determined set is refused before it runs, because a fit converges perfectly well onto the wrong state: a Bell state came back at fidelity 0.51 from five bases and 0.98 from all nine, at identical settings
- No fidelity bound exists for the reconstruction. A maximum-likelihood fit has no equivalent of the variational principle. What is reported is agreement with records held back from the fit, each with a Hoeffding interval, Bonferroni-corrected so they hold together
- CPU tier only. It is refused on
neural.tpu, where provisioning would dominate a fit that takes seconds


Qiskit
PennyLane
PyTorch