Skip to content
Felix Velasquez

01 / Simulation / Compilation

Adaptive Quantum Circuit Simulator

Four Rust-backed quantum-circuit simulators and a machine-learning selector that chooses a backend from static circuit features. The project explores how circuit structure, execution strategy and memory limits shape simulation performance.

  1. 01 / Circuit

    Import OpenQASM 2 and extract static circuit features.

  2. 02 / Rank & filter

    Rank four backends; apply support, allocation and learned memory checks.

  3. 03 / Select

    Return a backend, scores and exclusions—or abstain.

  4. 04 / Simulate

    Execute separately through the Python API backed by Rust.

Fig. 01 — Circuit-to-backend workflow. Selection uses static features without running competing simulators. Execution is a separate API call in the current prototype.

Abstract

I built four classical quantum-circuit simulators and a machine-learning pipeline that selects among them from circuit structure. The project combines a Rust simulation core, a Python interface and a gradient-boosted classifier to explore when different state representations offer a practical advantage.

The selector chooses between dense statevector, matrix product state (MPS), P-block and stabilizer simulation. On 370 resolved development comparisons, it selected the measured fastest backend in 336 cases (90.8%). The result is a research prototype with reproducible collection and training tools, a bundled model, and documented limitations.

Problem

No single representation is efficient for every quantum circuit. A dense statevector stores 2ⁿ complex amplitudes, while other methods exploit limited entanglement, independent qubit groups or Clifford structure. Qubit count alone does not reveal which representation will be fastest—or whether it will fit in memory.

Choosing a backend by running every simulator first defeats the purpose for expensive workloads. I wanted to make that choice from features available before execution, while accounting for unsupported operations, memory limits and the cost of a poor prediction.

Approach

I implemented four complementary backends behind a shared shot-simulation interface. Statevector handles general circuits with dense amplitudes. MPS represents the state as a tensor chain, with cost governed by bond growth. P-block maintains independent dense groups, merging them for interacting operations and splitting off measured or reset qubits. Stabilizer uses a packed binary tableau for supported Clifford circuits.

The execution layer includes gate compilation, terminal-measurement sampling and reuse of deterministic prefixes for dynamic circuits. Rust handles numerical execution, with PyO3 bindings exposing the simulators to Python. OpenQASM 2 import preserves supported measurements, resets and classical conditions.

The winner classifier uses 28 static circuit features plus task and shot context. A separate memory classifier adds 29 structural features; deterministic checks exclude unsupported stabilizer circuits and infeasible statevector allocations. The remaining candidates retain the winner model’s ranking. Prediction returns the selected backend and exclusion details, with explicit abstention when no candidate remains.

For training, I used circuits from QASMBench and the PennyLane-hosted MQT Bench dataset. The collector runs circuit/backend pairs in isolated processes, records repeated timings and resource failures, and supports resuming interrupted runs. Training deduplicates normalized circuits and keeps related algorithm families together during validation.

Personal contribution

My work spans the simulator implementations, performance optimizations, shared Python interface, circuit import, benchmark collection and ML selection pipeline. I connected the numerical backends to a common workflow so that their performance could be compared on the same circuits and execution tasks.

I also built validation around gate correctness, measurement correlations, classical feedback, compact result queries and memory-limited execution. The repository includes Qiskit comparisons, focused performance scripts, a model card and references to the simulation methods and upstream circuit datasets.

Evaluation and outcomes

The model card reports five-fold development validation with algorithm families held out together and tuning confined to training groups. The dataset contains 395 observed circuits, of which 370 have resolved winner labels. Measurements used macOS arm64, four native threads, 1,000 shots, five timed repetitions and a 1 GiB backend working-memory setting. This setting is not a total process-memory cap.

Among resolved comparisons, 336 of 370 selections matched the fastest backend, one selection failed, and six successful choices were at least 10× slower than the winner. Across all 395 circuits, there were 24 observed failed selections and one selection without an observation, at 100% selection coverage. The worst successful choice was about 526× slower than the fastest backend—an important limitation that average accuracy alone would hide.

The simulator tests also exercise structure-dependent scaling: MPS sampling on a 100-qubit GHZ state and P-block storage of 100 qubits as fifty independent Bell pairs under a 1 MiB allowance. These cases illustrate the value of compact representations; they do not imply efficient simulation of arbitrary 100-qubit circuits.

These figures describe development validation of backend selection, not a general speedup over Qiskit Aer or an independent final test. Fresh circuit families and measurements on the target hardware are needed before making stronger generalization claims.

Design decisions and lessons

I framed selection as direct classification of the fastest observed backend, rather than runtime prediction. The model returns a ranking with uncalibrated scores. Keeping that ranking separate from eligibility and memory checks makes it possible to inspect why a candidate was excluded.

Memory prediction introduced its own tradeoff: the filter incorrectly excluded P-block on 10 of 264 successful P-block runs, all QFT circuits where it was fastest. That exceeded the development target for false exclusions. The filter is therefore part of the experiment, not a guarantee that execution will succeed.

Compact state representations only help when execution and result queries preserve them. Terminal sampling avoids repeated evolution, while amplitude and expectation queries avoid unnecessary dense exports. Comparisons also need consistent output tasks and approximation settings, especially for MPS truncation.

The main lesson was to evaluate the cost of mistakes alongside winner accuracy: failures, coverage and extreme slowdowns all matter. The current prototype exposes selection and simulation separately; a combined circuit-to-results application remains a next step.

Back to selected work