About this project

PRISM-Q is a quantum circuit simulator written in Rust, with Python bindings. It automatically selects a backend from the circuit's structure, runs a multi-pass gate fusion pipeline, and uses AVX2, FMA and BMI2 kernels in the inner loop (NEON on ARM64). CUDA support is optional and covers the statevector, stabilizer (experimental) and density-matrix backends. Input is OpenQASM 3.0 with backward-compatible 2.0 syntax, and circuits can be exported back to OpenQASM 3.0. Installation is available via cargo (prism-q) or pip (prism-q). The default feature set includes Rayon parallelism and faer SVD; a no-default-features build is single-threaded with minimal dependencies. The gpu feature requires CUDA Toolkit 12.x or newer. Backends include Statevector (general circuits, O(2^n)), Stabilizer (Clifford only, O(n^2)), Factored Stabilizer (Clifford with independent blocks), Sparse (few live amplitudes), MPS (low entanglement or 1D, O(n*chi^2)), Product State (no entanglement), Tensor Network (low treewidth), Factored (partial entanglement), Density Matrix (exact noisy evolution, O(4^n), explicit dispatch only), and Distributed Statevector (beyond single-host memory, MPI ranks, exact results). BackendKind::Auto is the default; the statevector memory budget is half the machine's physical memory, overridable via PRISM_MAX_SV_QUBITS. The OpenQASM parser accepts the stdgates.inc set, common controlled and multi-controlled variants, Qiskit exporter gates, IonQ and Google/Cirq native gate names, decomposed multi-instruction gates, IBM legacy u1/u2/u3, and user-defined gate declarations. The inv @, ctrl @ and pow(k) @ modifiers chain on direct gates. qasm_export::to_qasm3 writes a Circuit back out as OpenQASM 3.0. Features include shots and sample counts, expectation values and marginals, parametric circuits with rebinding, and gradients via the adjoint method or the parameter-shift rule. The GPU feature compiles PTX at runtime through NVRTC for the device's compute capability; a size crossover keeps small sub-circuits on the CPU. Benchmarks are provided for circuit macrobenchmarks, gate microbenchmarks and GPU dispatch. The roadmap mentions mid-circuit branching, multi-GPU and distributed GPU execution, ROCm ports, and noisy shots on the distributed backend.