Skip to content

Replay-buffer engineering for noise-aware quantum circuit optimization

Conference: NeurIPS2026 (task-list assignment; source is arXiv v2)
arXiv: 2604.21863
Area: Reinforcement Learning
Keywords: experience replay, reliability-aware sampling, quantum architecture search, amortized evaluation, noiseless-to-noisy transfer

TL;DR

The paper shifts quantum circuit optimization from changing the agent to reusing experience more effectively: annealed reliability-aware replay improves sample efficiency, blockwise evaluation reduces quantum–classical calls, and buffer-only transfer accelerates noise adaptation, although accuracy, gate-count, and convergence advantages differ across tasks.

Background & Motivation

Quantum compilation approximates a target unitary with a finite gate library, while quantum architecture search seeks a parameterized circuit preparing a low-energy state. Reinforcement learning (RL) can model gate placement as sequential decisions, but an action is considerably more expensive than in an ordinary discrete environment: architecture search typically runs COBYLA to optimize continuous angles and then evaluates the Hamiltonian expectation with a quantum simulator. Larger circuits make each full evaluation expensive, and noise adds pressure to reduce error-prone two-qubit gates.

Existing replay methods have opposing weaknesses. Prioritized experience replay (PER) emphasizes large temporal-difference (TD) errors without distinguishing useful corrections from errors caused by unreliable downstream targets. Reliability-adjusted replay (ReaPER) discounts priorities using subsequent TD errors in the same trajectory, but may trust unstable reliability estimates too early, before the value network has matured. Moreover, promising circuits explored in noiseless simulation are often discarded when training moves to a noisy environment, forcing expensive search to restart.

The paper retains DQN/DDQN and existing circuit representations while changing how experience is sampled, how many architecture edits receive one evaluation, and how experience is reused across noise conditions. These are related but separately tested contributions; notably, transfer experiments use uniform replay, so their gains cannot all be attributed to annealed sampling. Core Idea: move replay gradually from TD-error emphasis toward trustworthy targets, reduce evaluation frequency, and initialize noisy-task replay with noiseless trajectories rather than inheriting source-policy weights.

Method

Overall Architecture

The input is a target unitary or a molecular/spin-system Hamiltonian, and the output is a circuit meeting a fidelity or energy-error target. The agent observes the current circuit, selects a gate action, and stores the resulting evaluated transition; training samples from this buffer, whereas executing the final circuit requires no replay buffer.

The three contributions are not a single model that every experiment must use end to end. “Amortized architecture evaluation” changes the expensive QAS environment; “annealed reliability-aware replay” changes training sampling; “buffer-only transfer” starts a corresponding noisy task after source training. The diagram orders these around collecting source experience and subsequently transferring it. Dashed edges represent training supervision or buffer reuse, not circuit inference data flow.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Target and current circuit"] --> B["Agent gate action"]
    B --> C["Amortized architecture<br/>evaluation"]
    C --> D["Transition buffer"]
    D --> E["Annealed reliability-aware<br/>replay"]
    E -. "TD training supervision" .-> B
    D -. "Copy after source training" .-> F["Buffer-only transfer"]
    F -. "Initialize target buffer" .-> G["Noisy-target training"]
    G --> H["Optimized circuit"]

Compilation observations contain the real and imaginary parts of the residual unitary between the current and target unitaries. The single-qubit small-rotation actions use positive/negative \(\pi/128\) rotations around X/Y/Z, with a separate HRC discrete-basis experiment. Two-qubit experiments include local Z rotations and XX/YY rotations. QAS records gate layers and connections in a tensor, together with a cost summary; its usual gate library contains RX, RY, RZ, and CNOT, with illegal actions masked.

Appendix Q further explains that an implemented encoding action can include both a two-qubit gate and a single-qubit rotation, so environment steps should not be equated directly with total gate count. The Heisenberg experiment instead uses gates such as RXX/RYY/RZZ and integer labels for multiple connection types. These are task-representation differences, not innovations of annealed sampling itself.

Key Designs

1. Amortized architecture evaluation: one expensive feedback call covers several circuit edits

Standard CRLQAS optimizes angles and evaluates energy after each gate placement. OptCRLQAS accumulates \(m\) architecture edits before a full evaluation, forcing a final evaluation when an episode ends early. For an episode of length \(T\), expensive calls decrease from approximately \(T\) to \(\lceil T/m\rceil\). This saves environment evaluations; it does not unconditionally reduce network updates to one tenth.

Between evaluations, the circuit structure keeps changing while the state's cost summary retains its previous evaluated value. The authors explicitly describe zero improvement reward during accumulation steps, with nontrivial feedback at evaluation steps. Termination must use the actual final energy to prevent the agent exploiting stale cached values to evade failure penalties. This also changes credit assignment: several gates jointly alter the architecture, rather than rewarding individual edits whose effects continuous-angle optimization can compensate away.

However, Appendix J.1 Equation (53) uses the difference between costs at consecutive evaluations. Its numerator is not necessarily zero merely because both costs remain fixed within a block, conflicting with the subsequent explanation of zero intermediate reward. This note follows the explicit prose description without treating that equation as an unambiguous executable specification. The main text defaults to \(m=10\), but Appendix P and Table 19 use \(m=15\) for 12 qubits; the interval must be checked per experiment.

2. Annealed reliability-aware replay: exploit large errors first, then increasingly trust downstream targets

TD error measures the gap between the current Q estimate and its bootstrapped target, but a large error need not indicate a correct target. ReaPER defines reliability from errors along an episode: the larger the fraction of absolute TD error located after a transition, the less trustworthy its target is. For an episode with \(n\) transitions, the idealized reliability and sampling rule are:

\[ \mathcal{R}_t=1-\frac{\sum_{i=t+1}^{n}|\delta_i|}{\sum_{i=1}^{n}|\delta_i|},\qquad \Psi_t^{(+,\tau)}=\mathcal{R}_t^{\omega_\tau}|\delta_t|^\alpha,\qquad \mu_t^{(+,\tau)}=\frac{\Psi_t^{(+,\tau)}}{\sum_j\Psi_j^{(+,\tau)}}. \]

Here \(\alpha\) controls TD prioritization and \(\omega_\tau\) controls reliability weighting. Setting \(\omega=0\) recovers PER; \(\omega=1\) gives fully reliability-adjusted replay. Fixed ReaPER can also use intermediate exponents rather than only 1. The proposed exponent increases linearly, from 0.1 to 0.7 in the main configuration, then stays constant instead of eventually reaching \(\omega=1\).

\[ \omega_\tau=\omega_{\min}+(\omega_{\max}-\omega_{\min})\min\!\left(\frac{\tau}{T_{\mathrm{ann}}},1\right). \]

An initially near-random Q network makes reliability a noisy proxy, so early replay preserves PER-style correction. Later, more mature estimates allow stronger suppression of high-error transitions with unreliable downstream targets. The main text and Table 20 specify \(T_{\mathrm{ann}}=5\times10^5\) steps for compilation and \(5\times10^4\) for LunarLander, with a rule of half the total training budget. These are steps, not compilation training episodes. Appendix N.2 instead identifies 0.2, 0.7, and 20000 as the best sweep configuration, then refers to 0.1 and 0.7 as the selected configuration; these should not silently be merged into one setting.

The implementation does not recompute exact episode suffix sums at every update. Appendix D.2 inserts new samples into a SumTree at maximum tree priority and initializes an error proxy with \(\max(|r_t|,0.1)\). Reliability is computed at episode completion; subsequent updates approximate the suffix sum by the episode error total minus the sampled transition's error. This approximation loses the strict temporal-suffix interpretation and must be distinguished from the ideal definition above. A fallback handles tiny total error, and priorities include a \(10^{-6}\) numerical-stability term.

The theoretical scope matters equally. Theorem D.1 studies tabular Q-learning on a fixed dataset and assumes bounded quantities, positive minimum reliability, reliability–variance alignment, and a monotone likelihood ratio (MLR) condition. The authors do not verify the latter two conditions on actual buffers. The analysis uses full importance correction, whereas the implementation anneals \(\beta\) toward 1. It is therefore a conditional argument for lowering the noise floor, not an unconditional convergence guarantee for deep networks, arbitrary environments, or arbitrary schedules.

3. Buffer-only transfer: retain source experience while relearning values under noise

After noiseless training, the entire source buffer initializes the corresponding noisy task's buffer. The copied objects are states, actions, source rewards, next states, and terminal flags, without filtering, reward relabeling, or Q-network weight transfer. Matching source/target state and action interfaces makes the data ingestible; it does not make source rewards equivalent to target rewards, and old transitions retain source-domain bias.

The new target network receives better initial coverage and subsequently collects real noisy-target transitions, progressively replacing old experience. The paper calls this online self-correction; it does not mean each stored source reward is automatically recomputed under noise. Appendix E also requires bounded task shift and source trajectories that remain informative in the target. Substantial changes to topology, gate vocabulary, or noise mechanisms cannot be handled merely by matching data shapes.

The main transfer experiments use OptCRLQAS with uniform replay, sourcing buffers from a fixed 12 GPU-hour training budget. To exploit the warm start, target exploration starts at 0.55 instead of 1.0, and the curriculum is tightened. Thus “weight-free transfer” does not mean changing only a buffer variable while leaving all hyperparameters unchanged, nor does it eliminate source-training cost. Appendix M separately studies the interaction with exploration and does not support reducing exploration under every noise condition.

A Worked Example

Consider an illustrative QAS episode with the main-text default \(m=10\). The agent makes ten consecutive edits, updating structural observations at every step, but reoptimizes angles and evaluates energy only on the tenth. If a step limit terminates the episode on step seven, step seven must also be evaluated. These counts illustrate the rule rather than introducing experimental results.

After these transitions enter the source buffer, annealed sampling initially emphasizes TD error and later increases reliability discounts. Once source training ends, the buffer can initialize a noisy environment for the same molecule while the network is initialized anew. Sampling a source transition still retrieves its source reward; only newly collected target transitions contain noisy evaluation feedback. The mechanism reuses promising architectures without treating noiseless Q estimates as the noisy task's answer.

Loss & Training

Appendix D.2 selects the next action with the online network and evaluates it with the target network, giving a DDQN-style target. Importance weights depend on sampling probability and are normalized by the maximum batch weight. Notably, SmoothL1 is not implemented as the usual per-sample loss multiplied by its weight; both predictions and targets are multiplied before computing the loss:

\[ Y_i=r_i+\gamma_{\mathrm{layer}}(1-d_i)Q_{\bar\theta}\!\left(s'_i,\arg\max_a Q_\theta(s'_i,a)\right),\qquad \mathcal{L}=\operatorname{SmoothL1}\!\left(W\odot Q_\theta(s)[a],W\odot Y\right). \]

This changes the effective Huber/SmoothL1 weighting, so reproductions should not substitute the conventional PER loss without qualification. QAS inherits CRLQAS rewards: 5 for success, -5 for failure at the step limit, and otherwise a normalized improvement relative to the target optimum, clipped below at -1. OptCRLQAS additionally delays evaluated feedback to block boundaries.

Table 20 gives compilation a two-layer network with 128 units per layer, learning rate \(3\times10^{-4}\), batch size 200, and replay capacity \(5\times10^5\). QAS uses three/four layers with 1000 units each, batch size 1000, capacity 20000, and 5/6-step returns; the two tasks should not share one hyperparameter table. LunarLander separately uses 5000 episodes and 6 seeds, not additional repetitions of the quantum experiments.

Key Experimental Results

Main Results

The following data come from main-text Table 1 for two-qubit \(ZZ(\pi)\) approximation at threshold 0.9914. They compare the best results obtained at the listed budgets, not final accuracy under equal episode counts. The 4-fold and 32-fold factors describe episode-budget ratios, not measured wall-clock speedups.

Method Training episodes Minimum gates Best fidelity
ReaPER+ (annealed method) \(2.5\times10^4\) 123 0.9920
Fixed ReaPER \(10^5\) 126 0.9931
PER \(10^5\) 127 0.9918
HER \(10^5\) NA \(<0.9914\)
PPO (Moro et al.) \(8\times10^5\) 122 0.9914

Annealing reaches comparable fidelity with fewer episodes, but fixed ReaPER has higher best fidelity and PPO has a slightly lower minimum gate count. The table's broad “outperforms all baselines” caption should not be interpreted as winning every column.

The embedded data in main-text Figure 4 for 12-qubit H2O report multi-seed means ± standard deviations, without stating the exact seed count locally. These are not chemically accurate results: every energy error exceeds \(1.6\times10^{-3}\) Ha.

Replay method Energy error (Ha) Total gates CNOT Steps
Fixed ReaPER \((2.00\pm0.26)\times10^{-2}\) \(221\pm57\) \(125\pm29\) \((2.5\pm2.3)\times10^4\)
ReaPER++ (annealed method) \((2.13\pm0.21)\times10^{-2}\) \(128\pm29\) \(75\pm35\) \((3.7\pm2.0)\times10^4\)
PER \((2.43\pm0.12)\times10^{-2}\) \(165\pm45\) \(95\pm28\) \((5.7\pm3.2)\times10^4\)
Uniform replay \((2.45\pm0.20)\times10^{-2}\) \(151\pm35\) \(91\pm41\) \((3.2\pm5.2)\times10^4\)

Fixed ReaPER has the lowest mean error and fewest steps; annealing stands out for circuit compactness. The main text reports a 67.5% reduction in average per-episode wall-clock time for OptCRLQAS versus CRLQAS. The theoretical call reduction at \(m=10\) should not be described as a tenfold end-to-end speedup.

Ablation Study

Table 3's 6-qubit BEH2 amplitude-damping experiment (\(p=0.001\)) includes both same-exploration transfer and reduced-exploration transfer, separating the buffer contribution from a hyperparameter change. The original table supplies neither seed means nor error bars, so it does not carry the same statistical strength as multi-seed results.

Buffer initialization Initial exploration \(\epsilon\) Energy error (Ha) CNOT ROT
No transfer 1.00 \(8.45\times10^{-5}\) 29 16
Transfer 1.00 \(5.81\times10^{-5}\) 26 15
Transfer 0.55 \(5.77\times10^{-5}\) 27 10

Reducing exploration gives slightly lower error and fewer ROT gates, but one more CNOT than same-exploration transfer. The source claims an 11.3% CNOT reduction for the reduced-exploration setting, inconsistent with 29→27 in the table; the table implies approximately 6.9%. This note does not alter the authors' table.

Appendix J.3's fixed-1000-episode sweep also shows that amortization does not always improve accuracy. At 10 qubits, CRLQAS reaches minimum error \(9.0\times10^{-4}\) Ha; \(m=7\) gives \(9.2\times10^{-4}\), and \(m=10\) gives \(1.0\times10^{-3}\), with classical-optimization times approximately 490, 110, and 75 seconds. Runtime falls reliably, while accuracy and gate counts require separate tradeoffs.

Key Findings

  • Single-qubit small-rotation annealed success rates are 89.30%, 85.30%, and 81.40% at the three tolerances, with the main text stating 40 seeds. Appendix G explicitly uses 5 training seeds for gate-count confidence intervals, which overlap substantially. These cannot be combined into a claim that every metric is significantly established with 40 seeds.
  • Table 2 uses 5 seeds for the proposed method, but most non-RL baselines are not rerun. On 8-qubit H2O, the proposed error \((2.2\pm1.4)\times10^{-4}\) Ha exceeds quantumDARTS's \(1.7\times10^{-4}\) Ha; the main advantage is total gates, 89.8 versus 219, not universally best energy.
  • Transfer improves speed and energy on 6/8-qubit tasks. At 12 qubits, the reported reductions are 88.2% in steps and 57.6% in CNOTs, without improved final energy. This should be described as reaching no-transfer accuracy faster, not generically reaching chemical accuracy faster.
  • The transfer score is \(S=0.4\Delta_{\mathrm{steps}}+0.1\Delta_{\mathrm{ROT}}+0.2\Delta_{\mathrm{CNOT}}+0.3\Delta_{\mathrm{err}}\), where terms are relative improvements over a no-transfer baseline. It is not physical accuracy and depends on metric weights.
  • The 15–20-qubit experiments rely on fixed MPS initial circuits; 20, 30, and 40 are only additional RL gates. Energy errors range from \(1.5\times10^{-2}\) to \(2.1\times10^{-2}\) Ha, with single-seed runs. This validates compatibility and structural scaling, not large-system chemical accuracy.

Highlights & Insights

  • Replay reliability uses within-trajectory error information rather than an additional noise-prediction network. Annealing recognizes that the proxy becomes more trustworthy as training matures, instead of imposing the same sampling preference at initialization and convergence.
  • Amortized evaluation saves computation while changing feedback granularity. A reusable lesson is to evaluate composite edits when an expensive optimizer masks individual changes, while specifying cached states, delayed rewards, and forced terminal evaluations precisely.
  • Transfer reuses exploration outcomes rather than source-policy parameters. This suits tasks with stable state/action interfaces and expensive target feedback, but retained source labels still require monitoring target behavior; weight-free transfer does not remove all distributional bias.

Limitations & Future Work

  • All quantum experiments use classical simulators, without real-device connectivity constraints, calibration drift, or richer noise. LunarLander is only one supplementary discrete-action validation.
  • A fixed backbone does not mean identical replay hyperparameters: compilation PER uses \(\alpha=0.6\), versus 0.4 for ReaPER/annealing. Results compare the reported configurations and cannot be attributed exclusively to changing \(\omega\).
  • The theoretical conditions are not empirically verified, and exact suffix reliability, online approximation, and importance-weighted losses differ. Future work should log buffer-bin statistics and independently ablate exact/approximate reliability and conventional/current losses.
  • Appendix M labels its empty-buffer rows noiseless while transfer rows correspond to noise conditions, preventing direct matched-noise causal attribution from those rows alone. Its strong-noise exploration comparison has an error ratio of approximately 7; the source's “order-of-magnitude” wording is only a coarse description.
  • ReaPER+ and ReaPER++ are used inconsistently, without sufficient evidence that they are two distinct new algorithms. This note treats them as the annealed method while retaining source labels in tables.
  • The HRC main text says tuned HER plateaus at 95%, but Appendix H Table 7 gives 100%; the claimed approximately 24% speedup over fixed ReaPER also conflicts with 1.56/1.60 million steps. Appendix N's sweep settings, main defaults, and intermediate-reward formula require code-level clarification.
  • The main text gives a 51.0% 12-qubit transfer score, whereas Appendix L's original-weight row gives 12.3%, without enough condition information to reconcile them. Composite scores should not replace individual metrics.
  • vs PER / fixed ReaPER: PER emphasizes information through error, while ReaPER adds downstream reliability. The main addition is a training-dependent reliability exponent, not a new reliability definition.
  • vs HER: HER relabels goals and corresponding rewards, whereas this buffer transfer preserves source transitions intact. They reuse different aspects of experience, and uniformly sampled HER is not equivalent to ordinary uniform replay.
  • vs CRLQAS / TensorRL-QAS: CRLQAS supplies curriculum-based circuit search, and this paper reduces evaluation frequency. TensorRL-QAS supplies MPS initialization; the proposed method refines its fixed base circuit, so added gates must not be presented as complete circuit resources.
  • vs offline-to-online RL: The paper extends buffer seeding with offline experience to noiseless-to-noisy quantum optimization. Selecting source experience through limited target evaluations is a promising direction, not reward relabeling already implemented here.

Rating

  • Novelty: 4/5 — Organizes scheduling, evaluation granularity, and instance transfer around replay, although reliability definitions and buffer seeding have precedents.
  • Experimental Thoroughness: 3/5 — Broad tasks and ablations, constrained by seed-count inconsistencies, unrepeated baselines, and single-run high-qubit experiments.
  • Writing Quality: 2/5 — Clear overall mechanism, but naming, budgets, reward explanations, and numerical conflicts impair reproducibility.
  • Value: 4/5 — Practical insights for expensive-feedback off-policy learning, with speed, accuracy, and circuit costs assessed separately.