Skip to content

Modeling Quantum Neural Network Gradient With Reinforcement Learning

Conference: NeurIPS2026 (main-track poster; supplied metadata, not verified online)
arXiv: 2609.31066v1
OpenReview: https://openreview.net/forum?id=lfrvu8xfsh
Code: https://doi.org/10.5281/zenodo.22956233
Area: Reinforcement Learning (a PPO-trained optimizer for quantum neural networks)
Keywords: learned optimizer, quantum neural network, barren plateau, PPO, spectral normalization

TL;DR

RLQ-Grad uses a classical PPO policy with spectral normalization to generate surrogate update signals from a quantum neural network's training state, improving validation accuracy and reducing additional differentiation overhead in the reported classification simulations without proving recovery of true gradients or universal avoidance of barren plateaus.

Background & Motivation

Quantum neural networks (QNNs) commonly encode classical features as quantum rotation angles, apply a trainable circuit, and pass measurements to a classical classification head. During training, back-propagation must retain or reconstruct quantum states, while parameter-shift repeatedly evaluates the circuit; these operations become expensive as the qubit count increases in statevector simulation. A separate difficulty is the barren plateau: for some circuit and observable combinations, the loss or its derivatives concentrate strongly over random parameter ensembles, making useful directions difficult to distinguish with a finite measurement budget. Expensive derivative computation and insufficiently distinguishable derivative signals are related but different problems.

Previous quantum reinforcement learning work searches circuit architectures or learns parameter controllers for problems such as QAOA and VQE. This paper keeps the hardware-efficient ansatz (HEA) used in its main experiments and turns supervised classification training into an RL environment: the policy proposes an update signal for quantum parameters, the environment applies it, and training loss and accuracy provide feedback. The attraction is avoiding explicit circuit derivatives at every quantum parameter update; the cost is introducing a classical optimizer that must be trained concurrently and depends on informative rewards. Actor inference, policy training, and the remaining QNN forward evaluations therefore require separate accounting.

Core idea: replace analytical differentiation of quantum parameters with a state-conditioned learned update rule, train the policy using training outcomes rather than true-gradient labels, and stabilize the feedback loop with PPO and spectral normalization.

Method

Overall Architecture

RLQ-Grad receives a summary of the optimization process rather than a single classification example: current quantum parameters, loss and accuracy on the current training batch, and the update signal produced by the policy at the previous step. A classical actor outputs a continuous action with the same dimension as the quantum parameters, places it in their gradient slots, and lets Adam apply the update; the remaining classical layers still use back-propagation according to the original algorithm. The updated QNN produces training-batch statistics for reward calculation and construction of the next state, while a critic learns state values to support PPO.

This loop has two distinct modes of use. Joint training collects state–action–reward trajectories and updates the actor and critic; when a learned optimizer is used to continue optimizing the QNN, QNN forward statistics are still needed to construct its state, but generating an action need not imply a PPO training update. Final classification inference uses only the trained QNN, without PPO rewards or the optimizer actor.

Solid edges below show the data flow during QNN optimization, while dashed edges show reward and policy-training feedback during joint training; classification deployment is separate from this optimization loop.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    I["Training batch + current QNN"] --> A["State-conditioned proposal"]
    A --> Q["QNN update + forward statistics"]
    Q --> B["Reward feedback"]
    B -. "training trajectories" .-> C["PPO + spectral normalization"]
    C -. "train actor and critic" .-> A
    Q --> O["Trained QNN: classification inference"]
    Q -->|"next optimization state"| A

Key Designs

1. State-conditioned proposal: learn an update rule rather than regress to true derivatives

The policy needs to know its current position and its recent movement; otherwise it can only produce random directions disconnected from training progress. Equation (6) concatenates quantum parameters, mean training-batch loss, mean accuracy, and the previous update signal. The shorthand below represents the same structure using batch statistics; the preceding signal comes from the actor and does not require an additional true circuit gradient.

\[ s_t=[\theta_t,\ell_t,\mathrm{acc}_t,g_{t-1}]. \]

The actor is a classical MLP whose action dimension equals the number of trainable quantum parameters. Current parameters indicate position, loss and accuracy summarize task performance, and the previous action supplies short-term optimization history. The policy can therefore acquire direction-selection heuristics through interaction rather than repeatedly querying the circuit for every parameter as parameter-shift does. The paper calls the action a surrogate gradient, but does not use true derivatives as regression targets or prove that it is an unbiased gradient estimator.

Equation (12) uses the following simplified update to explain the difference from analytical differentiation.

\[ \theta_{t+1}=\theta_t-\eta\,g_t,\qquad g_t\sim\pi_\phi(\cdot\mid s_t). \]

The implemented algorithm is not simply this fixed-step update: it writes the action into quantum parameter gradient slots and then uses Adam with a cosine annealing scheduler. Adam's momentum and second moments further alter the actual parameter displacement. Policy output, true derivative, and final parameter movement are consequently distinct objects; using a gradient interface does not make the action a derivative of the loss.

The QNN first reduces input dimensionality with PCA, maps features to angles through a classical input fully connected layer, and encodes them with Rx rotations. Each HEA layer contains single-qubit rotations and entangling operations, followed by Pauli Z measurements and an output fully connected layer that maps measurements to classes. Its quantum parameter count is \(P_{\mathrm{QNN}}=3nL\). The policy primarily takes over these quantum parameters; this does not imply that the classical input layer no longer requires gradients propagated through the quantum module. The algorithm explicitly retains back-propagation for other computation-graph components, and input gradients along that path may still decay.

2. Reward feedback: assess actions through classification outcomes, not derivative labels

A nonzero action is not necessarily useful; the policy needs training outcomes after execution to identify better directions. Equation (7) adds training-batch accuracy to reciprocal loss, rewarding both higher accuracy and lower loss. To preserve the source definition, the batch-expectation form is retained below rather than silently replacing the mean reciprocal loss with the reciprocal of mean loss.

\[ r_t=\mathbb{E}_{(x,y)\sim\mathcal{B}_t}\left[acc(f(x,\theta_t,\mathbf{W}_t,\mathbf{b}_t),y)+\frac{1}{\mathcal{L}(f(x,\theta_t,\mathbf{W}_t,\mathbf{b}_t),y)+\epsilon}\right]. \]

The small constant avoids division by zero but does not remove the sensitivity of reciprocal loss near small losses; accuracy is also a discrete statistic. The scaling between these terms and the effects of finite shots or noise need separate evaluation. Rewards here use training batches, not validation labels; the classification metrics reported later are validation accuracies.

Algorithm 1 evaluates post-update statistics but writes the reward formula using current-time parameter indices, leaving the timing notation unclear. This note follows the action–evaluation–transition feedback logic without claiming to have verified the precise pre-update or post-update reward implementation in code. Initialization descriptions also differ: the algorithm initializes quantum parameters and the previous signal at zero, while other sections discuss random initialization. These statements should not be merged into a single supposedly verified setting.

The defensible structural conclusion in Appendix C.1 is that a variance bound for true circuit derivatives does not directly constrain an external classical policy's output. Its stronger inference that absence of circuit differentiation implies statistical independence does not follow: parameters, loss, and accuracy in the state come from the circuit, and policy training depends on circuit-generated rewards. Actions may correlate with true derivatives, or have substantial variance without any descent effect. Appendix F.2 and the main limitations section further acknowledge that PPO loses its learning signal when loss and reward differences concentrate; this is more accurate than claiming universal barren plateau avoidance.

3. PPO + spectral normalization: stabilize an environment changed by its own actions

Every parameter update changes the next QNN and reward distribution, so old trajectories can become inconsistent with the current environment. The authors use on-policy PPO to collect trajectories from the current policy and a clipped objective to constrain policy changes; the critic supplies value estimates for advantage construction. The paper does not introduce a new PPO loss, and clipping range, GAE, entropy coefficient, and update epochs use SB3 defaults unless otherwise stated.

The authors do not establish that PPO achieves higher accuracy than every other agent here. In the preliminary comparison in Appendix D.1, TD3 and SAC are more accurate than PPO without spectral normalization; the choice of PPO emphasizes resource trade-offs involving a smaller model and rollout storage rather than large replay buffers. Its on-policy nature does not automatically guarantee unbiased update directions, global optimality, or polynomial sample complexity.

Spectral normalization acts on selected policy/value network layers, scaling weights by an estimated largest singular value to limit sensitivity to input changes. Equation (16) writes this as:

\[ \hat{W}_i=\frac{W_i}{\rho_i},\qquad\rho_i=\max\bigl(\rho(W_i),k\bigr). \]

The denominator uses stop-gradient, and each training update uses one power-iteration step to estimate the spectral norm. Algorithm 2 instead uses a normalization variant with a lower bound of 1. Appendix D.2 compares normalization of the final actor and critic layers, only the final critic layer, and no normalization; its explanation is that critic constraints reduce amplification of TD errors while actor constraints further stabilize update scales. These are mechanistic explanations with empirical support, not a convergence theorem; spectral norms need not grow monotonically along arbitrary training trajectories.

Loss & Training

The quantum classifier minimizes supervised loss, whereas PPO optimizes cumulative reward and trains a value function. These are different objectives, and there is no explicit true-gradient regression term. The main algorithm uses Adam with learning rate 1e-3, weight decay 1e-4, momentum (0.9, 0.999), and cosine annealing. Appendix D.1 separately describes SGD with learning rate 5e-4 for policy/value networks, while resource analysis includes Adam states for the actor/critic. The optimizer descriptions conflict and should not be silently reconciled.

The main classification comparison trains for 200 epochs with batch size 128 for BC and 256 for the other datasets. PCA selects features to retain approximately 95% cumulative explained variance; BC uses a fixed-seed split. A separate agent is trained for each dataset and circuit configuration, without demonstrated meta-training that transfers across tasks or qubit counts.

Appendix D.1 uses a single environment and a rollout covering the full supervised training step count: 400 for BC, 46800 for MNIST/F-MNIST, and 39000 for CIFAR10, with policy-update minibatches of 200. These settings clarify the work involved in joint training and show why actor parameter memory alone cannot explain total rollout storage. The full-pipeline time order in Appendix E.3 is analyzed per sample for a PPO update; a complete run also depends on trajectory length, update epochs, and environment evaluation count.

Key Experimental Results

Main Results

The first table comes from Table 2: BC top-1 validation accuracy (%), mean ± standard deviation over 4 runs. Configuration labels also specify the quantum parameter count.

Method 2q1d (6) 4q2d (24) 6q3d (54) 8q4d (96) 10q5d (150) 12q6d (216)
Backpropagation 80.62 ± 10.64 90.86 ± 1.998 91.88 ± 2.057 91.88 ± 1.126 91.96 ± 2.402 93.15 ± 2.942
Parameter-shift 85.71 ± 7.082 89.77 ± 1.316 89.60 ± 5.526 89.77 ± 2.010 91.99 ± 5.385 92.11 ± 5.526
Adjoint differentiation 83.01 ± 11.64 84.93 ± 14.82 85.04 ± 3.689 87.32 ± 8.476 92.75 ± 10.28 93.02 ± 8.926
RLQ-Grad 95.52 ± 0.080 95.20 ± 2.379 95.79 ± 3.418 95.96 ± 0.380 96.18 ± 4.178 96.58 ± 5.612

At 12q6d in the same Table 2, MNIST gives 95.27 ± 0.297 for RLQ-Grad and 92.48 ± 0.225 for Backpropagation; F-MNIST gives 87.62 ± 0.181 and 85.72 ± 0.198; CIFAR10 gives 41.15 ± 0.352 for RLQ-Grad and 40.08 ± 0.321 for the strongest analytical baseline, Parameter-shift. All are validation results, not performance on an independent held-out test set.

The second table comes from Table 3: top-1 validation accuracy (%) on larger CIFAR-10 circuits, mean ± standard deviation over 3 runs.

Method 14q7d (294) 16q8d (384) 18q9d (486) 20q10d (600)
Backpropagation 38.72 ± 3.21 40.45 ± 2.67 39.86 ± 3.48 40.13 ± 2.91
CMA-ES 10.10 ± 0.14 10.12 ± 0.21 10.11 ± 0.20 10.13 ± 0.15
sep-CMA-ES 10.09 ± 0.23 10.14 ± 0.17 10.11 ± 0.26 10.15 ± 0.19
SPSA 10.08 ± 0.27 10.15 ± 0.18 10.11 ± 0.32 10.14 ± 0.24
Layerwise 44.73 ± 3.21 48.58 ± 2.74 50.92 ± 3.47 53.16 ± 3.65
Gaussian Init 43.92 ± 3.15 46.87 ± 2.38 49.76 ± 3.02 52.41 ± 2.71
RLQ-Grad 45.86 ± 3.84 48.64 ± 2.46 51.29 ± 3.16 53.02 ± 2.83

RLQ-Grad is comparable to dedicated plateau-mitigation methods rather than uniformly better; Layerwise has the highest mean at 20q10d. Overlapping uncertainty ranges also do not justify calling small mean differences significant gains. CMA-ES, sep-CMA-ES, and SPSA are near chance in the reported settings, which does not establish failure under every budget, initialization, or circuit.

Ablation Study

The numerical spectral-normalization comparison comes from Appendix Table 6. PPO without normalization versus PPO (SN) gives 90.05 ± 4.760 and 95.52 ± 0.080 on BC; 59.52 ± 0.212 and 61.36 ± 2.216 on MNIST; 68.52 ± 0.534 and 73.64 ± 0.415 on F-MNIST; and 25.12 ± 0.961 and 27.17 ± 0.782 on CIFAR10. The table also changes parameter counts from 5575 to 11847 and MACs from 5440 to 10432, so it is not an architecture-matched pure SN ablation. Figure 4 supplies qualitative evidence about normalization placement; no numerical values are invented from figures absent from the text cache.

The third table extracts resource measurements from Appendix Tables 7–8. Times are seconds per iteration; RLQ-Grad includes what the authors describe as PPO rollout, actor/critic updates, gradient buffers, and optimizer states, rather than actor-only inference. The CPU is an AMD Ryzen 5 5600G, with averages over 12 measurements after 4 warm-up iterations.

Config Backpropagation time Parameter-shift time Adjoint time RLQ-Grad full time RLQ-Grad full memory
2q1d 0.0042 0.0067 0.0027 0.0455 161.22 KB
6q3d 0.0260 0.1333 0.0109 0.0641 306.35 KB
12q6d 0.1662 2.3931 0.0569 0.0858 796.14 KB
14q7d 0.2719 5.2019 0.0836 0.0842 1.02 MB
20q10d 219.41 693.96 59.3241 0.0881 1.95 MB

Key Findings

  • Small circuits show clear non-wins: full PPO is slower than all three analytical methods at 2q1d and remains slower than adjoint at 12q6d; 0.0842 versus 0.0836 at 14q7d is only an approximate tie. The largest CPU ratios concern additional optimization/differentiation cost, not the same speedup for total QNN training.
  • The GPU comparison includes only Backpropagation and RLQ-Grad on an NVIDIA RTX 3090, averaged over 15 trials; at 20q10d they take 1457.06 and 1190.83 ms. Speedups across the range are approximately 1.03–1.22 times, not the thousands-fold CPU ratios.
  • Table 1 explicitly excludes the common QNN forward cost, and Appendix E.1 retains the exponential dependence of statevector forward evaluation. Polynomial classical actor cost is therefore not a polynomial guarantee for end-to-end statevector simulation.
  • At 20q10d, reported memory is 6174 MB for Backpropagation, 22.55 MB for Parameter-shift, and 1.95 MB for RLQ-Grad full; reported adjoint values are below the 0.1 MB measurement resolution throughout. Adjoint retains a constant number of additional buffers, not constant total statevector memory. Higher parameter-shift memory also depends on batching implementation; serial queries need not retain a separate state for every parameter.

Highlights & Insights

  • The contribution changes how an update signal is generated rather than recovering an unmeasurable true derivative. The policy can encode experience from a particular training process, providing a concrete interface between learned optimizers and quantum circuits.
  • Separating differentiation cost from reward distinguishability is the most important lens for reading this paper. Computational savings can be real, while cheap actions with substantial variance remain unhelpful when rewards contain no information.
  • PPO and SN target the stability of policy training and can complement layerwise learning or structured initialization that changes QNN trainability. A useful next test combines these approaches and measures accuracy, interaction budgets, and total runtime rather than action variance alone.

Limitations & Future Work

  • Limited theoretical strength: absence of differentiation in Appendix C.1 does not imply statistical independence, and classical back-propagation in C.2 does not guarantee well-behaved policy gradients. Appendix F.1 invokes informative rewards, good local minima, and universal approximation to claim sufficient conditions for efficient training, but supplies no PPO sample-complexity or convergence guarantee. Universal useful updates or polynomial training cannot be inferred.
  • Reward concentration and poor local minima remain: the main text acknowledges vanishing reward differences on genuinely flat losses and local-minimum obstructions under underparameterization. The end of F.2 describes failure of “Condition 2” as a flat loss, conflicting with the earlier definitions of Conditions 1/2; the overparameterization ratio in G.1 includes a square root inconsistent with the earlier threshold. This note does not guess corrections to the authors' conditions.
  • Input-gradient and proof boundaries: classical input layers may still suffer quantum input-Jacobian decay. The exact 2-design assumption in Appendix G.2 does not follow simply from using an HEA; the tighter second-moment bound is not directly obtained from the stated trace bounds, and the cross-moment formula needs review. It is not treated here as a verified theorem for arbitrary supervised losses.
  • Conflicting resource statements: the paper calls 796.14 KB versus actor-only 102.28 KB at 12q6d approximately 4 times, although the ratio is approximately 7.8; 26% + 26% + 52% sums to 104%, not a consistent memory decomposition. Parameter/gradient/optimizer-state accounting also does not adequately explain the complete measurement boundary for the large rollout buffer.
  • Limited empirical scope: all experiments use shot-noise-free statevector simulation. The four-dataset comparison reaches only 12 qubits; larger scales cover only CIFAR-10 and resource evaluation up to 20 qubits. There is no new evidence on physical quantum hardware, realistic noise channels, cross-configuration agent transfer, or an independent test set.
  • Stricter future validation: provide network-size-matched SN ablations, measure action alignment with true gradients and actual descent rates, separate common forward evaluation, rollout, policy updates, and peak memory, and test reward distinguishability under finite measurement budgets. These are recommendations, not experiments already performed in the paper.
  • Versus analytical differentiation: Backpropagation, Parameter-shift, and Adjoint differentiation compute circuit derivatives; RLQ-Grad learns a surrogate signal consumable by Adam. The latter may be cheaper, but the former retain derivative semantics, so variance curves alone cannot establish which is more correct.
  • Versus quantum architecture search: earlier RL-QAS methods select gate sequences or circuit structures, whereas this paper keeps the HEA and learns parameter updates. PPO-trained optimization is the core method, and quantum computing provides the optimizer's application environment.
  • Versus learned optimizers: the classical-network quantum controllers of Verdon et al. and the learning-to-learn work of Andrychowicz et al. provide conceptual precedents. This is not yet a reusable cross-task meta-optimizer, and per-configuration agent training cost requires explicit reporting.
  • Versus landscape mitigation: Layerwise and Gaussian Init improve trainability, while this paper reduces reliance on explicit differentiation. Larger-scale results support comparable accuracy but do not establish that all three necessarily benefit through exactly the same mechanism.

Rating

  • Novelty: 4/5; connects the quantum update interface of supervised QNNs to PPO with a clear mechanism, while building on learned-optimizer ideas.
  • Experimental Thoroughness: 3/5; includes multiple configurations and stronger baselines, but remains limited by validation-only reporting, simulation conditions, and confounded ablations.
  • Writing Quality: 2/5; the limitations discussion is useful, but independence arguments, optimizer descriptions, and resource numbers contain clear inconsistencies.
  • Value: 3/5; motivates an alternative to expensive circuit differentiation, without establishing a general solution to barren plateaus.