Skip to content

๐ŸŽฎ Reinforcement Learning

๐Ÿง  NeurIPS2026 ยท 10 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (12) ยท ๐Ÿ“ท CVPR2026 (25) ยท ๐Ÿ”ฌ ICLR2026 (400) ยท ๐Ÿ’ฌ ACL2026 (46) ยท ๐Ÿงช ICML2026 (110) ยท ๐Ÿค– AAAI2026 (58)

๐Ÿ”ฅ Top topics: Reinforcement Learning ร—6

AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning

AlphaPareto encodes the current alpha pool with a frozen LLM, conditions token-by-token MaskPPO formula search on existing signals, and uses adaptive multi-objective rewards for predictive power, ranking stability, perturbation robustness, and diversity, achieving out-of-sample ICs of 3.92%, 5.70%, and 10.10% on CSI300, CSI800, and the full market.

From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning

QTPT replaces behavior-action prediction in a parameter-shared Transformer with context-conditioned Q-target regression, enabling a fixed-weight model to improve decisions using reward and transition history from the same task; gains are pronounced under weak behavior data, but the theory requires coverage and realizability rather than guaranteeing success on arbitrary low-coverage datasets.

HaM-World: Soft-Hamiltonian World Models with Selective Memory for Planning

HaM-World integrates selective history memory and a Soft-Hamiltonian prior on part of the latent coordinates into one planning model, ranking first on four and second on two state-based control tasks, but its imagined-prediction advantage is confined to planning-relevant horizons, and memory has a substantially larger observed ablation impact than geometry.

Learning Chance-Constrained MDPs with Bellman Distributional Certificates

The paper expresses cumulative-cost violation events through remaining-budget Bellman recursions, obtains nearly matching deterministic-policy sample complexity under bounded successor support and a certified planning oracle, and combines local approximate-KKT optimization with independent validation for stochastic policies; synthetic and power-system simulations distinguish statistical conservatism from conservatism induced by expected-cost surrogates.

Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching

LFIRL converts local action scores from a frozen diffusion policy into soft-Q gradient supervision, then sequentially fits values, calibrates state offsets, and regresses rewards, removing alternating rewardโ€“policy optimization and typically running about 2โ€“3 times faster than the fastest baseline when diffusion pretraining is included, without leading every reward-quality metric.

Modeling Quantum Neural Network Gradient With Reinforcement Learning

RLQ-Grad uses a classical PPO policy with spectral normalization to generate surrogate update signals from a quantum neural network's training state, improving validation accuracy and reducing additional differentiation overhead in the reported classification simulations without proving recovery of true gradients or universal avoidance of barren plateaus.

Replay-buffer engineering for noise-aware quantum circuit optimization

The paper shifts quantum circuit optimization from changing the agent to reusing experience more effectively: annealed reliability-aware replay improves sample efficiency, blockwise evaluation reduces quantumโ€“classical calls, and buffer-only transfer accelerates noise adaptation, although accuracy, gate-count, and convergence advantages differ across tasks.

Safe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement Learning

Safe Score Matching (SSM) uses learned Hamiltonโ€“Jacobi safety values to switch a diffusion policy between reward and recovery training targets, combining multimodal action modeling with low violation costs in online safe reinforcement learning, although its hard-constrained formulation does not establish a strict safety guarantee for the learned policy.

Trust Guided Decision Transformer

TGDT first filters trustworthy history suffixes using rolling next-state prediction errors on realized transitions, then ranks their candidate actions with a frozen IQL critic, improving return and persistent-error behavior in long-horizon navigation without guaranteeing closed-loop coverage or safety.

Verifying Neural Networks with Reinforcement Learning

Rsb uses graph-structured observations and actor-critic learning to reweight Fsb neuron-branching scores without replacing verification logic, increasing solved problems from 148 to 165 among 600 challenging test instances, although inconsistencies in rewards, feature masks, and some statistics need to be distinguished from the performance gains.