On-Policy Diffusion Reinforcement Learning Meets Off-Policy Quality Anchoring¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/HiDream-ai/DiffusionCompass
Area: Image Generation
Keywords: diffusion models, on-policy reinforcement learning, off-policy anchoring, reward overfitting, flow matching
TL;DR¶
DiffusionCompass breaks the self-limiting myopia of on-policy diffusion reinforcement learning by employing off-policy quality anchors as a stabilizing global compass—combining distribution-aware filtering, perceptual anchored rewards, and joint reward renormalization to simultaneously achieve superior reward optimization, enhanced visual diversity, and accelerated convergence across image and video generation tasks.
Background & Motivation¶
Reinforcement learning (RL) has emerged as a cornerstone post-training paradigm to align large-scale diffusion models and flow matching architectures with intricate human preferences and non-differentiable aesthetic targets. Predominant alignment approaches rely primarily on on-policy RL, where the generative model generates rollouts from its active policy and updates its vector field according to reward rankings. While this self-referential feedback loop supplies direct, capacity-aware optimization gradients in the immediate neighborhood of current model parameters, it inevitably suffers from severe intrinsic myopia.
Because exploration is strictly confined within the model's narrow output distribution, on-policy learning lacks an external and objective benchmark of visual quality. This isolation triggers two chronic failure modes: reward saturation, wherein the model quickly overfits to local optima or exploits vulnerabilities in the reward evaluator (reward hacking) to produce visually degraded artifacts; and mode collapse, wherein sample diversity sharply degrades as trajectories polarize onto repetitive, low-variance generation modes. Simply mixing static off-policy datasets into the online rollout buffer also fails, as severe distribution mismatches destabilize training and cause catastrophic gradient variance.
The key insight of this paper is that on-policy exploration can be effectively grounded and guided by an external coordinate system derived from static off-policy samples. Instead of treating off-policy data as rigid supervised targets, they should function as a global "compass" that calibrates directional optimization. The core idea is to establish DiffusionCompass, a dual-objective co-training framework that decouples on-policy exploration from off-policy quality anchoring, leveraging distribution-aware filtering, perceptual feature similarity modulation, and joint reward renormalization to provide stable, globally consistent guidance without sacrificing exploratory freedom.
Method¶
Overall Architecture¶
DiffusionCompass is formulated upon Rectified Flow (flow matching) models and coordinates a dual-model, dual-stream training architecture. The framework maintains a trainable exploration network \(v_\theta\), an Exponential Moving Average (EMA) snapshot network \(v_\theta^{\text{snap}}\) for reliable online rollout generation, and a frozen base model for reference KL regularization. In each training iteration, the on-policy stream uses the snapshot network to sample online rollouts that are subsequently evaluated by reward models; concurrently, the off-policy stream retrieves prompt-matched positive and negative exemplars from a pre-curated offline corpus. Both streams intersect during joint reward renormalization and perceptual feature similarity calculation, after which the on-policy contrastive objective, off-policy anchoring loss, and KL divergence regularizer jointly update the exploration policy.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Text Prompt c"] --> B["Snapshot Online Sampling<br/>Low-fidelity dynamic snapshots"]
A --> C["Offline Corpus Retrieval<br/>Positive set / Negative set"]
B --> D["Distribution-Aware Off-Policy Loss<br/>Velocity difference d with soft weighting w"]
C --> D
B --> E["Perceptual Anchored Reward<br/>DINOv2 max feature cosine similarity modulation"]
C --> E
B --> F["Joint Reward Renormalization<br/>Unified mean and variance over on/off samples"]
C --> F
E --> F
F --> G["Dynamic Sampling Exploration<br/>Flow matching contrastive on-policy loss LON"]
D --> H["Joint Objective Optimization<br/>LON + LOFF + LKL updates exploring policy"]
G --> H
H -->|EMA transfer| B
Key Designs¶
1. Distribution-Aware Off-Policy Loss: Soft-weighted positive alignment and negative concept-erasure To exploit off-policy static guidance without inducing distribution mismatch instability, offline samples are segregated into top \(D\%\) high-reward positive samples \(\mathcal{X}_{\text{off}}^+\) and bottom \(D\%\) low-reward negative samples \(\mathcal{X}_{\text{off}}^-\). For positive samples, the network directly minimizes the distance between predicted and target velocity fields. For negative samples, maximizing velocity distance directly is unbounded and causes gradient divergence; inspired by concept-erasure techniques, the model is guided to predict the velocity field of negative samples under a null prompt \(\emptyset\), enabling classifier-free guidance (CFG) during inference to naturally push generation away from low-quality modes. Crucially, to prevent distant positive outliers from destabilizing updates and avoid penalizing negative modes the model would never generate, a stop-gradient velocity distance metric is computed: $\(d = \text{sg}(\|\mathbf{v} - \mathbf{v}_\theta(\mathbf{x}_t, c)\|_2^2), \quad w = e^{-\alpha d}\)$ The resulting soft-filtering weight \(w\) adaptively modulates the off-policy objective based on distributional compatibility: $\(\mathcal{L}_{\text{OFF}} = \mathbb{E}_{\mathbf{x}_0\in\mathcal{X}_{\text{off}}^+} w \|\mathbf{v} - \mathbf{v}_\theta(\mathbf{x}_t, c)\| + \mathbb{E}_{\mathbf{x}_0\in\mathcal{X}_{\text{off}}^-} w \|\mathbf{v} - \mathbf{v}_\theta(\mathbf{x}_t, \emptyset)\|\)$
2. Perceptual Anchored Reward: External visual feature similarity to counter reward hacking Scalar reward models are susceptible to bias and exploitation, allowing on-policy policies to overfit to metric blind spots by generating unnatural high-frequency textures or distorted geometries while still receiving inflated scores. Because off-policy positive references embody verified perceptual fidelity, DiffusionCompass extracts deep visual representations from both online rollouts and corresponding offline positive anchors using DINOv2 (with temporal frame-averaging for video). The maximum cosine similarity between an on-policy rollout and the offline positive set modulates the raw reward multiplicatively: $\(r_p(\mathbf{x}_0, t) = r(\mathbf{x}_0, t) \cdot \max_{\hat{\mathbf{x}}\in\mathcal{X}_{\text{off}}^+} \text{similarity}(f_{\mathbf{x}_0}, f_{\hat{\mathbf{x}}})\)$ Even if an online rollout achieves an artificially high scalar reward, any significant deviation from realistic perceptual feature representations depresses the composite learning signal, curbing reward hacking and preventing collapse.
3. Joint Reward Renormalization: Breaking self-referential bias with a global quality baseline Standard on-policy RL algorithms (such as GRPO) normalize advantage estimates strictly within the mini-batch of current online rollouts. When the policy undergoes a quality dip or degenerates into a low-fidelity state, all generated samples in the group may be substandard; yet intra-group normalization arbitrarily assigns high positive advantage to the "least bad" rollout, perpetuating low-quality optimization. DiffusionCompass injects the offline reference pool \(\mathcal{X}_{\text{off}}\) into the normalizer alongside online rollouts \(\mathcal{X}_{\text{on}}\): $\(p \approx \frac{1}{2} \left[ \text{clip}\left( \frac{r(\mathbf{x}_0, c) - \mu_{\mathcal{X}_{\text{on}}\cup\mathcal{X}_{\text{off}}}}{\sigma_{\mathcal{X}_{\text{on}}\cup\mathcal{X}_{\text{off}}}}, -1, 1 \right) + 1 \right]\)$ By establishing a stable, globally anchored statistical baseline, online samples receive elevated probability weight only when their quality genuinely competes with or outstrips established off-policy benchmarks.
4. Dynamic Sampling Exploration: Decoupling spatial aesthetics and temporal dynamics overhead Iterative multi-step sampling during video policy exploration introduces severe latency and computational cost. Because the static off-policy corpus already provides a firm global guarantee of visual fidelity, on-policy exploration does not require high-resolution rendering at every optimization step. DiffusionCompass introduces specialized low-fidelity snapshots: for spatial quality rewards (e.g., aesthetics, fidelity), the online policy generates single-frame rollouts; for temporal dynamics rewards (e.g., motion smoothness, background consistency), it generates low spatial resolution sequences at high frame rates. This dynamic sampling drastically cuts training time while preserving optimization efficacy.
Loss & Training¶
The on-policy learning objective \(\mathcal{L}_{\text{ON}}\) adopts DiffusionNFT's vector field contrastive loss, steering perturbed velocity fields towards high-advantage targets. The comprehensive training objective unifies on-policy, off-policy, and KL terms: $\(\mathcal{L} = \lambda_1 \mathcal{L}_{\text{ON}} + \lambda_2 \mathcal{L}_{\text{OFF}} + \lambda_3 \mathcal{L}_{\text{KL}}\)$ Hyperparameters are set to \(\lambda_1 = 1\), \(\lambda_2 = 0.1\), and \(\lambda_3 = 1 \times 10^{-4}\). In image generation, the base model is SD3.5-Medium fine-tuned via LoRA (\(r=32, \alpha=64\)) using AdamW (learning rate \(3 \times 10^{-4}\)), with offline anchors provided by Qwen-Image or Flux; online rollout generation takes 10 sampling steps (evaluated at 40 steps). In video synthesis, Wan2.1-1.3B serves as the base model and Wan2.1-14B provides offline anchors, using 20 rollout sampling steps.
Key Experimental Results¶
Main Results¶
DiffusionCompass is comprehensively evaluated on compositional text-to-image synthesis (GenEval), visual text rendering (OCR accuracy), human preference alignment (PickScore), and 16-dimensional text-to-video generation (VBench).
| Task / Benchmark | Model / Method | Key Metric | Comparison vs. SOTA / Baseline | Iterations / Compute Cost |
|---|---|---|---|---|
| Compositional T2I (GenEval) | DiffusionCompass | Overall: 0.97 | FlowGRPO: 0.95, DiffusionNFT: 0.95 | 1K iters (1/5 of FlowGRPO 5K) |
| Visual Text Rendering (OCR) | DiffusionCompass (w/o CFG) | OCR Acc: 0.975 | DiffusionNFT: 0.970, FlowGRPO: 0.915 | 64 GPU hours (FlowGRPO: 600h) |
| Human Preference (PickScore) | DiffusionCompass | PickScore: 23.8 | DiffusionNFT: 23.7, FlowGRPO: 23.4 | Pixel Vendi: 1.87 vs. 1.46 |
| Video Alignment (VBench 16-dim) | Wan2.1-1.3B + Ours (RP) | Total Score: 85.00 | Wan2.1-14B: 83.69, CausVid: 83.88 | 1.3B model outperforming 14B models |
| Scaled Video Generation (VBench) | Wan2.1-14B + Ours (RP) | Total Score: 87.16 | DiffusionNFT (1.3B): 83.52 | Converged in only 160 iterations |
Ablation Study¶
Component-wise ablations on the visual text rendering benchmark (OCR) using SD3.5-Medium (baseline base model OCR: 0.550):
| Configuration | Reward Renorm | Anchored Reward | Anchoring Loss | Dist. Filter | OCR Acc | Aesthetic | CLIPScore | GPU Hours |
|---|---|---|---|---|---|---|---|---|
| Pure On-Policy (NFT) | - | - | - | - | 0.970 | 4.45 | 0.280 | 64 |
| + Reward Renorm only | ✓ | - | - | - | 0.969 | 4.48 | 0.282 | 56 |
| + Perceptual Anchored Reward | ✓ | ✓ | - | - | 0.963 | 5.03 | 0.309 | 56 |
| + Off-Policy Loss | ✓ | - | ✓ | - | 0.974 | 4.66 | 0.298 | 64 |
| Combined (w/o filtering) | ✓ | ✓ | ✓ | - | 0.972 | 4.88 | 0.302 | 64 |
| Full DiffusionCompass | ✓ | ✓ | ✓ | ✓ | 0.975 | 5.00 | 0.306 | 64 |
Key Findings¶
- Quality anchoring circumvents mode collapse: In out-of-domain human preference testing, pure on-policy DiffusionNFT exhibits severe structural collapse across random seeds despite high reward numbers (Group SSIM of 0.324, Pixel Vendi score 1.46). DiffusionCompass achieves a markedly lower Group SSIM of 0.262 and boosts Vendi diversity to 1.87, preserving broad visual variability.
- Benefits derive from relative quality contrast, not distillation: Replacing external off-policy teacher anchors (Qwen-Image) with pre-generated frozen samples from the base SD3.5-M itself (self-anchoring) still yields an OCR of 0.973 (vs. 0.975), proving the performance surge stems from relative quality contrasting rather than superficial model distillation.
- Agnostic to on-policy RL algorithms: Applying off-policy quality anchoring to FlowGRPO elevates its OCR from 0.915 to 0.939, aesthetic score from 5.04 to 5.08, and CLIPScore from 0.310 to 0.320, proving that off-policy anchoring is an architecture-agnostic enhancement.
- Accelerated convergence: Incorporating 4 to 6 off-policy exemplars per prompt allows the model to match or exceed pure on-policy 48-step performance in just 16 to 24 steps, cutting wall-clock training time substantially.
Highlights & Insights¶
- Negative concept erasure via null prompt velocity mapping: Penalizing low-reward samples directly in continuous vector fields often triggers catastrophic gradient explosion. Mapping negative samples to null prompt conditioning \(\emptyset\) leverages classifier-free guidance mechanics to steer sampling away from degraded regions at inference.
- Decoupled dynamic sampling for video post-training: Allocating spatial evaluations to single-frame rollouts and temporal evaluations to downscaled high-frame-rate sequences effectively bypasses the multi-minute rollout bottleneck of video diffusion models.
- Shattering the self-referential trap: DiffusionCompass bridges on-policy RL with static anchoring, offering an elegant blueprint for how generative models can actively explore without succumbing to myopic over-optimization.
Limitations & Future Work¶
- Dependency on offline sample curation: Generating and scoring offline reference sets introduces pre-computation overhead, which can become cumbersome when prompts scale dynamically or adapt interactively.
- Feature extraction latency: Incorporating real-time DINOv2 feature extraction during training adds computational and memory demands in distributed high-throughput clusters.
- Future directions: Exploring dynamic self-adjusting balancing ratios between on-policy and off-policy objectives, and extending this anchoring paradigm to autoregressive visual models and robot trajectory diffusion policies.
Related Work & Insights¶
- vs. FlowGRPO / DanceGRPO: FlowGRPO adapts GRPO to diffusion flow matching but relies strictly on on-policy ODE/SDE sampling, resulting in slow convergence (5K steps) and susceptibility to reward saturation. DiffusionCompass achieves higher accuracy and diversity within 1K steps.
- vs. DiffusionNFT: While DiffusionNFT enables probability-free contrastive velocity updates, its pure on-policy regime remains vulnerable to reward hacking and sample collapse. DiffusionCompass complements it with off-policy quality anchors and joint reward renormalization.
- vs. SFT / Offline DPO: Supervised fine-tuning or offline preference optimization is bounded by the static distribution of pre-collected datasets. DiffusionCompass preserves active on-policy self-exploration while utilizing static data strictly as an anchoring reference.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ An elegant fusion of off-policy quality anchoring, distribution-aware soft filtering, and dynamic sampling for diffusion RL.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive validation spanning compositional generation, OCR, human preferences, and full 16-dimension video VBench benchmarks.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous technical narrative, clear architectural formulation, and well-designed qualitative and quantitative evidence.
- Value: ⭐⭐⭐⭐⭐ Solves the critical challenges of reward hacking, mode collapse, and prohibitive video rollout costs in diffusion alignment; open-sourced and highly practical.