Rdm: Re-conceptualizing Distribution Matching as a Reward for Diffusion Distillation¶
Conference: ECCV 2026
Paper: ECCV 2026 Official
Code: https://github.com/Flq2002/Rdm
Area: Image Generation
Keywords: Diffusion Models, Diffusion Distillation, Distribution Matching, Reinforcement Learning, GRPO
TL;DR¶
This paper re-conceptualizes Distribution Matching Distillation (DMD) as an intrinsic policy reward (\(R_{\mathrm{dm}}\)), and proposes a unified Group Relative Policy Optimization framework (GNDMR) featuring group normalization and adaptive magnitude scaling that eliminates high-noise score variance and seamlessly integrates auxiliary aesthetic rewards for few-step diffusion models.
Background & Motivation¶
Diffusion models have become the cornerstone of modern visual synthesis by progressively denoising random latents, but their slow iterative sampling process fundamentally hinders real-time applications. To address this computational bottleneck, few-step distillation paradigms such as Distribution Matching Distillation (DMD/DMD2) have gained widespread adoption, training a student generator in 1 to 4 steps by minimizing the reverse-KL divergence between the pre-trained teacher's distribution and the student's output distribution. Nevertheless, traditional distillation objectives restrict the student's performance ceiling by strictly anchoring it to the teacher, making it difficult to surpass the teacher's aesthetic quality or align with complex human preferences.
Recent efforts attempt to break this ceiling by incorporating Reinforcement Learning (RL) techniques (such as DDPO, ReFL, and DPO) into the distillation pipeline, exemplified by DMDR which optimizes a linear combination of distillation loss and external reward objectives. However, treating distillation and RL as two isolated, competing optimization streams introduces acute optimization conflicts: first, as the diffusion timestep increases, the heavy Gaussian blur drastically amplifies the estimation variance of the score difference vector, destabilizing the distillation gradients; second, external reward models (e.g., HPS and CLIP Score) perform unconstrained scalar maximization, making the student vulnerable to reward hacking, while static loss weights fail to accommodate the dynamically shifting magnitudes of intermediate diffusion residuals.
The angle of attack in this paper is to eliminate the artificial dichotomy between distillation and reinforcement learning by recognizing that the mathematical formulations of DMD gradient and policy gradient are fundamentally isomorphic. Core idea: re-conceptualize distribution matching as an intrinsic policy reward \(R_{\mathrm{dm}}\), establishing a unified Group Relative Policy Optimization framework (GNDMR) that uses group normalization to cancel out high-noise score variance and employs an adaptive magnitude scale to harmoniously balance distillation with auxiliary downstream rewards.
Method¶
Overall Architecture¶
The core principle of Rdm is to map the multi-step distribution matching gradient directly into a Markov Decision Process (MDP) policy gradient, thereby bringing diffusion distillation under the broader umbrella of reinforcement learning. In each training iteration, the student generator \(G_\theta\) samples a group of few-step denoising trajectories for each given text prompt. For intermediate latents along these trajectories, an intrinsic distribution matching reward \(R_{\mathrm{dm}}\) is computed via score differences between the frozen real teacher \(\mu_{\mathrm{real}}\) and the online fake score estimator \(\mu_{\mathrm{fake}}\), while auxiliary aesthetic or alignment models compute sparse rewards on the final generated images. Group Normalization is applied separately to both rewards within the prompt group to eliminate shared estimation drifts, and an adaptive weighting factor synchronizes their gradient magnitudes. Finally, the generator policy is updated using clipped Importance Sampling for sample-efficient trajectory reuse.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Prompts and Initial Gaussian Latents"] --> B["Group Trajectory Sampling<br/>Generate M Few-Step Trajectories"]
B --> C["Fake Score Estimator Update<br/>Track Student Distribution"]
C --> D["Distribution Matching Reward Rdm<br/>Construct Intrinsic Dense Reward"]
D --> E["Group Normalized Distillation GNDM<br/>Shared Timesteps Mean Centering"]
E --> F["Adaptive Multi-Reward Dynamic Alignment<br/>Magnitude Matching Factor"]
F --> G["Importance Sampling Policy Update<br/>Multi-Step Generator Updates under GRPO"]
Key Designs¶
1. Distribution Matching Reward Rdm: Bridging Distillation and Policy Gradients
In multi-step distillation, DMD minimizes the distribution divergence by aligning the student generator's gradient with the real-versus-fake score difference \((s_{\mathrm{real}} - s_{\mathrm{fake}})\). By examining the Gaussian transition probability in the reverse SDE, its log-derivative \(\nabla_\theta \log p_\theta(x_{t-1} \mid x_t) = \frac{x_{t-1} - \mu_\theta(x_t)}{\sigma_{t-1}^2} \nabla_\theta \mu_\theta(x_t)\) connects directly with the generator output gradient. Formulating the denoising chain as a multi-step MDP where single-step transitions act as policy actions, the distribution matching reward is explicitly derived as:
Under this definition, the policy gradient \(\nabla_\theta \mathcal{J}_{\mathrm{DDPO}_{\mathrm{DM}}} = \mathbb{E}[R_{\mathrm{dm}} \nabla_\theta \log p_\theta(x_{t-1} \mid x_t)]\) strictly equals the negative DMD loss gradient \(-\nabla_\theta \mathcal{L}_{\mathrm{DMD}}\). Crucially, unlike unbounded rewards that cause over-optimization, \(R_{\mathrm{dm}}\) is a divergence measure optimized towards zero, functioning as an inherent safeguard against reward hacking. In practical implementation, to prevent numerical division-by-zero or sign flips caused by the stochastic residual \(x_{t-1} - \mu_\theta(x_t)\), a sign-based normalization is applied together with an \(L_1\) norm weighting derived from the real score estimator.
2. Group Normalized Distillation GNDM: Mitigating High Variance at Large Timesteps
At large diffused timesteps \(t'\), heavy Gaussian noise corrupts the samples, causing the score difference vector \(R_s(t') = s_{\mathrm{real}}(x_{t'}) - s_{\mathrm{fake}}(x_{t'})\) to fluctuate dramatically with extreme variance. This erratic signal misguides the distillation trajectory in early denoising stages. To overcome this limitation without introducing complex Critic networks, the authors introduce Group Normalization (GN) into \(R_{\mathrm{dm}}\), yielding Group Normalized Distribution Matching (GNDM).
For each prompt, the generator produces \(G\) trajectories where samples within the group strictly share the generation timestep \(t\) and diffused timestep \(t'\). The advantage of the \(i\)-th sample is obtained by subtracting the group-mean and dividing by the group standard deviation:
Mean subtraction effectively strips out the shared high-noise baseline bias, providing a clean and robust relative optimization signal that dramatically stabilizes the distillation process across noisy diffusion regimes.
3. Adaptive Multi-Reward Dynamic Alignment: Eliminating Scale Imbalance and Collapse
When incorporating external downstream rewards \(R_o\) (such as HPSv2.1 aesthetic quality or CLIP semantic alignment), the rewards are evaluated on the final clean images and normalized into advantage \(A_{o, t}^i\). Directly summing distillation and auxiliary advantages with fixed weights leads to unstable training or catastrophic collapse because the intrinsic magnitude of the distribution matching update changes dynamically over training steps. The authors extract the base scalar \(w_{\mathrm{dm}, t}\) from \(R_{\mathrm{dm}}\) and construct an adaptive scaling coefficient \(\beta_{\mathrm{dm}, t}\):
where \(C\) and \(S\) denote the channels and spatial dimensions. The total advantage is formulated as \(A_{\mathrm{sum}, t}^i = w_{\mathrm{dm}, t} A_{\mathrm{dm}, t}^{i, t'} + \beta_{\mathrm{dm}, t} \sum_j w_j A_{o_j, t}^i\). This ensures that auxiliary reward gradients adaptively match the dynamic magnitude of the distillation term throughout training, eliminating hyper-parameter brittleness and preventing aesthetic rewards from overwhelming distribution fidelity.
4. Importance Sampling Policy Update: Multi-Step Updates from Reused Trajectories
Applying standard On-Policy reinforcement learning to diffusion distillation is severely bottlenecked by sampling overhead, as generating multi-step trajectories across large diffusion models is computationally intensive. Because the reverse transitions are Gaussian, the likelihood ratio between current and old policy parameters \(r_t^i(\theta) = \frac{p_\theta(x_t^i \mid x_{t-1}^i)}{p_{\theta_{\mathrm{old}}}(x_t^i \mid x_{t-1}^i)}\) can be computed analytically in closed form. GNDMR integrates clipped Importance Sampling (with clipping range \(\eta = 0.5\)), enabling 2 to 5 generator gradient updates per sampling round. This allows the model to cut the total sampling demand in half while achieving fully comparable generation quality and aesthetic convergence.
Key Experimental Results¶
Main Results¶
Distillation experiments are conducted on SD3-Medium and SD3.5-Medium using prompts from LAION-AeS-6.5+, evaluated on 10K prompts from COCO2014 (Karpathy split). Evaluation metrics include Human Preference Score (HPS v2.1), PickScore (PS), Multi-dimensional Preference Score (MPS), CLIP Score (CS), Frechet Inception Distance (FID), and distribution deviation from the teacher (FID-SD).
| Method | Steps (NFE) | Res. | Img-Free | HPS ↑ | PS ↑ | MPS ↑ | CS ↑ | FID ↓ | FID-SD ↓ | Sampling Cost |
|---|---|---|---|---|---|---|---|---|---|---|
| Stable Diffusion 3 Medium | ||||||||||
| Base Model (CFG=7) | 50 | 1024 | - | 29.00 | 22.72 | 12.10 | 38.86 | 24.48 | - | - |
| Flow-GRPO (CFG=7) | 50 | 1024 | - | 30.35 | 22.90 | 12.56 | 38.52 | 26.63 | 12.39 | - |
| Hyper-SD (CFG=5) | 8 | 1024 | ✗ | 27.20 | 21.90 | 11.22 | 37.76 | 26.94 | 10.09 | - |
| LCM | 4 | 1024 | ✗ | 27.76 | 22.31 | 11.61 | 36.97 | 27.71 | 15.90 | - |
| DMD2 | 4 | 1024 | ✗ | 26.64 | 22.36 | 11.37 | 38.00 | 27.28 | 16.51 | - |
| Flash-SD3 | 4 | 1024 | ✗ | 27.47 | 22.65 | 11.98 | 38.07 | 26.01 | 12.21 | - |
| DMDR | 4 | 1024 | ✓ | 29.50 | 22.77 | 11.98 | 38.10 | 29.10 | 14.11 | - |
| GNDMR-IS (Ours) | 4 | 1024 | ✓ | 30.00 | 22.89 | 12.48 | 38.15 | 28.84 | 12.47 | 128 × 4k |
| GNDMR (Ours) | 4 | 1024 | ✓ | 30.37 | 22.88 | 12.53 | 38.20 | 28.02 | 12.21 | 128 × 8k |
| Stable Diffusion 3.5 Medium | ||||||||||
| Base Model (CFG=3.5) | 50 | 512 | - | 27.78 | 22.59 | 11.91 | 38.46 | 20.69 | - | - |
| Flow-GRPO (CFG=3.5) | 50 | 512 | - | 31.81 | 23.24 | 12.93 | 39.21 | 29.27 | 9.39 | - |
| DMD2 | 4 | 512 | ✗ | 30.44 | 22.92 | 12.73 | 38.59 | 26.64 | 14.63 | - |
| DMDR | 4 | 512 | ✓ | 30.83 | 23.07 | 12.80 | 38.22 | 26.05 | 16.73 | - |
| GNDMR-IS (Ours) | 4 | 512 | ✓ | 30.88 | 22.94 | 12.86 | 38.39 | 25.60 | 16.68 | 128 × 3k |
| GNDMR (Ours) | 4 | 512 | ✓ | 31.25 | 23.15 | 12.93 | 38.59 | 24.44 | 13.93 | 128 × 6k |
Ablation Study¶
Ablations are run on SD3-Medium at \(512 \times 512\) resolution, starting with 500 warmup distillation iterations followed by multi-reward training evaluated on COCO30K.
Table 1: Ablation on Group Normalization (GN) on Rdm
| Method | Only Rdm (500 iters) FID ↓ | +HPS: FID ↓ | +HPS: HPS ↑ | +PS: FID ↓ | +PS: PS ↑ |
|---|---|---|---|---|---|
| GNDMR (w/ GN) | 23.07 | 24.47 | 30.22 | 22.32 | 22.76 |
| w/o GN | 24.94 | 25.40 | 30.39 | 24.61 | 22.71 |
Table 2: Ablation on Timestep Sharing and Timestep Intervals for GN
| Sampling Strategy | Shared \(t'\) | Shared \(t\) | FID ↓ | GN Timestep Interval | FID ↓ |
|---|---|---|---|---|---|
| Vanilla DMD Baseline | ✗ | ✗ | 24.94 | \([0, 300]\) | 23.71 |
| Shared \(t'\) only | ✓ | ✗ | 24.34 | \([300, 600]\) | 23.46 |
| Shared \(t'\) & \(t\) (Ours) | ✓ | ✓ | 23.07 | \([600, 1000]\) (High Noise) | 23.43 |
Key Findings¶
- Group Normalization significantly cleans gradient estimates: In the first 500 pure distillation steps, GNDM reduces FID from 24.94 to 23.07 (-1.87), validating that subtracting group-mean stats removes destructive pseudo-gradient noise early in training.
- Large diffused timesteps introduce highest variance: Applying GN specifically on the high noise interval \([600, 1000]\) yields the most substantial FID improvement, confirming the analytical observation that score variance spikes when latents are heavily blurred.
- Dynamic scaling factor \(\beta_{\mathrm{dm}, t}\) prevents collapse: Under fixed static weights (\(w=10\) or \(100\)), training easily becomes unstable or suffers from premature score stagnation; in contrast, \(\beta_{\mathrm{dm}, t}\) dynamically matches auxiliary reward gradients to residual updates, enabling continuous, stable HPS growth.
- Importance Sampling halves sampling overhead: With two policy updates per trajectory sample, GNDMR-IS reduces total sampling iterations from \(128 \times 8\mathrm{k}\) to \(128 \times 4\mathrm{k}\) while maintaining competitive HPS (30.00 vs 30.37) and FID-SD (12.47 vs 12.21).
Highlights & Insights¶
- Unified Math Framework for Distillation and RL: Proves the mathematical equivalence between the multi-step DMD loss gradient and the reverse-SDE MDP policy gradient, bringing diffusion distillation natively into the standard reinforcement learning algorithmic toolkit.
- Critic-Free Variance Reduction via Mean Subtraction: Overcomes the notorious instability of score differences at high noise levels by applying group normalization without training extra Critic or value networks.
- Physics-Informed Magnitude Matching: Derives the adaptive scaling coefficient \(\beta_{\mathrm{dm}, t}\) from the closed-form residual denominator, harmonizing the optimization paces of dense diffusion distillation and sparse perceptual rewards.
Limitations & Future Work¶
- Author-Admitted Limitations: The pipeline still relies on tracking the student's output distribution using an auxiliary fake score estimator \(\mu_{\mathrm{fake}}\), which introduces extra memory consumption and training latency; the number of off-policy importance sampling iterations is constrained by the clipping boundary (\(\eta = 0.5\)).
- Open Directions: Shared timesteps within groups reduce the diversity of noise schedules per batch; extending this framework to large-scale text-to-video diffusion distillation where temporal consistency requires trajectory-level rewards represents an exciting future avenue.
Related Work & Insights¶
- vs DMD / DMD2 [44, 45]: DMD/DMD2 treats distribution matching purely as regression or GAN-based loss objectives anchored to the teacher. Rdm turns distribution matching into an intrinsic reward, achieving lower FID without real image data (Img-Free) while outperforming the teacher's aesthetic ceiling.
- vs DMDR [12]: DMDR naively sums distillation and RL loss terms, making it sensitive to weights and vulnerable to high-timestep variance. GNDMR unifies them from the ground up, providing group variance reduction and adaptive reward balancing.
- vs Flow-GRPO [19]: Flow-GRPO applies online RL to multi-step teacher models. GNDMR extends GRPO principles to 4-step distilled generators, matching the aesthetic quality of 50-step Flow-GRPO models while operating at a fraction of the inference latency.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Elegant mathematical unification of distribution matching distillation with policy gradient reinforcement learning.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive validation across SD3 and SD3.5 with fine-grained preference, fidelity, and ablation metrics.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous derivation, clean motivation, and insightful analysis of variance and dynamic weighting.
- Value: ⭐⭐⭐⭐⭐ Provides an efficient, robust, and state-of-the-art framework for real-time few-step generative modeling.