Skip to content

Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching

Conference: ECCV 2026
Paper: ECCV Official
Area: Video Generation
Keywords: Video Generation, Distribution Matching, Knowledge Distillation, Preference Alignment, RL-free Alignment

TL;DR

Proposes DM-Align, a unified single-stage sample-guided distribution matching framework that aggregates few-step distillation gradients and preference alignment gradients within a shared score space, completely circumventing multi-step sampling bottlenecks and post-distillation model collapse.

Background & Motivation

Visual generative models based on diffusion and flow matching architectures have achieved substantial progress, yet prolonged inference latency and alignment with human aesthetic preferences remain pivotal bottlenecks toward practical deployment. Existing efforts broadly bifurcate into two distinct research tracks: knowledge distillation, which compresses multi-step iterative denoising down to few-step (e.g., 4-step) generation, and human preference alignment, which typically leverages Reinforcement Learning (RL) techniques such as DPO and GRPO to fine-tune generative policies according to human evaluative criteria. Crucially, prevailing pipelines treat distillation and alignment as disconnected, sequential stages, inevitably precipitating an architectural dilemma.

Under the "alignment-before-distillation" paradigm, executing RL algorithms directly on full-step diffusion models incurs staggering computational overhead, requiring repetitive multi-step sampling along sampling paths for every optimization step. Fundamentally, adapting continuous Ordinary Differential Equation (ODE) generative trajectories into the discrete Markov Decision Process (MDP) formulation of standard RL algorithms requires complex reverse-time Stochastic Differential Equation (SDE) transformations and restrictive Gaussian assumptions to approximate trajectory log-probabilities. Conversely, adopting an "alignment-after-distillation" sequence to exploit accelerated sampling leads to severe instability: fine-tuning an intensively compressed few-step generator with sparse, noisy RL reward updates promptly ruptures its delicate continuous generative manifold, yielding visual artifacts, flickering motions, and catastrophic model collapse.

Departing from the conventional paradigm of forcing diffusion processes into rigid RL frameworks, this work examines the problem from the perspective of Distribution Matching (DM). Under the DM objective, distillation gradients and preference alignment gradients exhibit native compatibility in the score domain, as both can be rigorously cast as distribution-matching projections toward target manifolds. The core idea is: unify distillation and preference alignment into a single-stage sample-guided distribution matching framework (DM-Align), utilizing a shared dynamic fake model to provide robust score-field estimations that synergistically aggregate fidelity distillation gradients and preference guidance without any RL overhead.

Method

Overall Architecture

The optimization pipeline of DM-Align coordinates knowledge distillation and human preference alignment within a unified single-stage flow. The architecture comprises three interacting modules: a trainable few-step video generator \(G_\theta\), a frozen pre-trained teacher model \(\mu_{\text{real}}\) delivering natural high-fidelity visual priors, and an online continuously trained denoising fake model \(\mu^\phi_{\text{fake}}\) that tracks the generator's evolving output distribution and supplies robust score estimations.

During each training cycle, optimization alternates across two phases: in the fake model update phase, the generator produces few-step candidate samples, and the fake model minimizes a denoising score matching loss across arbitrary forward noise timesteps \(t\) to maintain precise score tracking over the generator manifold; in the generator update phase, the system simultaneously computes the distillation gradient \(\nabla \mathcal{L}_{\text{DMD}}\) against the teacher model and the sample-guided alignment gradient \(\nabla \mathcal{L}_{\text{align}}\) derived from human preference anchors or group exploration. These complementary vectors are linearly aggregated into a composite gradient \(\nabla \mathcal{L}_G\) to update generator parameters \(\theta\) in a single backward pass.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input text prompt c and Gaussian noise z"] --> B["Few-step Generator G_θ<br/>4-step denoising forward rollout"]
    B --> C["Fake Model Dynamic Update: Score Field Tracking<br/>Continuous denoising score matching on μ_fake"]
    C --> D["Distribution Matching Distillation: Fidelity Guidance<br/>DMD gradient via score gap between real teacher and fake model"]
    C --> E["Sample-Guided Alignment: Human Preference Guidance<br/>DM-Align gradient via preferred samples and fake score field"]
    D --> F["Synergistic Gradient Aggregation<br/>Weighted composite update of generator parameters θ"]
    E --> F
    F --> G["High-fidelity 4-step aligned video output"]

Key Designs

1. Fake Model Dynamic Update: Score Field Tracking

Conventional RL formulations convert generative trajectories into MDP state transitions, requiring cumbersome mathematical derivations to track trajectory probabilities via reverse-time SDEs. DM-Align completely eliminates discrete MDP abstractions by operating directly in continuous score space. To this end, the framework maintains a dedicated denoising fake model \(\mu_{\text{fake}}^\phi\) structurally mirroring the generator. In each iteration, the generator generates video samples \(x_0 = G_\theta(z)\) using few-step sampling, which are treated with a stop-gradient operation. By applying forward Gaussian perturbation to create overlapping support manifolds \(x_t = \alpha_t x_0 + \sigma_t \epsilon\), the fake model is updated via dense Mean Squared Error (MSE) supervision:

\[ \mathcal{L}_{\phi}^{\text{denoise}} = \mathbb{E}_{t, \epsilon} \left\| \mu_{\text{fake}}^\phi(x_t, t) - x_0 \right\|_2^2 \]

By virtue of Tweedie's formula, an optimally trained denoiser yields an unbiased estimator of the perturbed score function \(s_{\text{fake}}(x_t, t) = \nabla_{x_t} \log p_{\theta, t}(x_t)\). Consequently, the fake model acts as a dynamic high-dimensional vector field analyzer, serving simultaneously as the repulsive reference manifold for distillation and as the geometric metric projecting preference samples onto high-reward regimes, avoiding out-of-distribution variance typical of external RL critic networks.

2. Distribution Matching Distillation: Fidelity Guidance

To compress the multi-step diffusion trajectory into ultra-fast inference (e.g., 4 steps), DM-Align builds upon Distribution Matching Distillation (DMD) principles, minimizing the Kullback-Leibler (KL) divergence between the generator's distribution \(p_{\text{fake}}\) and the pre-trained teacher's distribution \(p_{\text{real}}\): \(\mathcal{D}_{\text{KL}}(p_{\text{fake}} \parallel p_{\text{real}})\). Because calculating exact probability densities is intractable in high-dimensional video space, the algorithm computes an analytical gradient by evaluating the score disparity between the frozen teacher \(\mu_{\text{real}}\) and the online fake model \(\mu_{\text{fake}}\) in noisy latent space:

\[ \nabla_\theta \mathcal{L}_{\text{DMD}} \approx - \mathbb{E}_{t, z} \left[ \left( s_{\text{real}}(F(G_\theta(z), t), t) - s_{\text{fake}}(F(G_\theta(z), t), t) \right) \frac{d G_\theta(z)}{d \theta} \right] \]

The operational mechanism is intuitive: \(s_{\text{real}}\) pulls generated latents toward the coherent, high-fidelity visual manifold captured by the teacher, while \(-s_{\text{fake}}\) penalizes blurriness and mode collapse in the current generator distribution. In early training phases where outputs are coarse, this distillation gradient naturally dominates optimization, establishing solid physical structures and spatial coherence that prevent the catastrophic collapse observed in post-distillation RL fine-tuning.

3. Sample-Guided Alignment: Human Preference Guidance

Under the canonical RLHF paradigm adhering to the Bradley-Terry preference model, the reward-maximizing policy constrained by reference regularization \(p_{\text{ref}}\) admits the closed-form density \(p^*(x) \propto p_{\text{ref}}(x) \exp(r(x) / \beta)\). Through variational reparameterization, aligning the generator equates to minimizing \(\mathcal{D}_{\text{KL}}(p_\theta \parallel p^*)\), yielding an exact update rule situated squarely in the score domain:

\[ \nabla_\theta \mathcal{L}_{\text{align}} \approx - \mathbb{E}_{t, z} \left[ \left( \nabla_{x_t} \log p_t^*(x_t) - \nabla_{x_t} \log p_{\theta, t}(x_t) \right) \frac{d G_\theta(z)}{d \theta} \right] \]

Since the exact target gradient \(\nabla_{x_t} \log p_t^*(x_t)\) cannot be evaluated in closed form, DM-Align leverages the fake model's score response on perturbed high-reward samples as an implicit manifold projection. This insight materializes as two complementary strategies: - DM-PairLoss (Pairwise Preference Anchor): Designed for paired high-quality video datasets (e.g., ConsistID), where ground-truth preferred videos \(x^+\) anchor the target manifold and generator rollouts act as negative counterparts: \(\nabla \mathcal{L}_{\text{align}} \approx - \mathbb{E} [ ( s_{\text{fake}}(F(x^+, t), t) - s_{\text{fake}}(F(G_\theta(z), t), t) ) \frac{d G_\theta}{d \theta} ]\). This constructs a direct contrastive gradient steering generation toward human-curated anchors. - DM-GroupLoss (Reward-Guided Group Exploration): Designed for open-ended prompt corpora (e.g., VidProM), where the generator samples \(K\) outputs \(\{x_1, \dots, x_K\}\) per prompt, scored by an external reward model to obtain relative advantages \(A_i\). Taking the geometric mean of the Top-\(k\) candidates as the dynamic anchor \(\bar{x}^+\), the group gradient is formulated as:

\[ \nabla_\theta \mathcal{L}_{\text{DM-Group}} \approx - \mathbb{E}_{t, z} \left[ \sum_{i=1}^K A_i \left( s_{\text{fake}}(F(x_i, t), t) - s_{\text{fake}}(F(\bar{x}^+, t), t) \right) \frac{d G_\theta(z)}{d \theta} \right] \]

Suboptimal outputs (\(A_i < 0\)) are pulled toward the high-quality group mean, while superior rollouts (\(A_i > 0\)) actively push probability mass toward peak preference regions, achieving adaptive exploration and alignment in a single step.

4. Synergistic Gradient Aggregation: Single-Stage Optimization

To ensure mutually reinforcing convergence, DM-Align linearly aggregates both objectives within the generator's latent space:

\[ \nabla \mathcal{L}_G = \lambda_{\text{DMD}} \nabla \mathcal{L}_{\text{DMD}} + \lambda_{\text{align}} \nabla \mathcal{L}_{\text{align}} \]

Generated samples \(x_{\text{final}}\) simultaneously serve as the negative distribution for distillation and as the evaluative inputs for preference alignment. Both components reuse intermediate feature representations and the fake model's score estimator, preventing distributional drift caused by multi-stage transitions. Distillation secures structural fidelity, while alignment provides vivid dynamic motion and fine-grained textual compliance, achieving superior generation quality at minimal compute.

Loss & Training

The framework is optimized with AdamW on 32 NVIDIA H100 GPUs using the Wan2.1-T2V-1.3B foundation model. Learning rates are set to \(2 \times 10^{-6}\) for generator \(G_\theta\) and \(4 \times 10^{-7}\) for fake model \(\mu_{\text{fake}}\). The fake model updates at every iteration, while the generator updates every \(N=5\) steps to stabilize score tracking. Distillation uses a 4-step schedule over timesteps \([1000, 750, 500, 250]\) with teacher classifier-free guidance (CFG) scale 6.0. Loss weights default to \(\lambda_{\text{DMD}} = \lambda_{\text{align}} = 0.5\). DM-Group uses group size \(K=8\) and averages the top-4 samples. The entire pipeline converges robustly within 1,000 steps without requiring a dedicated DMD-only warm-up phase.

Key Experimental Results

Main Results

Quantitative evaluations are conducted across two setups: pairwise preference alignment on ConsistID using DM-Align (Pair), and group reward alignment on VidProM prompts using DM-Align (Group). Metrics follow the standardized VBench suite, evaluating temporal consistency, motion quality (Motion Smoothness, Dynamic Degree), and visual quality (Aesthetic Quality, Imaging Quality).

Dataset / Setting Method NFE Average Score Dynamic Degree Aesthetic Quality Imaging Quality
ConsistID Dataset
(Pairwise Alignment)
Wan-T2V-1.3B (Raw Base) 100 78.20 50.78 56.09 72.29
ConsistID DMD2 (Distillation Baseline) 4 79.68 52.81 57.64 74.73
ConsistID Flow-DPO (Standalone RL) 100 81.54 67.13 56.32 72.91
ConsistID Flow-DPO + DMD2 (Sequential Pipeline) 4 79.89 53.43 58.26 74.35
ConsistID DM-Align (Pair) (Ours) 4 82.78 66.25 60.89 75.61
VidProM Dataset
(Group Reward Alignment)
Wan-T2V-1.3B (Raw Base) 100 77.85 56.40 57.33 66.20
VidProM DMD2 (Distillation Baseline) 4 78.88 51.64 63.02 69.30
VidProM DanceGRPO (Standalone RL) 100 82.76 75.78 60.44 70.15
VidProM DanceGRPO + DMD2 (Sequential Pipeline) 4 80.54 60.85 60.71 70.81
VidProM DM-Align (Group) (Ours) 4 84.40 80.19 62.17 71.45

Ablation Study

1. Training Efficiency Comparison The table below records sampling NFE and per-step wall-clock time on 32 NVIDIA H100 GPUs.

Optimization Method Sampling NFE per Step Time per Step (s) Acceleration & Characteristics
DanceGRPO (Flow-based RL) 1536 423 s Baseline (Prohibitive sampling over multi-step paths)
DMD2 (Standard DMD) 46 11 s ~38.5× speedup, lacks preference alignment
DM-Align (Ours) 51 15 s ~28.2× speedup, full alignment with only 4s overhead

2. Loss Weight Sensitivity (\(\lambda_{\text{DMD}} : \lambda_{\text{align}}\)) on CogVideoX-2B and Wan-1.3B Evaluates Textual Alignment (TA) and aesthetic scores (HPSv2) under varying balance ratios where \(\lambda_{\text{DMD}} + \lambda_{\text{align}} = 1\).

Ratio (\(\lambda_{\text{DMD}} : \lambda_{\text{align}}\)) Wan-1.3B TA (↑) Wan-1.3B HPSv2 (↑) CogVideoX TA (↑) CogVideoX HPSv2 (↑) Observations & Analysis
1:0 (Pure DMD) 0.75 20.50 0.22 18.71 Lacks alignment; poor text adherence and aesthetics
3:1 (Distillation Heavy) 1.34 24.01 0.58 20.92 Robust structural stability with noticeable gains
1:1 (Ours Default) 1.65 28.89 0.55 25.01 Peak performance across both models and metrics
1:3 (Alignment Heavy) 1.56 12.37 0.08 13.78 Insufficient distillation breaks base generation quality
Varying (Linear 1:0 to 1:3) 1.45 27.56 0.49 22.11 Annealing schedule shows high competitive stability

Key Findings

  • Elimination of Post-Distillation Collapse: Two-stage sequential pipelines suffer severe degradation. DanceGRPO achieves 75.78 Dynamic Degree at 100 steps, but subsequent DMD2 distillation causes it to plummet to 60.85 (-14.93). In contrast, DM-Align's unified single-stage optimization boosts Dynamic Degree to 80.19, fully resolving objective misalignment.
  • Nearly 30× Training Speedup: Bypassing reverse-time SDE trajectory simulation reduces optimization step time from 423 seconds to 15 seconds, rendering video alignment computationally tractable on standard clusters.
  • MoE Architectural Scalability: On the 14B Mixture-of-Experts Wan2.2-T2V-A14B model, applying DM-Align exclusively to the high-noise expert improves 4-step VBench score from 78.35 (DMD) to 83.45, outperforming the 100-step raw foundation model (80.00).

Highlights & Insights

  • Score-Domain Unification: Reveals that distillation and Bradley-Terry preference alignment share an identical score-matching formulation, unifying two disparate disciplines under a single denoising vector field.
  • Bypassing Discrete MDP Formulations: Avoids converting continuous flow matching into discrete RL MDPs, sidestepping trajectory probability estimation and SDE-ODE conversion errors.
  • Generalizable Alignment Strategy: The high-noise expert tuning strategy and dual-mode formulation (Pair/Group) provide a practical blueprint for low-cost human alignment across diverse video diffusion architectures.

Limitations & Future Work

  • Lack of Non-Convex Convergence Proofs: While single-objective gradients are variationally grounded, dynamic convergence of the weighted dual-objective formulation warrants deeper theoretical study.
  • Dependency on DMD Architecture: The formulation requires a continuously trained fake model, making it incompatible with pure discriminator-based one-step GAN architectures.
  • Video Reward Vulnerabilities: Existing video reward models frequently over-score high-frequency noise or blank frames under rapid motions, requiring more robust 3D spatio-temporal reward models.
  • vs Flow-DPO / DanceGRPO: Conventional RL requires 20-40 sampling steps per update and exhibits severe instability; DM-Align operates in score space, cutting training overhead by ~28× while preventing collapse.
  • vs DMD2 / FlashDMD / DMDR: DMD2 lacks human alignment; concurrent works like DMDR directly concatenate standard RL losses with distillation; DM-Align natively formulates preference matching within distribution matching, achieving superior prompt compliance and motion dynamics.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ First framework demonstrating gradient compatibility between distillation and preference alignment in score space.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Extensive evaluations on Wan2.1, Wan2.2 MoE, and CogVideoX across automated VBench metrics and human GSB studies.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear theoretical exposition, transparent mathematical formulations, and compelling empirical validations.
  • Value: ⭐⭐⭐⭐⭐ Provides an efficient, robust, single-stage solution for rapid 4-step video generation aligned with human preferences.