D2PO: Optimizing Diffusion Samplers via Dynamic Preference¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Area: Alignment & RLHF
Keywords: diffusion models, sampler optimization, direct preference optimization, energy-based model, dynamic preference
TL;DR¶
D2PO reformulates diffusion sampler parameter optimization from a rigid student-teacher regression task into a preference alignment problem using an energy-based surrogate and dynamic self-refining teachers, markedly enhancing perceptual fidelity under ultra-low NFE budgets.
Background & Motivation¶
Diffusion Probabilistic Models (DPMs) have achieved unprecedented synthesis quality across high-resolution image and video generation, but their iterative reverse denoising demands numerous function evaluations (NFE), incurring substantial computational costs that impede real-time deployment. Beyond accelerated numerical differential equation solvers and few-step distillation, directly optimizing the sampling policy parametersโspecifically the discrete timestep schedule \(S\) and per-step classifier-free guidance (CFG) weights \(\omega\)โpresents a lightweight, orthogonal acceleration frontier. However, existing sampler optimization techniques encounter a severe methodological bottleneck. Population-level search objectives (e.g., optimizing FID or KID over large batches) suffer from high gradient variance and provide weak training signals for low-dimensional sampler parameters. Conversely, instance-wise distillation methods (such as LD3) force a low-NFE student sampler to regress onto fixed high-NFE teacher trajectories using \(\ell_2\) or LPIPS losses.
This regression paradigm exhibits an inherent structural limitation when aggressive speedups are pursuedโthat is, when the gap between student and teacher NFE widens. When constrained by a minimal step budget, forcing the low-NFE student to trace the trajectory of an unyielding high-NFE teacher compromises generation fidelity. In practice, the student prioritizes coarse global consistency at the expense of high-frequency textural fidelity. Empirically, holding the student budget at 4 steps while scaling the teacher's steps actually exacerbates visual artifacts and blurriness, demonstrating that regressing onto an external fixed teacher enforces an insurmountable residual error floor.
To overcome this performance bottleneck, the optimization of sampling parameters is recast as a preference-based alignment problem rather than a trajectory regression task. The core idea is to model the deterministic sampler as a smooth energy-based surrogate parameterized by multi-scale score discrepancies, and guide it through a dynamic self-improving teacher that adaptively refines timesteps to minimize the sampler's own discretization error.
Method¶
Overall Architecture¶
D2PO (Dynamic Direct Preference Optimization) optimizes the low-dimensional sampler parameters \(\phi = \{S, \omega\}\) while keeping the pre-trained diffusion model backbone strictly frozen. Because deterministic numerical ODE solvers produce degenerate Dirac delta output distributions that preclude direct log-likelihood gradient estimation, D2PO introduces an energy-based model (EBM) surrogate. The energy is formulated directly through the multi-scale score field learned by the pre-trained diffusion network. During optimization, instead of mimicking an external immutable teacher, D2PO dynamically constructs winning samples by evaluating the current policy on an interpolated, denser timestep schedule (\(2N\) steps), paired with negative samples produced by low-pass filtering the student's output. Under a unified DPO objective, the sampler iteratively drives down its own numerical truncation error.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Initial Noise & Prompt<br/>x_T ~ N(0, I), c"] --> B["Energy-Based Surrogate<br/>converts deterministic ODE to smooth EBM"]
B --> C["Score-Based Energy Metric<br/>multi-scale weighted denoising score residual"]
C --> D["Dynamic Teacher Mechanism<br/>generates 2x refined higher-fidelity winner"]
D --> E["Dynamic Preference Optimization<br/>DPO objective enforcing discretization error decay"]
E --> F["Optimized Sampler Policy<br/>discrete schedule S & step CFG weights omega"]
Key Designs¶
1. Energy-Based Surrogate: enabling differentiable preference optimization for deterministic ODE solvers
Given that a deterministic ODE solver produces an output \(x_\phi = \Phi_\phi(x_T, c; \theta)\) whose conditional probability density collapses to a non-differentiable Dirac delta function \(q_\phi(x|x_T, c) = \delta(x - x_\phi)\), standard likelihood-based DPO cannot be directly applied. D2PO bypasses this degeneracy by formulating a smooth surrogate policy \(\pi_\phi(x|c, x_T) = \frac{1}{Z(\phi, c, x_T)} \exp(-\alpha E(x; \phi, c, x_T))\), where the energy is defined as the distance to the generated output \(E(x; \phi, c, x_T) \equiv d(x, x_\phi)\). Because candidate images share the identical initial noise \(x_T\) and condition \(c\), the intractable partition function \(Z\) cancels out identically in the log-probability ratio, transforming the ill-posed preference alignment problem into a tractable, differentiable energy-difference classification objective.
2. Score-Based Energy Metric: leveraging multi-scale manifold geometry for high-frequency detail
Conventional \(\ell_2\) or LPIPS perceptual metrics lack sensitivity to subtle numerical truncation errors along the reverse trajectory. D2PO derives an intrinsic energy metric directly from the pre-trained score network \(s_\theta(x_t, t)\), which approximates \(\nabla_{x_t} \log p_t(x_t)\). To resolve the numerical explosion of score differences as \(t \to 0\) (caused by the \(\sigma_t^{-2}\) scaling in score-to-noise conversion), a weighting function \(w(t) = \sigma_t^2\) is applied, yielding a balanced, multi-scale noise-prediction distance over continuous noise levels:
Large noise levels regularize global semantic geometry, whereas small noise levels preserve high-frequency textural details, supplying a substantially more informative alignment signal than generic perceptual losses.
3. Dynamic Teacher Mechanism: overcoming the static teacher residual error floor via self-improvement
In standard teacher-student distillation, a static teacher policy \(\pi^{fix}\) imposes a fixed asymptotic error bound \(\epsilon^{fix}\) below which the student cannot improve. D2PO introduces a dynamic teacher whose policy \(\pi_{\phi'}\) is generated on the fly from the current student parameters \(\phi\). Specifically, given an \(N\)-step schedule \(S\), a refined \(2N\)-step schedule \(S'\) is constructed via intermediate linear interpolation and solved using the same ODE solver to yield the winning candidate \(x_w\). Meanwhile, the losing candidate \(x_l\) is synthesized by passing the student output through a degradation operator (e.g., a low-pass filter). For the reference sampler \(\phi_{ref}\), its timestep schedule \(S_{ref}\) is periodically copied from the student at the end of each epoch, while CFG weights \(\omega_{ref}\) are updated at each step via an exponential moving average. Theoretically, for a solver of order \(k\), the dynamic teacher satisfies \(\epsilon_{\phi'}^{true} \approx 2^{-k} \epsilon_\phi^{true}\), guaranteeing that the dynamic training loss forms a rigorous upper bound on the student's true error that monotonically vanishes as the policy refines.
Loss & Training¶
Combining the energy-based surrogate with the score-based distance, the practical D2PO objective optimizes the expectation over uniform noise timesteps \(t \sim \mathcal{U}(0, T)\):
where the differential term is:
Only the low-dimensional sampler schedule and CFG parameters \(\phi = \{S, \omega\}\) receive gradients during training, leaving all base model weights untouched.
Key Experimental Results¶
Main Results¶
Quantitative evaluation on text-to-image synthesis using Stable Diffusion v1.5 across 30k MS-COCO prompts shows that D2PO consistently outperforms state-of-the-art discretization methods (DMN, GITS, and LD3) across three representative high-order ODE solvers:
| Solver | Steps | Method | HPSv2 โ | Aesthetic โ | FID โ |
|---|---|---|---|---|---|
| iPNDM | 4 | DMN | 0.2030 | 5.0936 | 21.39 |
| iPNDM | 4 | GITS | 0.2128 | 5.1413 | 18.12 |
| iPNDM | 4 | LD3 | 0.2191 | 5.1756 | 17.60 |
| iPNDM | 4 | D2PO | 0.2237 | 5.2024 | 15.69 |
| iPNDM | 5 | LD3 | 0.2346 | 5.2463 | 13.59 |
| iPNDM | 5 | D2PO | 0.2385 | 5.2701 | 13.38 |
| UniPC | 4 | LD3 | 0.2180 | 5.1755 | 18.34 |
| UniPC | 4 | D2PO | 0.2185 | 5.1761 | 16.97 |
| DPM-Solver++ | 4 | LD3 | 0.2191 | 5.1736 | 17.46 |
| DPM-Solver++ | 4 | D2PO | 0.2216 | 5.1854 | 16.84 |
On ImageNet-256 (latent space, 3rd-order iPNDM), D2PO achieves an FID of 7.28 at 4 steps, significantly outperforming LD3 (9.19) and uniform spacing (13.86). When evaluated on the flow-matching model InstaFlow at 2 steps, D2PO achieves an FID of 40.68 and an Aesthetic score of 5.0924, surpassing both uniform sampling (44.27 / 5.0060) and LD3 (63.55 / 4.6270).
Ablation Study¶
Component ablations on Stable Diffusion v1.5 with the iPNDM solver demonstrate the critical role of each design element:
| Steps | Config | Aesthetic โ | FID โ | Note |
|---|---|---|---|---|
| 4 | D2PO (Full) | 5.2024 | 15.69 | full model |
| 4 | w/o dynamic preference (fixed \(N+1\) teacher) | 5.1810 | 16.91 | bounded by static error floor |
| 4 | w/o score-based energy (replaced with LPIPS) | 5.1796 | 17.88 | largest performance drop |
| 4 | w/ EMA reference timestep | 5.1937 | 15.75 | continuous smoothing destabilizes step boundaries |
| 5 | D2PO (Full) | 5.2701 | 13.38 | full model |
| 5 | w/o dynamic preference (fixed \(N+1\) teacher) | 5.2615 | 13.70 | ceiling limited |
| 5 | w/o score-based energy (replaced with LPIPS) | 5.2630 | 14.59 | texture fidelity degraded |
| 5 | w/ EMA reference timestep | 5.2639 | 13.42 | minor schedule oscillation |
Key Findings¶
- Score-based energy metric is foundational: Replacing the score-based energy formulation with LPIPS causes the sharpest performance deterioration across all metrics (e.g., FID worsens from 15.69 to 17.88 at 4 steps), proving that multi-scale score field differences capture high-frequency numerical errors that standard perceptual features miss.
- Regime transition across step budgets: At 4-5 steps, D2PO acts primarily as an effective discretization error corrector, dominating all baselines on both FID and perceptual metrics. At 6-7 steps, as truncation errors naturally recede, D2PO shifts seamlessly toward perceptual preference alignment, maintaining peak Aesthetic/HPSv2 scores while exhibiting expected quality-diversity trade-offs.
- Asynchronous update stability: While exponential moving averages suit continuous CFG weights, direct epoch-wise copying of discrete timestep schedules provides a far more stable reference anchor, avoiding harmful gradient flutter across discrete step locations.
Highlights & Insights¶
- Extending DPO to deterministic generative solvers: By reformulating deterministic ODE integration through an energy-based surrogate, D2PO allows direct preference optimization without requiring stochastic policies or density approximations.
- Self-refinement replaces external static supervision: Constructing dynamic winning targets via \(2N\) interpolated steps allows the model to continuously minimize its own discretization errors, mathematically dismantling the performance ceiling imposed by fixed teacher regression.
- Frozen-backbone efficiency: Achieving dramatic perceptual fidelity and FID gains solely by optimizing a compact set of schedule and guidance parameters avoids costly denoiser fine-tuning and preserves the underlying model distribution.
Limitations & Future Work¶
- Online computational cost of dynamic teachers: Generating \(2N\)-step winning candidates on the fly requires real-time solver evaluations during training, which is slower than training against pre-computed offline static targets.
- Heuristic negative sample construction: The current reliance on low-pass filtering to synthesize negative samples may not fully encapsulate complex structural failures such as unnatural anatomy or prompt-image misalignment.
- Expansion to broader solver parameters: The framework can be extended to jointly optimize high-order solver integration coefficients and multi-step extrapolation factors alongside timesteps and CFG weights.
Related Work & Insights¶
- vs LD3 (Learning to Discretize Denoising Diffusion ODEs): LD3 enforces relaxed matching regression against a fixed high-NFE teacher, suffering from severe blurring when the compression gap widens; D2PO eliminates the static error floor by combining preference alignment with dynamic self-improving teachers.
- vs Diffusion-DPO & Direct Preference Fine-Tuning: Traditional preference alignment for diffusion models updates large denoiser backbones (U-Net or DiT), risking catastrophic degradation; D2PO keeps all weights frozen, optimizing only the low-dimensional sampling policy as an orthogonal post-training acceleration method.
Rating¶
- Novelty: โญโญโญโญโญ Pioneering adaptation of DPO to diffusion sampler schedule optimization via energy-based surrogates and dynamic self-refinement.
- Experimental Thoroughness: โญโญโญโญโญ Validated across 3 high-order ODE solvers, Latent Diffusion, Flow Matching, and multiple benchmarks with formal theoretical bounds.
- Writing Quality: โญโญโญโญโญ Elegant mathematical derivations, sharp identification of regression bottlenecks, and thorough ablation analyses.
- Value: โญโญโญโญโญ Provides an efficient, training-free-backbone framework for ultra-fast diffusion inference with significant practical and academic utility.