Skip to content

Beyond Linear Shortcuts: Rectifying Diffusion Preference Optimization with Intrinsic Generative Geometry

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/wangpin-aier/TCPO
Area: Image Generation
Keywords: diffusion models, direct preference optimization, trajectory geometry, text-to-image generation, alignment fine-tuning

TL;DR

Addressing the "Linear Shortcut Fallacy" in diffusion Direct Preference Optimization (DPO)—caused by ignoring the intrinsic curvature of the generative trajectory—this paper proposes Trajectory-Consistent Preference Optimization (TCPO), which dynamically anchors on-manifold geometry via Curvature-Adaptive Sampling and fuses semantic preference targets via hyperspherical interpolation, boosting visual quality and achieving a 4.3× training speedup.

Background & Motivation

Text-to-image diffusion models have achieved remarkable success in synthesizing high-fidelity visual content. To align these models with nuanced human aesthetics and fine-grained textual prompts, fine-tuning via Direct Preference Optimization (DPO) and its derivatives has become the dominant paradigm. Unlike reinforcement learning frameworks that rely on explicit external reward models—which often suffer from severe reward hacking and optimization instability—diffusion DPO directly leverages paired winner-loser images in the latent space to optimize the denoising network via implicit reward formulation, significantly streamlining the alignment pipeline.

However, existing diffusion DPO fine-tuning methods share a critical, unaddressed geometric blind spot: they uniformly enforce endpoint-directed supervision, driving a straight-line update direction from intermediate noisy latents directly toward fully clean target images. This linear assumption directly contradicts the intrinsic physical dynamics of pretrained diffusion models. Due to spectral evolution bias—where coarse, low-frequency scene layouts emerge substantially earlier than fine, high-frequency textural details—the true generative trajectory in latent space is inherently a curved, non-linear manifold. Forcing a direct linear shortcut (which the authors term the "Linear Shortcut Fallacy") inevitably pushes intermediate latent representations off the valid data manifold into low-density regions where pretrained score estimates are unreliable, ultimately causing structural collapse, concept distortion, and optimization instability.

To resolve this issue, fine-tuning must move beyond naive endpoint targeting and respect the model's intrinsic non-linear generation dynamics. The core idea is to propose Trajectory-Consistent Preference Optimization (TCPO), which dynamically locates on-manifold geometric anchors along DDIM-inverted trajectories via Curvature-Adaptive Sampling (CAS), and blends them with clean preference endpoints using hyperspherical interpolation (TRF) to enforce trajectory-consistent supervision on the valid data manifold.

Method

Overall Architecture

TCPO aims to align diffusion models with human preferences while strictly constraining gradient updates to remain on the pretrained data manifold. The pipeline first reconstructs the sample's generation path via deterministic DDIM inversion and identifies an intermediate geometric anchor based on the rate of change of local tangent vectors. Next, it fuses this on-manifold geometric anchor with the clean semantic endpoint using spherical linear interpolation to construct a rectified target noise vector. Finally, this trajectory-rectified target is integrated into the diffusion DPO objective for end-to-end policy updates.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Paired samples (x_w, x_l) and noisy latent x_t"] --> B["Curvature-Adaptive Sampling (CAS)<br/>Reconstruct DDIM path and locate geometric anchor x_geo"]
    B --> C["Trajectory-Rectified Fusion (TRF)<br/>Fuse geometric and semantic anchors via SLERP on hypersphere"]
    C --> D["Preference Optimization Objective & Backpropagation<br/>Optimize diffusion DPO likelihood ratio with rectified margin"]
    D --> E["Output: Trajectory-consistent and human-aligned diffusion model"]

Key Designs

1. Curvature-Adaptive Sampling (CAS): Dynamically Locating Trajectory Geometric Anchors Conventional preference optimization uses the clean endpoint \(x_0\) to supervise denoising gradients, implicitly assuming that the velocity field remains constant throughout the entire time horizon—an assumption that introduces severe secant errors over curved manifolds. To maximize lookahead distance while preserving local linearity, CAS uses the discrete sequence obtained via DDIM inversion \(\mathcal{T} = \{x_{t_0}, x_{t_1}, \dots, x_{t_{n-1}}\}\) as a proxy for the generative trajectory. Let the velocity field at step \(i\) be \(v_i = \epsilon_\theta(x_{t_i}, t_i)\) and its unit tangent vector be \(u_i = v_i / \|v_i\|\). In continuous dynamics, curvature corresponds to the angular derivative \(\|du/dt\|^2\). Under uniform timestep spacing \(\Delta t\), the finite-difference approximation reveals that the squared norm of this derivative is strictly proportional to the cosine distance between consecutive velocity vectors: $\(\kappa_i = 1 - \frac{v_i^\top v_{i-1}}{\|v_i\|\|v_{i-1}\|} \in [0, 2]\)$ Equipped with this scale-invariant curvature metric, CAS abandons rigid fixed-step lookahead horizons. Instead, for a current noisy latent \(x_t\), it dynamically searches for the furthest valid anchor state \(x_{\text{geo}} := x_{t_{i^*}}\) such that the accumulated curvature along the subpath does not exceed an adaptive budget \(\eta\): $\(i^* = \min \left\{ 1 < \tau < n \;\Bigg|\; \sum_{j=\tau}^{n-1} \kappa_j \le \eta \right\}, \quad \text{where } \eta = \frac{1}{\alpha} \sum_{j=0}^{n-1} \kappa_j\)$ The hyperparameter \(\alpha\) controls the granularity of linearization. In trajectory intervals with sharp turning and high curvature, the lookahead window automatically contracts to honor the budget; in smooth, flat intervals, the window expands to provide long-range guidance while ensuring the anchor stays strictly on the intrinsic manifold.

2. Trajectory-Rectified Fusion (TRF): Hyperspherical Interpolation Balancing Geometry and Semantics Supervising exclusively with the geometric anchor \(x_{\text{geo}}\) makes the model overly conservative, merely encouraging the reconstruction of the unaligned reference model while failing to inject preference signals, and remains susceptible to accumulated inversion errors. Conversely, relying solely on the clean semantic anchor \(x_0\) (\(x_{\text{sem}}\)) repeats the linear shortcut fallacy. TRF constructs a unified target by fusing the implicit noise vectors pointing toward these two anchors: \(\epsilon_{\text{geo}} = \epsilon(x_{\text{geo}})\) and \(\epsilon_{\text{sem}} = \epsilon(x_{\text{sem}})\). High-dimensional Gaussian distributions exhibit the Concentration of Measure phenomenon, meaning noise vectors concentrate tightly on the surface of a hypersphere. Standard Euclidean linear interpolation (LERP) shrinks the norm toward the origin, pulling intermediate states off-manifold. To respect this spherical geometry, TRF performs Spherical Linear Interpolation (SLERP): $\(\epsilon_{\text{TRF}} = \text{SLERP}(\epsilon_{\text{geo}}, \epsilon_{\text{sem}}; \gamma) = \frac{\sin((1-\gamma)\Theta)}{\sin\Theta}\epsilon_{\text{geo}} + \frac{\sin(\gamma\Theta)}{\sin\Theta}\epsilon_{\text{sem}}\)$ where \(\Theta = \arccos\left(\frac{\epsilon_{\text{geo}}^\top \epsilon_{\text{sem}}}{\|\epsilon_{\text{geo}}\|\|\epsilon_{\text{sem}}\|}\right)\) is the angle between the two noise anchors, and \(\gamma \in [0, 1]\) balances geometric fidelity with preference learning. This formulation guarantees that the supervision noise vector remains on the hypersphere, combining semantic attraction toward the preferred endpoint with strong geometric anchoring back to the valid trajectory.

3. Preference Optimization Objective & Backpropagation During paired fine-tuning on winning images \(x^w\) and losing images \(x^l\), the inversion and TRF fusion processes produce rectified target noise vectors \(\epsilon_w^{\text{TRF}}\) and \(\epsilon_l^{\text{TRF}}\) for each sample. Substituting these targets into the diffusion DPO log-likelihood ratio under the mean squared error (MSE) formulation yields the end-to-end objective: $\(\mathcal{L}_{\text{TCPO}}(\theta) = -\mathbb{E}_{(x^w, x^l)\sim\mathcal{D}, t\sim\mathcal{U}[0, T]} \left[ \log \sigma \left( -\beta \left( \|\epsilon_\theta(x_t^w) - \epsilon_w^{\text{TRF}}\|^2 - \|\epsilon_\theta(x_t^l) - \epsilon_l^{\text{TRF}}\|^2 - \|\epsilon_{\text{ref}}(x_t^w) - \epsilon_w^{\text{TRF}}\|^2 + \|\epsilon_{\text{ref}}(x_t^l) - \epsilon_l^{\text{TRF}}\|^2 \right) \right) \right]\)$ This objective reweights the denoising error gradients via preference margins while evaluating errors against trajectory-consistent noise targets, widening the implicit reward gap between preferred and dispreferred outputs without destabilizing the pretrained generative geometry.

Key Experimental Results

Main Results

Experiments were conducted using Stable Diffusion XL-base-1.0 (SDXL) trained on the Pick-a-Pic v2 dataset (~851k pairs after filtering ties) across 8 NVIDIA A100 GPUs for 400 steps (effective batch size 1024). Evaluated across DrawBench, Parti-Prompts, HPDv2, and the Pick-a-Pic v2 test set against Base-SDXL, Diffusion-DPO, InPO, and MaPO using Aesthetic, PickScore, HPS, and MPS metrics.

Dataset Model Aesthetic (Mean) PickScore (Mean) HPS (Mean) MPS (Mean)
DrawBench Base-SDXL 5.5886 22.4952 28.5959 11.4993
DrawBench Diffusion-DPO 5.6428 22.8330 29.0824 12.1187
DrawBench InPO 5.6435 22.9671 29.1010 12.0145
DrawBench MaPO 5.7401 22.5230 28.8760 11.6293
DrawBench TCPO (Ours) 5.7302 23.0146 29.2670 12.1738
Parti-Prompts Base-SDXL 5.7693 22.6274 28.4241 11.2538
Parti-Prompts Diffusion-DPO 5.7946 22.9307 28.9085 11.8097
Parti-Prompts InPO 5.8332 22.9952 28.8443 11.6527
Parti-Prompts MaPO 5.8979 22.5946 28.5551 11.2751
Parti-Prompts TCPO (Ours) 5.9001 23.1618 29.1819 11.8487
HPDv2 Base-SDXL 6.1338 22.7835 28.6278 14.2736
HPDv2 Diffusion-DPO 6.1124 23.1330 29.1650 14.6354
HPDv2 InPO 6.1797 23.2761 29.2413 14.6645
HPDv2 MaPO 6.2416 22.8488 28.9939 14.3608
HPDv2 TCPO (Ours) 6.1917 23.3950 29.5000 14.8954
Pick-a-Pic v2 Base-SDXL 6.0040 22.1655 27.9776 11.5371
Pick-a-Pic v2 Diffusion-DPO 6.0127 22.6397 28.5933 11.9866
Pick-a-Pic v2 InPO 6.0416 22.5820 28.5171 11.9247
Pick-a-Pic v2 MaPO 6.2037 22.3179 28.5381 11.6570
Pick-a-Pic v2 TCPO (Ours) 6.0744 22.8045 28.8740 12.2160

Ablation Study

Ablations on the Pick-a-Pic v2 test set evaluate the fusion weight \(\gamma\), interpolation scheme (SLERP vs LERP), and curvature budget factor \(\alpha\) (\(\gamma=0\) reduces to unanchored endpoint supervision; \(\alpha=1\) collapses the anchor to the endpoint):

Config Aesthetic (Mean) PickScore (Mean) HPS (Mean) MPS (Mean) Note
Base-SDXL 6.0040 22.1655 27.9776 11.5371 unaligned baseline
InPO-SDXL 6.0416 22.5820 28.5171 11.9247 strong reparameterization baseline
\(\gamma = 0.0\) (pure endpoint) 6.0107 22.7554 28.6238 10.2044 no geometric anchor; MPS plummets
\(\gamma = 0.4\) (overly geometric) 6.0071 22.5730 28.5480 11.9249 insufficient preference injection
w/o SLERP (standard LERP) 6.0942 22.7661 28.8214 12.2125 norm collapse toward origin; scores drop
\(\alpha = 1.43\) 6.0654 22.8019 28.8297 12.1869 coarse anchor segmentation; solid gains
\(\alpha = 3.33\) 6.0742 22.7939 28.7977 12.2424 finer geometric budget; balanced
TCPO Full Model (\(\alpha=2.0, \gamma=0.2\)) 6.0744 22.8045 28.8740 12.2160 optimal balance of geometry and alignment

Training Efficiency Analysis

TCPO significantly improves training dynamics and gradient quality:

Method HPS Score (Pick-a-Pic) A100 GPU Training Hours Speedup vs DPO
Diffusion-DPO 28.5017 ~776 1.00×
Diffusion-InPO 28.6445 ~221 3.51×
TCPO (Ours) 28.9950 ~180 4.31×

Key Findings

  • Destructive Off-Manifold Drift: Setting \(\gamma=0\) leads to a catastrophic drop in MPS (from 12.2160 to 10.2044), confirming that unconstrained linear shortcuts force latents off the data manifold; conversely, \(\gamma=0.4\) stalls preference gains, showing that balance is essential.
  • Superiority of Hyperspherical Interpolation: SLERP consistently outperforms standard Euclidean linear interpolation (LERP) across benchmarks (+0.05 to +0.08 in HPS and PickScore), proving that preserving the constant norm of high-dimensional Gaussian noise prevents structural degradation.
  • Substantial Training Acceleration: By aligning gradient updates with the model's high-density pretrained trajectory, TCPO prevents wasteful gradient steps spent correcting off-manifold latents, converging in just 180 GPU hours (4.31× faster than DPO and 1.23× faster than InPO).
  • Attribute Binding & Disentanglement: Qualitative evaluations reveal that TCPO completely eliminates subject hijacking (e.g., rendering a queen with cat traits rather than an outright cat) and resolves complex multi-entity interactions without feature bleeding.

Highlights & Insights

  • Differential Geometry Perspective on Alignment: Unpacks the overlooked Linear Shortcut Fallacy in diffusion DPO, deriving a clean, scale-invariant discrete curvature metric from the cosine distance of velocity vectors.
  • Concentration of Measure & SLERP Fusion: Grounded in high-dimensional probability geometry, TRF uses SLERP rather than naive Euclidean interpolation to avoid vector norm shrinkage toward the origin.
  • High-Efficiency Post-Training Paradigm: Delivers state-of-the-art alignment across four benchmark datasets while reducing compute overhead by over 75%, making preference fine-tuning significantly more accessible.

Limitations & Future Work

  • Reliance on DDIM Inversion Accuracy: The geometric anchor relies on a 5-step DDIM inversion trajectory; for highly noisy or intricate inputs, inversion approximation errors could slightly bias anchor selection.
  • Static Global Hyperparameters: The parameters \(\alpha=2.0\) and \(\gamma=0.2\) are currently kept constant throughout the entire denoising schedule; exploring timestep-dependent adaptive schedules (e.g., higher geometric weighting at coarse stages and higher semantic weighting at fine stages) is a promising future direction.
  • Evaluation on Broader Modalities: While the authors verified generalizability on Flow Matching (SD3.5) and compact architectures (SSD-1B) in the appendix, scaling to high-dimensional spatiotemporal video diffusion models remains to be explored.
  • vs Diffusion-DPO: Diffusion-DPO imposes an unconstrained linear shortcut from noisy latents to the clean endpoint; TCPO identifies this cause of off-manifold drift and introduces CAS and TRF to anchor updates to the curved trajectory.
  • vs InPO / SPO: InPO leverages DDIM reparameterization to circumvent intractable marginal integrations; TCPO goes further by analyzing the intrinsic differential geometry of the trajectory, correcting the supervision vector via spherical interpolation and delivering superior quality and efficiency.
  • Takeaway: Post-training alignment should not treat generative dynamics as a disposable black box; respecting and preserving the intrinsic geometric manifold formed during pretraining is vital for stable and efficient fine-tuning.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneers the identification of the Linear Shortcut Fallacy in diffusion DPO and devises a mathematically grounded geometric rectification framework]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across four major benchmarks, four quantitative evaluators, thorough ablations, user studies, and compute efficiency metrics]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear geometric intuition, clean mathematical derivations, and coherent narrative structure]
  • Value: ⭐⭐⭐⭐⭐ [Provides an impactful, highly efficient post-training solution that improves visual quality while cutting training cost by more than 75%]