Beyond Linear Shortcuts: Rectifying Diffusion Preference Optimization with Intrinsic Generative Geometry¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/wangpin-aier/TCPO
Area: Image Generation
Keywords: diffusion models, direct preference optimization, trajectory geometry, text-to-image generation, alignment fine-tuning
TL;DR¶
Addressing the "Linear Shortcut Fallacy" in diffusion Direct Preference Optimization (DPO)—caused by ignoring the intrinsic curvature of the generative trajectory—this paper proposes Trajectory-Consistent Preference Optimization (TCPO), which dynamically anchors on-manifold geometry via Curvature-Adaptive Sampling and fuses semantic preference targets via hyperspherical interpolation, boosting visual quality and achieving a 4.3× training speedup.
Background & Motivation¶
Text-to-image diffusion models have achieved remarkable success in synthesizing high-fidelity visual content. To align these models with nuanced human aesthetics and fine-grained textual prompts, fine-tuning via Direct Preference Optimization (DPO) and its derivatives has become the dominant paradigm. Unlike reinforcement learning frameworks that rely on explicit external reward models—which often suffer from severe reward hacking and optimization instability—diffusion DPO directly leverages paired winner-loser images in the latent space to optimize the denoising network via implicit reward formulation, significantly streamlining the alignment pipeline.
However, existing diffusion DPO fine-tuning methods share a critical, unaddressed geometric blind spot: they uniformly enforce endpoint-directed supervision, driving a straight-line update direction from intermediate noisy latents directly toward fully clean target images. This linear assumption directly contradicts the intrinsic physical dynamics of pretrained diffusion models. Due to spectral evolution bias—where coarse, low-frequency scene layouts emerge substantially earlier than fine, high-frequency textural details—the true generative trajectory in latent space is inherently a curved, non-linear manifold. Forcing a direct linear shortcut (which the authors term the "Linear Shortcut Fallacy") inevitably pushes intermediate latent representations off the valid data manifold into low-density regions where pretrained score estimates are unreliable, ultimately causing structural collapse, concept distortion, and optimization instability.
To resolve this issue, fine-tuning must move beyond naive endpoint targeting and respect the model's intrinsic non-linear generation dynamics. The core idea is to propose Trajectory-Consistent Preference Optimization (TCPO), which dynamically locates on-manifold geometric anchors along DDIM-inverted trajectories via Curvature-Adaptive Sampling (CAS), and blends them with clean preference endpoints using hyperspherical interpolation (TRF) to enforce trajectory-consistent supervision on the valid data manifold.
Method¶
Overall Architecture¶
TCPO aims to align diffusion models with human preferences while strictly constraining gradient updates to remain on the pretrained data manifold. The pipeline first reconstructs the sample's generation path via deterministic DDIM inversion and identifies an intermediate geometric anchor based on the rate of change of local tangent vectors. Next, it fuses this on-manifold geometric anchor with the clean semantic endpoint using spherical linear interpolation to construct a rectified target noise vector. Finally, this trajectory-rectified target is integrated into the diffusion DPO objective for end-to-end policy updates.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Paired samples (x_w, x_l) and noisy latent x_t"] --> B["Curvature-Adaptive Sampling (CAS)<br/>Reconstruct DDIM path and locate geometric anchor x_geo"]
B --> C["Trajectory-Rectified Fusion (TRF)<br/>Fuse geometric and semantic anchors via SLERP on hypersphere"]
C --> D["Preference Optimization Objective & Backpropagation<br/>Optimize diffusion DPO likelihood ratio with rectified margin"]
D --> E["Output: Trajectory-consistent and human-aligned diffusion model"]
Key Designs¶
1. Curvature-Adaptive Sampling (CAS): Dynamically Locating Trajectory Geometric Anchors Conventional preference optimization uses the clean endpoint \(x_0\) to supervise denoising gradients, implicitly assuming that the velocity field remains constant throughout the entire time horizon—an assumption that introduces severe secant errors over curved manifolds. To maximize lookahead distance while preserving local linearity, CAS uses the discrete sequence obtained via DDIM inversion \(\mathcal{T} = \{x_{t_0}, x_{t_1}, \dots, x_{t_{n-1}}\}\) as a proxy for the generative trajectory. Let the velocity field at step \(i\) be \(v_i = \epsilon_\theta(x_{t_i}, t_i)\) and its unit tangent vector be \(u_i = v_i / \|v_i\|\). In continuous dynamics, curvature corresponds to the angular derivative \(\|du/dt\|^2\). Under uniform timestep spacing \(\Delta t\), the finite-difference approximation reveals that the squared norm of this derivative is strictly proportional to the cosine distance between consecutive velocity vectors: $\(\kappa_i = 1 - \frac{v_i^\top v_{i-1}}{\|v_i\|\|v_{i-1}\|} \in [0, 2]\)$ Equipped with this scale-invariant curvature metric, CAS abandons rigid fixed-step lookahead horizons. Instead, for a current noisy latent \(x_t\), it dynamically searches for the furthest valid anchor state \(x_{\text{geo}} := x_{t_{i^*}}\) such that the accumulated curvature along the subpath does not exceed an adaptive budget \(\eta\): $\(i^* = \min \left\{ 1 < \tau < n \;\Bigg|\; \sum_{j=\tau}^{n-1} \kappa_j \le \eta \right\}, \quad \text{where } \eta = \frac{1}{\alpha} \sum_{j=0}^{n-1} \kappa_j\)$ The hyperparameter \(\alpha\) controls the granularity of linearization. In trajectory intervals with sharp turning and high curvature, the lookahead window automatically contracts to honor the budget; in smooth, flat intervals, the window expands to provide long-range guidance while ensuring the anchor stays strictly on the intrinsic manifold.
2. Trajectory-Rectified Fusion (TRF): Hyperspherical Interpolation Balancing Geometry and Semantics Supervising exclusively with the geometric anchor \(x_{\text{geo}}\) makes the model overly conservative, merely encouraging the reconstruction of the unaligned reference model while failing to inject preference signals, and remains susceptible to accumulated inversion errors. Conversely, relying solely on the clean semantic anchor \(x_0\) (\(x_{\text{sem}}\)) repeats the linear shortcut fallacy. TRF constructs a unified target by fusing the implicit noise vectors pointing toward these two anchors: \(\epsilon_{\text{geo}} = \epsilon(x_{\text{geo}})\) and \(\epsilon_{\text{sem}} = \epsilon(x_{\text{sem}})\). High-dimensional Gaussian distributions exhibit the Concentration of Measure phenomenon, meaning noise vectors concentrate tightly on the surface of a hypersphere. Standard Euclidean linear interpolation (LERP) shrinks the norm toward the origin, pulling intermediate states off-manifold. To respect this spherical geometry, TRF performs Spherical Linear Interpolation (SLERP): $\(\epsilon_{\text{TRF}} = \text{SLERP}(\epsilon_{\text{geo}}, \epsilon_{\text{sem}}; \gamma) = \frac{\sin((1-\gamma)\Theta)}{\sin\Theta}\epsilon_{\text{geo}} + \frac{\sin(\gamma\Theta)}{\sin\Theta}\epsilon_{\text{sem}}\)$ where \(\Theta = \arccos\left(\frac{\epsilon_{\text{geo}}^\top \epsilon_{\text{sem}}}{\|\epsilon_{\text{geo}}\|\|\epsilon_{\text{sem}}\|}\right)\) is the angle between the two noise anchors, and \(\gamma \in [0, 1]\) balances geometric fidelity with preference learning. This formulation guarantees that the supervision noise vector remains on the hypersphere, combining semantic attraction toward the preferred endpoint with strong geometric anchoring back to the valid trajectory.
3. Preference Optimization Objective & Backpropagation During paired fine-tuning on winning images \(x^w\) and losing images \(x^l\), the inversion and TRF fusion processes produce rectified target noise vectors \(\epsilon_w^{\text{TRF}}\) and \(\epsilon_l^{\text{TRF}}\) for each sample. Substituting these targets into the diffusion DPO log-likelihood ratio under the mean squared error (MSE) formulation yields the end-to-end objective: $\(\mathcal{L}_{\text{TCPO}}(\theta) = -\mathbb{E}_{(x^w, x^l)\sim\mathcal{D}, t\sim\mathcal{U}[0, T]} \left[ \log \sigma \left( -\beta \left( \|\epsilon_\theta(x_t^w) - \epsilon_w^{\text{TRF}}\|^2 - \|\epsilon_\theta(x_t^l) - \epsilon_l^{\text{TRF}}\|^2 - \|\epsilon_{\text{ref}}(x_t^w) - \epsilon_w^{\text{TRF}}\|^2 + \|\epsilon_{\text{ref}}(x_t^l) - \epsilon_l^{\text{TRF}}\|^2 \right) \right) \right]\)$ This objective reweights the denoising error gradients via preference margins while evaluating errors against trajectory-consistent noise targets, widening the implicit reward gap between preferred and dispreferred outputs without destabilizing the pretrained generative geometry.
Key Experimental Results¶
Main Results¶
Experiments were conducted using Stable Diffusion XL-base-1.0 (SDXL) trained on the Pick-a-Pic v2 dataset (~851k pairs after filtering ties) across 8 NVIDIA A100 GPUs for 400 steps (effective batch size 1024). Evaluated across DrawBench, Parti-Prompts, HPDv2, and the Pick-a-Pic v2 test set against Base-SDXL, Diffusion-DPO, InPO, and MaPO using Aesthetic, PickScore, HPS, and MPS metrics.
| Dataset | Model | Aesthetic (Mean) | PickScore (Mean) | HPS (Mean) | MPS (Mean) |
|---|---|---|---|---|---|
| DrawBench | Base-SDXL | 5.5886 | 22.4952 | 28.5959 | 11.4993 |
| DrawBench | Diffusion-DPO | 5.6428 | 22.8330 | 29.0824 | 12.1187 |
| DrawBench | InPO | 5.6435 | 22.9671 | 29.1010 | 12.0145 |
| DrawBench | MaPO | 5.7401 | 22.5230 | 28.8760 | 11.6293 |
| DrawBench | TCPO (Ours) | 5.7302 | 23.0146 | 29.2670 | 12.1738 |
| Parti-Prompts | Base-SDXL | 5.7693 | 22.6274 | 28.4241 | 11.2538 |
| Parti-Prompts | Diffusion-DPO | 5.7946 | 22.9307 | 28.9085 | 11.8097 |
| Parti-Prompts | InPO | 5.8332 | 22.9952 | 28.8443 | 11.6527 |
| Parti-Prompts | MaPO | 5.8979 | 22.5946 | 28.5551 | 11.2751 |
| Parti-Prompts | TCPO (Ours) | 5.9001 | 23.1618 | 29.1819 | 11.8487 |
| HPDv2 | Base-SDXL | 6.1338 | 22.7835 | 28.6278 | 14.2736 |
| HPDv2 | Diffusion-DPO | 6.1124 | 23.1330 | 29.1650 | 14.6354 |
| HPDv2 | InPO | 6.1797 | 23.2761 | 29.2413 | 14.6645 |
| HPDv2 | MaPO | 6.2416 | 22.8488 | 28.9939 | 14.3608 |
| HPDv2 | TCPO (Ours) | 6.1917 | 23.3950 | 29.5000 | 14.8954 |
| Pick-a-Pic v2 | Base-SDXL | 6.0040 | 22.1655 | 27.9776 | 11.5371 |
| Pick-a-Pic v2 | Diffusion-DPO | 6.0127 | 22.6397 | 28.5933 | 11.9866 |
| Pick-a-Pic v2 | InPO | 6.0416 | 22.5820 | 28.5171 | 11.9247 |
| Pick-a-Pic v2 | MaPO | 6.2037 | 22.3179 | 28.5381 | 11.6570 |
| Pick-a-Pic v2 | TCPO (Ours) | 6.0744 | 22.8045 | 28.8740 | 12.2160 |
Ablation Study¶
Ablations on the Pick-a-Pic v2 test set evaluate the fusion weight \(\gamma\), interpolation scheme (SLERP vs LERP), and curvature budget factor \(\alpha\) (\(\gamma=0\) reduces to unanchored endpoint supervision; \(\alpha=1\) collapses the anchor to the endpoint):
| Config | Aesthetic (Mean) | PickScore (Mean) | HPS (Mean) | MPS (Mean) | Note |
|---|---|---|---|---|---|
| Base-SDXL | 6.0040 | 22.1655 | 27.9776 | 11.5371 | unaligned baseline |
| InPO-SDXL | 6.0416 | 22.5820 | 28.5171 | 11.9247 | strong reparameterization baseline |
| \(\gamma = 0.0\) (pure endpoint) | 6.0107 | 22.7554 | 28.6238 | 10.2044 | no geometric anchor; MPS plummets |
| \(\gamma = 0.4\) (overly geometric) | 6.0071 | 22.5730 | 28.5480 | 11.9249 | insufficient preference injection |
| w/o SLERP (standard LERP) | 6.0942 | 22.7661 | 28.8214 | 12.2125 | norm collapse toward origin; scores drop |
| \(\alpha = 1.43\) | 6.0654 | 22.8019 | 28.8297 | 12.1869 | coarse anchor segmentation; solid gains |
| \(\alpha = 3.33\) | 6.0742 | 22.7939 | 28.7977 | 12.2424 | finer geometric budget; balanced |
| TCPO Full Model (\(\alpha=2.0, \gamma=0.2\)) | 6.0744 | 22.8045 | 28.8740 | 12.2160 | optimal balance of geometry and alignment |
Training Efficiency Analysis¶
TCPO significantly improves training dynamics and gradient quality:
| Method | HPS Score (Pick-a-Pic) | A100 GPU Training Hours | Speedup vs DPO |
|---|---|---|---|
| Diffusion-DPO | 28.5017 | ~776 | 1.00× |
| Diffusion-InPO | 28.6445 | ~221 | 3.51× |
| TCPO (Ours) | 28.9950 | ~180 | 4.31× |
Key Findings¶
- Destructive Off-Manifold Drift: Setting \(\gamma=0\) leads to a catastrophic drop in MPS (from 12.2160 to 10.2044), confirming that unconstrained linear shortcuts force latents off the data manifold; conversely, \(\gamma=0.4\) stalls preference gains, showing that balance is essential.
- Superiority of Hyperspherical Interpolation: SLERP consistently outperforms standard Euclidean linear interpolation (LERP) across benchmarks (+0.05 to +0.08 in HPS and PickScore), proving that preserving the constant norm of high-dimensional Gaussian noise prevents structural degradation.
- Substantial Training Acceleration: By aligning gradient updates with the model's high-density pretrained trajectory, TCPO prevents wasteful gradient steps spent correcting off-manifold latents, converging in just 180 GPU hours (4.31× faster than DPO and 1.23× faster than InPO).
- Attribute Binding & Disentanglement: Qualitative evaluations reveal that TCPO completely eliminates subject hijacking (e.g., rendering a queen with cat traits rather than an outright cat) and resolves complex multi-entity interactions without feature bleeding.
Highlights & Insights¶
- Differential Geometry Perspective on Alignment: Unpacks the overlooked Linear Shortcut Fallacy in diffusion DPO, deriving a clean, scale-invariant discrete curvature metric from the cosine distance of velocity vectors.
- Concentration of Measure & SLERP Fusion: Grounded in high-dimensional probability geometry, TRF uses SLERP rather than naive Euclidean interpolation to avoid vector norm shrinkage toward the origin.
- High-Efficiency Post-Training Paradigm: Delivers state-of-the-art alignment across four benchmark datasets while reducing compute overhead by over 75%, making preference fine-tuning significantly more accessible.
Limitations & Future Work¶
- Reliance on DDIM Inversion Accuracy: The geometric anchor relies on a 5-step DDIM inversion trajectory; for highly noisy or intricate inputs, inversion approximation errors could slightly bias anchor selection.
- Static Global Hyperparameters: The parameters \(\alpha=2.0\) and \(\gamma=0.2\) are currently kept constant throughout the entire denoising schedule; exploring timestep-dependent adaptive schedules (e.g., higher geometric weighting at coarse stages and higher semantic weighting at fine stages) is a promising future direction.
- Evaluation on Broader Modalities: While the authors verified generalizability on Flow Matching (SD3.5) and compact architectures (SSD-1B) in the appendix, scaling to high-dimensional spatiotemporal video diffusion models remains to be explored.
Related Work & Insights¶
- vs Diffusion-DPO: Diffusion-DPO imposes an unconstrained linear shortcut from noisy latents to the clean endpoint; TCPO identifies this cause of off-manifold drift and introduces CAS and TRF to anchor updates to the curved trajectory.
- vs InPO / SPO: InPO leverages DDIM reparameterization to circumvent intractable marginal integrations; TCPO goes further by analyzing the intrinsic differential geometry of the trajectory, correcting the supervision vector via spherical interpolation and delivering superior quality and efficiency.
- Takeaway: Post-training alignment should not treat generative dynamics as a disposable black box; respecting and preserving the intrinsic geometric manifold formed during pretraining is vital for stable and efficient fine-tuning.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneers the identification of the Linear Shortcut Fallacy in diffusion DPO and devises a mathematically grounded geometric rectification framework]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across four major benchmarks, four quantitative evaluators, thorough ablations, user studies, and compute efficiency metrics]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear geometric intuition, clean mathematical derivations, and coherent narrative structure]
- Value: ⭐⭐⭐⭐⭐ [Provides an impactful, highly efficient post-training solution that improves visual quality while cutting training cost by more than 75%]