Skip to content

DiffRGD: An Inference-Time Diffusion Guidance Through Riemannian Gradient Descent

Conference: ECCV 2026
Paper: ECCV Official
Project: https://diffrgd.github.io/
Area: Image Generation
Keywords: diffusion models, inference-time guidance, Riemannian gradient descent, polar decomposition, inverse problems

TL;DR

Addressing sample quality degradation caused by the destruction of latent Gaussian distributions during inference-time guidance, DiffRGD leverages the polar decomposition of isotropic Gaussians to construct distribution-induced spherical manifolds, achieving high-fidelity conditional generation via Riemannian Gradient Descent.

Background & Motivation

Diffusion models have become the dominant foundation for both unconditional and conditional generative modeling. However, adapting pre-trained foundation models to specific downstream tasks via fine-tuning or re-training incurs prohibitive computational overhead and maintenance costs. To achieve plug-and-play controllability without modifying model weights, inference-time guidance techniques have received considerable attention. Pioneering frameworks such as DPS and FreeDoM steer the reverse denoising trajectory by directly injecting gradients computed from task-specific loss functions. Yet, these methods perform gradient descent in the unconstrained ambient Euclidean space. Due to distributional drift induced by the Jensen gap, naive gradient updates perturb the step-wise Gaussian marginal distributions, pulling latent trajectories off the true data manifold and causing severe visual artifacts and perceptual degradation. Conversely, artificially shrinking the step size to stabilize generation leads to inadequate steerability, failing to align samples with conditional inputs.

Recent attempts to mitigate this off-manifold issue exhibit substantial shortcomings. For instance, MPGD relies on a local linear manifold hypothesis and projects gradients onto the tangent space of the data manifold via a pre-trained VAE encoder; however, its reliance on an ideal encoder assumption frequently induces numerical instability. DSG proposes projecting guided latents onto a fixed-radius hypersphere, but this heuristic breaks the intrinsic radial distribution of the latent noise, effectively degrading it into a uniform distribution. Meanwhile, ADMMDiff formalizes guidance as a constrained optimization problem via the Alternating Direction Method of Multipliers, but its dependence on proximal operators and inner-loop iterations results in sluggish convergence and burdensome hyperparameter sensitivity.

Faced with this fundamental tension between sample fidelity and guidance controllability, this paper returns to the statistical and geometric principles of the diffusion process itself: each DDIM sampling step naturally forms an isotropic Gaussian distribution centered at the predicted mean. Core idea: by mathematically decoupling the isotropic Gaussian latent via polar decomposition into statistically independent chi-distributed radial and uniform directional components, each sampling step is formulated as a constrained optimization problem on a distribution-induced spherical manifold and solved efficiently via Riemannian Gradient Descent (RGD) with tangent projection and retraction.

Method

Overall Architecture

DiffRGD models inference-time guidance at each reverse sampling step \(t\) as a constrained optimization problem defined on a specific spherical manifold \(\mathcal{S}_{t,r}\). Given a conditional input \(y\) and a noisy latent \(x_t\), the pre-trained diffusion model estimates the clean sample \(\hat{x}_0(x_t, t)\) and the reverse step mean \(\mu_t\). Using the polar decomposition property, the algorithm independently samples a chi-distributed radius \(r\) and an initial spherical direction, constructing a spherical manifold \(\mathcal{S}_{t,r}\) that strictly conforms to the statistical geometry of the isotropic Gaussian. On this manifold, Euclidean gradients derived from the task guidance loss are orthogonally projected onto the tangent space to eliminate components that perturb the radial variance. The latent is then updated along the negative Riemannian gradient and pulled back to the manifold via a retraction operator, yielding the distribution-preserving guided latent \(x_{t-1}\) after \(K\) inner iterations.

The complete workflow is illustrated in the diagram below:

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: observation y and noisy latent xt"] --> B["Polar decomposition geometric modeling<br/>sample chi radius r and spherical slice St,r"]
    B --> C["Orthogonal tangent space projection<br/>project Euclidean gradient to eliminate radial component"]
    C --> D["Manifold retraction and convergence guarantee<br/>update along geodesic tangent and retract to sphere"]
    D --> E["Output: prior-preserving guided sample xt-1"]

Key Designs

1. Polar decomposition geometric modeling: projecting prior constraints onto a spherical manifold To resolve the core issue where unconstrained Euclidean updates deviate latents from the prior distribution, the paper establishes the polar decomposition of an isotropic Gaussian (Proposition 1). At any diffusion timestep \(t\), the latent variable \(x_t \sim \mathcal{N}(\mu_t, \sigma_t^2 I_n)\) admits an exact decomposition into statistically independent radial and directional components: $\(x_t = \mu_t + \sigma_t r u, \quad r \sim \chi(n), \quad u \sim \text{Unif}(\mathbb{S}^{n-1})\)$ where \(r\) is a chi-distributed random variable with \(n\) degrees of freedom, and \(u\) is a unit vector distributed uniformly over the unit hypersphere \(\mathbb{S}^{n-1}\). While earlier methods like DSG enforce an empirical fixed-radius sphere that destroys the radial probability structure, DiffRGD samples a chi variable \(r\) at each reverse step, restricting the optimization to the exact feasible set centered at \(\mu_t\) with radius \(\sigma_t r\): $\(\mathcal{S}_{t,r} = \{ x \in \mathbb{R}^n \mid \| x - \mu_t \|_2 = \sigma_t r \}\)$ This design preserves the marginal probability density of the diffusion latent, reformulating guidance as identifying the optimal direction vector \(u\) that minimizes the conditioning loss while strictly retaining prior statistics.

2. Orthogonal tangent space projection: removing radial components to prevent distributional drift Once the spherical manifold \(\mathcal{S}_{t,r}\) is defined, directly applying Euclidean gradients derived from the guidance objective \(\mathcal{L}(\hat{x}_0, y)\) would alter the radial norm, compounding distributional drift across sampling steps. The tangent space \(T_x \mathcal{S}_{t,r}\) comprises all vectors orthogonal to the radial direction \(x - \mu_t\). DiffRGD establishes an orthogonal projection operator onto the tangent space using the standard Riemannian metric: $\(\Pi_{T_x \mathcal{S}_{t,r}}(g) = g - \frac{\langle g, x - \mu_t \rangle}{\| x - \mu_t \|_2^2} (x - \mu_t)\)$ Given the Euclidean gradient \(g = \nabla_{x} \mathcal{L}(\hat{x}_0(x, t-1), y)\), this operator strips away the radial component that expands or contracts the variance, preserving solely the tangent component along the manifold surface (\(\text{grad}_{\mathcal{S}_{t,r}} \mathcal{L}\)). This mechanism prevents variance distortion, ensuring that the step-wise guidance focuses exclusively on conditional alignment without compromising Gaussian fidelity.

3. Manifold retraction and convergence guarantee: ensuring strict return to the spherical manifold Because the tangent space is a flat local approximation, stepping along the tangent direction \(x - \eta_t \text{grad}_{\mathcal{S}_{t,r}} \mathcal{L}\) inevitably causes the latent to leave the curved spherical manifold. To map the updated point back onto the feasible set, DiffRGD applies a retraction operator: $\(R_x(v) = \mu_t + \sigma_t r \frac{x + v - \mu_t}{\| x + v - \mu_t \|_2}\)$ This operation displaces the latent along the tangent direction and rescales it radially to the exact radius \(\sigma_t r\), strictly confining the trajectory to \(\mathcal{S}_{t,r}\). Theoretically, assuming the objective function is \(L\)-smooth on the manifold, the authors prove that Algorithm 1 converges to a first-order stationary point at a sublinear rate of \(\mathcal{O}(1/\sqrt{k})\). In practice, setting the inner iteration count to \(K = 3\) is sufficient to diminish the Riemannian gradient norm, avoiding the complex dual updates of ADMMDiff while ensuring rapid convergence.

Key Experimental Results

Main Results

Quantitative evaluations on the FFHQ 256×256 validation set (150 images) across four standard linear inverse problems using 1,000 DDIM sampling steps (averaged across 100 bootstrap runs) demonstrate the consistent superiority of DiffRGD:

Task Method Venue PSNR ↑ SSIM ↑ LPIPS ↓ FID ↓
Inpainting (70% mask) DPS ICLR 2023 30.44 0.863 0.153 42.68
MPGD ICLR 2024 27.51 0.724 0.256 68.24
DSG ICML 2024 31.03 0.866 0.144 36.30
ADMMDiff CVPR 2025 32.38 0.899 0.119 29.52
DiffRGD (Ours) ECCV 2026 34.04 0.926 0.096 21.88
Super-Resolution 4× DPS ICLR 2023 26.03 0.727 0.260 80.05
MPGD ICLR 2024 24.40 0.614 0.354 101.50
DSG ICML 2024 26.71 0.737 0.256 74.67
ADMMDiff CVPR 2025 26.48 0.712 0.297 96.69
DiffRGD (Ours) ECCV 2026 27.77 0.783 0.220 63.94
Gaussian Deblurring DPS ICLR 2023 25.88 0.721 0.237 69.38
MPGD ICLR 2024 24.07 0.576 0.328 95.12
DSG ICML 2024 27.45 0.751 0.259 75.85
ADMMDiff CVPR 2025 26.57 0.757 0.226 79.30
DiffRGD (Ours) ECCV 2026 26.80 0.757 0.218 63.83
Motion Deblurring DPS ICLR 2023 24.47 0.685 0.271 80.75
MPGD ICLR 2024 23.15 0.569 0.357 106.99
DSG ICML 2024 26.80 0.709 0.290 87.96
ADMMDiff CVPR 2025 27.26 0.778 0.222 72.92
DiffRGD (Ours) ECCV 2026 25.84 0.736 0.250 72.26

On conditional human face generation tasks on CelebA-HQ 256×256 with 100 DDIM steps, DiffRGD achieves state-of-the-art alignment while preserving perceptual quality:

Task / Metric FreeDoM (ICCV 2023) DSG (ICML 2024) ADMMDiff (CVPR 2025) DiffRGD (Ours)
Segmentation Map mIoU ↑ 0.622 0.750 0.758 0.804
Segmentation Map FID ↓ 156.02 117.48 101.86 96.10
Sketch Guidance \(\ell_2\) ↓ 30.85 21.36 30.82 19.48
Sketch Guidance FID ↓ 101.90 107.00 97.52 87.82
FaceID Guidance \(\ell_2\) ↓ 0.557 0.340 0.346 0.303
FaceID Guidance FID ↓ 127.05 95.27 100.81 93.80

Ablation Study

The choice of guidance strength \(\eta_t\) is critical in inference-time steering. The table below analyzes the sensitivity of DSG and DiffRGD across varying guidance scales under segmentation map guidance on CelebA-HQ:

Method Guidance Strength \(\eta_t\) Alignment mIoU ↑ Quality FID ↓ Observation & Analysis
DSG 0.05 0.53 95.69 Weak condition adherence at small step size
0.10 0.76 110.00 Improved control accompanied by degrading FID
0.20 0.90 157.82 Drifts off manifold, generating severe artifacts
0.30 0.93 248.77 Total sample collapse due to unchecked distortion
DiffRGD (Ours) 10 0.74 80.73 High perceptual fidelity with solid baseline control
30 0.82 89.03 Optimal balance of visual quality and alignment
50 0.84 93.47 Strong structural adherence with well-controlled FID
100 0.86 108.33 Stable even under extreme guidance strengths

Key Findings

  • Robustness under few sampling steps: When reducing DDIM steps from 1,000 to 100, methods like DPS and ADMMDiff suffer severe performance degradation. DiffRGD maintains superior metrics under 100 steps (e.g., reaching 31.16 dB PSNR on ImageNet 256×256 inpainting compared to 27.97 dB for ADMMDiff).
  • Broad guidance strength tolerance: As shown in the ablation study, DSG exhibits catastrophic failure when \(\eta_t\) reaches 0.2 (FID rises to 157.82). Thanks to Riemannian projection and retraction, DiffRGD remains exceptionally stable across order-of-magnitude parameter scaling (\(\eta_t = 10\) to \(100\)), with FID rising only modestly to 108.33.
  • Fidelity trade-off at higher scaling factors: In super-resolution tasks with Stable Diffusion (8× to 12×), 12× upsampling synthesizes richer fine textures, yielding higher SSIM (0.709 vs 0.700), while minor pixel misalignments with ground truth lead to a slight drop in PSNR (26.39 dB vs 26.41 dB).

Highlights & Insights

  • Theoretical alignment of Gaussian statistics and Riemannian geometry: By leveraging the polar decomposition of isotropic Gaussian random variables, DiffRGD elegantly maps the complex distributional constraint to a tractable spherical Riemannian manifold with full theoretical convergence guarantees.
  • Computational simplicity through analytic projection and retraction: The method avoids nested proximal optimizations, relying strictly on standard inner products and vector normalization, reaching effective stationarity within just \(K=3\) iterations.
  • Universal plug-and-play adaptability: Requiring no auxiliary networks or structural re-training, DiffRGD seamlessly integrates into any pre-trained diffusion model following DDIM sampling schedules.

Limitations & Future Work

  • Reliance on isotropic Gaussian assumptions: The mathematical formulation assumes isotropic Gaussian marginals at each reverse step in DDIM. Extending this framework to non-isotropic noise schedules or Continuous Flow Matching models requires reformulating the underlying geometric manifold.
  • Suboptimal performance in severe motion blur: In severe motion deblurring, DiffRGD achieves 25.84 dB PSNR, trailing ADMMDiff (27.26 dB), indicating that purely local first-order Riemannian gradients may face limitations when inverting highly non-local degradation operators.
  • Future directions: Integrating adaptive Riemannian momentum optimization and generalizing the manifold constraint framework to Diffusion Transformers (DiTs) and Flow Matching architectures represent promising avenues.
  • vs DPS (ICLR 2023) / FreeDoM (ICCV 2023): DPS and FreeDoM apply unconstrained gradients in ambient Euclidean space, leading to off-manifold artifacts and hyperparameter fragility; DiffRGD eliminates off-manifold drift via tangent projection and retraction onto the Gaussian-induced manifold.
  • vs MPGD (ICLR 2024): MPGD attempts data manifold projection using an external VAE encoder, which introduces instability and local linearity errors; DiffRGD operates directly on the latent reverse sampling manifold without auxiliary networks.
  • vs DSG (ICML 2024): DSG fixes a heuristic spherical radius that degrades the radial distribution into a uniform distribution; DiffRGD accurately preserves Gaussian statistics by sampling radial lengths via the chi distribution.
  • vs ADMMDiff (CVPR 2025): ADMMDiff applies operator splitting but suffers from slow proximal convergence and multiple sensitive hyperparameters; DiffRGD converges rapidly in \(K=3\) steps via lightweight Riemannian operations.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneering use of isotropic Gaussian polar decomposition to formulate diffusion guidance on spherical Riemannian manifolds]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Comprehensive benchmarking across 4 restoration tasks, 3 conditional generation tasks, and multiple resolutions]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Rigorous theoretical derivations, intuitive conceptual visualizations, and clearly structured experimental presentation]
  • Value: ⭐⭐⭐⭐⭐ [Provides an elegant, plug-and-play paradigm for mitigating off-manifold artifacts in training-free diffusion control]