DiffRGD: An Inference-Time Diffusion Guidance Through Riemannian Gradient Descent¶
Conference: ECCV 2026
Paper: ECCV Official
Project: https://diffrgd.github.io/
Area: Image Generation
Keywords: diffusion models, inference-time guidance, Riemannian gradient descent, polar decomposition, inverse problems
TL;DR¶
Addressing sample quality degradation caused by the destruction of latent Gaussian distributions during inference-time guidance, DiffRGD leverages the polar decomposition of isotropic Gaussians to construct distribution-induced spherical manifolds, achieving high-fidelity conditional generation via Riemannian Gradient Descent.
Background & Motivation¶
Diffusion models have become the dominant foundation for both unconditional and conditional generative modeling. However, adapting pre-trained foundation models to specific downstream tasks via fine-tuning or re-training incurs prohibitive computational overhead and maintenance costs. To achieve plug-and-play controllability without modifying model weights, inference-time guidance techniques have received considerable attention. Pioneering frameworks such as DPS and FreeDoM steer the reverse denoising trajectory by directly injecting gradients computed from task-specific loss functions. Yet, these methods perform gradient descent in the unconstrained ambient Euclidean space. Due to distributional drift induced by the Jensen gap, naive gradient updates perturb the step-wise Gaussian marginal distributions, pulling latent trajectories off the true data manifold and causing severe visual artifacts and perceptual degradation. Conversely, artificially shrinking the step size to stabilize generation leads to inadequate steerability, failing to align samples with conditional inputs.
Recent attempts to mitigate this off-manifold issue exhibit substantial shortcomings. For instance, MPGD relies on a local linear manifold hypothesis and projects gradients onto the tangent space of the data manifold via a pre-trained VAE encoder; however, its reliance on an ideal encoder assumption frequently induces numerical instability. DSG proposes projecting guided latents onto a fixed-radius hypersphere, but this heuristic breaks the intrinsic radial distribution of the latent noise, effectively degrading it into a uniform distribution. Meanwhile, ADMMDiff formalizes guidance as a constrained optimization problem via the Alternating Direction Method of Multipliers, but its dependence on proximal operators and inner-loop iterations results in sluggish convergence and burdensome hyperparameter sensitivity.
Faced with this fundamental tension between sample fidelity and guidance controllability, this paper returns to the statistical and geometric principles of the diffusion process itself: each DDIM sampling step naturally forms an isotropic Gaussian distribution centered at the predicted mean. Core idea: by mathematically decoupling the isotropic Gaussian latent via polar decomposition into statistically independent chi-distributed radial and uniform directional components, each sampling step is formulated as a constrained optimization problem on a distribution-induced spherical manifold and solved efficiently via Riemannian Gradient Descent (RGD) with tangent projection and retraction.
Method¶
Overall Architecture¶
DiffRGD models inference-time guidance at each reverse sampling step \(t\) as a constrained optimization problem defined on a specific spherical manifold \(\mathcal{S}_{t,r}\). Given a conditional input \(y\) and a noisy latent \(x_t\), the pre-trained diffusion model estimates the clean sample \(\hat{x}_0(x_t, t)\) and the reverse step mean \(\mu_t\). Using the polar decomposition property, the algorithm independently samples a chi-distributed radius \(r\) and an initial spherical direction, constructing a spherical manifold \(\mathcal{S}_{t,r}\) that strictly conforms to the statistical geometry of the isotropic Gaussian. On this manifold, Euclidean gradients derived from the task guidance loss are orthogonally projected onto the tangent space to eliminate components that perturb the radial variance. The latent is then updated along the negative Riemannian gradient and pulled back to the manifold via a retraction operator, yielding the distribution-preserving guided latent \(x_{t-1}\) after \(K\) inner iterations.
The complete workflow is illustrated in the diagram below:
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: observation y and noisy latent xt"] --> B["Polar decomposition geometric modeling<br/>sample chi radius r and spherical slice St,r"]
B --> C["Orthogonal tangent space projection<br/>project Euclidean gradient to eliminate radial component"]
C --> D["Manifold retraction and convergence guarantee<br/>update along geodesic tangent and retract to sphere"]
D --> E["Output: prior-preserving guided sample xt-1"]
Key Designs¶
1. Polar decomposition geometric modeling: projecting prior constraints onto a spherical manifold To resolve the core issue where unconstrained Euclidean updates deviate latents from the prior distribution, the paper establishes the polar decomposition of an isotropic Gaussian (Proposition 1). At any diffusion timestep \(t\), the latent variable \(x_t \sim \mathcal{N}(\mu_t, \sigma_t^2 I_n)\) admits an exact decomposition into statistically independent radial and directional components: $\(x_t = \mu_t + \sigma_t r u, \quad r \sim \chi(n), \quad u \sim \text{Unif}(\mathbb{S}^{n-1})\)$ where \(r\) is a chi-distributed random variable with \(n\) degrees of freedom, and \(u\) is a unit vector distributed uniformly over the unit hypersphere \(\mathbb{S}^{n-1}\). While earlier methods like DSG enforce an empirical fixed-radius sphere that destroys the radial probability structure, DiffRGD samples a chi variable \(r\) at each reverse step, restricting the optimization to the exact feasible set centered at \(\mu_t\) with radius \(\sigma_t r\): $\(\mathcal{S}_{t,r} = \{ x \in \mathbb{R}^n \mid \| x - \mu_t \|_2 = \sigma_t r \}\)$ This design preserves the marginal probability density of the diffusion latent, reformulating guidance as identifying the optimal direction vector \(u\) that minimizes the conditioning loss while strictly retaining prior statistics.
2. Orthogonal tangent space projection: removing radial components to prevent distributional drift Once the spherical manifold \(\mathcal{S}_{t,r}\) is defined, directly applying Euclidean gradients derived from the guidance objective \(\mathcal{L}(\hat{x}_0, y)\) would alter the radial norm, compounding distributional drift across sampling steps. The tangent space \(T_x \mathcal{S}_{t,r}\) comprises all vectors orthogonal to the radial direction \(x - \mu_t\). DiffRGD establishes an orthogonal projection operator onto the tangent space using the standard Riemannian metric: $\(\Pi_{T_x \mathcal{S}_{t,r}}(g) = g - \frac{\langle g, x - \mu_t \rangle}{\| x - \mu_t \|_2^2} (x - \mu_t)\)$ Given the Euclidean gradient \(g = \nabla_{x} \mathcal{L}(\hat{x}_0(x, t-1), y)\), this operator strips away the radial component that expands or contracts the variance, preserving solely the tangent component along the manifold surface (\(\text{grad}_{\mathcal{S}_{t,r}} \mathcal{L}\)). This mechanism prevents variance distortion, ensuring that the step-wise guidance focuses exclusively on conditional alignment without compromising Gaussian fidelity.
3. Manifold retraction and convergence guarantee: ensuring strict return to the spherical manifold Because the tangent space is a flat local approximation, stepping along the tangent direction \(x - \eta_t \text{grad}_{\mathcal{S}_{t,r}} \mathcal{L}\) inevitably causes the latent to leave the curved spherical manifold. To map the updated point back onto the feasible set, DiffRGD applies a retraction operator: $\(R_x(v) = \mu_t + \sigma_t r \frac{x + v - \mu_t}{\| x + v - \mu_t \|_2}\)$ This operation displaces the latent along the tangent direction and rescales it radially to the exact radius \(\sigma_t r\), strictly confining the trajectory to \(\mathcal{S}_{t,r}\). Theoretically, assuming the objective function is \(L\)-smooth on the manifold, the authors prove that Algorithm 1 converges to a first-order stationary point at a sublinear rate of \(\mathcal{O}(1/\sqrt{k})\). In practice, setting the inner iteration count to \(K = 3\) is sufficient to diminish the Riemannian gradient norm, avoiding the complex dual updates of ADMMDiff while ensuring rapid convergence.
Key Experimental Results¶
Main Results¶
Quantitative evaluations on the FFHQ 256×256 validation set (150 images) across four standard linear inverse problems using 1,000 DDIM sampling steps (averaged across 100 bootstrap runs) demonstrate the consistent superiority of DiffRGD:
| Task | Method | Venue | PSNR ↑ | SSIM ↑ | LPIPS ↓ | FID ↓ |
|---|---|---|---|---|---|---|
| Inpainting (70% mask) | DPS | ICLR 2023 | 30.44 | 0.863 | 0.153 | 42.68 |
| MPGD | ICLR 2024 | 27.51 | 0.724 | 0.256 | 68.24 | |
| DSG | ICML 2024 | 31.03 | 0.866 | 0.144 | 36.30 | |
| ADMMDiff | CVPR 2025 | 32.38 | 0.899 | 0.119 | 29.52 | |
| DiffRGD (Ours) | ECCV 2026 | 34.04 | 0.926 | 0.096 | 21.88 | |
| Super-Resolution 4× | DPS | ICLR 2023 | 26.03 | 0.727 | 0.260 | 80.05 |
| MPGD | ICLR 2024 | 24.40 | 0.614 | 0.354 | 101.50 | |
| DSG | ICML 2024 | 26.71 | 0.737 | 0.256 | 74.67 | |
| ADMMDiff | CVPR 2025 | 26.48 | 0.712 | 0.297 | 96.69 | |
| DiffRGD (Ours) | ECCV 2026 | 27.77 | 0.783 | 0.220 | 63.94 | |
| Gaussian Deblurring | DPS | ICLR 2023 | 25.88 | 0.721 | 0.237 | 69.38 |
| MPGD | ICLR 2024 | 24.07 | 0.576 | 0.328 | 95.12 | |
| DSG | ICML 2024 | 27.45 | 0.751 | 0.259 | 75.85 | |
| ADMMDiff | CVPR 2025 | 26.57 | 0.757 | 0.226 | 79.30 | |
| DiffRGD (Ours) | ECCV 2026 | 26.80 | 0.757 | 0.218 | 63.83 | |
| Motion Deblurring | DPS | ICLR 2023 | 24.47 | 0.685 | 0.271 | 80.75 |
| MPGD | ICLR 2024 | 23.15 | 0.569 | 0.357 | 106.99 | |
| DSG | ICML 2024 | 26.80 | 0.709 | 0.290 | 87.96 | |
| ADMMDiff | CVPR 2025 | 27.26 | 0.778 | 0.222 | 72.92 | |
| DiffRGD (Ours) | ECCV 2026 | 25.84 | 0.736 | 0.250 | 72.26 |
On conditional human face generation tasks on CelebA-HQ 256×256 with 100 DDIM steps, DiffRGD achieves state-of-the-art alignment while preserving perceptual quality:
| Task / Metric | FreeDoM (ICCV 2023) | DSG (ICML 2024) | ADMMDiff (CVPR 2025) | DiffRGD (Ours) |
|---|---|---|---|---|
| Segmentation Map mIoU ↑ | 0.622 | 0.750 | 0.758 | 0.804 |
| Segmentation Map FID ↓ | 156.02 | 117.48 | 101.86 | 96.10 |
| Sketch Guidance \(\ell_2\) ↓ | 30.85 | 21.36 | 30.82 | 19.48 |
| Sketch Guidance FID ↓ | 101.90 | 107.00 | 97.52 | 87.82 |
| FaceID Guidance \(\ell_2\) ↓ | 0.557 | 0.340 | 0.346 | 0.303 |
| FaceID Guidance FID ↓ | 127.05 | 95.27 | 100.81 | 93.80 |
Ablation Study¶
The choice of guidance strength \(\eta_t\) is critical in inference-time steering. The table below analyzes the sensitivity of DSG and DiffRGD across varying guidance scales under segmentation map guidance on CelebA-HQ:
| Method | Guidance Strength \(\eta_t\) | Alignment mIoU ↑ | Quality FID ↓ | Observation & Analysis |
|---|---|---|---|---|
| DSG | 0.05 | 0.53 | 95.69 | Weak condition adherence at small step size |
| 0.10 | 0.76 | 110.00 | Improved control accompanied by degrading FID | |
| 0.20 | 0.90 | 157.82 | Drifts off manifold, generating severe artifacts | |
| 0.30 | 0.93 | 248.77 | Total sample collapse due to unchecked distortion | |
| DiffRGD (Ours) | 10 | 0.74 | 80.73 | High perceptual fidelity with solid baseline control |
| 30 | 0.82 | 89.03 | Optimal balance of visual quality and alignment | |
| 50 | 0.84 | 93.47 | Strong structural adherence with well-controlled FID | |
| 100 | 0.86 | 108.33 | Stable even under extreme guidance strengths |
Key Findings¶
- Robustness under few sampling steps: When reducing DDIM steps from 1,000 to 100, methods like DPS and ADMMDiff suffer severe performance degradation. DiffRGD maintains superior metrics under 100 steps (e.g., reaching 31.16 dB PSNR on ImageNet 256×256 inpainting compared to 27.97 dB for ADMMDiff).
- Broad guidance strength tolerance: As shown in the ablation study, DSG exhibits catastrophic failure when \(\eta_t\) reaches 0.2 (FID rises to 157.82). Thanks to Riemannian projection and retraction, DiffRGD remains exceptionally stable across order-of-magnitude parameter scaling (\(\eta_t = 10\) to \(100\)), with FID rising only modestly to 108.33.
- Fidelity trade-off at higher scaling factors: In super-resolution tasks with Stable Diffusion (8× to 12×), 12× upsampling synthesizes richer fine textures, yielding higher SSIM (0.709 vs 0.700), while minor pixel misalignments with ground truth lead to a slight drop in PSNR (26.39 dB vs 26.41 dB).
Highlights & Insights¶
- Theoretical alignment of Gaussian statistics and Riemannian geometry: By leveraging the polar decomposition of isotropic Gaussian random variables, DiffRGD elegantly maps the complex distributional constraint to a tractable spherical Riemannian manifold with full theoretical convergence guarantees.
- Computational simplicity through analytic projection and retraction: The method avoids nested proximal optimizations, relying strictly on standard inner products and vector normalization, reaching effective stationarity within just \(K=3\) iterations.
- Universal plug-and-play adaptability: Requiring no auxiliary networks or structural re-training, DiffRGD seamlessly integrates into any pre-trained diffusion model following DDIM sampling schedules.
Limitations & Future Work¶
- Reliance on isotropic Gaussian assumptions: The mathematical formulation assumes isotropic Gaussian marginals at each reverse step in DDIM. Extending this framework to non-isotropic noise schedules or Continuous Flow Matching models requires reformulating the underlying geometric manifold.
- Suboptimal performance in severe motion blur: In severe motion deblurring, DiffRGD achieves 25.84 dB PSNR, trailing ADMMDiff (27.26 dB), indicating that purely local first-order Riemannian gradients may face limitations when inverting highly non-local degradation operators.
- Future directions: Integrating adaptive Riemannian momentum optimization and generalizing the manifold constraint framework to Diffusion Transformers (DiTs) and Flow Matching architectures represent promising avenues.
Related Work & Insights¶
- vs DPS (ICLR 2023) / FreeDoM (ICCV 2023): DPS and FreeDoM apply unconstrained gradients in ambient Euclidean space, leading to off-manifold artifacts and hyperparameter fragility; DiffRGD eliminates off-manifold drift via tangent projection and retraction onto the Gaussian-induced manifold.
- vs MPGD (ICLR 2024): MPGD attempts data manifold projection using an external VAE encoder, which introduces instability and local linearity errors; DiffRGD operates directly on the latent reverse sampling manifold without auxiliary networks.
- vs DSG (ICML 2024): DSG fixes a heuristic spherical radius that degrades the radial distribution into a uniform distribution; DiffRGD accurately preserves Gaussian statistics by sampling radial lengths via the chi distribution.
- vs ADMMDiff (CVPR 2025): ADMMDiff applies operator splitting but suffers from slow proximal convergence and multiple sensitive hyperparameters; DiffRGD converges rapidly in \(K=3\) steps via lightweight Riemannian operations.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneering use of isotropic Gaussian polar decomposition to formulate diffusion guidance on spherical Riemannian manifolds]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Comprehensive benchmarking across 4 restoration tasks, 3 conditional generation tasks, and multiple resolutions]
- Writing Quality: ⭐⭐⭐⭐⭐ [Rigorous theoretical derivations, intuitive conceptual visualizations, and clearly structured experimental presentation]
- Value: ⭐⭐⭐⭐⭐ [Provides an elegant, plug-and-play paradigm for mitigating off-manifold artifacts in training-free diffusion control]