Spectral Prior for Reducing Exposure Bias in Diffusion Models¶
Conference: ECCV 2026
Paper: ECCV 2026
Area: Image Generation
Keywords: Diffusion Models, Exposure Bias, Spectral Prior, Guided Sampling, Frequency-dependent SNR
TL;DR¶
To alleviate exposure bias (distributional shift) in multi-step diffusion sampling that manifests as frequency- and timestep-dependent SNR distortions, this paper proposes Spectral Alignment (SPA)—a method that fits an offline power-law spectral prior and applies lightweight inference-time FFT-based guidance, yielding consistent generation gains across pixel, latent, and flow-matching models with only 3–4% computational overhead.
Background & Motivation¶
Diffusion models achieve state-of-the-art visual synthesis through iterative reverse denoising. However, during sampling, small prediction errors at each timestep accumulate, progressively shifting the intermediate trajectory away from the true marginal distributions encountered during forward diffusion training—a failure mode termed exposure bias. Prior research primarily attributed exposure bias to a mismatch between the pre-scheduled noise levels and the empirical noise levels during inference, leading to remedies such as predicted-noise rescaling (\(\epsilon\)-rescaling), time-shifted sampling schedules, and wavelet-based frequency reweighting. Nevertheless, these strategies either assume uniform scalar variance shifts across all spatial frequencies or rely on heuristic, hand-crafted rules that fail to generalize across differing image resolutions and sampling budgets.
A deeper analysis of the intermediate clean prediction \(\hat{x}_{0|t}\) reveals that exposure bias fundamentally corresponds to frequency-dependent effective SNR errors. Crucially, empirical investigation demonstrates that the direction and structure of this spectral discrepancy differ sharply across architectures, latent representations, and timesteps: for instance, ADM exhibits severe high-frequency attenuation, Stable Diffusion 2.0 suffers from low-frequency power loss, while modern models like SDXL, SD3.5, and FLUX display intricate, channel-wise, time-varying distortions. This wide architectural diversity indicates that fixed correction heuristics (such as globally boosting high frequencies or uniform scaling) cannot universally succeed, creating an acute demand for a data-driven framework capable of characterizing the ideal spectral profile for any given model and correcting it adaptively.
This paper tackles the challenge by steering the power spectrum of the single-step Tweedie prediction toward an error-free reference power spectrum derived from the forward process. Core idea: model the radially averaged power spectrum of single-step forward predictions as a continuous timestep-dependent power-law prior, and calibrate inference trajectories via an asymmetric spectral guidance loss computed with efficient inverse FFTs without extra neural network evaluations.
Method¶
Overall Architecture¶
Spectral Alignment (SPA) is structured into two decoupled stages: offline target power spectrum modeling (executed once per model) and inference-time spectral alignment guidance. In Stage 1, clean training samples are corrupted via the forward diffusion process, fed into the pretrained denoiser for a single-step Tweedie prediction, and transformed into 2D channel-wise power spectra. Radially averaged power spectra (RAPS) are computed and fitted to a parametric power-law function with offset and linear terms across discrete timesteps, followed by cubic spline interpolation over continuous time \(t \in [0, T]\) to yield \(S(t, f)\). In Stage 2, during iterative reverse sampling, the denoiser evaluates CFG noise to form the current Tweedie estimate \(\hat{x}_{0|t}\). Its RAPS is extracted and compared against the target prior \(S(t, f)\) using an asymmetric penalty loss \(\phi_a\) that penalizes spectral deficits more severely than surpluses. The analytical gradient with respect to \(x_t\) is backpropagated through the FFT, updating \(x_t\) before the standard solver takes the final denoising step.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input noisy state x_t and condition c"] --> B["Denoising step with CFG<br/>Compute ε_CFG and Tweedie estimate x̂_0|t"]
B --> C["Target power spectrum modeling<br/>Compute 2D FFT, RAPS y_t(f), and target prior S(t, f)"]
C --> D["Asymmetric guidance loss & state calibration<br/>Compute asymmetric log-spectral loss and update x_t via iFFT"]
D --> E["Standard denoising update<br/>Advance calibrated state to x_{t-1}"]
Key Designs¶
1. Target power spectrum modeling: data-driven continuous-time power-law fitting Because the spectral mismatch exhibits divergent patterns across architectures and timesteps, and ground-truth \(x_0\) is inaccessible during sampling, the method uses the single-step Tweedie denoiser prediction \(\hat{x}_{0|\tau}\) from the forward process as an accumulation-free reference proxy. For each channel, the 2D power spectrum \(\mathbf{Y}_\tau = |\mathcal{F}[\hat{x}_{0|\tau}]|^2\) is transformed into a 1D radially averaged power spectrum (RAPS) \(y_\tau(f) = \text{RadialAvg}(\mathbf{Y}_\tau)\), where \(f \in (0, 0.5]\) denotes normalized spatial frequency. Leveraging the empirical fact that natural and blurry intermediate images adhere to power-law frequency decay, the target spectrum is parameterized as: $\(S(t, f) = p_t \cdot f^{-q_t} + r_t \cdot f + s_t\)$ where \(p_t, q_t, r_t, s_t\) are independent parameters per channel and timestep. The primary term captures power-law decay, while the bias \(s_t\) and linear term \(r_t \cdot f\) capture deviations from a pure power law (\(r_t=0\) by default, activated for complex models such as SDXL and FLUX). After fitting parameters across discrete timesteps via least-squares regression, cubic spline interpolation is applied across time, providing a continuous reference function \(S(t, f)\) compatible with any sampling step count.
2. Asymmetric guidance loss and state calibration: penalizing energy deficits via fast spectral operators To steer sampling toward the target prior without prohibitive computational overhead, SPA adapts Diffusion Posterior Sampling (DPS) principles by constraining Tweedie's estimate and detaching \(\epsilon_\theta\) from gradient computation. Statistical observation reveals that intermediate spectra falling below the target cause severe blurring and structural collapse, whereas exceeding the target is comparatively benign. SPA introduces an asymmetric piecewise penalty function \(\phi_a(x)\) (where \(a > 1\) amplifies negative log-errors when \(x < 0\)), formulating the guidance loss across channels and frequency bins as: $\(\mathcal{L}_{\text{spec}}(x_t, t) = \frac{1}{N} \sum_{c=1}^{N} \sum_{k=1}^{K} \left( \phi_a\left( \log_{10} y_t(f_k^c) - \log_{10} S(t, f_k^c) \right) \right)^2\)$ Because gradient computation requires only element-wise derivatives, radial unpooling, and an inverse 2D FFT backpropagated through Tweedie's formula, its complexity is strictly \(\mathcal{O}(D \log D)\) without requiring additional neural network forward or backward passes. The latent is directly shifted via \(x_t \leftarrow x_t - \eta \nabla_{x_t} \mathcal{L}_{\text{spec}}\).
Loss & Training¶
The offline modeling stage requires only 10K training samples (e.g., ImageNet, CelebA-HQ, or LAION-Aesthetics V2) with single forward-pass evaluations, incurring computational cost comparable to calculating standard FID; the coefficient of determination \(R^2\) exceeds 0.9 across nearly all configurations. During inference, hyperparameters consist solely of guidance strength \(\eta\) and asymmetric penalty intensity \(a\): DDPM uses \((\eta, a) = (0.01, 1.0)\), ADM uses \((0.05, 2.0)\), SD2.0 and SDXL share \((0.2, 5.0)\), SD3.5 uses \((0.75, 3.0)\), and FLUX.1 [dev] uses \((0.2, 2.0)\) at \(w=2.5\).
Key Experimental Results¶
Main Results¶
SPA was evaluated across unconditional/class-conditional pixel models (ADM and DDPM) and text-to-image models (latent architectures SD2.0, SDXL, and flow-matching models SD3.5, FLUX.1), compared against Vanilla sampling, \(\epsilon\)-rescaling, Time-shift sampling, and Wavelet-based frequency regulation.
| Model & Dataset | Method | FID ↓ | KID (×1000) ↓ | Density ↑ | Coverage ↑ |
|---|---|---|---|---|---|
| ADM (ImageNet 256×256) | Vanilla | 9.29 | 19.48 | 1.133 | 0.5939 |
| ADM (ImageNet 256×256) | \(\epsilon\)-rescaling | 8.00 | 11.01 | 1.146 | 0.5976 |
| ADM (ImageNet 256×256) | Time-shift | 15.35 | 77.47 | 0.761 | 0.4321 |
| ADM (ImageNet 256×256) | Wavelet reg. | 8.35 | 12.80 | 1.122 | 0.5929 |
| ADM (ImageNet 256×256) | SPA (Ours) | 7.81 | 10.25 | 1.226 | 0.6184 |
| DDPM (CelebA-HQ) | Vanilla | 49.44 | 43.57 | 0.8186 | 0.0169 |
| DDPM (CelebA-HQ) | \(\epsilon\)-rescaling | 48.53 | 40.98 | 0.8791 | 0.0169 |
| DDPM (CelebA-HQ) | SPA (Ours) | 48.63 | 41.36 | 0.8344 | 0.0167 |
For text-to-image synthesis, human-preference reward models HPSv3 and ImageReward serve as primary quality benchmarks, with CLIP Score (ViT-G/14) monitored to confirm prompt alignment (5K MS-COCO validation prompts):
| Model Setting | Method | HPSv3 ↑ | ImageReward ↑ | CLIP Score ↑ | Improvement Note |
|---|---|---|---|---|---|
| SD2.0 (\(w=7.5\)) | Vanilla | 7.043 | 0.393 | 0.333 | Baseline |
| SD2.0 (\(w=7.5\)) | Wavelet reg. | 7.193 | 0.429 | 0.332 | Heuristic frequency weighting |
| SD2.0 (\(w=7.5\)) | SPA (Ours) | 7.238 | 0.421 | 0.333 | Best overall generation quality |
| SDXL (\(w=7.5\)) | Vanilla | 8.426 | 0.791 | 0.341 | Baseline |
| SDXL (\(w=7.5\)) | \(\epsilon\)-rescaling | 8.528 | 0.800 | 0.340 | Scalar noise rescaling |
| SDXL (\(w=7.5\)) | Wavelet reg. | 8.576 | 0.807 | 0.340 | Wavelet band regulation |
| SDXL (\(w=7.5\)) | SPA (Ours) | 8.829 | 0.829 | 0.341 | Substantial ~4.5% quality gain |
| SD3.5 (\(w=5.0\)) | Vanilla | 9.782 | 0.939 | 0.335 | Flow matching baseline |
| SD3.5 (\(w=5.0\)) | SPA (Ours) | 9.975 | 0.976 | 0.334 | Consistent preference boost |
| FLUX.1 dev (\(w=2.5\)) | Vanilla | 12.53 (11.55†) | 1.044 (0.989) | 0.328 (0.332) | Parentheses: un-distilled true CFG |
| FLUX.1 dev (\(w=2.5\)) | SPA (Ours) | 12.57 (11.76) | 1.054 (1.005) | 0.328 (0.332) | Marked gains on de-distilled model |
Overhead and Tail Distribution Analysis¶
On SDXL, per-step timing breakdown reveals that standard denoiser inference (neural network evaluation and CFG composition) takes 2.47 s, whereas SPA's operations (Tweedie extraction, FFT/RAPS calculation, and iFFT gradient calibration) take only 0.0952 s—adding just +3.86% latency overhead.
Analysis on the guidance-distilled FLUX.1 [dev] model demonstrates that SPA acts as a selective error-correction mechanism for degraded long-tail samples:
| Evaluated Subset | Metric | Vanilla | SPA (Ours) | Net Advantage |
|---|---|---|---|---|
| FLUX.1 full set (\(w=2.5\)) | Win-rate | 50.0% | 53.3 ± 1.5% | Statistically significant (\(p = 2.3 \times 10^{-5}\)) |
| FLUX.1 bottom 20% samples (\(w=3.5\)) | Win-rate | 50.0% | 54.1 ± 3.3% | Significant correction (\(p = 0.02\)) |
| FLUX.1 bottom 5% tail samples (\(w=2.5\)) | HPSv3 Gain | Baseline | +0.70 pts | Substantial repair of malformed structures |
| FLUX.1 bottom 5% tail samples (\(w=3.5\)) | HPSv3 Gain | Baseline | +0.19 pts | Persistent refinement on severe failures |
Key Findings¶
- Broad architectural generalizability: SPA delivers consistent improvements across pixel-space architectures, latent diffusions, and flow-matching models. Unlike heuristic wavelet regulation, which introduces ringing artifacts, SPA preserves high visual fidelity without edge distortion.
- Physical object repair: Calibrating intermediate SNR enables models to exploit early semantic signals more effectively. Furthermore, VAE latent channels disentangle shape, layout, and lighting, allowing channel-wise alignment to repair structural defects (e.g., deformed sports equipment).
Highlights & Insights¶
- Lucid diagnostic insight: Unveils that exposure bias in diffusion processes operates as a model- and timestep-dependent frequency SNR distortion, overturning previous assumptions of uniform scalar shifts.
- Remarkable efficiency and deployability: Imposing guidance on Tweedie estimates using analytical FFT gradients avoids backpropagation through large diffusion backbones, introducing only 3–4% latency overhead.
- Plug-and-play modularity: Operates downstream of CFG combination, maintaining full compatibility with diverse ODE/SDE solvers (DDIM, Euler) and CFG variants without requiring model retraining.
Limitations & Future Work¶
- Context-agnostic spectral prior: The current formulation fits an aggregate target spectrum over the entire dataset, ignoring distinct domain-specific distributions (e.g., line illustrations or medical imaging); conditioning priors on prompts or styles remains an open path.
- Uniform frequency weighting: The guidance loss utilizes unweighted MSE across log-frequencies rather than integrating human perceptual contrast sensitivity functions (CSF).
- First-order gradient dynamics: Optimization relies on first-order gradient updates (DPS-style); higher-order or momentum-based trajectory corrections could offer additional stability.
Related Work & Insights¶
- vs \(\epsilon\)-rescaling (Ning et al., ICLR 2024): Employs scalar multiplication to match global noise norm, assuming frequency-uniform errors; SPA demonstrates that errors are sharply frequency-dependent and model-specific, achieving superior quality across all metrics.
- vs Wavelet Regulation (Yu & Zhan, ACM MM 2025): Applies discrete wavelet transforms with manual high/low-band reweighting that requires sensitive tuning for each resolution; SPA leverages continuous physical power-law priors with minimal artifacts.
- vs Discriminator Guidance (Kim et al., ICML 2023) / MPGD (He et al., ICLR 2024): Discriminator methods require adversarial training and heavy inference overhead, while MPGD backpropagates through autoencoders; SPA provides training-free, analytical FFT guidance at a fraction of the compute.
Rating¶
- Novelty: ⭐⭐⭐⭐ [Identifies frequency-dependent SNR distortion as the core of exposure bias and designs a lightweight RAPS-based guidance scheme]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensively validated across pixel, latent, and flow-matching models, combining distribution metrics, reward models, and tail-sample analyses]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, rigorous mathematical formulation, and intuitive empirical visualizations]
- Value: ⭐⭐⭐⭐⭐ [Training-free, minimal overhead (~3-4%), and highly practical for eliminating structural glitches in modern generative pipelines]