Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/huggingface/diffusers
Area: Image Generation
Keywords: diffusion models, training-free refinement, internal-latent analysis, anomaly detection, artifact suppression
TL;DR¶
By identifying that abrupt early-stage fluctuations in deep low-noise internal latents directly cause visual artifacts, this paper proposes DUNE, a training-free framework that detects outlier activations via score-consistent EMA tracking and applies backbone-specific suppression across U-Net and DiT models, simultaneously improving visual fidelity and suppressing hallucinations without extra training.
Background & Motivation¶
Diffusion models have achieved remarkable success across text-to-image synthesis, super-resolution, and multimodal generation, driven by denoising backbones such as U-Net and Diffusion Transformers (DiT) that parameterize the score function. To enhance visual fidelity and suppress synthesis defects, recent efforts have introduced training-free inference-time interventions such as feature reweighting in FreeU. However, directly manipulating internal activations or raw score outputs often causes severe semantic drift, resulting in oversaturated textures, blurred contours, unnatural contrasts, and exacerbated anatomical hallucinations such as deformed limbs or distorted facial features.
The fundamental tension stems from an indiscriminate treatment of different network components and sampling stages. Within a U-Net, skip connections and upsampling blocks govern fine details and low-frequency structures respectively, while Transformer self-attention blocks process tokens across diverse spatial scales. Naively scaling component activations across arbitrary timesteps or shallow layers disrupts the normal distribution of the score network output, exaggerating localized anomalies rather than correcting them. Consequently, the key challenge lies in pinpointing exactly when and where artifact-inducing dynamics emerge within the backbone, and designing targeted interventions that preserve underlying semantic trajectories.
Through systematic component-wise and phase-aware analysis, this paper reveals that artifact-associated score accelerations are heavily concentrated in the early denoising stages and are prominently reflected within deep low-noise internal latents (such as U-Net h-space). Core idea: partition the sampling trajectory into an early EMA-guided detect-and-suppress phase and an optional late detail phase, performing surgical anomaly suppression in deep, noise-attenuated internal latents to eliminate artifacts while preserving semantic consistency.
Method¶
Overall Architecture¶
DUNE decouples the reverse diffusion trajectory into two distinct operational intervals: the core detect phase (\(1.0T \sim 0.6T\)) and an optional detail phase (\(0.6T \sim 0.0T\)). During the early detect phase, where global semantic structure is established, DUNE normalizes internal latents to match score dimensions and tracks their temporal trajectory via Exponential Moving Average (EMA). Outlier features with high log-ratio deviations are isolated into a binary mask \(M_t\). A backbone-specific suppressor then modifies the anomalous activations in deep, low-noise representation spaces (U-Net bottleneck h-space or deep Transformer self-attention blocks) to stabilize score acceleration.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Denoising Timestep Latent<br/>t ∈ [1.0T, 0.6T]"] --> LNorm["Score-Consistent Latent Normalization<br/>z_t = -h_t / σ_t"]
LNorm --> EMA["EMA State Tracking & Log-Deviation<br/>Δ = |log|z_t / z_bar_{t+1}||"]
EMA --> TopK["Quantile Thresholding & Masking<br/>M_t = (Δ > λ)"]
TopK --> BBCheck{"Backbone Architecture"}
BBCheck -->|U-Net h-space| UNetS["Channel-Aware Masked Scaling<br/>S_UNet = κ (n_t ⊙ h_t)"]
BBCheck -->|Transformer DiT| DiTS["Masked EMA Feature Blending<br/>S_Tr = κ h_t + (1-κ) h_bar_t"]
UNetS --> GatedMerge["Gated Masked Feature Fusion<br/>h_hat_t = (1-M_t) ⊙ h_t + M_t ⊙ S_bb"]
DiTS --> GatedMerge
GatedMerge --> DPhase["Optional Detail Phase Modulation<br/>t ∈ [0.6T, 0.0T] Style & Contrast Control"]
Key Designs¶
1. Deep Low-Noise Intervention Point: Leveraging Convolutional Downsampling and Attention Denoising
Standard feature modification applied to shallow layers or skip connections tends to induce volatile score perturbations. The paper theoretically establishes that deep representation spaces serve as optimal intervention sites due to intrinsic noise attenuation. For U-Nets, Proposition 1 shows that convolutional downsampling scales the standard deviation of injected Gaussian noise by the kernel 2-norm \(\|K\|_2 = \sqrt{\sum \lambda_{i,j}^2}\) (empirically evaluated at 0.2036 and 0.1509 in SDXL downsampling layers), meaning the deepest bottleneck (h-space) naturally concentrates semantic signal while attenuating noise. For Diffusion Transformers, Proposition 2 establishes that self-attention outputs \(y = \sum_{j=1}^N w_j x_j\) scale isotropic noise standard deviation by \(\sqrt{\sum_{j=1}^N w_j^2} \in [1/\sqrt{N}, 1]\). As attention weights become diffuse across deep stacked layers (specifically at the 3/4 depth mark), self-attention aggregates token semantics and suppresses stochastic fluctuations, providing a noise-attenuated sweet spot for anomaly detection.
2. Score-Consistent EMA Anomaly Detection: Log-Ratio Quantile Thresholding
Because latent magnitudes undergo wide dynamic shifts across timestep schedules \(\sigma_t\), detecting anomalies on raw features causes severe temporal drift. DUNE projects internal latents onto a score-consistent space: \(z_t = -h_t / \sigma_t\). An exponential moving average (EMA) reference is maintained across the detect phase:
with \(\gamma = 0.7\) held constant across all backbones. To ensure detection remains invariant to per-image brightness or prompt magnitude variations, DUNE evaluates the log-ratio deviation:
Using a quantile threshold \(p = 0.9\), DUNE isolates the top \((1-p)\%\) anomalous coordinates into a binary mask \(M_t = (\Delta > \lambda)\), precisely tagging regions suffering from abrupt score acceleration spikes.
3. Backbone-Specific Masked Suppression: Channel-Aware Shrinkage and Attention Blending
To rectify the detected outliers, DUNE formulates a gated suppression template: \(\hat{h}_t = (1 - M_t) \odot h_t + M_t \odot \mathcal{S}_{\text{bb}}(h_t, \bar{h}_t; \kappa)\), where the operator \(\mathcal{S}_{\text{bb}}\) is specialized according to backbone structure. In U-Net models, h-space comprises low-resolution spatial feature maps with specialized semantic channels (\(1280 \times 32 \times 32\) in SDXL). Directly smoothing these features with EMA references washes out channel-specific semantics. DUNE computes the absolute channel mean \(n_t \in \mathbb{R}^{C \times 1 \times 1}\) and applies channel-aware scaling: \(\mathcal{S}_{\text{UNet}}(h_t, \bar{h}_t; \kappa) = \kappa (n_t \odot h_t)\). In Transformer architectures, where tokens are patchified and globally mixed via attention, channel scaling causes patchy artifacts; DUNE instead applies linear EMA blending: \(\mathcal{S}_{\text{Tr}}(h_t, \bar{h}_t; \kappa) = \kappa h_t + (1 - \kappa) \bar{h}_t\). Theorem 3 formally proves that whenever noise concentration satisfies \((1-\kappa^2)\eta_t > (1-\kappa)\varepsilon_t\), this operator guarantees a strict increase in latent signal-to-noise ratio \(\text{SNR}(\hat{z}_t) > \text{SNR}(z_t)\).
Loss & Training¶
DUNE operates entirely at inference time as a training-free plug-and-play algorithm, requiring no parameter backpropagation or auxiliary teacher models. Hyperparameters are largely universal: EMA momentum \(\gamma = 0.7\), detection quantile \(p = 0.9\), and intervention duration strictly confined to the early \(40\%\) timesteps (\(1.0T \sim 0.6T\)). For Transformer models, intervention is positioned at 3/4 layer depth. The suppression intensity \(\kappa \in [0, 1]\) is tuned once per backbone family and kept fixed across all evaluation prompts.
Key Experimental Results¶
Main Results¶
Evaluation is performed across 5,000 images generated from COCO prompts for FID, and 5,000 human-annotated prompts for HADM-L (a detection-based metric quantifying anatomical flaws like deformed hands and extra digits). Semantic consistency is tracked via cosine CLIP similarity between original and refined images.
| Backbone Model | Method | FID ↓ | HADM-L ↓ | CLIP Sim. ↑ |
|---|---|---|---|---|
| SDXL | Original | 18.86 | 0.227 | 1.000 |
| SDXL | FreeU | 31.22 | 0.351 | 0.781 |
| SDXL | ASCED | 22.07 | 0.227 | 0.945 |
| SDXL | PAG | 21.72 | 0.210 | 0.821 |
| SDXL | DUNE (Ours) | 18.75 | 0.074 | 0.925 |
| LCM | Original | 22.53 | 0.152 | 1.000 |
| LCM | FreeU | 27.57 | 0.237 | 0.756 |
| LCM | ASCED | 27.60 | 0.151 | 0.995 |
| LCM | PAG | 35.87 | 0.123 | 0.362 |
| LCM | DUNE (Ours) | 22.41 | 0.052 | 0.940 |
| Kandinsky 3 | Original | 21.36 | 0.250 | 1.000 |
| Kandinsky 3 | ASCED | 22.44 | 0.252 | 0.924 |
| Kandinsky 3 | DUNE (Ours) | 20.76 | 0.095 | 0.933 |
| PixArt-Σ | Original | 28.75 | 0.316 | 1.000 |
| PixArt-Σ | ASCED | N/A (pure noise) | N/A | N/A |
| PixArt-Σ | PAG | 32.61 | 0.284 | 0.840 |
| PixArt-Σ | DUNE (Ours) | 27.80 | 0.254 | 0.952 |
| Hunyuan-DiT | Original | 30.82 | 0.294 | 1.000 |
| Hunyuan-DiT | DUNE (Ours) | 30.67 | 0.283 | 0.988 |
Ablation Study¶
Ablations examine intervention locations, random mask baselines, and architectural choices on score acceleration and fidelity:
| Configuration / Variant | Target Benchmark | Metric | Value | Note / Conclusion |
|---|---|---|---|---|
| DUNE (h-space) | SDXL | Excess Accel. Gap Red. (EAGR) ↑ | 24.9% | Effectively flattens abnormal score acceleration spikes |
| DUNE (3/4 depth SA) | PixArt-Σ | EAGR ↑ | 18.1% | Confirms deep self-attention suppresses acceleration in DiT |
| Random Mask Control | SDXL | EAGR | ~0.0% | Verifies gains stem from accurate outlier localization |
| Shallow Skip Suppression | U-Net | Score Deviation | Increased | Intervening in high-frequency paths amplifies hallucinations |
| Shallow Up-Block Suppression | U-Net | Score Deviation | Increased | Causes tonal distortions and brightness drift |
| Full Model (p=0.9, optimal κ) | LCM | FID ↓ | 22.41 | Superior to low-quantile choices (p=0.3/0.5/0.7) |
Key Findings¶
- Eliminating the Artifact vs. Fidelity Trade-off: On SDXL, all prior training-free methods degrade FID (FreeU 31.22, ASCED 22.07, PAG 21.72), whereas DUNE simultaneously improves FID to 18.75 while dramatically reducing anatomical hallucinations from 0.227 down to 0.074 (a 67.4% reduction).
- Robustness Across Diverse Architectures: On high-resolution Transformer backbones like PixArt-Σ, trajectory perturbation in ASCED causes total sampling collapse into pure noise, whereas DUNE smoothly improves FID to 27.80 and HADM to 0.254.
- Minimal Sampling Latency Overhead: On SDXL, standard sampling requires 9.55 seconds. DUNE incurs almost negligible overhead (9.71 seconds), vastly outperforming gradient-guided methods like PAG (14.38 seconds) and iterative score trajectory analysis in ASCED (13.37 seconds).
Highlights & Insights¶
- Root-Cause Attribution via Latent Acceleration: Connects generative hallucinations directly to early temporal score acceleration spikes within deep low-noise latents, establishing a rigorous dynamical perspective for diffusion interpretability.
- Unified Dual-Backbone Formulation: Distinguishes between channel-specialized convolutional features and token-mixed attention representations, formulating channel-aware scaling for U-Nets and EMA blending for DiTs under a single SNR-improving framework.
- Phase Decoupling Paradigm: Restricting structural corrections to the early 40% timesteps prevents downstream semantic distortion, while offering an optional late-stage playground for user-guided contrast and tone adjustment.
Limitations & Future Work¶
- Static Backbone Hyperparameters: While the detection quantile \(p = 0.9\) generalizes across models, suppression factor \(\kappa\) still relies on per-backbone calibration rather than prompt-adaptive dynamic selection.
- Subjective Detail Phase Tuning: Modifications during the latter 60% of sampling primarily serve stylistic preferences (contrast, brightness), lacking an automated objective metric for optimal detail parameter selection.
Related Work & Insights¶
- vs FreeU: FreeU uniformly scales backbone and skip features throughout the entire sampling schedule, frequently causing oversaturation and severe FID degradation; DUNE intervenes only during early timesteps on detector-selected anomalous regions.
- vs ASCED: ASCED modifies the outer score trajectory directly and injects corrective noise, which causes sampling collapse on modern Diffusion Transformers; DUNE operates inside deep representations to ensure smooth transitions.
- vs PAG (Perturbed-Attention Guidance): PAG perturbs self-attention maps to construct guidance branches, which inflates inference latency by ~50% and sharpens edges unnaturally; DUNE performs in-place latent correction with only ~1.6% overhead.
Rating¶
- Novelty: ⭐⭐⭐⭐ [Pinpoints latent score acceleration as the root of artifacts; presents a unified U-Net and DiT refinement framework]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive verification across 5 major diffusion backbones, fine-grained anatomical metrics, and user preference studies]
- Writing Quality: ⭐⭐⭐⭐⭐ [Mathematically rigorous propositions and theorems paired with clear empirical justifications]
- Value: ⭐⭐⭐⭐⭐ [Highly practical, training-free, and computationally lightweight drop-in refiner for generative deployment]