Skip to content

LUA: Latent Upscaling Adapter for Diffusion-Based Image Synthesis

Conference: ECCV 2026
Paper: ECCV Official
Area: Image Generation
Keywords: latent super-resolution / diffusion models / high-resolution image synthesis / latent upscaling adapter

TL;DR

Addressing the heavy computational cost of high-resolution diffusion denoising and the artifacts of post-hoc super-resolution, LUA inserts a lightweight feed-forward latent upscaling adapter before the frozen VAE decoder, combining a shared Swin backbone with a three-stage latent-pixel curriculum to achieve native-level 2K/4K synthesis via a single decode with multi-fold acceleration.

Background & Motivation

Latent diffusion models (LDMs) have transformed modern visual synthesis by shifting iterative generative computation into compact latent representations. Nevertheless, deployed models are essentially bounded by the spatial resolutions encountered during their training phase (typically \(512\times 512\) or \(1024\times 1024\)). When prompted to sample directly beyond these native scales, generators frequently suffer from severe visual pathologies, including duplicated subjects, fragmented anatomy, and collapsed textures. While retraining or high-resolution full-model fine-tuning can mitigate such flaws, the computational expenditure and curated data requirements are prohibitively demanding.

Post-hoc super-resolution presents a practical alternative to avoid full retraining, but existing paradigms enforce steep trade-offs. Pixel-space super-resolution processes decoded RGB images; however, its computational complexity grows quadratically with output dimensions, which incurs high latency and causes oversmoothing or semantic drift. Conversely, existing latent-space super-resolution frameworks (such as DemoFusion and LSRNA) upscale representations in latent space but depend on multi-stage re-diffusion pipelines with auxiliary noise schedules and guided reverse steps, severely burdening throughput. Meanwhile, naïve latent interpolation departs from the generative manifold, producing distorted grid patterns and blurred outputs upon decoding.

The central challenge lies in upscaling latent features while rigorously preserving manifold geometry and high-frequency latent details, all without triggering an additional diffusion process. Core idea: insert a lightweight, plug-and-play Latent Upscaling Adapter (LUA) between the generator and the frozen VAE decoder to enlarge latents via a single feed-forward pass, governed by a three-stage progressive latent-pixel curriculum that achieves single-decode high-fidelity synthesis at substantially lower latency.

Method

Overall Architecture

The core inference pipeline of LUA operates at the latent interface between the pretrained generator and the frozen VAE decoder. First, the base diffusion generator takes text conditioning and noise to produce a low-resolution latent code \(z \in \mathbb{R}^{h \times w \times C}\) (where \(C=16\) for FLUX and SD3, and \(C=4\) for SDXL). Next, LUA processes \(z\) through a deterministic feed-forward pass to predict an upscaled latent representation \(\hat{z} \in \mathbb{R}^{\alpha h \times \alpha w \times C}\) for a target scale factor \(\alpha \in \{2, 4\}\). Finally, the frozen VAE decoder performs a single decoding pass on \(\hat{z}\) to render the high-resolution image \(\hat{x}\). Because the VAE decoder incorporates a spatial downsampling stride of \(s=8\), upscaling the latent by \(\times 2\) yields a \(16\times\) enlargement in pixel count, while LUA operates on only \(1/s^2 = 1/64\) of the spatial positions required by pixel-space super-resolution models.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Low-Resolution Latent<br/>generator output z ∈ R^(h×w×C)"] --> B["Channel Alignment Conv<br/>1×1 convolution matching target VAE channels"]
    B --> C["Shared Swin Feature Backbone<br/>windowed self-attention capturing local context"]
    C --> D["Scale-Specific Pixel-Shuffle Heads<br/>sub-pixel convolution predicting upscaled z_hat"]
    D --> E["Single VAE Decoding<br/>frozen decoder outputting HR image x_hat"]

Key Designs

1. Latent-domain single-pass upscaling: circumventing re-diffusion and quadratic cost Conventional high-resolution generation relies on secondary diffusion passes that multiply sampling steps, whereas image-space super-resolution processes full-resolution pixel grids with quadratic computational overhead. LUA establishes a deterministic upscaling operator \(U_\alpha(z)\), confining all generative stochasticity to the base generator while LUA focuses exclusively on latent manifold interpolation and high-frequency statistical restoration. For a decoder with spatial stride \(s=8\), pixel-space operators must calculate across \((sh) \times (sw)\) positions, whereas LUA operates over the compact \(h \times w\) space, reducing spatial elements by \(O(hw) / O((sh)(sw)) \approx 1/s^2 = 1/64\). Consequently, LUA delivers multi-megapixel outputs while avoiding multi-step denoising loops and memory bottlenecks, introducing only \(+0.42\) s to \(+2.21\) s overhead on an NVIDIA L40S.

2. Shared backbone with scale-specific heads: multi-scale unity and cross-VAE transfer Rather than training separate networks for different scale factors (\(\times 2\) and \(\times 4\)), LUA adopts a shared feature backbone paired with dedicated upscaling heads. The backbone leverages a SwinIR-style windowed self-attention structure whose shifted-window mechanics naturally align with the spatial-statistical characteristics of VAE latent patches. Following feature extraction, independent lightweight convolutions and pixel-shuffle layers (\(U_{\times 2}\) and \(U_{\times 4}\)) process the features, enabling the shared backbone to learn scale-invariant representations while the heads specialize in scale-specific aliasing profiles. Furthermore, this architectural separation facilitates cross-VAE reuse: when transferring between 16-channel models (FLUX/SD3) and the 4-channel SDXL, the Swin backbone and prediction heads remain intact, requiring only the replacement of the initial \(1\times 1\) convolution and brief fine-tuning on target latents, eliminating the need for retraining from scratch.

3. Three-stage progressive curriculum: aligning latent geometry with perceptual fidelity Single-domain loss functions prove inadequate for latent super-resolution: latent-only optimization preserves macro structures but leaves decoded images with residual grid artifacts and high-frequency noise, whereas optimizing purely in pixel space via backpropagation through a frozen VAE decoder is unstable due to unnormalized latent gradients. LUA resolves this through a three-stage training curriculum: - Stage I (Latent-domain structural alignment): Learns the primary mapping \(\hat{z} = U_\alpha(z)\) using element-wise spatial \(\ell_1\) loss and 2D fast Fourier transform (FFT) magnitude alignment: $\(\mathcal{L}_{\mathrm{SI}} = \alpha_1 \mathcal{L}_{\mathrm{L1}}^z + \beta_1 \mathcal{L}_{\mathrm{FFT}}^z\)$ where \(\mathcal{L}_{\mathrm{FFT}}^z = \| \mathcal{F}(\hat{z}) - \mathcal{F}(z_{\mathrm{HR}}) \|_1\), \(\mathcal{F}\) computes channel-wise 2D FFT magnitudes, and coefficients are set to \(\alpha_1=1.0, \beta_1=0.1\). - Stage II (Joint latent–pixel consistency): Introduces the frozen decoder to enforce cross-domain coherence by combining the latent losses with bicubic downsampled consistency \(\mathcal{L}_{\mathrm{DS}}^x\) and Gaussian-blurred high-frequency residual matching \(\mathcal{L}_{\mathrm{HF}}^x\): $\(\mathcal{L}_{\mathrm{SII}} = \alpha_2 \mathcal{L}_{\mathrm{L1}}^z + \beta_2 \mathcal{L}_{\mathrm{FFT}}^z + \gamma_2 \mathcal{L}_{\mathrm{DS}}^x + \delta_2 \mathcal{L}_{\mathrm{HF}}^x\)$ where \(\mathcal{L}_{\mathrm{DS}}^x = \| \downarrow_d(\hat{x}) - \downarrow_d(x_{\mathrm{HR}}) \|_1\) and \(\mathcal{L}_{\mathrm{HF}}^x = \| (\hat{x} - G_\sigma(\hat{x})) - (x_{\mathrm{HR}} - G_\sigma(x_{\mathrm{HR}})) \|_1\), with weights \(\alpha_2=1.0, \beta_2=0.1, \gamma_2=0.1, \delta_2=0.05\). This stage stabilizes latent properties under decoding. - Stage III (Edge-aware image refinement): Omits latent loss supervision entirely, training in pixel space with image \(\ell_1\), image FFT, and the edge-aware gradient localization (EAGLE) loss: $\(\mathcal{L}_{\mathrm{SIII}} = \alpha_3 \mathcal{L}_{\mathrm{L1}}^x + \beta_3 \mathcal{L}_{\mathrm{FFT}}^x + \gamma_3 \mathcal{L}_{\mathrm{EAGLE}}^x\)$ with weights \(\alpha_3=10.0, \beta_3=1.0, \gamma_3=5 \times 10^{-5}\), sharpening fine boundaries and eliminating staircase artifacts without secondary denoising.

Loss & Training

Training is conducted on OpenImages photos (\(\ge 1440\) px on both dimensions), cropped into non-overlapping \(512\times 512\) patches and downsampled bicubically to produce 3.8M pairs encoded via the FLUX VAE (\(s=8, C=16\)). The network is optimized using Adam (\(\text{lr}=2\times 10^{-4}\), weight decay 0, gradient clipping at 0.4), EMA 0.999, and a MultiStepLR schedule. Effective batch sizes are 2,048 for Stage I and 32 for Stages II–III, each stage running for 125k steps, totaling \(\sim 270\) GPU-hours across 8 NVIDIA H100 GPUs. Cross-model adaptation to SDXL and SD3 requires only 500k latent pairs and a brief fine-tuning phase.

Key Experimental Results

Main Results

Quantitative evaluations are conducted on 1,000 held-out high-resolution OpenImages scenes. Table 1 benchmarks SDXL-anchored configurations across 1024, 2048, and 4096 resolutions against state-of-the-art pipelines on a single H100 GPU (batch size 1).

Resolution Method FID ↓ pFID ↓ KID ↓ pKID ↓ CLIP ↑ Time (s)
1024×1024 HiDiffusion 232.55 230.39 0.0211 0.0288 0.695 1.54
DemoFusion 195.82 193.99 0.0153 0.0229 0.725 2.04
LSRNA–DemoFusion 194.55 192.73 0.0151 0.0228 0.734 3.09
SDXL (Direct) 194.53 192.71 0.0151 0.0225 0.731 1.61
SDXL + SwinIR 210.40 204.23 0.0313 0.0411 0.694 2.47
SDXL + LUA (Ours) 209.80 191.75 0.0330 0.0426 0.738 1.42
2048×2048 HiDiffusion 200.72 114.30 0.0030 0.0090 0.738 4.97
DemoFusion 184.79 177.67 0.0030 0.0100 0.750 28.99
LSRNA–DemoFusion 181.24 98.09 0.0019 0.0066 0.762 20.77
SDXL (Direct) 202.87 116.57 0.0030 0.0086 0.741 7.23
SDXL + SwinIR 183.16 100.09 0.0020 0.0077 0.757 6.29
SDXL + LUA (Ours) 180.80 97.90 0.0018 0.0065 0.764 3.52
4096×4096 HiDiffusion 233.65 95.95 0.0158 0.0214 0.698 122.62
DemoFusion 185.36 177.89 0.0043 0.0113 0.749 225.77
LSRNA–DemoFusion 177.95 62.07 0.0023 0.0071 0.757 91.64
SDXL (Direct) 280.42 101.89 0.0396 0.0175 0.663 148.71
SDXL + SwinIR 183.15 65.71 0.0018 0.0103 0.756 7.29
SDXL + LUA (Ours) 176.90 61.80 0.0015 0.0152 0.759 6.87

Table 2 highlights LUA performance across different diffusion backbones (FLUX, SDXL, SD3) and scale factors (\(\times 2\) and \(\times 4\)):

Scale Diffusion Model FID ↓ pFID ↓ KID ↓ pKID ↓ CLIP ↑ Time (s)
×2 FLUX + LUA 180.99 100.40 0.0020 0.0079 0.773 29.83
SDXL + LUA 183.15 101.18 0.0020 0.0087 0.753 3.52
SD3 + LUA 184.58 103.94 0.0022 0.0083 0.768 20.29
×4 FLUX + LUA 181.06 62.30 0.0018 0.0085 0.772 31.91
SDXL + LUA 182.42 71.92 0.0015 0.0152 0.754 6.87
SD3 + LUA 183.34 67.25 0.0016 0.0095 0.769 21.84

Ablation Study

Table 3 validates the contribution of each stage within the progressive curriculum under \(\times 2\) and \(\times 4\) upscaling:

Configuration ×2 PSNR ↑ ×2 LPIPS ↓ ×4 PSNR ↑ ×4 LPIPS ↓ Note
Latent \(\ell_1\) only 28.53 0.198 26.16 0.236 Lacks spectral and perceptual constraints; blurry boundaries
Stages I + II (w/o III) 28.96 0.172 26.67 0.213 Omits pixel-level edge sharpening; slight visual artifacts persist
Stages I + III (w/o II) 31.05 0.150 27.10 0.198 Abrupt transition between domains; latent-decoder mismatch
Stages II + III (w/o I) 31.60 0.145 27.40 0.192 Lacks latent-domain warm-up; suboptimal training stability
Full (I + II + III) 32.54 0.138 27.94 0.184 Superior reconstruction across both metrics and resolutions

Table 4 compares architectural upsampling choices:

Variant ×2 PSNR ↑ ×2 LPIPS ↓ ×4 PSNR ↑ ×4 LPIPS ↓ Note
LIIF (Implicit coordinate decoder) 29.10 0.210 26.10 0.235 Continuous coordinates struggle with fine high-frequency textures
Separated specialist models 31.92 0.150 27.71 0.189 Separate networks miss cross-scale feature regularization
Joint Multi-Head (Ours) 32.54 0.138 27.94 0.184 Shared representations yield stronger generalization and fidelity

Key Findings

  • LUA defines the Pareto frontier at 2048 and 4096 resolutions: at 4K, SDXL+LUA achieves 176.90 FID and 61.80 pFID, outperforming LSRNA–DemoFusion (177.95) while slashing runtime from 91.64 s to 6.87 s (over \(13\times\) speedup).
  • Direct high-resolution sampling undergoes catastrophic degradation at 4K (FID collapses to 280.42), exhibiting object duplication in 9–19% of samples, whereas LUA remains immune to structural repetition.
  • Cross-VAE adaptation exhibits remarkable sample efficiency: fine-tuning on only 50k target pairs matches established baselines, and scaling to 250k pairs surpasses LSRNA, confirming the transferable nature of latent spatial priors.
  • At 1024 resolution, the \(64\times 64\) latent bottleneck limits global FID (209.80 vs. 194.53 for native generation), yet LUA preserves superior patch-level fidelity (pFID 191.75) and text alignment (CLIP 0.738).

Highlights & Insights

  • Targeting structural efficiency: Bypassing recursive diffusion loops and executing super-resolution directly in latent space harnesses the VAE's spatial stride to reduce compute by \(64\times\) relative to pixel-space models.
  • Progressive cross-domain curriculum: The staged progression from latent spectral alignment to hybrid consistency and pixel edge sharpening overcomes training divergence under frozen non-linear decoders.
  • Minimalist cross-model transfer: A single shared backbone adapts across distinct generator families (FLUX, SD3, SDXL) merely by swapping the input convolution and fine-tuning briefly on target latents.

Limitations & Future Work

  • Preserved base hallucinations: As a deterministic operator, LUA upscales existing latent representations without generative re-sampling; structural or semantic errors generated by the base model are magnified rather than corrected.
  • Dead pixel anomalies: Rare out-of-distribution latent outputs can trigger localized visual artifacts upon decoding, currently requiring empirical latent value clamping.
  • Latent capacity ceiling: Because current decoders downsample by stride 8, initial latents at \(512\text{px}\) retain only \(64\times 64\) tokens; future exploration into smaller-stride or continuous latent representations could unlock higher fidelity.
  • vs DemoFusion / LSRNA: Reference-guided diffusion methods rely on multi-step noise-denoise loops for upscaling; LUA achieves comparable or better fidelity with a single feed-forward pass, reducing runtime by \(6\times\) to \(13\times\).
  • vs SwinIR / SeeSR (Pixel-Space SR): Pixel-domain models process huge spatial grids and incur quadratic compute scaling, often introducing haloing and oversmoothed textures; LUA operates compactly in latent space before decoding, avoiding pixel-level ringing.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Introduces a single-pass latent upscaling paradigm that eliminates secondary diffusion stages, supported by an effective multi-scale curriculum.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarks across 1K/2K/4K scales, three foundation models (FLUX/SD3/SDXL), detailed ablations, and a 2AFC human study.
  • Writing Quality: ⭐⭐⭐⭐⭐ Crisp mathematical formulation, clear narrative structure, and high information density across figures and tables.
  • Value: ⭐⭐⭐⭐⭐ Provides an immediately deployable, production-ready path for low-latency 2K/4K generative pipelines.