LUA: Latent Upscaling Adapter for Diffusion-Based Image Synthesis¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Image Generation
Keywords: latent super-resolution / diffusion models / high-resolution image synthesis / latent upscaling adapter
TL;DR¶
Addressing the heavy computational cost of high-resolution diffusion denoising and the artifacts of post-hoc super-resolution, LUA inserts a lightweight feed-forward latent upscaling adapter before the frozen VAE decoder, combining a shared Swin backbone with a three-stage latent-pixel curriculum to achieve native-level 2K/4K synthesis via a single decode with multi-fold acceleration.
Background & Motivation¶
Latent diffusion models (LDMs) have transformed modern visual synthesis by shifting iterative generative computation into compact latent representations. Nevertheless, deployed models are essentially bounded by the spatial resolutions encountered during their training phase (typically \(512\times 512\) or \(1024\times 1024\)). When prompted to sample directly beyond these native scales, generators frequently suffer from severe visual pathologies, including duplicated subjects, fragmented anatomy, and collapsed textures. While retraining or high-resolution full-model fine-tuning can mitigate such flaws, the computational expenditure and curated data requirements are prohibitively demanding.
Post-hoc super-resolution presents a practical alternative to avoid full retraining, but existing paradigms enforce steep trade-offs. Pixel-space super-resolution processes decoded RGB images; however, its computational complexity grows quadratically with output dimensions, which incurs high latency and causes oversmoothing or semantic drift. Conversely, existing latent-space super-resolution frameworks (such as DemoFusion and LSRNA) upscale representations in latent space but depend on multi-stage re-diffusion pipelines with auxiliary noise schedules and guided reverse steps, severely burdening throughput. Meanwhile, naïve latent interpolation departs from the generative manifold, producing distorted grid patterns and blurred outputs upon decoding.
The central challenge lies in upscaling latent features while rigorously preserving manifold geometry and high-frequency latent details, all without triggering an additional diffusion process. Core idea: insert a lightweight, plug-and-play Latent Upscaling Adapter (LUA) between the generator and the frozen VAE decoder to enlarge latents via a single feed-forward pass, governed by a three-stage progressive latent-pixel curriculum that achieves single-decode high-fidelity synthesis at substantially lower latency.
Method¶
Overall Architecture¶
The core inference pipeline of LUA operates at the latent interface between the pretrained generator and the frozen VAE decoder. First, the base diffusion generator takes text conditioning and noise to produce a low-resolution latent code \(z \in \mathbb{R}^{h \times w \times C}\) (where \(C=16\) for FLUX and SD3, and \(C=4\) for SDXL). Next, LUA processes \(z\) through a deterministic feed-forward pass to predict an upscaled latent representation \(\hat{z} \in \mathbb{R}^{\alpha h \times \alpha w \times C}\) for a target scale factor \(\alpha \in \{2, 4\}\). Finally, the frozen VAE decoder performs a single decoding pass on \(\hat{z}\) to render the high-resolution image \(\hat{x}\). Because the VAE decoder incorporates a spatial downsampling stride of \(s=8\), upscaling the latent by \(\times 2\) yields a \(16\times\) enlargement in pixel count, while LUA operates on only \(1/s^2 = 1/64\) of the spatial positions required by pixel-space super-resolution models.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Low-Resolution Latent<br/>generator output z ∈ R^(h×w×C)"] --> B["Channel Alignment Conv<br/>1×1 convolution matching target VAE channels"]
B --> C["Shared Swin Feature Backbone<br/>windowed self-attention capturing local context"]
C --> D["Scale-Specific Pixel-Shuffle Heads<br/>sub-pixel convolution predicting upscaled z_hat"]
D --> E["Single VAE Decoding<br/>frozen decoder outputting HR image x_hat"]
Key Designs¶
1. Latent-domain single-pass upscaling: circumventing re-diffusion and quadratic cost Conventional high-resolution generation relies on secondary diffusion passes that multiply sampling steps, whereas image-space super-resolution processes full-resolution pixel grids with quadratic computational overhead. LUA establishes a deterministic upscaling operator \(U_\alpha(z)\), confining all generative stochasticity to the base generator while LUA focuses exclusively on latent manifold interpolation and high-frequency statistical restoration. For a decoder with spatial stride \(s=8\), pixel-space operators must calculate across \((sh) \times (sw)\) positions, whereas LUA operates over the compact \(h \times w\) space, reducing spatial elements by \(O(hw) / O((sh)(sw)) \approx 1/s^2 = 1/64\). Consequently, LUA delivers multi-megapixel outputs while avoiding multi-step denoising loops and memory bottlenecks, introducing only \(+0.42\) s to \(+2.21\) s overhead on an NVIDIA L40S.
2. Shared backbone with scale-specific heads: multi-scale unity and cross-VAE transfer Rather than training separate networks for different scale factors (\(\times 2\) and \(\times 4\)), LUA adopts a shared feature backbone paired with dedicated upscaling heads. The backbone leverages a SwinIR-style windowed self-attention structure whose shifted-window mechanics naturally align with the spatial-statistical characteristics of VAE latent patches. Following feature extraction, independent lightweight convolutions and pixel-shuffle layers (\(U_{\times 2}\) and \(U_{\times 4}\)) process the features, enabling the shared backbone to learn scale-invariant representations while the heads specialize in scale-specific aliasing profiles. Furthermore, this architectural separation facilitates cross-VAE reuse: when transferring between 16-channel models (FLUX/SD3) and the 4-channel SDXL, the Swin backbone and prediction heads remain intact, requiring only the replacement of the initial \(1\times 1\) convolution and brief fine-tuning on target latents, eliminating the need for retraining from scratch.
3. Three-stage progressive curriculum: aligning latent geometry with perceptual fidelity Single-domain loss functions prove inadequate for latent super-resolution: latent-only optimization preserves macro structures but leaves decoded images with residual grid artifacts and high-frequency noise, whereas optimizing purely in pixel space via backpropagation through a frozen VAE decoder is unstable due to unnormalized latent gradients. LUA resolves this through a three-stage training curriculum: - Stage I (Latent-domain structural alignment): Learns the primary mapping \(\hat{z} = U_\alpha(z)\) using element-wise spatial \(\ell_1\) loss and 2D fast Fourier transform (FFT) magnitude alignment: $\(\mathcal{L}_{\mathrm{SI}} = \alpha_1 \mathcal{L}_{\mathrm{L1}}^z + \beta_1 \mathcal{L}_{\mathrm{FFT}}^z\)$ where \(\mathcal{L}_{\mathrm{FFT}}^z = \| \mathcal{F}(\hat{z}) - \mathcal{F}(z_{\mathrm{HR}}) \|_1\), \(\mathcal{F}\) computes channel-wise 2D FFT magnitudes, and coefficients are set to \(\alpha_1=1.0, \beta_1=0.1\). - Stage II (Joint latent–pixel consistency): Introduces the frozen decoder to enforce cross-domain coherence by combining the latent losses with bicubic downsampled consistency \(\mathcal{L}_{\mathrm{DS}}^x\) and Gaussian-blurred high-frequency residual matching \(\mathcal{L}_{\mathrm{HF}}^x\): $\(\mathcal{L}_{\mathrm{SII}} = \alpha_2 \mathcal{L}_{\mathrm{L1}}^z + \beta_2 \mathcal{L}_{\mathrm{FFT}}^z + \gamma_2 \mathcal{L}_{\mathrm{DS}}^x + \delta_2 \mathcal{L}_{\mathrm{HF}}^x\)$ where \(\mathcal{L}_{\mathrm{DS}}^x = \| \downarrow_d(\hat{x}) - \downarrow_d(x_{\mathrm{HR}}) \|_1\) and \(\mathcal{L}_{\mathrm{HF}}^x = \| (\hat{x} - G_\sigma(\hat{x})) - (x_{\mathrm{HR}} - G_\sigma(x_{\mathrm{HR}})) \|_1\), with weights \(\alpha_2=1.0, \beta_2=0.1, \gamma_2=0.1, \delta_2=0.05\). This stage stabilizes latent properties under decoding. - Stage III (Edge-aware image refinement): Omits latent loss supervision entirely, training in pixel space with image \(\ell_1\), image FFT, and the edge-aware gradient localization (EAGLE) loss: $\(\mathcal{L}_{\mathrm{SIII}} = \alpha_3 \mathcal{L}_{\mathrm{L1}}^x + \beta_3 \mathcal{L}_{\mathrm{FFT}}^x + \gamma_3 \mathcal{L}_{\mathrm{EAGLE}}^x\)$ with weights \(\alpha_3=10.0, \beta_3=1.0, \gamma_3=5 \times 10^{-5}\), sharpening fine boundaries and eliminating staircase artifacts without secondary denoising.
Loss & Training¶
Training is conducted on OpenImages photos (\(\ge 1440\) px on both dimensions), cropped into non-overlapping \(512\times 512\) patches and downsampled bicubically to produce 3.8M pairs encoded via the FLUX VAE (\(s=8, C=16\)). The network is optimized using Adam (\(\text{lr}=2\times 10^{-4}\), weight decay 0, gradient clipping at 0.4), EMA 0.999, and a MultiStepLR schedule. Effective batch sizes are 2,048 for Stage I and 32 for Stages II–III, each stage running for 125k steps, totaling \(\sim 270\) GPU-hours across 8 NVIDIA H100 GPUs. Cross-model adaptation to SDXL and SD3 requires only 500k latent pairs and a brief fine-tuning phase.
Key Experimental Results¶
Main Results¶
Quantitative evaluations are conducted on 1,000 held-out high-resolution OpenImages scenes. Table 1 benchmarks SDXL-anchored configurations across 1024, 2048, and 4096 resolutions against state-of-the-art pipelines on a single H100 GPU (batch size 1).
| Resolution | Method | FID ↓ | pFID ↓ | KID ↓ | pKID ↓ | CLIP ↑ | Time (s) |
|---|---|---|---|---|---|---|---|
| 1024×1024 | HiDiffusion | 232.55 | 230.39 | 0.0211 | 0.0288 | 0.695 | 1.54 |
| DemoFusion | 195.82 | 193.99 | 0.0153 | 0.0229 | 0.725 | 2.04 | |
| LSRNA–DemoFusion | 194.55 | 192.73 | 0.0151 | 0.0228 | 0.734 | 3.09 | |
| SDXL (Direct) | 194.53 | 192.71 | 0.0151 | 0.0225 | 0.731 | 1.61 | |
| SDXL + SwinIR | 210.40 | 204.23 | 0.0313 | 0.0411 | 0.694 | 2.47 | |
| SDXL + LUA (Ours) | 209.80 | 191.75 | 0.0330 | 0.0426 | 0.738 | 1.42 | |
| 2048×2048 | HiDiffusion | 200.72 | 114.30 | 0.0030 | 0.0090 | 0.738 | 4.97 |
| DemoFusion | 184.79 | 177.67 | 0.0030 | 0.0100 | 0.750 | 28.99 | |
| LSRNA–DemoFusion | 181.24 | 98.09 | 0.0019 | 0.0066 | 0.762 | 20.77 | |
| SDXL (Direct) | 202.87 | 116.57 | 0.0030 | 0.0086 | 0.741 | 7.23 | |
| SDXL + SwinIR | 183.16 | 100.09 | 0.0020 | 0.0077 | 0.757 | 6.29 | |
| SDXL + LUA (Ours) | 180.80 | 97.90 | 0.0018 | 0.0065 | 0.764 | 3.52 | |
| 4096×4096 | HiDiffusion | 233.65 | 95.95 | 0.0158 | 0.0214 | 0.698 | 122.62 |
| DemoFusion | 185.36 | 177.89 | 0.0043 | 0.0113 | 0.749 | 225.77 | |
| LSRNA–DemoFusion | 177.95 | 62.07 | 0.0023 | 0.0071 | 0.757 | 91.64 | |
| SDXL (Direct) | 280.42 | 101.89 | 0.0396 | 0.0175 | 0.663 | 148.71 | |
| SDXL + SwinIR | 183.15 | 65.71 | 0.0018 | 0.0103 | 0.756 | 7.29 | |
| SDXL + LUA (Ours) | 176.90 | 61.80 | 0.0015 | 0.0152 | 0.759 | 6.87 |
Table 2 highlights LUA performance across different diffusion backbones (FLUX, SDXL, SD3) and scale factors (\(\times 2\) and \(\times 4\)):
| Scale | Diffusion Model | FID ↓ | pFID ↓ | KID ↓ | pKID ↓ | CLIP ↑ | Time (s) |
|---|---|---|---|---|---|---|---|
| ×2 | FLUX + LUA | 180.99 | 100.40 | 0.0020 | 0.0079 | 0.773 | 29.83 |
| SDXL + LUA | 183.15 | 101.18 | 0.0020 | 0.0087 | 0.753 | 3.52 | |
| SD3 + LUA | 184.58 | 103.94 | 0.0022 | 0.0083 | 0.768 | 20.29 | |
| ×4 | FLUX + LUA | 181.06 | 62.30 | 0.0018 | 0.0085 | 0.772 | 31.91 |
| SDXL + LUA | 182.42 | 71.92 | 0.0015 | 0.0152 | 0.754 | 6.87 | |
| SD3 + LUA | 183.34 | 67.25 | 0.0016 | 0.0095 | 0.769 | 21.84 |
Ablation Study¶
Table 3 validates the contribution of each stage within the progressive curriculum under \(\times 2\) and \(\times 4\) upscaling:
| Configuration | ×2 PSNR ↑ | ×2 LPIPS ↓ | ×4 PSNR ↑ | ×4 LPIPS ↓ | Note |
|---|---|---|---|---|---|
| Latent \(\ell_1\) only | 28.53 | 0.198 | 26.16 | 0.236 | Lacks spectral and perceptual constraints; blurry boundaries |
| Stages I + II (w/o III) | 28.96 | 0.172 | 26.67 | 0.213 | Omits pixel-level edge sharpening; slight visual artifacts persist |
| Stages I + III (w/o II) | 31.05 | 0.150 | 27.10 | 0.198 | Abrupt transition between domains; latent-decoder mismatch |
| Stages II + III (w/o I) | 31.60 | 0.145 | 27.40 | 0.192 | Lacks latent-domain warm-up; suboptimal training stability |
| Full (I + II + III) | 32.54 | 0.138 | 27.94 | 0.184 | Superior reconstruction across both metrics and resolutions |
Table 4 compares architectural upsampling choices:
| Variant | ×2 PSNR ↑ | ×2 LPIPS ↓ | ×4 PSNR ↑ | ×4 LPIPS ↓ | Note |
|---|---|---|---|---|---|
| LIIF (Implicit coordinate decoder) | 29.10 | 0.210 | 26.10 | 0.235 | Continuous coordinates struggle with fine high-frequency textures |
| Separated specialist models | 31.92 | 0.150 | 27.71 | 0.189 | Separate networks miss cross-scale feature regularization |
| Joint Multi-Head (Ours) | 32.54 | 0.138 | 27.94 | 0.184 | Shared representations yield stronger generalization and fidelity |
Key Findings¶
- LUA defines the Pareto frontier at 2048 and 4096 resolutions: at 4K, SDXL+LUA achieves 176.90 FID and 61.80 pFID, outperforming LSRNA–DemoFusion (177.95) while slashing runtime from 91.64 s to 6.87 s (over \(13\times\) speedup).
- Direct high-resolution sampling undergoes catastrophic degradation at 4K (FID collapses to 280.42), exhibiting object duplication in 9–19% of samples, whereas LUA remains immune to structural repetition.
- Cross-VAE adaptation exhibits remarkable sample efficiency: fine-tuning on only 50k target pairs matches established baselines, and scaling to 250k pairs surpasses LSRNA, confirming the transferable nature of latent spatial priors.
- At 1024 resolution, the \(64\times 64\) latent bottleneck limits global FID (209.80 vs. 194.53 for native generation), yet LUA preserves superior patch-level fidelity (pFID 191.75) and text alignment (CLIP 0.738).
Highlights & Insights¶
- Targeting structural efficiency: Bypassing recursive diffusion loops and executing super-resolution directly in latent space harnesses the VAE's spatial stride to reduce compute by \(64\times\) relative to pixel-space models.
- Progressive cross-domain curriculum: The staged progression from latent spectral alignment to hybrid consistency and pixel edge sharpening overcomes training divergence under frozen non-linear decoders.
- Minimalist cross-model transfer: A single shared backbone adapts across distinct generator families (FLUX, SD3, SDXL) merely by swapping the input convolution and fine-tuning briefly on target latents.
Limitations & Future Work¶
- Preserved base hallucinations: As a deterministic operator, LUA upscales existing latent representations without generative re-sampling; structural or semantic errors generated by the base model are magnified rather than corrected.
- Dead pixel anomalies: Rare out-of-distribution latent outputs can trigger localized visual artifacts upon decoding, currently requiring empirical latent value clamping.
- Latent capacity ceiling: Because current decoders downsample by stride 8, initial latents at \(512\text{px}\) retain only \(64\times 64\) tokens; future exploration into smaller-stride or continuous latent representations could unlock higher fidelity.
Related Work & Insights¶
- vs DemoFusion / LSRNA: Reference-guided diffusion methods rely on multi-step noise-denoise loops for upscaling; LUA achieves comparable or better fidelity with a single feed-forward pass, reducing runtime by \(6\times\) to \(13\times\).
- vs SwinIR / SeeSR (Pixel-Space SR): Pixel-domain models process huge spatial grids and incur quadratic compute scaling, often introducing haloing and oversmoothed textures; LUA operates compactly in latent space before decoding, avoiding pixel-level ringing.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Introduces a single-pass latent upscaling paradigm that eliminates secondary diffusion stages, supported by an effective multi-scale curriculum.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarks across 1K/2K/4K scales, three foundation models (FLUX/SD3/SDXL), detailed ablations, and a 2AFC human study.
- Writing Quality: ⭐⭐⭐⭐⭐ Crisp mathematical formulation, clear narrative structure, and high information density across figures and tables.
- Value: ⭐⭐⭐⭐⭐ Provides an immediately deployable, production-ready path for low-latency 2K/4K generative pipelines.