Extreme Face Super-Resolution through Identity Fitting and Decoupling¶
Conference: ECCV 2026
Paper: ECCV Official Page
Area: Image Restoration
Keywords: Face Super-Resolution, Identity Disentanglement, Identity Fitting, Diffusion Models, Reference-Based Restoration
TL;DR¶
Addressing severe identity loss and pervasive facial hallucination in extreme degradation regimes (8ร/16ร scaling factors), IDFSR masks corrupted low-resolution facial regions to eliminate misleading cues, injects landmark-warped reference features via style-adaptive normalization alongside ArcFace identity embeddings into diffusion attention layers, and achieves unprecedented identity consistency and perceptual fidelity via lightweight test-time identity embedding optimization.
Background & Motivation¶
Face super-resolution (FSR) serves as a cornerstone technology for downstream applications including biometric identification, automated facial attribute analysis, and historical media restoration. Under standard, moderate degradation regimes, modern generative restoration frameworks (such as CodeFormer and DifFace) effectively leverage learned codebook dictionaries or diffusion-driven stochastic trajectories to synthesize realistic facial textures while maintaining satisfactory identity consistency. However, when images undergo extreme degradationโsuch as scaling factors exceeding 8ร or 16ร compounded by heavy sensor noise, severe motion blur, and aggressive JPEG compressionโthe intrinsic fine-grained facial geometry, gaze direction, and delicate identity features within the low-resolution (LR) input are almost completely obliterated. From an information-theoretic standpoint, the inverse reconstruction problem becomes severely underdetermined, yielding an expansive solution space of visually plausible yet identity-discrepant facial hypotheses. Consequently, conventional reference-free generative models inevitably fabricate unrealistic, hallucinated facial details that fail to retain true individual identityโfor instance, turning an unresolvable blurred tattoo into an unnatural dark artifact.
Reference-based face restoration methods provide a natural avenue to bypass this fundamental limit by harvesting high-frequency facial priors from same-identity high-resolution reference portraits. Nonetheless, traditional single-reference paradigms depend heavily on explicit, rigid or dense pixel-level alignment between the LR target and the reference exemplar; substantial discrepancies in head pose, camera lighting, or facial expressions routinely cause optical flow and warping algorithms to fail, propagating catastrophic visual artifacts. Conversely, multi-reference schemes aggregate broader identity distributions but uniformly fail to disentangle the identity cues inherent to the LR input from those provided by reference exemplars. Because super-resolution remains a pixel-level reconstruction task conditioned on the LR substrate, un-disentangled networks face a destructive trade-off: they either over-rely on the corrupted, misleading low-frequency artifacts surviving in the LR input, or blindly clone misaligned high-frequency textures from reference images, failing to faithfully restore authentic identity signatures.
The core angle of attack in this work is that severely degraded LR facial pixels act as corrupting noise rather than helpful constraints, warranting active elimination; simultaneously, reference portraits must be cleanly decoupled into structural style guidance and compact identity manifolds. Core idea: mask out corrupted low-resolution facial regions to prevent hallucination from degraded cues, decouple reference inputs into style embeddings injected via adaptive group normalization and identity embeddings infused through cross-attention, and freeze the pretrained diffusion backbone to perform lightweight personalized fitting on a compact target identity embedding.
Method¶
Overall Architecture¶
The input to IDFSR (Identity Decoupling and Fitting Face Super-Resolution) comprises a severely degraded low-resolution face \(I_L\), its paired ground-truth high-resolution image \(I_G\) (available during training), and a same-identity high-resolution reference image \(I_R\). The backbone is built upon a parameterized Denoising Diffusion Probabilistic Model (DDPM) following the Guided Diffusion U-Net topology, comprising residual blocks (ResBlocks), spatial self-attention and cross-attention blocks (AttnBlocks), an auxiliary style encoder \(E_s\), and a frozen ArcFace identity feature extractor \(E_{id}\).
During data preprocessing, a fine-tuned RetinaFace detector localizes the facial region of \(I_L\) and applies a rectangular/contour mask, generating the corrupted face-masked image \(I_M\); simultaneously, landmark-based affine transformation warps the facial region of \(I_R\) onto the spatial coordinates of \(I_L\), yielding the warped reference \(I_W\). In the pretraining phase, the masked image \(I_M\) is concatenated channel-wise with the noisy latent \(x_t\) to serve as the U-Net input, completely preventing the model from drawing corrupted identity cues from \(I_L\). The warped reference \(I_W\) passes through style encoder \(E_s\) to yield a style embedding \(z_s\), which modulates ResBlocks via Adaptive Group Normalization (AdaGN). Concurrently, the ground-truth image \(I_G\) is processed by \(E_{id}\) into a 512-dimensional compact ID embedding \(z_{id}\), injected into AttnBlocks via cross-attention. In the subsequent personalized fine-tuning phase, all parameters of the U-Net and style encoder are frozen; \(z_{id}\) is replaced with a learnable continuous vector, which is rapidly optimized using a handful of target identity portraits.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
InL["Degraded LR Input IL<br/>Extreme 8ร / 16ร Scaling"] --> PrepM["Corrupted ID Masking<br/>RetinaFace Face Occlusion"]
PrepM --> IM["Masked Condition IM<br/>Preserves Background Only"]
InR["HQ Reference Image IR<br/>Same ID / Varying Pose"] --> PrepW["Landmark-Based Warping<br/>Affine Alignment to IL Space"]
PrepW --> IW["Warped Reference IW"]
IW --> Es["Style Encoder Es"]
Es --> zs["Style Embedding zs"]
InG["HQ Ground-Truth IG<br/>(Target ID Set in FT Phase)"] --> Eid["ArcFace ID Encoder Eid<br/>Extracts 512-D Identity Priors"]
Eid --> zid["Identity Embedding zid<br/>(Learnable Vector in FT Phase)"]
IM --> Concat["Channel-wise Concatenation<br/>[xt, IM]"]
Noisy["Noisy Diffusion Latent xt<br/>Timestep t"] --> Concat
Concat --> UNet["Diffusion U-Net Backbone<br/>(Frozen During FT Stage)"]
zs -->|AdaGN Affine Modulation| UNet
zid -->|Cross-Attention Injection| UNet
UNet --> Pred["Noise Prediction ฮตฮธ<br/>DDIM 20-Step Sampling"]
Pred --> Out["Restored HR Face<br/>High-Fidelity & ID-Consistent"]
Key Designs¶
1. Corrupted ID Masking: Eradicating Degraded Artifacts and Enforcing Background Consistency Under extreme 8ร or 16ร downsampling and noise, the remaining facial pixels in \(I_L\) lose all resolvable biometric structure and become deceptive noise. Conditioning directly on unmasked LR pixels causes diffusion models to overfit to blurred low-frequency patterns, muting the impact of reference guidance. IDFSR addresses this bottleneck by identifying the face bounding box via a specialized RetinaFace detector and replacing the interior facial area with zeros to produce \(I_M\), which is concatenated with noisy latent \(x_t\). This effectively reformulates face super-resolution into an inpainting-inspired generation task: hair contours, shoulder silhouettes, and peripheral background scene textures are preserved with pristine spatial coordinates, while the blanked face interior forces the diffusion backbone to synthesize anatomical facial geometry strictly guided by external style and identity embeddings.
2. Fault-Tolerant Warping & Style-Condition Embedding: Coarse Structure Prior without Alignment Rigidity Relying solely on a 1D global identity vector fails to capture head pose, illumination gradients, and skin tone. IDFSR performs facial landmark detection to affinely warp the reference image \(I_R\) onto \(I_L\), creating \(I_W\). Crucially, the authors recognize that real-world pose discrepancies inevitably make this warping imperfect and prone to spatial stretching. Instead of viewing this as a flaw, IDFSR treats warping imperfections as a form of natural data augmentation that discourages trivial pixel copy-pasting. \(I_W\) is processed by style encoder \(E_s\) into a style embedding \(z_s = E_s(I_W)\), which is injected into the diffusion backbone via Adaptive Group Normalization (AdaGN): $$ \text{AdaGN}(h, t, z_s) = z_f \cdot \left( t_s \cdot \text{GroupNorm}(h) + t_b \right) $$ where \(h \in \mathbb{R}^{c \times h \times w}\) denotes the intermediate feature map, \(z_f = \text{MLP}_{\text{style}}(z_s)\) provides channel-wise scale modulation, and \((t_s, t_b) = \text{MLP}_{\text{time}}(\psi(t))\) parameterizes the sinusoidal time embedding \(\psi(t)\). AdaGN dynamically modulates global tone, skin illumination, and structural plausibility without enforcing rigid spatial correspondences.
3. Disentangled ID Pretraining and Personalized ID Fitting: Two-Stage Identity Customization To cleanly decouple general human facial distribution modeling from subject-specific identity signatures, IDFSR implements a rigorous two-stage training scheme. During pretraining across diverse subjects on CelebRef-HQ, a frozen ArcFace encoder maps ground-truth images \(I_G\) into compact identity representations \(z_{id} = E_{id}(I_G)\), conditioned into U-Net attention blocks via cross-attention. This equips the diffusion network with a rich understanding of how 1D identity codes steer 2D facial anatomy. In the inference and personalized restoration stage, extracting an ID code from a single degraded or off-pose reference image risks introducing severe viewpoint bias. IDFSR completely freezes all weights of the pretrained U-Net and style encoder \(E_s\), eliminating the need for ArcFace at test time. Instead, a lightweight continuous embedding vector \(z_{id}^*\) is instantiated and fine-tuned using a small gallery (e.g., ~5 images) of the target subject. Because the backbone is locked, the model retains its robust generative prior without catastrophic forgetting, while the fitted embedding precisely locks onto the subject's distinctive ocular contours, nasal bridge, lip lines, and jaw shape.
Loss & Training¶
Both pretraining and personalized fine-tuning stages employ the standard simplified DDPM mean-squared error objective. With forward noise schedule \(T = 1000\) and linear \(\beta\) parameters, the noisy state \(x_t\) at timestep \(t\) is denoised via: $$ \mathcal{L}{\text{sim}} = \mathbb{E}) |_2^2 \right] $$ Pretraining is conducted on 4 NVIDIA RTX 3090 GPUs for 2 days, initialized from DiffAE weights pretrained on FFHQ. Personalized fine-tuning requires only ~20 minutes on a single RTX 3090 GPU per identity, optimizing exclusively the }, I_M} \left[ | \epsilon - \epsilon_\theta([x_t, I_M], t, z_s, z_{id\(z_{id}\) vector. During inference, DDIM deterministic sampling with only 20 steps is employed to produce final high-resolution restorations.
Key Experimental Results¶
Main Results¶
Quantitative evaluations are conducted on the CelebRef-HQ benchmark, utilizing a curated test split of 56 identities and 586 images under severe \(8\times\) and \(16\times\) downscaling factors. Metrics assess reconstruction fidelity (PSNR, SSIM), perceptual distance (LPIPS, FID, CLIPIQA, MUSIQ), and facial identity distance measured via DeepFace (IDS, lower is better). As shown below, IDFSR achieves state-of-the-art results across both degradation scales.
| Method | Scale | PSNR (dB)โ | SSIMโ | LPIPSโ | FIDโ | CLIPIQAโ | MUSIQโ | IDS (DeepFace)โ |
|---|---|---|---|---|---|---|---|---|
| CodeFormer | 8ร | 24.42 | 0.7147 | 0.1489 | 31.88 | 0.6873 | 68.98 | 0.3865 |
| DifFace | 8ร | 24.41 | 0.7255 | 0.1694 | 32.94 | 0.6132 | 60.54 | 0.3815 |
| ASFFNet (Ref-based) | 8ร | 21.82 | 0.6624 | 0.1764 | 26.25 | 0.6158 | 65.81 | 0.3169 |
| DMDNet (Ref-based) | 8ร | 22.36 | 0.6984 | 0.1733 | 29.76 | 0.6570 | 68.28 | 0.4169 |
| IDFSR (Pretrained) | 8ร | 24.29 | 0.6996 | 0.1554 | 27.89 | 0.6798 | 69.30 | 0.3173 |
| IDFSR (Fine-tuned) | 8ร | 28.85 | 0.7604 | 0.1031 | 26.69 | 0.7092 | 73.73 | 0.2242 |
| CodeFormer | 16ร | 20.72 | 0.5947 | 0.2379 | 43.89 | 0.6834 | 67.85 | 0.6673 |
| PGDiff | 16ร | 20.47 | 0.6202 | 0.2727 | 55.70 | 0.5008 | 52.19 | 0.7355 |
| DR2 | 16ร | 22.35 | 0.6356 | 0.2481 | 50.84 | 0.6102 | 62.60 | 0.7082 |
| DifFace | 16ร | 22.23 | 0.6737 | 0.2195 | 38.65 | 0.5957 | 57.84 | 0.5948 |
| ASFFNet (Ref-based) | 16ร | 18.23 | 0.5732 | 0.2457 | 40.64 | 0.5827 | 63.72 | 0.7955 |
| DMDNet (Ref-based) | 16ร | 17.38 | 0.5590 | 0.2497 | 46.34 | 0.6051 | 64.05 | 0.8291 |
| IDFSR (Pretrained) | 16ร | 23.32 | 0.6863 | 0.2062 | 38.58 | 0.6723 | 68.56 | 0.5227 |
| IDFSR (Fine-tuned) | 16ร | 24.09 | 0.7158 | 0.1845 | 35.50 | 0.6992 | 71.83 | 0.3625 |
Ablation Study¶
The evaluation of biometric verification, facial attribute preservation under \(16\times\) degradation (Table A), and computational resource profiling on a single RTX 3090 GPU (Table B) are summarized below.
Table A: Quantitative Comparison of Face ID Verification (IDV) and Attribute Consistency at 16ร Scale
| Method / Configuration | ID Verification (IDV)โ | Gender Accuracy (ACCGen)โ | Emotion Accuracy (ACCEmo)โ | Race Accuracy (ACCRa)โ | Age Difference (DiffAge)โ |
|---|---|---|---|---|---|
| CodeFormer | 50.3% | 95.2% | 64.3% | 64.8% | 4.6 ยฑ 4.3 |
| PGDiff | 32.9% | 88.7% | 51.0% | 68.8% | 5.2 ยฑ 4.6 |
| DR2 | 38.9% | 91.9% | 57.6% | 71.5% | 5.2 ยฑ 5.1 |
| DifFace | 73.0% | 95.4% | 67.4% | 77.3% | 4.6 ยฑ 4.4 |
| ASFFNet | 21.4% | 72.1% | 39.6% | 45.2% | 7.5 ยฑ 7.2 |
| DMDNet | 18.7% | 66.3% | 36.8% | 42.5% | 7.6 ยฑ 7.3 |
| IDFSR (Fine-tuned) | 89.6% | 98.7% | 66.2% | 96.6% | 3.3 ยฑ 1.7 |
Table B: Complexity and Runtime Profile on a Single NVIDIA RTX 3090 GPU
| Method | Parameters (M) | FLOPs (G) | NFEs (Sampling Steps) | Step Latency (ms) | Total Inference Time (s) |
|---|---|---|---|---|---|
| PGDiff | 159.7 | 185.95 | 100 | 71 | 7 |
| DR2 | 179.31 | 918.85 | 100 | 89 | 9 |
| DifFace | 175.42 | 272.67 | 100 | 70 | 7 |
| StableSR | 918.93 | >2000 | 50 | 362 | 18 |
| DPI | 145.53 | 136.57 | 20 | 50 | 1 |
| IDFSR (Ours) | 160.7 | 168.68 | 20 | 54 | 1 |
Key Findings¶
- Breakthrough in Extreme Identity Retention: Under 16ร degradation, baseline reference-driven models like ASFFNet and DMDNet experience total alignment failure, with their ID verification rates plummeting to 21.4% and 18.7% and identity distances soaring above 0.79. In contrast, fine-tuned IDFSR sustains an 89.6% ID verification rate and an IDS of 0.3625, proving that decoupling reference style from identity vectors effectively insulates the model from alignment degradation.
- Component Contributions and Saturation Dynamics: Progressive ablations demonstrate that directly concatenating unmasked LR inputs (Case 1) enforces excessive rigid constraints that degrade perceptual realism, whereas masking alone (Case 2) increases generative freedom without identity grounding. Omitting ID embeddings (Case 4) severely undermines facial recognition, while dropping style embeddings (Case 5) compromises texture realism. Furthermore, reference image scaling experiments show that personalization performance plateaus at approximately 5 reference portraits, highlighting high sample efficiency.
- Superior Practical Efficiency: Requiring only 20 DDIM sampling steps and consuming 168.68 GFLOPs, IDFSR achieves full-resolution reconstruction within 1 second on a single consumer GPU, outperforming multi-step diffusion baselines requiring 7 to 18 seconds.
Highlights & Insights¶
- Active Masking as an Antidote to Corrupted Cues: Rather than treating degraded low-resolution inputs as sacrosanct constraints, IDFSR recognizes that severely corrupted facial pixels operate as adversarial noise. Masking the face interior forces the diffusion network to draw deterministic identity features strictly from decoupled reference spaces.
- Zero-Backbone-Drift Personalization: Freezing the generative core and optimizing only a low-dimensional identity vector circumvents the catastrophic forgetting and heavy computational overhead typical of full-parameter fine-tuning, requiring only 20 minutes on a single GPU to master a new identity.
Limitations & Future Work¶
- Subtle Expression Discrepancies: Because the interior facial region is masked out, momentary micro-expressions (such as subtle smiles or eye narrows) are reconstructed predominantly from reference style and learned priors. Consequently, emotion consistency (66.2%) slightly lags behind unmasked models like DifFace (67.4%).
- Dependency on Personal Reference Sets for Maximum Fidelity: While the pretrained model provides robust generic restoration, the leap to 89.6% verification accuracy depends on acquiring 3โ5 clean reference portraits for fine-tuning, which may not always be accessible in strictly unconstrained surveillance scenarios.
Related Work & Insights¶
- vs CodeFormer / DifFace: These blind restoration models rely on discrete codebook priors or diffused error contraction. While competitive under 8ร scaling, their lack of explicit identity conditioning leads to hallucinated identities under 16ร degradation. IDFSR raises ID verification from 50.3% / 73.0% to 89.6%.
- vs ASFFNet / DMDNet: Prior reference-based models depend on optical flow or moving least squares to deform exemplar features. Under extreme degradation, alignment crashes and introduces catastrophic artifacts. IDFSR absorbs alignment errors as data augmentation, decoupling style via AdaGN from identity attention.
Rating¶
- Novelty: โญโญโญโญโญ [Pioneering integration of corrupted face masking, AdaGN style modulation, and compact identity vector fitting for extreme FSR]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive comparisons across 8ร/16ร scales, biometric attribute consistency, computational latency, and progressive degradation]
- Writing Quality: โญโญโญโญโญ [Clear motivation, technically sound formulations, and insightful architectural justifications]
- Value: โญโญโญโญโญ [High practical utility for forensic face enhancement, long-range surveillance, and legacy portrait restoration]