D²R²OSR: Degradation-Disentangled Representation for Real-World Omnidirectional Image Super-Resolution¶
Conference: ECCV 2026
Paper: ECCV Official Link
Full Text Cache: ../paper_cache/ECCV2026/eccv-5315.txt
Code: Pending Release
Area: Image Restoration
Keywords: Omnidirectional Image Super-Resolution, Real-World Degradation Disentanglement, Perspective Projection Representation, Degradation-Specific Modulation, Cross-Domain Projection Fusion
TL;DR¶
Addressing the compounded challenges of fisheye capture degradations and Equirectangular Projection (ERP) geometric distortions in omnidirectional imaging, D²R²OSR introduces a dual-branch network coupled with a Perspective Projection Representation (PPR) that losslessly decouples viewport representations, achieving SOTA real-world ODI-SR performance with superior computational efficiency.
Background & Motivation¶
Omnidirectional images (ODIs), capturing a full 360° field of view (FoV), have become foundational for immersive virtual reality and augmented reality systems. Delivering genuine immersion requires rendering at ultra-high resolutions (4K, 8K, or even 16K). However, in practical deployment, ODIs are severely bandwidth- and sensor-constrained, yielding low-resolution inputs. Crucially, the end-to-end omnidirectional imaging pipeline comprises fisheye lens acquisition, multi-view stitching, Equirectangular Projection (ERP) mapping, compression, and network streaming. Across this pipeline, real-world degradations—such as blur, sensor noise, downsampling, and codec artifacts—are non-linearly entangled with severe latitude-dependent geometric stretching in ERP space, rendering conventional super-resolution frameworks ill-suited.
Existing super-resolution paradigms face an acute dilemma when applied to real-world ODIs. On one hand, conventional ODI-SR approaches (e.g., LAU-Net, OSRT, BPOSR) rely on idealized Bicubic downsampling assumptions and focus strictly on mitigating latitude distortions via geometric priors; they degrade drastically when confronted with unknown, compound real-world degradations. On the other hand, general Real-SR models (e.g., Real-ESRGAN, BSRGAN, AdaIR) construct realistic degradation pipelines but treat inputs as planar, Euclidean images. Consequently, they remain oblivious to non-uniform spherical sampling densities, inducing severe geometric aliasing, boundary seams, and structural deformation in local viewports.
Human visual perception in immersive environments is inherently viewpoint-centric, typically focusing on local viewports spanning 100° to 120° FoV, where projection distortions and sensor-level degradations exhibit distinct physical priors. The core idea of this paper is to decouple omnidirectional super-resolution into an ERP global branch and a Perspective Projection Representation (PPR) viewport branch, leveraging second-order curvature-guided continuous implicit Fourier mapping to disentangle geometric warping from blind real-world degradations, followed by degradation-specific modulation and cross-domain gated fusion for high-fidelity restoration.
Method¶
Overall Architecture¶
Given a low-quality ERP image \(I_{ERP}^{LQ}\) corrupted by mixed geometric and real-world degradations, D²R²OSR aims to reconstruct an ultra-high-resolution ERP image \(I_{ERP}^{HQ}\). The network adopts a symmetric dual-branch architecture. The low-quality input is partitioned into ERP patches \(P_{ERP}^{LQ}\) and simultaneously transformed into perspective patches \(P_{PPR}^{LQ}\) via the proposed Perspective Projection Representation (PPR) with minimal projection loss. The two streams are fed into parallel feature extraction branches composed of \(N=6\) ERP-Adaptive Modulation Blocks (EAMBs) and PPR-Adaptive Modulation Blocks (PAMBs). Within each block, Degradation-Specific Modules (DSMs) adaptively extract latitude-dependent geometric priors for the ERP branch and real-world degradation descriptors for the PPR branch. Finally, Projection Fusion Attention Modules (PFAMs) perform cross-stage, cross-projection feature gating and aggregation to reconstruct the final high-quality ODI.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input: Low-Quality ERP<br/>(ERP Patches)"] --> PPR["Perspective Projection Representation (PPR)<br/>Jacobian/Hessian Priors + Implicit Fourier Mapping"]
In --> EAMB["ERP-Adaptive Modulation Branch<br/>(EAMB)"]
PPR --> PAMB["PPR-Adaptive Modulation Branch<br/>(PAMB)"]
subgraph S1 ["Degradation Disentanglement & Modulation (DSM)"]
direction TB
EAMB --> DSM1["Degradation-Specific Module (DSM)<br/>Latitude-Quantized Geometric Calibration"]
PAMB --> DSM2["Degradation-Specific Module (DSM)<br/>Adaptive Pooled Descriptors + Prompt Injection"]
end
DSM1 --> PFAM["Projection Fusion Attention Module (PFAM)<br/>Spatial-Channel Gating & Depth-wise Aggregation"]
DSM2 --> PFAM
PFAM --> Out["Patch Aggregation & Reconstruction<br/>(High-Quality ERP Output)"]
Key Designs¶
1. Perspective Projection Representation (PPR): High-Order Geometry-Guided Continuous Viewport Mapping Conventional conversion from ERP to local perspective viewports relies on intermediate spherical remapping and interpolation (e.g., OpenCV Remap), which inevitably suffers from interpolation blur and high-frequency textural attenuation (exhibiting over 4.5 dB loss in round-trip reconstruction). PPR establishes a continuous direct mapping between ERP coordinates and target perspective plane coordinates \(p'(u', v')\). To model the non-uniform, latitude-dependent warping accurately, PPR incorporates both the first-order Jacobian matrix \(J_f\) and the second-order Hessian matrix \(H_f\) to capture local orientation and surface curvature. The six independent elements from \(J_f\) and \(H_f\) form a spatial prior vector \(s(\mathcal{P})\). Combined with a local implicit Fourier formulation, a coordinate transformation estimator \(\mathcal{E}\) predicts amplitude, frequency, and phase components: $\(E(\mathcal{E}(p), \delta p, s(\mathcal{P})) = E_a(\mathcal{E}(p)) \otimes \begin{bmatrix} \cos\{\pi(\langle E_f(\mathcal{E}(p)), \delta p \rangle + E_p(s(\mathcal{P})))\} \\ \sin\{\pi(\langle E_f(\mathcal{E}(p)), \delta p \rangle + E_p(s(\mathcal{P})))\} \end{bmatrix}\)$ A lightweight 4-layer MLP with local ensemble coefficients synthesizes perspective pixel distributions. By bypassing discrete spherical resampling, PPR preserves sharp high-frequency structures, allowing the perspective branch to focus purely on restoring real-world degradations within natural image distributions.
2. Degradation-Specific Module (DSM): Dual-Domain Prior Calibration and Attention-Level Prompt Injection Because feature statistics vary significantly across projections—the ERP domain is dominated by latitude-dependent stretching, whereas the PPR domain suffers from isotropic blur, noise, and compression—DSM is embedded into both EAMB and PAMB to perform tailored modulation. For the ERP branch, a pre-computed latitude quantization map (\(H \times W \times 1\)) calibrates spatial distortion. Within both branches, features undergo \(3 \times 3\) convolutions, LeakyReLU, and Adaptive Average Pooling (AAP) to generate compact, learnable degradation descriptors \(D_{ERP}^i\) and \(D_{PPR}^i\). To avoid heavy computation or corrupting query-key semantic correspondences, authors introduce Attention-Level Injection (ALI). The degradation descriptor is projected into a multi-head-aligned prompt tensor \(P_v\) and injected exclusively into the Value component of window self-attention (\(V' = V + P_v\)). This retains precise query-key structural matching while dynamically adapting feature aggregation intensity based on the local degradation level.
3. Projection Fusion Attention Module (PFAM): Cross-Projection Spatial-Channel Gating and Inter-Stage Depth Attention ERP and PPR features reside in heterogeneous coordinate spaces and have asymmetric channel capacities (\(C_e = 180\), \(C_p = 60\)). Direct concatenation induces severe spatial misalignments. PFAM first aligns PPR channel dimensions to \(C_e\) using a learnable \(1 \times 1\) convolution. Concatenated features then pass through convolutional layers with LeakyReLU and Sigmoid activations to derive adaptive spatial gate maps \(S_e\) and \(S_p\), executing residual complementary fusion: \(F^i = S_e \odot F_e^i + S_p \odot \tilde{F}_p^i\). Furthermore, to model long-range hierarchical dependencies across different network depths, PFAM reshapes multi-stage features into query and value tensors and computes depth-wise attention via \(\text{Softmax}(QQ^\top)\). This resolves inter-stage feature inconsistency and ensures global geometric continuity across the reconstructed panorama.
Loss & Training¶
D²R²OSR is trained end-to-end without explicit degradation category or severity annotations. The training objective relies strictly on the pixel-level \(\ell_1\) reconstruction loss between the predicted high-resolution panorama and the ground truth: $\(\mathcal{L} = \| I_{ERP}^{HQ} - \hat{I}_{ERP}^{HQ} \|_1\)$ Implemented in PyTorch, the network is optimized on a single NVIDIA A6000 GPU using the Adam optimizer (\(\beta_1=0.9, \beta_2=0.99\)) for 500K iterations with an initial learning rate of \(1 \times 10^{-4}\) and a batch size of 4. Low-quality inputs are simulated using a realistic fisheye-to-ERP compound degradation pipeline, cropped into \(256 \times 256\) patches with random horizontal flips during training.
Key Experimental Results¶
Main Results¶
Quantitative evaluations are conducted across Flickr360, ODI-SR, and SUN360 benchmarks under \(\times 4\), \(\times 8\), and extreme \(\times 16\) scaling factors. Metrics include Weighted Spherical Peak Signal-to-Noise Ratio (WS-PSNR) and Weighted Spherical Structural Similarity (WS-SSIM) on the luminance (Y) channel.
| Dataset | Scale | 360-SS | Real-ESRGAN | SwinIR | HAT | OSRT | AdaIR | D²R²OSR (Ours) |
|---|---|---|---|---|---|---|---|---|
| Flickr360 | \(\times 4\) | 24.20 / 0.6198 | 24.30 / 0.6171 | 25.16 / 0.6827 | 25.19 / 0.6838 | 25.41 / 0.6886 | 25.49 / 0.6912 | 25.64 / 0.6934 |
| \(\times 8\) | 16.98 / 0.5728 | 21.85 / 0.6037 | 22.83 / 0.6347 | 22.93 / 0.6354 | 22.93 / 0.6354 | 22.96 / 0.6377 | 23.24 / 0.6433 | |
| \(\times 16\) | 15.91 / 0.5844 | 19.84 / 0.5934 | 21.19 / 0.6202 | 21.36 / 0.6206 | 21.41 / 0.6211 | 21.45 / 0.6238 | 21.64 / 0.6249 | |
| ODI-SR | \(\times 4\) | 23.30 / 0.6031 | 23.24 / 0.6001 | 23.97 / 0.6537 | 23.97 / 0.6541 | 24.22 / 0.6569 | 24.26 / 0.6591 | 24.37 / 0.6626 |
| \(\times 8\) | 17.11 / 0.5562 | 21.52 / 0.5904 | 22.32 / 0.6129 | 22.41 / 0.6139 | 22.39 / 0.6135 | 22.42 / 0.6158 | 22.62 / 0.6210 | |
| \(\times 16\) | 15.97 / 0.5667 | 19.61 / 0.5785 | 20.77 / 0.5996 | 20.92 / 0.6000 | 20.96 / 0.6004 | 21.00 / 0.6036 | 21.18 / 0.6053 | |
| SUN360 | \(\times 4\) | 23.02 / 0.5943 | 23.07 / 0.5922 | 23.66 / 0.6439 | 23.67 / 0.6444 | 23.82 / 0.6484 | 23.85 / 0.6502 | 24.00 / 0.6528 |
| \(\times 8\) | 16.64 / 0.5728 | 21.23 / 0.5926 | 22.04 / 0.6296 | 22.08 / 0.6300 | 22.13 / 0.6301 | 22.14 / 0.6306 | 22.16 / 0.6328 | |
| \(\times 16\) | 15.14 / 0.5875 | 19.34 / 0.5793 | 20.48 / 0.6176 | 20.55 / 0.6177 | 20.58 / 0.6180 | 20.59 / 0.6185 | 20.69 / 0.6184 |
Ablation Study¶
Ablations on Flickr360 under real-world \(\times 4\) settings demonstrate the progressive gains contributed by each design:
| Config / Stage | Branch Structure | Representation | WS-PSNR / WS-SSIM | Note |
|---|---|---|---|---|
| Single Baseline | Single | ERP | 25.20 / 0.6858 | Standard single ERP branch baseline |
| Perspective Only | Single | Perspective (Remap) | 25.09 / 0.6820 | Perspective alone with naive remapping degrades |
| Dual-Baseline | Dual | ERP + ERP | 25.35 / 0.6865 | Pure channel expansion yields limited gain (+0.15 dB) |
| Baseline + Pers | Dual | ERP + Perspective (Remap) | 25.47 / 0.6882 | Adding remapped perspective branch (+0.27 dB) |
| Baseline + PPR | Dual | ERP + PPR (Ours) | 25.50 / 0.6892 | Implicit Fourier-based PPR representation (+0.30 dB) |
| + \(DSM_{ERP\&PPR}\) | Dual | ERP + PPR | 25.58 / 0.6910 | Dual-domain degradation modulation (+0.08 dB) |
| + ALI (Prompt Injection) | Dual | ERP + PPR | 25.57 / 0.6918 | Value-only injection enhances structure (SSIM +0.0026) |
| Full D²R²OSR (+PFAM) | Dual | ERP + PPR | 25.64 / 0.6934 | Full model with cross-domain gated fusion (+0.07 dB) |
Computational efficiency benchmarks on SUN360 (\(256 \times 512\) input, NVIDIA A6000): - Parameters: D²R²OSR requires only 3.50M parameters (drastically lower than AdaIR's 28.90M and OSRT's 6.02M); - FLOPs: Requires 0.55 TFLOPs (vs. HAT's 2.84T and SwinIR's 1.70T); - Inference Time: Achieves 0.34 s per image, nearly \(4\times\) faster than OSRT (1.33 s).
Round-trip reconstruction fidelity across 1,200 images (ERP \(\to\) Viewport \(\to\) ERP): - Equator (0° latitude): OpenCV Remap achieves 29.28 dB, while PPR reaches 34.89 dB (+5.61 dB); - Mid-to-high latitude (60° latitude): OpenCV Remap reaches 36.21 dB, while PPR achieves 40.07 dB (+3.86 dB); - Overall Average: PPR reaches 38.17 dB / 0.9728 compared to OpenCV Remap's 33.63 dB / 0.9295, delivering an average gain of +4.54 dB.
Key Findings¶
- PPR Decoupling Resolves Severe Projection Loss: Relying solely on conventional interpolation remapping drops reconstruction fidelity by 4.54 dB. The Hessian and Jacobian guided continuous implicit representation preserves structural fine details (e.g., license plate alphanumeric text "AUSS") while enabling the perspective branch to focus on blind real-world degradation.
- Attention-Level Prompt Injection (ALI) Maintains Topology: Injecting degradation prompts exclusively into self-attention Value tensors avoids distorting Query-Key spatial affinity matrices while dynamically scaling restoration intensity, preserving crisp object boundaries with zero extra attention heads.
- Extreme Upscaling Robustness: Under extreme \(\times 16\) scaling on Flickr360, standard Real-ESRGAN plummets to 19.84 dB, whereas D²R²OSR retains 21.64 dB (+1.80 dB), proving exceptional resilience in low-bandwidth scenarios.
Highlights & Insights¶
- Embedding High-Order Geometry into Fourier Continua: Rather than utilizing Jacobian and Hessian matrices purely for geometric warping, the paper integrates them as continuous spatial priors within an implicit neural Fourier framework, conquering the classic trade-off between projection fidelity and interpolation blur.
- Value-Only Modulation for Lightweight Attention: Modulating exclusively the Value path (\(V' = V + P_v\)) establishes an elegant, plug-and-play paradigm for multi-degradation vision transformers, bypassing burdensome cross-attention networks.
- Domain Disentanglement Simplifies Blind Restoration: By breaking an ill-posed blind omnidirectional problem into an ERP geometric prior calibration task and a perspective natural-image restoration task, the architecture delivers SOTA performance with an ultra-compact 3.50M parameter budget.
Limitations & Future Work¶
- Static Viewports Without Dynamic Gaze Tracking: The current PPR samples fixed tangent viewports rather than dynamically adjusting resolution or rendering density according to real-time VR foveated eye-tracking signals.
- Absence of Generative Diffusion Priors: While extraordinarily fast and lightweight (0.34 s runtime), under extreme information loss (e.g., heavily corrupted polar regions), discriminative \(\ell_1\) optimization inevitably tends toward blurred statistical averages; integrating lightweight latent diffusion priors represents a promising frontier.
Related Work & Insights¶
- vs OSRT / BPOSR (Conventional ODI-SR): Prior ODI-SR models focus heavily on ERP latitude deformation (e.g., continuous coordinate offsets or cubemap branches) but assume synthetic Bicubic downsampling. D²R²OSR is the first to formalize realistic fisheye-to-ERP composite degradations and decouple them via dual-branch PPR.
- vs Real-ESRGAN / AdaIR (Real-World Image Restoration): Generic real-world SR models lack spherical and latitude consciousness, frequently mistaking ERP geometric elongation for camera blur and generating severe ringing artifacts around polar latitudes. D²R²OSR injects geometric quantization priors via DSM to ensure latitude-consistent restoration.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ [Pioneering degradation disentanglement for real-world ODIs; clever integration of second-order differential geometry into implicit neural representations]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Rigorous evaluations across 3 benchmarks, 3 scaling factors (\(\times 4/\times 8/\times 16\)), comprehensive round-trip fidelity tests, and runtime complexity analyses]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation narrative, disciplined mathematical formulation, and self-consistent structural alignment]
- Value: ⭐⭐⭐⭐☆ [Compact 3.5M parameters with 0.34s inference; highly deployable for client-side VR/AR viewport rendering and streaming enhancement]