SVI360: Spherical Video Interpolation¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://icb-vision-ai.github.io/video360_interpolation/
Area: Video Understanding
Keywords: spherical video interpolation, omnidirectional optical flow estimation, dual-branch architecture, equivariance priors, multi-field fusion
TL;DR¶
To tackle severe polar distortions and correspondence failures under large displacements in equirectangular projection, SVI360 introduces a dual-branch coarse-to-fine framework that leverages rotated orthogonal views to iteratively refine optical flow and intermediate features, achieving state-of-the-art omnidirectional video frame interpolation.
Background & Motivation¶
Spherical (omnidirectional) video is vital for virtual reality, immersive telepresence, and robotics. In head-mounted displays, ultra-high frame rates (such as 90โ120 fps) are indispensable for delivering smooth visual transitions and mitigating cybersickness. However, streaming uncompressed or high-frame-rate 360ยฐ videos demands enormous network bandwidth. Video frame interpolation (VFI) offers an appealing solution by synthesizing intermediate frames on local client devices from lower frame rate streams. Nevertheless, directly deploying state-of-the-art perspective VFI models (e.g., FILM, RIFE, AMT) onto equirectangular projection (ERP) videos causes severe performance degradation, producing prominent ghosting artifacts and blurring near the poles.
This failure stems from a fundamental conflict: standard 2D VFI architectures rely on planar projection assumptions with spatially uniform pixel distributions. In contrast, equirectangular projection exhibits extreme spatial non-uniformityโequatorial areas preserve geometric proportions with low distortion, whereas polar latitudes suffer from massive lateral stretching and disproportionately large flow displacements. While 360VFI introduced distortion maps and learned deformable convolutions to mitigate polar stretch, its single-view mechanism degrades rapidly under medium-to-large motions. Furthermore, although recent optical flow works such as PriOr-Flow demonstrated the benefits of orthogonal rotations for motion estimation, frame synthesis demands cross-view feature alignment and distortion-aware texture synthesis far beyond isolated optical flow calculation.
The entry point of this paper is to exploit the rotational equivariance of spherical geometry by introducing a 90ยฐ rotated orthogonal view that maps heavily distorted polar regions onto the distortion-free equatorial plane. Core idea: construct an orthogonal dual-branch coarse-to-fine architecture that iteratively refines flow and intermediate features via dual-cost collaborative lookup and a spherical refiner, combines multi-candidate frame synthesis from both views, and optimizes reconstruction using a latitude-weighted Charbonnier loss.
Method¶
Overall Architecture¶
Given two consecutive spherical ERP frames \(I_0^p, I_1^p\), SVI360 synthesizes high-fidelity intermediate frames \(I_t\) at arbitrary timesteps \(t \in (0, 1)\). The overall pipeline operates in a coarse-to-fine multi-stage decoder: the input pair is first projected via a 90ยฐ rotation around the \(x\)-axis into orthogonal views \(I_0^o, I_1^o\). Both views are processed by separate multi-scale feature encoders. Across decoder stages, dense all-pairs correlation volumes are queried using Dual-Cost Collaborative Lookup (DCCL) within an AMT-style recurrent block. An aligned Spherical Refiner module subsequently injects orthogonal flow and intermediate feature cues into the primitive branch. Finally, at the finest scale, both branches output multiple candidate synthesis fields that are rotated and adaptively fused through a lightweight convolutional head.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input Consecutive Spherical Frames<br/>I0_p, I1_p"] --> Rot["Orthogonal View Projection<br/>90-deg rotation around x-axis to I0_o, I1_o"]
Rot --> Feat["Dual-Branch Feature Encoding & Correlation Pyramid"]
Feat --> Upd["AMT-style DCCL Recurrent Update<br/>Joint flow and feature refinement"]
Upd --> SR["Spherical Refiner<br/>Inverse rotation alignment & cross-view residual injection"]
SR --> MF["Spherical Multi-field Refinement<br/>Dual-branch N-candidate synthesis & fusion"]
MF --> Out["Output High-Fidelity Interpolated Frame It"]
Key Designs¶
1. Orthogonal View Projection: Relocating High-Latitude Distortions to Low-Distortion Zones Equirectangular projection incurs severe longitudinal expansion near the poles (\(v \to 0\) or \(v \to H\)). Standard correlation volumes computed directly on polar regions suffer from local minima and mismatched correspondences. To establish a robust geometric prior, input pixel coordinates \((u, v)\) are mapped through spherical coordinates \((\theta, \phi)\) to 3D Cartesian coordinates \((x, y, z)\), followed by a rigid 90ยฐ rotation around the \(x\)-axis denoted by \(\mathbf{T}_p^o\). This transformation shifts the distorted polar contents to the equatorial region of the orthogonal frame, enabling feature matching under minimal distortion and yielding complementary input pairs \(\{\mathbf{I}_0^o, \mathbf{I}_1^o\}\).
2. AMT-style DCCL Recurrent Update: Dual-Branch Joint Flow and Feature Decoupling To reliably track large-displacement motions, the model builds 4-level correlation pyramids in both primitive and orthogonal branches. In each iteration, Dual-Cost Collaborative Lookup (DCCL) retrieves correlation responses from both the current branch and its counterpart rotated branch according to the current flow estimate. Lightweight recurrent update blocks \(\mathcal{U}_1\) take the current flow, multi-view correlation features, and intermediate representations as input, predicting residual updates for both optical flow and intermediate latent features simultaneously rather than updating flow in isolation.
3. Spherical Refiner: Cross-View Equivariant Motion and Feature Alignment To directly rectify distortions in the primitive branch across pyramid levels, the Spherical Refiner (SR) acts as a cross-view residual injection unit. Estimated flows \(\mathbf{f}^o\) and intermediate features \(\mathbf{X}_t^o\) from the orthogonal branch are transformed back to the primitive coordinate frame via inverse projection: $\(\mathbf{f}^{o\rightarrow p} = \mathbf{T}_o^p(\mathbf{f}^o), \qquad \mathbf{X}_t^{o\rightarrow p} = \mathbf{T}_o^p(\mathbf{X}_t^o)\)$ Update network \(\mathcal{U}_2\) processes primitive representations, aligned orthogonal representations, and freshly looked-up correlation maps to predict fine residual corrections \(\Delta \mathbf{f}^p\) and \(\Delta \mathbf{X}_t^p\) solely for the primitive branch, achieving hierarchical geometric error suppression across stages.
4. Spherical Multi-field Refinement: Multi-Candidate Synthesis Mitigating Singular Occlusions In severe occlusion or polar wrap-around areas, a single bilateral flow cannot adequately resolve ambiguous pixel trajectories. SVI360 equips the finest decoder stage of each branch to predict \(N\) candidate sets of bilateral flow fields, occlusion masks \(\mathbf{M}\), and residual images \(\mathbf{R}\) (set to \(N = 5\)). The \(N\) candidate frames synthesized in the orthogonal branch are re-projected onto the primitive sphere: $\(\mathbf{I}_t^{o2p, n} = \mathbf{T}_o^p(\mathbf{I}_t^{o, n}), \quad n = 1, \dots, N\)$ A lightweight two-layer convolutional fusion head \(\mathcal{F}_{\text{fuse}}\) aggregates all \(2N\) candidates into the final frame, combining the structural fidelity of the orthogonal view at the poles with the equatorial sharpness of the primitive branch.
Loss & Training¶
The framework is supervised end-to-end with Census structural loss and spherical-weighted Charbonnier loss: $\(\mathcal{L}_{r} = \mathcal{L}_{\text{cen}} + \lambda \mathcal{L}_{\text{char}}\)$ The Census term \(\mathcal{L}_{\text{cen}}\) computes soft Hamming distance over \(7\times 7\) patches to preserve structural boundaries under illumination shifts. To prevent stretched polar pixels from overpowering the loss function at the expense of visually salient equatorial regions, the Charbonnier loss is weighted by latitude: $\(\mathcal{L}_{\text{char}} = \sum_{\mathbf{p}} \rho \Big( \boldsymbol{\omega}(\mathbf{p}) \odot (\hat{\mathbf{I}}_t(\mathbf{p}) - \mathbf{I}_t^{gt}(\mathbf{p})) \Big), \qquad \rho(x) = (x^2 + \epsilon^2)^\alpha\)$ where \(\boldsymbol{\omega}(\mathbf{p}) = \cos(\boldsymbol{\theta}_{\mathbf{p}})\) with latitude \(\boldsymbol{\theta}_{\mathbf{p}} = \pi (v_{\mathbf{p}} / H - 0.5)\). This weight smoothly decays from 1 at the equator to 0 at the poles, ensuring balanced gradient optimization across the entire sphere.
Training is performed on two NVIDIA H100 GPUs (80 GB) using AdamW with cosine learning rate decay from \(2 \times 10^{-4}\) to \(2 \times 10^{-5}\) over 300 epochs with batch size 16 (\(\lambda = 1, \alpha = 0.5, \epsilon = 10^{-3}\)). In addition to standard flipping and random erasing, random 3D rotations (roll, pitch, yaw) are applied to expose the network to diverse spherical wrap-around boundaries.
Key Experimental Results¶
Main Results¶
SVI360 was evaluated on synthetic benchmarks FlowScape (\(t=0.5\) middle-frame interpolation with ground-truth flow) and Flow360 (8x arbitrary-timestep interpolation), as well as real-world benchmarks ODV360 and 360VFI. Metrics include standard PSNR/SSIM, spherical-weighted WS-PSNR/WS-SSIM, and optical flow metrics EPE, SEPE, and AE.
| Dataset / Setting | Metric | SVI360 (Ours) | Second Best (AMT-G) | Gain |
|---|---|---|---|---|
| FlowScape (Middle frame) | PSNR (dB) | 38.68 | 38.05 | +0.63 dB |
| FlowScape (Middle frame) | WS-PSNR (dB) | 37.73 | 37.40 | +0.33 dB |
| FlowScape (Middle frame) | SSIM / WS-SSIM | 0.9754 / 0.9647 | 0.9704 / 0.9622 | +0.0050 / +0.0025 |
| FlowScape (Optical flow) | SEPE / AE | 15.22 / 19.74 | 25.91 / 20.90 | -10.69 / -1.16 |
| Flow360 (Arbitrary timestep) | PSNR / WS-PSNR (dB) | 34.02 / 33.30 | 33.77 / 33.22 | +0.25 / +0.08 dB |
| Flow360 (Arbitrary timestep) | SSIM / WS-SSIM | 0.9752 / 0.9615 | 0.9708 / 0.9594 | +0.0044 / +0.0021 |
| ODV360 (Real-world) | PSNR / WS-PSNR (dB) | 29.25 / 29.93 | 29.07 / 29.75 | +0.18 / +0.18 dB |
| ODV360 (Real-world) | SSIM / WS-SSIM | 0.9332 / 0.9105 | 0.9303 / 0.9067 | +0.0029 / +0.0038 |
On the 360VFI benchmark stratified across motion difficulty levels (Easy, Middle, Hard, Extreme), SVI360 consistently achieved the highest WS-PSNR and WS-SSIM:
| Difficulty (360VFI) | 360VFI [15] (WS-PSNR/SSIM) | AMT-G [13] (WS-PSNR/SSIM) | SVI360 (Ours) |
|---|---|---|---|
| Easy (flow < 2) | 33.95 / 0.9537 | 34.84 / 0.9726 | 34.81 / 0.9731 |
| Middle | 28.96 / 0.9060 | 29.74 / 0.9376 | 29.98 / 0.9409 |
| Hard | 27.81 / 0.8879 | 28.58 / 0.9209 | 28.95 / 0.9264 |
| Extreme (large motion) | 25.63 / 0.8517 | 25.55 / 0.8817 | 25.92 / 0.8880 |
Ablation Study¶
Ablations conducted on FlowScape substantiate the individual contributions of the loss terms and architectural components:
| Configuration | PSNR (dB) | WS-PSNR (dB) | SSIM | WS-SSIM | Description |
|---|---|---|---|---|---|
| Full model | 38.68 | 37.73 | 0.9754 | 0.9647 | Complete pipeline with all loss functions |
| w/o Spherical Charbonnier Loss | 33.99 | 33.85 | 0.9099 | 0.9471 | Drastic drop (-4.69 dB); polar errors dominate optimization |
| w/o Census Loss | 38.34 | 37.63 | 0.9734 | 0.9644 | Loss of structural coherence under lighting variations |
| w/o Spherical Multi-field Refinement | 38.08 | 37.29 | 0.9715 | 0.9627 | Notable drop (-0.60 dB); shows need for orthogonal candidates |
| w/o Spherical Refiner (SR) | 38.59 | 37.65 | 0.9748 | 0.9640 | Intermediate features lack cross-view geometric guidance |
Key Findings¶
- Latitude-weighted loss is fundamental: Omitting the spherical Charbonnier loss causes a catastrophic performance collapse of 4.69 dB in PSNR. Without cosine latitude weighting, gradient updates are dominated by extreme polar pixel stretching, compromising representation learning across the central visual field.
- Dual utility of orthogonal views: The orthogonal view provides complementary value at two distinct levels: cross-scale feature refinement via the SR module, and direct image candidate generation yielding a 0.60 dB gain in final multi-field fusion.
- Robustness under extreme displacement: While competing perspective algorithms degrade to 23โ24 dB under the Extreme motion split of 360VFI, SVI360 maintains 25.92 dB / 0.8880, proving its resilience against large motion fields.
Highlights & Insights¶
- Geometric unwarping via orthogonal rotation: By rotating equirectangular panoramas by 90ยฐ, polar singularities are mapped onto the distortion-free equator, resolving high-latitude correspondence ambiguities within standard 2D convolutional operators.
- End-to-end synthesis with cross-view equivariance: Extending beyond prior optical-flow-only methods (e.g., PriOr-Flow), SVI360 integrates cross-view equivariance into flow residual updates, intermediate latent feature propagation, and multi-candidate image blending.
- Broad paradigm applicability: The orthogonal dual-branch interaction strategy can be naturally extended to other spherical visual restoration and reconstruction tasks, such as 360ยฐ video super-resolution, omnidirectional depth estimation, and inpainting.
Limitations & Future Work¶
- Computational latency: SVI360 has 50.1M parameters and an inference latency of 365 ms/frame (~2.7 fps), which cannot yet meet the interactive real-time frame rates required by VR headsets (>120 fps). It is primarily targeted at offline video enhancement and cloud/edge pre-encoding workflows.
- Future optimization: Future efforts can explore weight-sharing between orthogonal and primitive branches, or employ knowledge distillation to compress dual-branch equivariant features into a compact single-stream network for real-time mobile deployment.
Related Work & Insights¶
- vs 360VFI [15]: 360VFI relies on learned deformable convolution offsets derived from distortion maps, which fail under fast, large motions; SVI360 establishes explicit orthogonal view projections and dual-cost correlation volumes, providing a mathematically grounded correspondence space.
- vs PriOr-Flow [14]: PriOr-Flow confines orthogonal view exploitation strictly to single-direction optical flow estimation; SVI360 systematically extends the orthogonal prior to flow-feature co-refinement and multi-candidate frame synthesis.
- vs AMT-G [13]: AMT is engineered for planar perspective videos and deteriorates under ERP polar distortion; SVI360 adapts its correlation volume concepts to spherical geometry through orthogonal cross-branch interactions and latitude-weighted losses.
Rating¶
- Novelty: โญโญโญโญ [Systematic formulation of orthogonal-view geometric equivariance for omnidirectional video interpolation]
- Experimental Thoroughness: โญโญโญโญโญ [Evaluated across four public benchmarks, multiple motion tiers, and rigorous ablations]
- Writing Quality: โญโญโญโญโญ [Well-structured narrative, disciplined mathematical formulation, and clear figures]
- Value: โญโญโญโญ [Sets a solid benchmark and architectural paradigm for immersive 360ยฐ video processing]