OmniLife360: A Benchmark for 3D Reconstruction from In-the-Wild 360° Captures¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/AutoLab-SAI-SJTU/Omni4DGS
Area: 3D Vision
Keywords: Panoramic 3D Reconstruction, 4D Gaussian Splatting, In-the-Wild Benchmark, Spherical Motion Decomposition, Dynamic-Static Decoupling
TL;DR¶
To tackle severe camera ego-motion and dynamic interference in long-horizon captures from consumer 360° cameras, this paper presents OmniLife360—a benchmark of 2,407 4K minute-level in-the-wild panoramic sequences—along with Omni4DGS, the first spherical 4D Gaussian Splatting framework that decomposes radial/tangential velocities and decouples dynamic foreground via mask supervision.
Background & Motivation¶
The rapid proliferation of consumer 360° cameras has transformed how people record and share real-world experiences. In daily lifelogging, travel documentation, and outdoor action sports, omnidirectional video captures continuous and globally coherent context within a single exposure, effectively eliminating viewpoint fragmentation and multi-camera stitching seams. However, transforming these unconstrained in-the-wild panoramic videos into persistent, navigable 3D digital experiences remains notoriously difficult. Everyday captures inherently entail aggressive camera rotations, unpredictable ego-motion, and cumulative drift across extended trajectories, compounded by the severe nonlinear distortions of Equirectangular Projection (ERP) in polar regions.
A critical gap exists between current panoramic reconstruction methods and the realistic deployment domain. Existing panoramic benchmarks remain severely constrained: scanned or synthetic datasets such as HM3D, Replica, and OmniBlender lack natural camera shake, authentic motion blur, and illumination shifts; meanwhile, real-world collections like OmniPhotos, 360Loc, or KITTI-360 either comprise brief clips of only a few seconds or rely on specialized vehicular sensor rigs that produce smooth, artificial trajectories. On the algorithmic side, scene-specific optimization methods (e.g., ODGS) and generalizable feed-forward architectures (e.g., PanSplat, OmniSplat) predominantly rely on static-scene assumptions. When confronted with moving pedestrians, vehicles, or the camera operator, these methods inevitably entangle dynamic components into static representations, triggering severe blur, ghosting, and floating artifacts.
Systematically addressing panoramic reconstruction under realistic capture conditions requires advancing both empirical benchmarks and dynamic representation models. Core idea: build OmniLife360, the first large-scale in-the-wild 4K panoramic benchmark integrating minute-level continuity, diverse real-world behaviors, and verified multi-view geometric consistency; and propose Omni4DGS, a spherical-aware 4D Gaussian Splatting framework that decouples radial and tangential motion with latitude-adaptive weighting and enforces mask-supervised dynamic-static separation.
Method¶
Overall Architecture¶
Omni4DGS aims to synthesize photo-realistic novel views across arbitrary viewpoints and timestamps from unconstrained panoramic video sequences. The input consists of ERP video frames captured by a single panoramic camera together with initial poses and sparse point clouds from COLMAP, and the output is a continuous dynamic omnidirectional radiance field. The pipeline augments each 3D Gaussian with periodic harmonic displacement and temporal opacity decay, parameterizes motion via spherical velocity decomposition with latitude-aware weighting, and leverages offline panoramic segmentation masks to decouple dynamic foreground Gaussians from the static background during rasterization.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Monocular 360° Video Frames + Poses"] --> B["Initialize Panoramic Gaussians<br/>COLMAP / OpenMVG Sparse Points"]
B --> C["Time-Varying Gaussian Parameterization<br/>Harmonic displacement + Opacity decay"]
C --> D["Spherical Velocity Decomposition & Latitude Weighting<br/>Radial/tangential split with adaptive blending"]
D --> E["Mask-Guided Dynamic-Static Decoupling<br/>OmniSAM masks + Feature rasterization"]
E --> F["Spherical Motion & Velocity Regularization<br/>Zero velocity in static regions + Ray smoothness"]
F --> G["Dynamic & Static Panoramic Rendering Output"]
Key Designs¶
1. Spherical Velocity Decomposition & Latitude-Aware Weighting: eliminating motion distortion under equirectangular projection In conventional Cartesian space, 4D Gaussian Splatting typically assigns a single Euclidean velocity vector \(\mathbf{v} \in \mathbb{R}^3\) to each Gaussian. In panoramic imaging based on unit-sphere projection, however, pixel positions correspond directly to directional rays, and spatial motion manifests as spherical angular displacement rather than linear Cartesian translation. Optimizing a single Euclidean velocity cannot disentangle motion along the viewing ray (radial) from lateral drift perpendicular to it (tangential), and equatorial versus polar regions experience drastically different distortion rates under ERP. Omni4DGS decomposes the panoramic velocity \(\mathbf{v}_{pano}\) into two learnable components—radial velocity \(\mathbf{v}_r\) and tangential velocity \(\mathbf{v}_t\)—combined via latitude-adaptive weights: $\(\mathbf{v}_{pano} = w_r \mathbf{v}_r + w_t \mathbf{v}_t\)$ where \(\hat{\mu}_y\) denotes the normalized vertical coordinate of the Gaussian center, and the latitude angle \(\theta\) satisfies \(\cos\theta = \hat{\mu}_y\). The weights are defined as: $\(w_t = \left(\sqrt{1 - \hat{\mu}_y^2}\right)^\gamma, \quad w_r = 1 - w_t\)$ The exponent \(\gamma\) shapes the nonlinear weighting curve such that tangential motion smoothly dominates in equatorial regions and decays rapidly near the poles into radial-dominant behavior. This formulation respects spherical geometry and prevents numerical instability caused by severe polar stretching.
2. Mask-Guided Dynamic-Static Decoupling: resolving dynamic-static ambiguity under photometric supervision Relying solely on RGB reconstruction loss causes self-supervised optimization to misinterpret moving objects as floating semi-transparent Gaussian artifacts across multiple viewpoints. Omni4DGS defines a staticness coefficient \(\rho = \frac{\beta}{l}\) (the ratio of lifespan \(\beta\) to oscillation period \(l\)), identifying Gaussians with \(\rho < \theta_{dyn}\) as dynamic Gaussians \(\mathcal{G}_d\). To establish explicit geometric boundaries, OmniSAM—a foundation segmentation model specialized for panoramic imagery—extracts moving foreground objects (such as pedestrians and vehicles) as an offline 2D motion mask prior \(M_d\). The ODGS rasterizer is extended to support general feature map rendering: $\(F = \sum_{i=1}^N f_i \alpha_i \prod_{j=1}^{i-1}(1 - \alpha_j)\)$ Setting feature \(f_i\) to Gaussian instantaneous opacity produces the rendered dynamic opacity map \(\hat{O}_d\), supervised via an \(L_1\) objective \(\mathcal{L}_o = \|\hat{O}_d - M_d\|_1\). Meanwhile, the rendered dynamic color \(\hat{C}_d\) is supervised against masked ground-truth pixels \(M_d \odot C_{gt}\) using a joint photometric loss: $\(\mathcal{L}_d = (1 - \lambda_d)\|\hat{C}_d - M_d \odot C_{gt}\|_1 + \lambda_d (1 - \text{SSIM}(\hat{C}_d, M_d \odot C_{gt}))\)$ Combining both terms yields the dynamic objective \(\mathcal{L}_{dynamic} = \mathcal{L}_d + \mathcal{L}_o\), strictly constraining dynamic Gaussians to actual moving regions and preventing them from contaminating static background reconstruction.
3. Spherical Motion & Velocity Regularization: suppressing angular perturbation amplification in polar and static areas Because equirectangular projection is highly sensitive to directional perturbations, minor angular displacements in 3D space can trigger multi-pixel jumps across the ERP image, causing background Gaussians to oscillate unnaturally during optimization. Omni4DGS introduces dual regularizers to stabilize the scene. First, feature rasterization aggregates instantaneous velocities into an average 2D velocity map \(\hat{V}\), penalizing nonzero velocities across static background regions (indicated by \(1 - M_d\)): $\(\mathcal{L}_v = \|(1 - M_d) \odot \hat{V}\|_1\)$ Second, for static Gaussians (\(\rho \ge \theta_{dyn}\)), the normalized ray direction of the center coordinate at time \(t\) is computed as \(d(t) = \frac{\boldsymbol{\mu}(t)}{\|\boldsymbol{\mu}(t)\|_2}\), and adjacent time steps are constrained via a direction-smoothness regularizer: $\(\mathcal{L}_m = \|d(t + \Delta t) - d(t)\|_2^2\)$ This constraint firmly stabilizes static Gaussians along their viewing rays, ensuring that buildings, terrain, and distant backgrounds remain rigid, substantially eliminating background blur and temporal flickering.
Loss & Training¶
The framework is optimized end-to-end using a compound objective combining static scene rendering, dynamic object supervision, and motion regularizations: $\(\mathcal{L} = \lambda_{rgb}\mathcal{L}_{rgb} + \lambda_{dyn}\mathcal{L}_{dynamic} + \lambda_v\mathcal{L}_v + \lambda_m\mathcal{L}_m\)$ where \(\mathcal{L}_{rgb} = (1 - \lambda_{rgb})\|\hat{C} - C_{gt}\|_1 + \lambda_{rgb}(1 - \text{SSIM}(\hat{C}, C_{gt}))\) ensures overall visual fidelity. Training runs for 30,000 iterations using the Adam optimizer on an NVIDIA RTX 4090 GPU, with alternating updates between spatial and temporal parameters ensuring stable convergence in highly dynamic scenarios.
Key Experimental Results¶
Main Results¶
Evaluations are conducted on 111 representative scenes selected from OmniLife360, comparing against the scene-specific optimization baseline ODGS and four leading feed-forward architectures (Splatter-360, PanSplat, PanoSplatt3R, OmniSplat). Standard metrics include PSNR, SSIM, and LPIPS.
| Evaluation Subset | Method | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|---|---|---|---|
| Reconstruction Fidelity (Input View) | ODGS [16] | 23.727 | 0.792 | 0.385 |
| Reconstruction Fidelity (Input View) | Ours (Omni4DGS) | 29.655 | 0.896 | 0.146 |
| Novel View Synthesis (Novel View) | ODGS [16] | 20.145 | 0.712 | 0.433 |
| Novel View Synthesis (Novel View) | Ours (Omni4DGS) | 20.640 | 0.624 | 0.335 |
Under the feed-forward protocol (uniform groups of 10 frames, predicting intermediate 8 frames from the first and last frames), PanoSplatt3R achieves the highest overall scores (PSNR 17.646 dB, SSIM 0.521, LPIPS 0.450), followed by OmniSplat (PSNR 17.270 dB, SSIM 0.543, LPIPS 0.472), whereas PanSplat and Splatter-360 suffer more noticeable performance degradation under wide unposed baselines.
Quantitative evaluation across four challenging dynamic scene types (Market, Lake, Mountain, Urban Street) demonstrates consistent advantages across all environments:
| Method | Metric | Avg. | Market | Lake | Mountain | Urban Street |
|---|---|---|---|---|---|---|
| PanSplat | PSNR ↑ / LPIPS ↓ | 15.988 / 0.469 | 13.183 / 0.567 | 17.136 / 0.412 | 17.923 / 0.413 | 15.711 / 0.485 |
| Splatter-360 | PSNR ↑ / LPIPS ↓ | 15.017 / 0.517 | 12.660 / 0.595 | 15.347 / 0.477 | 17.439 / 0.460 | 14.621 / 0.536 |
| OmniSplat | PSNR ↑ / LPIPS ↓ | 17.356 / 0.454 | 14.354 / 0.568 | 19.000 / 0.379 | 19.391 / 0.395 | 16.677 / 0.474 |
| PanoSplatt3R | PSNR ↑ / LPIPS ↓ | 17.797 / 0.437 | 13.507 / 0.573 | 21.443 / 0.318 | 20.816 / 0.361 | 15.421 / 0.496 |
| ODGS | PSNR ↑ / LPIPS ↓ | 20.492 / 0.416 | 16.789 / 0.505 | 22.596 / 0.333 | 22.048 / 0.404 | 20.535 / 0.422 |
| Ours | PSNR ↑ / LPIPS ↓ | 21.519 / 0.302 | 17.849 / 0.381 | 23.727 / 0.248 | 22.776 / 0.264 | 21.722 / 0.313 |
Ablation Study¶
Ablations on four representative dynamic scenes isolate the impact of individual architectural components. In addition, adapting pinhole 4DGS baselines (Deformable-GS and PVG) to equirectangular projection demonstrates the necessity of spherical-aware motion formulation.
| Config | PSNR ↑ | SSIM ↑ | LPIPS ↓ | Note |
|---|---|---|---|---|
| w/o temporal opacity | 19.865 | 0.631 | 0.324 | Cannot handle objects entering/exiting view; PSNR drops by 1.608 dB |
| w/o temporal motion | 20.999 | 0.666 | 0.299 | Lacks periodic position compensation, causing dynamic blurring |
| w/o velocity decom. | 21.191 | 0.664 | 0.300 | Cartesian velocity degrades near poles, lowering perceptual quality |
| w/o dynamic mask | 20.753 | 0.651 | 0.302 | Sole RGB self-supervision leads to severe static-dynamic entanglement |
| w/o motion reg. | 21.023 | 0.659 | 0.302 | Static background Gaussians exhibit random directional drift |
| Ours (full model) | 21.473 | 0.679 | 0.285 | All components combined achieve the best overall fidelity |
Comparison with adapted 4DGS methods on dynamic sequences (reported in PSNR ↑): - Bike scene: ODGS 10.137, Deformable-GS 11.183, PVG 12.379, Ours 12.930 - Community scene: ODGS 15.154, Deformable-GS 16.641, PVG 17.168, Ours 18.063 - Road scene: ODGS 15.064, Deformable-GS 15.618, PVG 17.484, Ours 17.533 - Square scene: ODGS 17.063, Deformable-GS 19.510, PVG 21.020, Ours 21.941
Key Findings¶
- Temporal opacity decay is critical for long sequences: Removing time-varying opacity triggers the sharpest performance drop (>1.6 dB PSNR). Across extended user paths, dynamic objects continuously enter and leave camera visibility; without opacity decay, static Gaussians are forced to permanently memorize transient objects.
- Dramatic perceptual gains (LPIPS): Omni4DGS slashes LPIPS from 0.385 to 0.146 on input views and from 0.433 to 0.335 on novel views. Explicit static-dynamic decoupling eliminates high-frequency floating floaters and boundary tearing.
- Feed-forward models struggle under unconstrained dynamics: In outdoor environments with vigorous ego-motion (e.g., Market and Road), feed-forward baselines drop below 16 dB PSNR, highlighting the need for benchmark-scale fine-tuning on diverse in-the-wild datasets.
Highlights & Insights¶
- Spherical motion parameterization tailored for ERP: Splitting velocity into radial and tangential components and modulating them via latitude cosine weights addresses the mathematical friction between Cartesian velocity vectors and spherical projections.
- Collaborative human-AI benchmark construction: Sourcing through LLM-expanded behavioral taxonomies combined with ShotTransNetV2 cut-detection and manual auditing ensures both wild diversity and strict multi-view geometric integrity.
- Versatile feature rasterization operator: Generalizing the ODGS rasterizer to render multi-channel feature buffers (opacity, velocity, depth) enables seamless backpropagation from 2D segmentation priors into 3D Gaussian attributes.
Limitations & Future Work¶
- Reliance on precomputed poses and offline segmentation: The pipeline depends on accurate COLMAP sparse reconstructions and OmniSAM masks. Extreme camera shake or heavy occlusions can still cause SfM track failure, and mask edge jitter propagates into Gaussian geometry.
- Optimization efficiency constraints: Per-scene optimization over 30K steps at 1K-2K resolution requires substantial training time, precluding real-time interactive reconstruction on edge or mobile hardware.
- Future directions: Integrating camera pose estimation directly into a 4D panoramic dynamic SLAM system and leveraging lightweight feed-forward models for rapid Gaussian field initialization.
Related Work & Insights¶
- vs ODGS: ODGS introduces tangent-plane rasterization for static panoramic scenes but fails under moving pedestrians and rapid camera translation. Omni4DGS extends the rasterization interface with temporal modeling and dynamic decoupling, dramatically boosting dynamic visual fidelity.
- vs PVG: While PVG introduces periodic vibration Gaussians for pinhole camera urban driving, applying Cartesian velocities directly to 360° projections induces severe polar distortion. Omni4DGS resolves this mismatch via radial-tangential velocity decomposition and latitude adaptation.
- vs PanoSplatt3R / OmniSplat: Feed-forward panoramic splatting offers rapid inference but exhibits high generalization error under complex long-horizon user-generated captures. OmniLife360 provides an indispensable training and benchmarking bed for enhancing the robustness of future feed-forward models.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Pioneers spherical motion decomposition and dynamic-static decoupling tailored for 360° 4DGS alongside an in-the-wild benchmark.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Evaluates 2,407 sequences, 111 detailed test scenes, across per-scene, feed-forward, and 4DGS variants with exhaustive ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Well-structured narrative with tight coupling between geometric intuition and mathematical formulation.
- Value: ⭐⭐⭐⭐⭐ Fills a long-standing void in long-horizon in-the-wild panoramic 3D reconstruction benchmarks; open-sourced dataset and codebase offer strong community utility.