Skip to content

StratoSplat: Taming Layered Regularities for Sparse Aerial 3D Gaussian Splatting

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/keloee/StratoSplat
Area: 3D Vision
Keywords: Sparse-view 3D Gaussian Splatting, Aerial Photogrammetry, Layered Depth Images, Neural Multi-Plane Gaussians, Test-Time Adaptation

TL;DR

StratoSplat exploits the inherent gravity-aligned stratification of aerial scenes, coupling layered depth image-guided test-time diffusion adaptation for dense point initialization with a hybrid neural multi-plane Gaussian representation to suppress view-aligned degeneracies under sparse drone captures.

Background & Motivation

Unmanned aerial vehicles (UAVs) play a pivotal role in urban planning, large-scale 3D mapping, environmental monitoring, and autonomous navigation. 3D Gaussian Splatting (3DGS) has rapidly emerged as a prominent foundation for novel view synthesis owing to its real-time rendering speeds and photorealistic fidelity. Nonetheless, real-world drone missions often face operational constraints such as limited battery endurance, flight path regulations, volatile weather conditions, and sparse viewing geometry, preventing the acquisition of dense multi-view imagery. When trained on sparse, anisotropic aerial captures dominated by oblique perspectives, standard radiance fields and 3DGS suffer catastrophic structural collapse.

Two critical failure modes plague sparse aerial 3DGS. First, initialization sparsity in textureless regions: aerial imagery frequently contains expansive, homogeneous surfaces such as building rooftops, asphalt roads, and uniform vegetation. In these regions, classical Structure-from-Motion (SfM, e.g., COLMAP) feature matching and triangulation fail completely, leaving vast voids and holes in the sparse point clouds that destabilize Gaussian initialization. Although drone GNSS/IMU systems provide reliable camera poses, the lack of geometric anchor points cannot be recovered by standard pipelines. Second, view-aligned degeneracies: sparse, anisotropic observations lack sufficient multi-view parallax constraints. Consequently, explicit 3D Gaussians overfit the 2D appearance of the few training views by elongating and aligning along the training ray directions. Instead of recovering true 3D surface geometry, they form needle-like and sheet-like floating artifacts that produce severe distortion under novel viewpoints. Generic monocular depth priors (e.g., DNGaussian) and generative video diffusion models (e.g., Guidedvd-3DGS) fail due to the acute domain gap between ground-level training data and top-down aerial perspectives. Furthermore, feed-forward Gaussian predictors exhibit severe view-aligned overfitting—for instance, TranSplat reaches nearly 50 dB on training views but collapses to 17 dB on novel views.

These failure modes share a fundamental root cause: conventional 3DGS completely ignores the strong layered structural regularities intrinsic to aerial environments. Unlike generic indoor or object-centric setups, aerial scenes are shaped by gravity: terrain, roads, building facades, and rooftops naturally stratify into distinct parallel horizontal layers perpendicular to the ground plane. Core idea: exploit the inherent gravity-aligned stratification of aerial scenes as a strong geometric regularizer by utilizing a Layered Depth Image (LDI) memory to guide test-time depth diffusion adaptation for multi-view consistent dense point initialization, and introducing neural multi-plane virtual Gaussians anchored to learned parallel planes to rigidly penalize view-aligned collapse during joint optimization.

Method

Overall Architecture

StratoSplat operates through a two-stage synergistic pipeline designed for sparse aerial imagery with known onboard GNSS/IMU poses: "Layered Memory Guided Initialization" followed by "Neural Multi-Plane Hybrid Optimization." In the initialization phase, starting from incomplete triangulated SfM points, the framework sequentially processes views ordered by coverage density. It adapts a pretrained depth diffusion model at test time, accumulating multi-view depth estimates into a globally coherent Layered Depth Image (LDI) anchored to a reference camera. LDI's explicit depth-ordered merging prevents uncontrolled point explosion while providing occlusion-aware visibility guidance back into the diffusion denoising loop, ultimately unprojecting a dense, hole-free colored point cloud \(P_{\text{dense}}\). In the optimization phase, singular value decomposition (SVD) on \(P_{\text{dense}}\) extracts the dominant stratification normal aligned with gravity, defining a stack of globally shared parallel planes spanning the scene depth. Virtual Gaussians generated at camera ray-plane intersections query a continuous coordinate neural field for low-frequency attributes, rasterizing jointly with explicit Gaussians initialized from \(P_{\text{dense}}\) via \(\alpha\)-compositing. This hybrid planar anchoring eliminates the degrees of freedom required for Gaussians to collapse into view-aligned minima.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Sparse Aerial Images + GNSS/IMU Poses"] --> B["Layered Memory Guided Initialization<br/>LDI Ordered Storage + Sparse/Visibility Guided Adaptation"]
    B --> C["Dense Colored Point Cloud P_dense"]
    C --> D["Neural Multi-Plane Gaussian Representation<br/>SVD Stratification Normal + Continuous Neural Field"]
    D --> E["Hybrid Joint Optimization & Geometric Regularization<br/>Explicit Detail Capture + Planar Anchored Regularization"]
    E --> F["High-Fidelity Aerial Novel View Synthesis"]

Key Designs

1. Layered Memory Guided Initialization: Eliminating Textureless Voids and Multi-View Inconsistencies

Standard SfM triangulation leaves wide holes in textureless rooftops and roads, whereas naive monocular depth predictors suffer from severe ground-to-aerial domain gaps and multi-view contradictions. StratoSplat introduces sequential test-time depth adaptation conditioned on a compact Layered Depth Image (LDI) memory. Anchored to reference camera \(C_{\text{ref}}\), each pixel \((u, v)\) stores an ordered sequence of depth layers \(\{\ell_k\}_{k=1}^{K_{u,v}}\) containing depth \(d_k\), world coordinates \(\mathbf{p}_k\), and color \(\mathbf{c}_k\) sorted by strictly increasing depth (\(d_1 < d_2 < \dots < d_{K_{u,v}}\)). At diffusion denoising timestep \(t\), the clean latent is estimated via the diffusion posterior, decoded into intermediate depth \(\hat{D}_i\), and scale-shifted to metric scale \(\hat{D}_i^{(m)} = \hat{D}_i \cdot \hat{a} + \hat{b}\). Multi-view geometric consistency is enforced via a dual-objective loss: $$ \mathcal{L} = \sum_{(u,v) \in \mathcal{P}i^{\text{SfM}}} |\hat{D}_i^{(m)}(u,v) - d|_1 $$ where the sparse loss pins depth predictions to known SfM triangulated points, while the visibility loss penalizes discrepancies against the projected guidance depth }^{\text{SfM}}| + \lambda_{\text{vis}} |(\hat{D}_i^{(m)} - D_i^{\text{guide}}) \odot M_i^{\text{vis}\(D_i^{\text{guide}}\) rendered from the current LDI state with visibility mask \(M_i^{\text{vis}}\). When updating memory, new depth samples within threshold \(\epsilon_d\) of an existing layer are fused via weighted averaging, whereas distinct depths instantiate new layers. This depth-ordered merging intrinsically suppresses point cloud explosion by \(3.6\times\) compared to unstructured accumulation, generating a dense, multi-view coherent point cloud \(P_{\text{dense}}\).

2. Neural Multi-Plane Gaussian Representation: Planar Anchoring Against View-Aligned Degeneracies

Even when initialized from dense point clouds, unconstrained explicit 3D Gaussians tend to overfit sparse oblique views by aligning their primitives along camera viewing rays. Recognizing that aerial structures primarily organize into horizontal ground and vertical elevations, StratoSplat performs SVD on centered points \(P_{\text{dense}} - \bar{\mathbf{p}}\) and selects the singular vector corresponding to the minimal singular value as the stratification plane normal \(\mathbf{n} = V[:, \arg\min_j \Sigma_{jj}]\), which physically aligns with the gravity axis. A stack of \(K\) globally shared parallel planes \(\pi_k: \mathbf{n}^\top \mathbf{p} + d_k = 0\) is uniformly distributed over the scene depth range \([d_{\min}, d_{\max}]\). During rendering, camera ray \(\mathbf{r}(t) = \mathbf{o} + t \mathbf{d}_p\) intersects plane \(\pi_k\) at: $$ \mathbf{p}_k = \mathbf{o} + t_k \mathbf{d}_p, \quad t_k = -\frac{\mathbf{n}^\top \mathbf{o} + d_k}{\mathbf{n}^\top \mathbf{d}_p} $$ Rather than optimizing unconstrained discrete primitive parameters, virtual Gaussian attributes \((\mathbf{c}_k, \tilde{\alpha}_k, \mathbf{s}_k, \mathbf{q}_k)\) are queried from a continuous coordinate neural field \(\Phi(\mathbf{p}_k)\). The neural parameterization provides smooth spatial inductive bias, while primitive centers remain strictly anchored to planar intersections. To prevent front layers from saturating early and occluding background strata during \(\alpha\)-compositing, an opacity distance bias modulates the activation: \(\alpha_k = \sigma(\tilde{\alpha}_k + b(t_k))\).

3. Hybrid Joint Optimization & Geometric Regularization: Complementary Frequency Decomposition

The final scene representation integrates explicit Gaussians \(\mathcal{G}_{\text{exp}}\) and neural multi-plane virtual Gaussians \(\mathcal{G}_{\text{vir}}\), rasterized jointly via unified front-to-back depth sorting. Crucially, the two representations follow an asymmetric optimization division. Explicit Gaussians \(\mathcal{G}_{\text{exp}}\) (initialized from \(P_{\text{dense}}\)) employ spherical harmonics (SH) and standard adaptive clone/split/prune densification to reconstruct high-frequency local textures, sharp facade boundaries, and specular variations. In contrast, virtual Gaussians \(\mathcal{G}_{\text{vir}}\) remain locked to planar coordinates and output view-independent Lambertian colors, acting as a low-frequency geometric scaffolding that absorbs structural ambiguity. Because virtual primitives cannot translate in 3D space to overfit oblique rays, they regularize the optimization space and prevent the explicit Gaussians from collapsing into view-aligned floaters.

Loss & Training

The pipeline commences with offline test-time depth adaptation taking approximately 12 seconds per view on a single NVIDIA RTX 4090 GPU to produce \(P_{\text{dense}}\). Subsequent hybrid Gaussian training converges in approximately 9 minutes: - Photometric Objective: Jointly rendered images \(\hat{I}\) are optimized against ground-truth training views \(I\) using a combined \(\mathcal{L}_1\) and D-SSIM loss: \(\mathcal{L}_{\text{rgb}} = (1 - \lambda_{\text{ssim}}) \|\hat{I} - I\|_1 + \lambda_{\text{ssim}} (1 - \text{SSIM}(\hat{I}, I))\) with \(\lambda_{\text{ssim}} = 0.2\). - Opacity Bias: The distance-dependent opacity bias function \(b(t_k)\) maps into logit range \([-4.6, 1.4]\) to balance transmittance across layers. - Hyperparameters: Consistency weight \(\lambda_{\text{vis}} = 0.1\); layer merging threshold \(\epsilon_d = 0.02\) for 3DAS and \(\epsilon_d = 0.55\) for LEVIR-NVS.

Key Experimental Results

Main Results

StratoSplat is evaluated on the standard aerial benchmark LEVIR-NVS (3 training views across 16 diverse scenes comprising urban buildings, campuses, stadiums, towns, mountains, and parks) and the large-scale 3DAS benchmark (7 training views across 9 complex scenes spanning city, country, and port environments).

Table 1: Quantitative comparison on LEVIR-NVS with 3 training views (16-scene average)

Category Method PSNR ↑ (dB) SSIM ↑ LPIPS ↓ Note
NeRF-based NeRF [34] 15.13 0.20 0.58 Dense baseline, catastrophic failure under 3 views
NeRF-based DietNeRF [23] 15.71 0.24 0.57 High-level semantic CLIP consistency
NeRF-based RegNeRF [36] 13.40 0.18 0.63 Unobserved view regularization fails on aerial data
NeRF-based FreeNeRF [52] 15.54 0.27 0.60 Frequency regularization baseline
NeRF-based MPNeRF [18] 21.72 0.80 0.19 Aerial multiplane NeRF; slow volume rendering
3DGS-based Vanilla 3DGS [25] 20.20 0.72 0.23 Severe view-aligned degeneracies and floaters
3DGS-based FSGS [64] 20.93 0.71 0.27 Soft Pearson correlation depth supervision
3DGS-based DNGaussian [29] 17.48 0.49 0.44 Hard monocular depth prior suffers from domain gap
3DGS-based DropoutGS [51] 19.31 0.56 0.43 Generic primitive dropout lacks aerial structure
3DGS-based Guidedvd-3DGS [62] 17.05 0.47 0.49 Video diffusion + DUSt3R aerial domain shift
Feed-forward NoPoSplat ft [53] 19.75 0.57 0.24 Jointly fine-tuned pose-free predictor
Feed-forward MVSplat ft [7] 17.30 0.39 0.43 Pairwise cost volume predictor
Feed-forward DepthSplat ft [50] 15.75 0.30 0.53 Generalizable depth-gaussian model
Feed-forward TranSplat ft [58] 16.79 0.29 0.56 Overfits training views (49.76 dB) vs novel views
Ours StratoSplat (Ours) 24.48 0.83 0.15 Surpasses strongest 3DGS baseline by +3.55 dB

Ablation Study

The ablation experiments isolate the impact of initialization schemes, structural representations, and regularization designs on LEVIR-NVS under the 3-view setting.

Table 2: Ablation of initialization and representation components on LEVIR-NVS (16-scene mean)

Ablation Category Configuration Variant PSNR ↑ (dB) SSIM ↑ LPIPS ↓ Core Finding & Mechanism
Baseline Vanilla 3DGS (SfM, Explicit) 20.20 0.72 0.23 Sparse SfM seeds leave large holes in geometry
Initialization + DUSt3R init 18.65 0.56 0.34 Severe aerial domain shift generates distorted points
Initialization + VGGT init 16.43 0.30 0.46 Multi-view visual geometry transformer breaks down
Initialization + Depth Anything V3 (DA3) init 20.49 0.71 0.23 Monocular scale inconsistency across views
Initialization + COLMAP MVS init 20.45 0.75 0.21 Stereo matching fails in textureless rooftops
Initialization + Marigold diffusion init 20.31 0.73 0.23 Independent single-view diffusion lacks 3D consistency
Initialization + Layered Memory guided init (LDI) 21.00 0.77 0.18 Robust dense initialization, +0.80 dB over SfM
Representation Explicit 3DGS (on LDI points) 21.00 0.77 0.18 Without planar regularizer, view-aligned floaters linger
Representation Fixed multi-plane GS 21.81 0.79 0.16 Discrete primitives lack continuous spatial coherence
Representation Neural multi-plane (w/o opacity bias) 23.64 0.79 0.18 Continuous neural field enforces smoothness
Full Model StratoSplat (Full Model) 24.48 0.83 0.15 Synergistic combination achieves +4.28 dB over baseline

Table 3: Ablation on geometric memory structure design (3DAS City Scene 0, 7 views)

Memory Structure Variant PSNR ↑ (dB) SSIM ↑ Points Count ↓ Processing Time (s) ↓ Design Advantage
Unstructured Point Cloud Accumulation 22.99 0.76 778,646 98 Uncontrolled accumulation and redundant conflicts
Layered Memory (LDI, Ours) 23.21 0.79 215,607 84 3.6× fewer points (-72.3%), 14% faster, +0.22 dB PSNR

Key Findings

  • Layered Memory Prevents Redundancy and Enhances Geometry: Directly stacking unprojected points across views causes uncontrolled memory growth (778k points). In contrast, LDI-guided depth merging exploits natural depth stratification to resolve multi-view contradictions, using only 215k points (\(3.6\times\) reduction) while boosting PSNR by 0.22 dB.
  • Continuous Neural Field vs. Discrete Primitives: Placing discrete Gaussians on fixed planes (Fixed multi-plane GS) yields 21.81 dB. Moving to a continuous coordinate neural field \(\Phi\) elevates PSNR to 23.64 dB, demonstrating that shared network weights provide essential spatial smoothness and prevent discrete primitive overfitting.
  • Crucial Role of Depth-Dependent Opacity Bias: Introducing the distance-aware opacity bias \(b(t_k)\) yields an additional +0.84 dB improvement (23.64 dB \(\to\) 24.48 dB) by preventing frontmost planes from prematurely extinguishing viewing rays during volume accumulation.
  • Robustness to UAV Pose Noise: When controlled SE(3) perturbations are added during depth adaptation (rotation \(1.0^\circ\), translation \(0.002\) extent), PSNR drops modestly from 24.48 dB to 23.21 dB. The system retains solid reconstruction fidelity, matching the precision standards of consumer-grade UAV GNSS/IMU hardware.

Highlights & Insights

  • Scene-Specific Structural Regularities over External Black-Box Foundation Models: Rather than fine-tuning massive generic video or depth foundation models that suffer severe aerial domain gaps, StratoSplat anchors its design in the simple physical reality that man-made and natural aerial environments naturally stratify along gravity.
  • Frequency-Decomposed Asymmetric Hybrid Optimization: Explicit Gaussians capture sharp high-frequency details, while planar-anchored neural Gaussians enforce rigid low-frequency geometry. This division allows 3DGS to achieve NeRF-like geometric stability without relinquishing real-time rasterization efficiency.
  • Broad Transferability: The LDI-guided test-time adaptation paradigm is representation-agnostic; it can be readily plugged into sparse neural implicit surfaces (SDF), Instant-NGP, or automated mesh completion pipelines for low-altitude UAV mapping.

Limitations & Future Work

  • Extreme Non-Planar Topography: The global stratification assumption relies on dominant horizontal planes aligned with gravity. In scenes with extreme vertical relief such as steep gorges, overhangs, or sheer cliffs, a single set of parallel planes may struggle to capture slanted surfaces.
  • Lack of Joint Pose Optimization: The framework assumes fixed, accurate GNSS/IMU poses. In GNSS-denied urban canyons where flight telemetry drifts, pose errors degrade LDI projection consistency. Integrating joint bundle adjustment into the diffusion loop remains a promising direction.
  • Dynamic Transient Objects: Moving vehicles and pedestrians can introduce conflicting depth layers across views, necessitating semantic transient filtering during LDI memory fusion.
  • vs. MPNeRF [18]: While MPNeRF pioneered multiplane priors in aerial NeRF, its volumetric representation requires hours of optimization and suffers from slow rendering. StratoSplat bridges multiplane regularities with explicit/neural Gaussian splatting, executing training in 9 minutes, delivering real-time rendering, and beating MPNeRF by +2.76 dB.
  • vs. FSGS [64] & DNGaussian [29]: Prior sparse 3DGS methods impose monocular depth maps as continuous loss constraints during optimization, propagating estimation noise and scale drift directly into gradients. StratoSplat restricts depth diffusion strictly to initial point cloud generation, decoupling subsequent 3DGS training from biased 2D depth supervision.
  • vs. Guidedvd-3DGS [62] & DUSt3R [45]: External generative priors trained on ground-level imagery suffer acute domain mismatch under aerial perspectives. StratoSplat demonstrates that exploiting intrinsic scene stratification provides far more reliable geometric constraints.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Elegant formalization of gravity-aligned stratification combining dynamic LDI memory with neural multi-plane primitives.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarks across LEVIR-NVS and 3DAS datasets with extensive ablations uncovering feed-forward overfitting traps.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical formulation and insightful diagnosis of view-aligned degeneracies and initialization voids.
  • Value: ⭐⭐⭐⭐⭐ Highly practical framework for real-world drone inspection, urban digital twins, and low-cost aerial photogrammetry.