Skip to content

Geometry-Propagated Gaussian Splatting for Aerial Sparse Novel View Synthesis

Conference: ECCV 2026
Paper: ECCV 2026 Official
Code: https://github.com/GeoProp-GS/GeoProp-GS
Area: 3D Vision
Keywords: 3D Gaussian Splatting, Sparse Novel View Synthesis, Aerial Reconstruction, Geometric Initialization, Anchor-constrained Optimization

TL;DR

Addressing the dual bottlenecks of SfM feature matching collapse and extensive unobserved areas in aerial sparse-view settings, GeoProp-GS proposes Depth-guided Geometric Initialization (DGI) and Anchor-constrained Gaussian Optimization (AGO), which densify and complete geometry before optimization and propagate reliable geometric priors into under-supervised regions via an anchor-residual parameterization to achieve state-of-the-art sparse-view rendering.

Background & Motivation

3D Gaussian Splatting (3DGS) has rapidly become the dominant paradigm for novel view synthesis (NVS), enabling photorealistic, real-time rendering via explicit, differentiable Gaussian primitives. Its explicit geometry and efficient inference make it an ideal candidate for aerial mapping, infrastructure inspection, and UAV autonomous navigation. However, aerial deployment inherently faces strict platform limitations—limited drone battery life, satellite revisit orbits, weather windows, and airspace regulations—which inevitably force viewpoint sampling to be extremely sparse. Standard 3DGS relies heavily on dense photometric supervision across overlapping views; under sparse aerial observations, it rapidly suffers from severe floater artifacts, geometric holes, and catastrophic rendering degradation.

Existing sparse-view NVS methods broadly follow two paradigms: optimization-based frameworks that regularize scene geometry using sparse point clouds generated by Structure-from-Motion (SfM) such as COLMAP, and feed-forward architectures that directly regress Gaussian primitives or cost volumes from image features. However, aerial imagery presents unique challenges: repetitive roof patterns, uniform highways, and homogeneous building facades offer few discriminative keypoints, leading to severe correspondence matching collapse. Furthermore, wide-baseline flight paths severely limit viewpoint overlap. Consequently, COLMAP yields severely incomplete or empty point clouds, stripping optimization methods of their geometric scaffolds. Feed-forward methods similarly fail to jointly estimate camera poses and scene geometry on low-texture aerial surfaces. Experiments verify that the core bottleneck of sparse-view aerial 3DGS lies in the collapse of geometric initialization rather than downstream optimization strategies.

To overcome this fundamental limitation, the authors rethink the dependency on feature matching and propose to derive complete, metric-consistent geometric scaffolds directly from monocular depth priors, while explicitly populating unobserved blind zones prior to optimization. The core idea is to introduce a unified framework combining Depth-guided Geometric Initialization (DGI) and Anchor-constrained Gaussian Optimization (AGO): DGI synthesizes dense, metrically calibrated point clouds and inpainting-driven pseudo-view proposals, while AGO parameterizes proposal Gaussians into reliable geometric anchors and learnable residuals to propagate stable geometric knowledge from observed into under-supervised regions.

Method

Overall Architecture

GeoProp-GS is built upon two core principles: establishing complete geometric initialization across the scene and propagating geometric knowledge from observed to unobserved areas. The pipeline first uses DGI to construct an initial point set \(C_{\text{init}}\) with full scene coverage: for observed viewpoints, Hierarchical Depth-guided Densification (HDD) fuses monocular depth predictions with sparse SfM points to generate dense, metric-consistent points \(C_{\text{dense}}\); for unobserved regions, Geometry-aware Viewpoint Augmentation (GVA) projects current geometry into pseudo-viewpoints, inpaints missing textures with a diffusion model, and back-projects pseudo-depth to generate proposal points \(C_{\text{proposal}}\). During optimization, points are partitioned into reliable Gaussians \(G_{\text{reliable}}\) and proposal Gaussians \(G_{\text{prop}}\). AGO decomposes proposal Gaussians into nearest reliable anchors plus learnable residuals, ensuring robust convergence in under-constrained regions without cumbersome regularization terms.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Sparse Aerial Images Input"] --> B["Stage 1: Hierarchical Depth-guided Densification (HDD)<br/>Monocular Depth + Hierarchical Metric Calibration"]
    B --> C["Stage 2: Geometry-aware Viewpoint Augmentation (GVA)<br/>Pseudo-view Projection + Diffusion Inpainting Proposal"]
    C --> D["Gaussian Attribute Decomposition<br/>Reliable Gaussians G_reliable & Proposal Gaussians G_prop"]
    D --> E["Stage 3: Anchor-constrained Gaussian Optimization (AGO)<br/>Anchor Base + Learnable Residual Adaptation"]
    E --> F["Photorealistic 3D Scene Representation & Novel View Rendering"]

Key Designs

1. Hierarchical Depth-guided Densification (HDD): Bypassing feature matching collapse with metric depth priors

In aerial environments with repetitive textures, COLMAP produces sparse and incomplete point clouds. HDD aims to generate dense, metrically accurate point clouds for observed views without dense keypoint correspondence. Because monocular depth estimators (e.g., Depth Anything V2) output scale-ambiguous relative depth \(\hat{D}_i\) subject to local distortions, HDD introduces a Hierarchical Correction Strategy (HCS). At the global level, a scale factor \(\alpha\) and shift \(\beta\) are fitted against sparse SfM depth observations \(D_{\text{sparse}}\) with a valid mask \(M_{\text{sparse}}\) via repeated least-squares:

\[\arg\min_{\alpha, \beta} \left\| M_{\text{sparse}} \odot (\alpha \hat{D}_i + \beta - D_{\text{sparse}}) \right\|_2^2\]

producing a globally aligned depth \(D_{\text{global}} = \alpha \hat{D}_i + \beta\). To correct perspective-induced local geometric warps, a distance-weighted nearest-neighbor interpolation computes a residual correction field that decays exponentially with Euclidean distance from the nearest sparse SfM point. Pixel depths near sparse references closely adhere to SfM accuracy, while distant regions smoothly transition to the global solution. Median filtering suppresses boundary extrapolation artifacts. Back-projecting the corrected depth maps via camera intrinsics and extrinsics produces a dense point cloud \(C_{\text{dense}}\) that replaces flawed SfM points.

2. Geometry-aware Viewpoint Augmentation (GVA): Inpainting-driven geometry completion for unobserved regions

Sparse aerial viewpoints leave extensive regions unobserved by any training camera. GVA extends the geometric scaffold into these blind spots prior to training. By interpolating camera trajectories, pseudo-view poses \(P_h\) outside the training frustums are generated. Existing dense geometry \(C_{\text{dense}}\) is projected onto \(P_h\) to render a partial RGB image \(I_{\text{partial}}\) alongside an occlusion mask \(M\). A latent diffusion inpainting model (FLUX.1 Kontext) fills in missing regions, generating a complete pseudo-image \(I_{\text{complete}}\).

While diffusion inpainting may hallucinate local textures, state-of-the-art monocular depth models infer global 3D geometric layout from high-level visual perspective and spatial context rather than fine texture fidelity. Consequently, applying Depth Anything V2 and the HCS pipeline to \(I_{\text{complete}}\) yields consistent pseudo-depth maps. Back-projecting these depths generates geometric proposal points \(C_{\text{proposal}}\). Filtering out spatial overlaps with \(C_{\text{dense}}\) yields a complete geometric initialization \(C_{\text{init}} = C_{\text{dense}} \cup C_{\text{proposal}}\).

3. Anchor-constrained Gaussian Optimization (AGO): Architectural anchor-residual decomposition preventing blind-zone drift

Initialization yields two distinct groups: reliable Gaussians \(G_{\text{reliable}}\) (from \(C_{\text{dense}}\)) supported by real training views, and proposal Gaussians \(G_{\text{prop}}\) (from \(C_{\text{proposal}}\)) located in unobserved regions lacking ground-truth photometric supervision. Standard gradient optimization on \(G_{\text{prop}}\) inevitably causes severe geometric drift or degenerate shapes. Rather than tuning complex regularizer losses, AGO embeds geometric priors directly into the architectural parameterization. Each proposal Gaussian \(g_k \in G_{\text{prop}}\) identifies its nearest reliable anchor \(g_a \in G_{\text{reliable}}\) based on initial center coordinates:

\[a = \arg\min_{j \in G_{\text{reliable}}} \|\mathbf{x}_k - \mathbf{x}_j\|_2\]

All Gaussian attributes—position \(\mathbf{x}\), spherical harmonics coefficients \(\mathbf{c}\), scale \(\mathbf{s}\), rotation \(\mathbf{q}\), and opacity \(\mathbf{o}\)—are decomposed into anchor values and learnable residuals:

\[\mathbf{x}_k = \mathbf{x}_a + \Delta\mathbf{x}_k, \quad \mathbf{c}_k = \mathbf{c}_a + \Delta\mathbf{c}_k, \quad \mathbf{s}_k = \mathbf{s}_a + \Delta\mathbf{s}_k, \quad \mathbf{q}_k = \mathbf{q}_a + \Delta\mathbf{q}_k, \quad \mathbf{o}_k = \mathbf{o}_a + \Delta\mathbf{o}_k\]

Applied in pre-activation parameter space, proposal attributes are reconstructed on the fly during forward rendering. Under sparse or nonexistent photometric gradients, proposal residuals undergo minimal updates, ensuring proposals inherit the well-optimized geometric orientation and appearance of reliable anchors. When occasional multi-view consistency signals are received, the residuals fine-tune local details. This structural design reliably propagates geometric knowledge across boundaries into unobserved zones.

Loss & Training

The optimization backbone builds on FSGS with pruning and densification disabled, relying entirely on the high-quality initialized points. Supervision uses standard photometric \(L_1\) loss and D-SSIM loss computed solely on real training views. To prevent color drift in weakly supervised proposal Gaussians, a proposal color statistical regularizer aligns the mean and variance of proposal DC color coefficients with the global Gaussian distribution. Training converges in 10,000 iterations within 6.1 minutes on an RTX 3090, using only 2.73 GB of GPU memory.

Key Experimental Results

Main Results

Experiments were conducted on LEVIR-NVS (16 nadir vertical aerial scenes, evaluated under extreme 3-view training settings) and 3D-AS (9 oblique orbital aerial scenes spanning City, Country, and Port environments, evaluated under 7-view training settings).

Dataset / Scene Category Method PSNR (dB) ↑ SSIM ↑ LPIPS ↓
LEVIR-NVS (3 views, Mean) NeRF Baseline FreeNeRF 15.54 0.27 0.60
LEVIR-NVS (3 views, Mean) Aerial NeRF MPNeRF 21.72 0.80 0.19
LEVIR-NVS (3 views, Mean) Vanilla 3DGS 3DGS 20.20 0.72 0.23
LEVIR-NVS (3 views, Mean) Sparse 3DGS FSGS 20.93 0.71 0.27
LEVIR-NVS (3 views, Mean) Sparse 3DGS DNGaussian 17.48 0.49 0.44
LEVIR-NVS (3 views, Mean) Sparse 3DGS DropoutGS 19.31 0.56 0.43
LEVIR-NVS (3 views, Mean) Diffusion 3DGS Guidedvd-3DGS 17.05 0.47 0.49
LEVIR-NVS (3 views, Mean) Feed-Forward 3DGS NoPoSplat (ft) 19.75 0.57 0.24
LEVIR-NVS (3 views, Mean) Feed-Forward 3DGS MVSplat (ft) 17.30 0.39 0.43
LEVIR-NVS (3 views, Mean) Ours GeoProp-GS 24.35 0.83 0.15
3D-AS (7 views, City Scene 0) Sparse 3DGS FSGS 19.67 0.64 0.26
3D-AS (7 views, City Scene 0) Ours GeoProp-GS 22.20 0.75 0.18
3D-AS (7 views, Country Scene 2) Vanilla 3DGS 3DGS 19.43 0.42 0.45
3D-AS (7 views, Country Scene 2) Ours GeoProp-GS 26.16 0.74 0.25
3D-AS (7 views, Port Scene 2) Sparse 3DGS FSGS 24.07 0.86 0.14
3D-AS (7 views, Port Scene 2) Ours GeoProp-GS 26.49 0.89 0.10

On LEVIR-NVS under 3 input views, GeoProp-GS surpasses the previous best aerial-specific method MPNeRF by 2.63 dB PSNR (from 21.72 to 24.35 dB) and reduces LPIPS from 0.19 to 0.15. Compared with baseline FSGS, PSNR improves by 3.42 dB, effectively eliminating floating artifacts and blurred geometries.

Ablation Study

The ablation investigates progressive component integration on LEVIR-NVS, alongside plug-and-play evaluations when integrating HDD into baseline 3DGS methods.

Config Observed Region Strategy Unobserved Region Strategy PSNR (dB) ↑ SSIM ↑ LPIPS ↓ Note
A1 (Baseline) COLMAP - 21.10 0.75 0.23 FSGS default with incomplete SfM points
A2 HDD (Ours) - 21.57 0.82 0.15 Dense depth calibration improves observed fidelity
A3 (Full) HDD (Ours) GVA + AGO (Ours) 24.35 0.83 0.15 Proposal extension + anchor regularization achieves SOTA
B1 (Plug-and-Play) Vanilla 3DGS COLMAP 20.20 0.72 0.23 Standard 3DGS baseline
B2 (Plug-and-Play) 3DGS + HDD Replaced with HDD 21.80 (+1.60) 0.83 (+0.11) 0.14 (-0.09) Consistent gain by upgrading initialization alone
B3 (Plug-and-Play) FSGS COLMAP 20.93 0.71 0.27 Standard FSGS baseline
B4 (Plug-and-Play) FSGS + HDD Replaced with HDD 21.75 (+0.82) 0.80 (+0.09) 0.19 (-0.08) Significant artifact reduction across all metrics

Testing across different depth backbones (MiDas, Marigold, Depth Anything V2) demonstrates robustness, achieving 23.79 dB, 23.93 dB, and 24.35 dB PSNR respectively, with all choices substantially outperforming prior arts.

Key Findings

  • Crucial role of unobserved region modeling: Enhancing observed regions via HDD alone (A2) yields a modest +0.47 dB PSNR gain over baseline A1, although perceptual LPIPS improves noticeably. Incorporating GVA and AGO to complete unobserved regions (A3) triggers a dramatic +2.78 dB PSNR surge, confirming that handling blind frustums is the decisive factor in sparse aerial NVS.
  • Superiority over feed-forward point predictors: Modern feed-forward architectures like VGGT (16.43 dB) and DUSt3R (18.65 dB) collapse on low-texture aerial surfaces, generating oversized, distorted Gaussians. In contrast, HDD produces uniformly distributed, geometrically metric points, reaching 21.57 dB out of the box.
  • High training and resource efficiency: By encoding geometric constraints into the parameterization rather than computing per-iteration online diffusion sampling or complex depth distillation losses, GeoProp-GS trains in 6.1 minutes using 2.73 GB VRAM, operating nearly 25x faster than Guidedvd-3DGS (2.5 hours, 21.3 GB VRAM).

Highlights & Insights

  • Architectural embedding of geometric constraints: Instead of fragile hyperparameter balancing across multiple regularization losses, AGO decomposes proposal parameters into anchors and residuals, providing natural structural rigidity while preserving fine-tuning adaptability.
  • Decoupled utilization of diffusion inpainting: The authors recognize that while diffusion models hallucinate color textures, monocular depth estimators accurately extract perspective geometry from inpainted layouts. Using inpainting strictly for geometric scaffold expansion avoids injecting hallucinated artifacts into the radiance field.
  • Plug-and-play practicality: HDD operates as a standalone dense point generator that consistently boosts existing 3DGS frameworks (e.g., +1.60 dB on vanilla 3DGS) without modifying their downstream rendering pipelines.

Limitations & Future Work

  • Offline preprocessing overhead: Although GPU training takes only 6.1 minutes, offline pseudo-view synthesis and inpainting in GVA require 8 to 15 minutes per scene on an RTX 4090, leaving room for acceleration in emergency mapping workflows.
  • Dependency on minimal initial poses: The hierarchical calibration strategy (HCS) relies on valid sparse COLMAP points to compute initial scale and shift. In featureless regions like open water bodies or sand dunes, supplementary GPS/IMU navigation priors remain necessary.
  • Extreme extrapolation viewpoints: When novel viewpoints deviate drastically from the flight path, monocular depth models may face degraded perspective accuracy, suggesting future integration with satellite Digital Elevation Models (DEM).
  • vs FSGS / DNGaussian: While FSGS uses unobserved ray regularization and DNGaussian applies depth normalization, both rely on COLMAP's flawed initial point clouds in sparse aerial settings. GeoProp-GS tackles the root cause at initialization and stabilizes learning via anchor propagation.
  • vs Guidedvd-3DGS / RI3D: These methods run expensive diffusion models during iterative optimization, leading to heavy computational overhead. GeoProp-GS decouples generative inpainting into an offline initialization step, maintaining light 2.73 GB VRAM training.
  • vs DUSt3R / VGGT: Feed-forward transformers struggle with repetitive aerial textures, resulting in severe registration failures. GeoProp-GS adopts a hybrid strategy, leveraging metric monocular calibration to construct robust scene scaffolds.

Rating

  • Novelty: ⭐⭐⭐⭐ [Innovative combination of hierarchical monocular depth alignment, generative view extension, and anchor-residual decomposition specifically addressing aerial NVS challenges]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluated extensively on LEVIR-NVS (3 views) and 3D-AS (7 views) with comprehensive ablations, visualizations, and plug-and-play verifications]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear structural organization, thorough bottleneck analysis, and coherent technical descriptions]
  • Value: ⭐⭐⭐⭐⭐ [Solves a practical vulnerability of 3DGS in sparse aerial remote sensing, combining state-of-the-art fidelity with high computational efficiency]