Skip to content

Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Project Page: https://focusgs.github.io/
Area: Autonomous Driving
Keywords: 3D scene reconstruction, Gaussian Splatting, geometric ambiguity, sparse-view reconstruction, targeted structure completion

TL;DR

Addressing the substantial computational redundancy of uniform volumetric densification in sparse surround-view driving scenarios (<15% overlap), FocusGS decouples deterministic surfaces from ambiguous boundaries via a 3D geometric ambiguity manifold, instantiating and optimizing queries exclusively within this sparse topological subspace to cut Gaussian count by ~74% and rendering time by ~34% while establishing state-of-the-art visual quality.

Background & Motivation

In ego-centric autonomous driving scenarios, reconstructing high-fidelity 3D scene representations from sparse multi-camera inputs is fundamental for simulation, downstream perception, and generative world modeling. Feed-forward single-pass frameworks built on 3D Gaussian Splatting (3DGS) have demonstrated exceptional inference efficiency and generalizability. However, existing pixel-aligned feed-forward models heavily rely on substantial cross-view visual overlap to enforce multi-view geometric consistency. In production autonomous driving rigs, surround-view camera suites are designed for maximal spatial coverage with minimal overlap (often less than 15%). Under such limited overlap and severe frustum truncation, standard pixel-based unprojection breaks down, producing pronounced geometric collapse and empty holes near occlusion boundaries.

To alleviate the reliance on cross-view overlap, recent dual-branch methods such as Omni-Scene incorporate voxel-based Gaussian branches that lift 2D image features into 3D voxel space to infer occluded structures. Nevertheless, this uniform volumetric processing strategy indiscriminately instantiates and optimizes millions of Gaussians across the entire 3D space, introducing massive memory and rendering overhead. Quantitative empirical analysis indicates that over 82% of typical driving scenes comprise continuous deterministic surfaces (such as flat asphalt roads and open skies) where simple pixel-aligned Gaussians already achieve high visual fidelity (26.50 dB PSNR); severe artifacts and geometric collapse are confined to a mere 17.93% of the scene characterized by visibility transitions and occlusion boundaries (where baseline PSNR drops to 21.90 dB). Processing the whole volume with dense voxel grids essentially squanders enormous computation to fix localized, sparse defects.

This stark efficiency-completeness tension suggests that structural completion should not be uniformly scattered across space, but decoupled from deterministic regions. The core idea of this paper is to explicitly derive a "3D geometric ambiguity manifold" from depth gradients and visibility transitions, selectively instantiating and optimizing a fixed budget of Gaussian queries strictly within this sparse topological subspace for targeted structure completion (FocusGS).

Method

Overall Architecture

FocusGS comprises two tightly coordinated components: a pixel-aligned base representation branch and a Targeted Structure Completion (TSC) branch. Given sparse multi-view surround RGB images, the model first extracts downsampled image features enhanced with Plรผcker ray embeddings and depth priors to predict pixel-aligned base 3D Gaussians that efficiently cover deterministic visible surfaces. Concurrently, a geometric ambiguity localization module detects depth boundaries and expands them into a 2D ambiguity mask, which is lifted along camera rays into a sparse 3D uncertainty subspace. Finally, the targeted structure completion module samples a fixed budget of continuous queries exclusively within this subspace, enriches them via 3D sparse convolution and multi-view deformable cross-attention, and decodes compensatory Gaussian attributes through an MLP to compose the full scene with the base Gaussians.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Sparse Multi-View Input<br/>6-view surround RGB images"] --> B["Multi-View Feature Extraction<br/>ResNet-50 + Plรผcker ray priors"]
    B --> C["Pixel-Aligned Base Gaussian Predictor<br/>Efficiently covers 82% deterministic surfaces"]
    B --> D["2D Geometric Ambiguity Manifold Extraction<br/>Central-difference depth gradients + dilation"]
    D --> E["3D Uncertainty Subspace Lifting<br/>Depth-adaptive longitudinal ray segments"]
    E --> F["Targeted Structure Completion Module<br/>Fixed budget queries + sparse conv + deformable attention"]
    C --> G["Gaussian Aggregation & Differentiable Rendering<br/>G_final = G_base โˆช G_comp"]
    F --> G
    G --> H["Novel-View Synthesized Images & Depth"]

Key Designs

1. Geometric ambiguity localization and 3D subspace lifting: isolating occlusion zones from depth discontinuities To bypass indiscriminate volumetric densification, the framework must reliably isolate regions suffering from occlusion and view transitions. In surround-view images, depth varies smoothly across continuous surfaces but jumps abruptly at occlusion boundaries where camera sightlines shift from foreground obstacles to distant backgrounds. The method computes horizontal and vertical gradients on predicted depth maps using central-difference convolution kernels: $\(D_x = Z_i \times K_x, \quad D_y = Z_i \times K_y, \quad E_i = \sqrt{D_x^2 + D_y^2}\)$ A threshold \(\tau_g\) is applied to isolate sharp foreground-background silhouettes, followed by morphological dilation using a \(k \times k\) square structuring element to encompass continuous uncertainty bands, yielding the 2D geometric ambiguity manifold \(\mathcal{O}_d\). Next, ambiguous pixels are projected along camera rays into world coordinates, using the predicted depth \(d_p\) as the midpoint of a 3D line segment \([p_0, p_1]\) with thickness \(\Delta_p = \kappa_{\text{rel}} d_p + \kappa_{\text{abs}}\). This formulation dynamically scales longitudinal uncertainty with distance, allocating wider search margins for distant objects. The union of these lifted segments across all surround views forms the compact, irregular 3D uncertainty subspace \(\mathcal{O}_{3d}\).

2. Uncertainty-aware query sampling: guaranteeing constant memory footprint via a fixed token budget To handle the irregular and sparse geometry of the derived manifold efficiently without dense grid interpolation, FocusGS samples a fixed global token budget of \(N_q\) (default 60k) discrete 3D coordinates uniformly from the 3D uncertainty subspace \(\mathcal{O}_{3d}\): $\(\mathcal{P} = \left\{\mathbf{p}_m \in \mathcal{O}_{3d} \;\middle|\; m = 1, \dots, N_q\right\}, \quad \mathbf{p}_m \sim \mathcal{U}(\mathcal{O}_{3d})\)$ Each sampled coordinate serves as a localized spatial anchor coupled with a learnable high-dimensional feature embedding. By decoupling query generation from image resolution and scene volume dimensions, this mechanism provides a strict guarantee of constant, lightweight GPU memory consumption regardless of scene size.

3. Dual context aggregation and structural compensation: sparse geometry reasoning and cross-view feature injection The sampled compensatory queries are iteratively updated through four cascaded completion blocks. First, for intra-query spatial reasoning, 3D sparse convolution (spconv) is executed exclusively on occupied voxels, completely avoiding the cubic computational complexity \(O(X \times Y \times Z)\) of dense 3D grids. Second, to resolve depth ambiguity from single-view occlusion, deformable cross-attention projects the 3D queries onto 2D multi-view image feature maps, harvesting unoccluded contextual visual cues from neighboring cameras. Finally, an MLP decodes residual spatial offsets relative to initial positions along with opacity, scale, rotation quaternion, and RGB color parameters, producing the targeted compensatory Gaussian set \(\mathcal{G}_{\text{comp}}\).

Loss & Training

The scene representation combines base and compensatory Gaussians as \(\mathcal{G}_{\text{final}} = \mathcal{G}_{\text{base}} \cup \mathcal{G}_{\text{comp}}\) for global differentiable rendering and optimization with photometric \(L_1\) and LPIPS losses. To ensure that compensatory Gaussians remain dedicated to resolving local ambiguities rather than duplicating road surfaces, the 2D manifold \(\mathcal{O}_d\) acts as a spatial mask supervising independent renders of \(\mathcal{G}_{\text{comp}}\) with masked \(L_1\), LPIPS, and depth losses: $\(\mathcal{L} = \mathcal{L}_{\text{full}}^{l_1} + \lambda_1 \mathcal{L}_{\text{full}}^{\text{lpips}} + \lambda_2 \mathcal{L}_{\text{comp}}\)$ where \(\mathcal{L}_{\text{comp}} = \mathcal{L}_{\text{comp}}^{l_1} + \lambda_{c1} \mathcal{L}_{\text{comp}}^{\text{lpips}} + \lambda_{c2} \mathcal{L}_{\text{comp}}^{\text{dpt}}\). The model is optimized using AdamW with cosine learning rate decay on two 80GB GPUs over 100k iterations on nuScenes.

Key Experimental Results

Main Results

Evaluated on the nuScenes ego-centric benchmark (6 surround cameras, ~15% overlap) and the RealEstate10K scene-centric benchmark.

Dataset Method Latency (s) โ†“ PSNR โ†‘ SSIM โ†‘ LPIPS โ†“ PCC โ†‘
nuScenes AttnRend 9.980 20.96 0.533 0.467 N/A
nuScenes MuRF 0.672 20.34 0.504 0.433 -0.332
nuScenes pixelSplat 0.508 21.51 0.616 0.372 0.001
nuScenes MVSplat 0.174 21.61 0.658 0.295 0.181
nuScenes SCube - 23.85 0.721 0.258 0.651
nuScenes DrivingForward - 24.32 0.732 0.229 0.766
nuScenes STORM - 24.56 0.752 0.217 0.788
nuScenes Omni-Scene (Prev. SOTA) 0.088 24.27 0.736 0.237 0.800
nuScenes FocusGS (Ours) 0.058 24.65 0.754 0.220 0.837
RealEstate10K Omni-Scene - 26.19 0.865 0.131 0.368
RealEstate10K FocusGS - 26.32 0.872 0.123 0.365
RealEstate10K MVSplat + FocusGS - 27.02 0.892 0.118 0.374

Ablation Study

Table 1 examines the progressive addition of framework components; Table 2 validates internal aggregation layers and ambiguity mask strategies.

Config / Component PSNR โ†‘ SSIM โ†‘ LPIPS โ†“ PCC โ†‘ Note
(a) Pixel-aligned Baseline 22.89 0.698 0.290 0.780 Base pixel-aligned Gaussians only
(b) (a) + Depth Init 23.14 0.703 0.320 0.802 Marginal gain from prior depth unprojection
(c) (b) + Random Structure Completion 23.40 0.708 0.306 0.810 Uniform random queries without spatial focus
(d) Full Model (with G.A. Localization) 24.65 0.754 0.220 0.837 Targeted to manifold, PSNR jumps +1.25 dB
TSC w/o Self-context Aggregation (w/o S.A.) 24.04 0.738 0.234 0.827 Missing 3D sparse conv neighborhood smoothing
TSC w/o Cross-view Aggregation (w/o C.A.) 23.30 0.723 0.265 0.802 Missing complementary views drops PSNR by 1.35 dB
Ambiguity Mask: Learned Mask 24.03 0.732 0.261 - High precision (0.86) but low recall (0.55)
Ambiguity Mask: Depth Gradient (Ours) 24.65 0.754 0.220 - High recall (0.82) boosts PSNR by +0.62 dB

Key Findings

  • Targeted completion vastly outperforms uniform random completion: Randomly scattering supplementary queries throughout the entire volume yielded an insignificant gain of only 0.26 dB PSNR. In contrast, anchoring completion to the geometric ambiguity manifold surged performance by +1.25 dB, proving that spatial targeting is essential.
  • Cross-view attention provides the primary disambiguation signal: Ablating cross-view deformable attention caused a major drop of 1.35 dB PSNR, significantly exceeding the 0.61 dB decrease from removing 3D sparse convolutions. This highlights that recovering occluded structures relies fundamentally on ingesting visual evidence from complementary camera angles.
  • High recall is more critical than high precision for structural repair: While a learned neural mask achieved higher precision (0.86 vs 0.46) and IoU (0.50 vs 0.42), its lower recall (0.55 vs 0.82) missed critical depth boundaries, trailing the gradient-based manifold by 0.62 dB PSNR. In geometric completion, false negatives are far more damaging than false positives.
  • Query budget and block depth saturation: Performance scales steadily from 10k to 60k queries (23.79 to 24.65 dB PSNR) before saturating at 120k (24.67 dB). Similarly, 4 completion blocks strike an optimal balance, whereas 6 blocks slightly degrade quality (24.56 dB) due to optimization hurdles in sparse feature spaces.

Highlights & Insights

  • Spatial decoupling over brute-force volumetric densification: Recognizes that over 80% of autonomous driving scenes are simple planar surfaces, decoupling structural completion from deterministic regions to prune ~74% of redundant Gaussians and achieve a 34% speedup.
  • Physics-grounded gradient heuristics beating black-box learned masks: Leveraging classical central-difference depth gradients and depth-adaptive ray expansion achieves an 82% recall of structural errors without requiring neural mask supervision, demonstrating the elegance of geometric domain priors.
  • Constant token budget for predictable deployment: Formulating structural completion as a fixed-size query sampling process within a sparse 3D manifold delivers strictly bounded memory and predictable execution latency.

Limitations & Future Work

  • Vulnerability to weather-induced optical artifacts: Under adverse weather like heavy rain, water droplets on camera lenses produce severe optical blur and distortions. The network cannot easily separate 2D lens-level noise from genuine 3D geometry, mistakenly generating blurred artifacts in reconstruction.
  • Absence of temporal dynamics in the spatial manifold: The geometric ambiguity manifold is currently derived per frame independently; rapidly moving dynamic vehicles and pedestrians can exhibit temporal flickering due to the lack of motion compensation.
  • Future Directions: Future research could incorporate multi-frame spatio-temporal attention and explicit optical artifact modeling to extend FocusGS into dynamic 4D world models.
  • vs Omni-Scene: Omni-Scene allocates millions of voxel Gaussians uniformly throughout the entire 3D space to fill blind spots; FocusGS restricts completion to the geometric ambiguity manifold, surpassing Omni-Scene's visual quality with 74% fewer Gaussians and 34% lower latency.
  • vs MVSplat / pixelSplat: Feed-forward pixel-unprojection methods degrade significantly in autonomous driving settings due to near-zero view overlap (<15%); FocusGS preserves their lightweight base path while eliminating boundary hole artifacts through targeted compensatory queries.
  • vs SCube / VoxSplat: SCube relies on dense voxel grids that scale cubically in memory; FocusGS relies on continuous queries sampled in sparse manifolds and processed via sparse convolutions, maintaining a constant memory budget.

Rating

  • Novelty: โญโญโญโญ [Elegant spatial decoupling of structural completion from deterministic surfaces via geometric manifolds]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive comparisons on nuScenes and RealEstate10K across visual, geometric, and latency metrics with rigorous ablations]
  • Writing Quality: โญโญโญโญโญ [Clear motivation supported by empirical region analysis, coherent formulations, and intuitive structure]
  • Value: โญโญโญโญโญ [Provides a practical, highly efficient representation paradigm for autonomous driving 3D reconstruction and simulation]