Skip to content

ReInGS: Re-Initializing 3D Gaussians against Sparsity Discrepancy in Few-Shot Novel View Synthesis

Conference: ECCV 2026
Paper: ECCV Official
Full Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-5651.txt
Area: 3D Vision
Keywords: 3D Gaussian Splatting, Few-Shot Novel View Synthesis, Sparsity Discrepancy, Gaussian Re-Initialization, Hybrid Point Sampling

TL;DR

ReInGS tackles the overlooked sparsity discrepancy and data leakage problem in few-shot 3DGS caused by dense SfM points from extended views, proposing a two-stage coarse-to-fine Gaussian re-initialization framework combining expansion-based hybrid point sampling and view-dominant details-aware sampling to achieve state-of-the-art novel view synthesis purely from sparse training views.

Background & Motivation

Novel view synthesis (NVS) powered by 3D Gaussian Splatting (3DGS) has achieved remarkable progress by replacing dense neural coordinate networks with explicit ellipsoidal primitives and tile-based rasterization, delivering photorealistic real-time rendering. However, when deployed in the challenging few-shot regime with only 2 to 4 training views, standard Structure-from-Motion (SfM) pipelines inevitably produce degraded, unstable, or near-empty point clouds due to the lack of sufficient multi-view parallax. To counteract this instability, recent representative few-shot 3DGS methods, such as FSGS and CoR-GS, initialize 3D Gaussians using dense fused stereo points generated from an extended set of views across the entire dataset—in fact incorporating all available viewpoints beyond the training set.

This convention introduces a critical flaw termed sparsity discrepancy: while evaluated as few-shot methods, these systems secretly leak dense spatial geometry and surface distributions from test views during initialization. Once stripped of extended-view points and initialized purely from sparse random points or actual few-shot views, these prior methods suffer catastrophic performance drops (e.g., plunging by 2 to 4 dB in PSNR on LLFF under 3-view settings). A pilot diagnostic shows that the benefit of extended SfM points stems primarily from their high spatial density rather than strict ground truth positioning; under few-shot conditions, sparse initialization forces Adaptive Density Control (ADC) to transport Gaussians across long, erratic gradient trajectories, rapidly getting trapped in local optima and producing oversized floating artifacts or severe cracks.

Because true few-shot reconstruction strictly prohibits test-view leakage, the key challenge is how to synthesize high-density, geometrically faithful initial Gaussian distributions relying exclusively on the given sparse training views. Core idea: establish a coarse-to-fine 3D Gaussian re-initialization framework that first deploys expansion-based hybrid point sampling to capture the global scene geometry and pre-optimize macro topology, followed by view-dominant details-aware sampling leveraging depth residuals and cross-view transmittance coverage to reconstruct high-fidelity micro details without any extended-view leakage.

Method

Overall Architecture

The ReInGS pipeline is structured into three progressive phases: coarse-level initialization, fine-level re-initialization, and final joint optimization. During coarse initialization, the bounded scene volume estimated from sparse views is scaled outward, and a hybrid combination of uniform grid and random sampling is deployed to generate a space-filling point distribution. These points instantiate the initial Gaussian set \(\Theta_1\), which undergoes coarse pre-optimization using photometric loss and hard depth regularization to anchor macro geometric structure. Subsequently, recognizing that \(\Theta_1\) captures global layout but misses high-frequency local textures, the fine-level re-initialization stage is triggered. It computes pixel-wise importance scores via rendering depth residuals and transmittance, back-projects high-residual pixels into 3D candidate Gaussians \(\Theta_2\), and refines them through a cross-view visibility selection to form the final re-initialized Gaussian set \(\Theta\).

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Sparse Training Views & Camera Poses"] --> B["Expansion-based Hybrid Point Sampling<br/>Volume scaling + uniform grid & random point sampling"]
    B --> C["Coarse-level Pre-optimization<br/>Photometric loss + hard depth regularization"]
    C --> D["Intra-view Point Sampling<br/>Depth residual & transmittance importance back-projection"]
    D --> E["Fine-level Pre-optimization<br/>Soft depth regularization to suppress floaters"]
    E --> F["Cross-view Point Sampling<br/>Multi-view ray intersection coverage probabilistic re-filtering"]
    F --> G["Final Global Optimization & Novel View Rendering"]

Key Designs

1. Expansion-based Hybrid Point Sampling: Alleviating Long-Distance Transport via Space Expansion

Under few-shot constraints, the scene extent \(V\) reconstructed by SfM is typically truncated and undersampled, failing to cover background regions and peripheral geometry. From the perspective of Gaussian Mixture Model (GMM) maximum likelihood estimation, each Gaussian centroid \(\mu_i\) is heavily influenced by neighbors within its three-standard-deviation sphere (\(3\sigma_i\)). Sparse initialization leaves Gaussians isolated; moving them to correct surfaces requires long-range transportation paths where directional gradient accumulation in ADC frequently breaks monotonic loss descent, causing Gaussians to blow up into diffuse floaters.

To shorten optimization paths and eliminate boundary undersampling, this design first applies a spatial scaling factor \(\gamma = 1.4\) along each coordinate axis to broaden the bounding cuboid into an expanded space \(V'\). Within \(V'\), an expansion-based hybrid point sampling (EHPS) strategy combines uniform regular grid sampling and uniform random sampling. Uniform sampling guarantees structured, non-redundant spatial coverage across depth layers, effectively disambiguating overlapping multi-view ray intersections; random sampling breaks grid-induced aliasing artifacts and injects spatial diversity. Controlled by sampling parameters \(\beta_1 = 200, \beta_2 = 50\) over the expanded dimensions \([\Delta'_x, \Delta'_y, \Delta'_z]\), EHPS provides dense, well-dispersed initial coordinates for the coarse Gaussian set \(\Theta_1\).

2. Intra-View Point Sampling: Local Detail Completion via Depth & Transmittance Residuals

While coarse pre-optimized Gaussians \(\Theta_1\) establish the global skeleton, their view-irrelevant initialization leaves them dominated by large-scale primitives, yielding smoothed edges and missing micro-structures. Inspired by image-order rendering, intra-view point sampling (IVPS) identifies 2D pixels on rendered training views that correspond to under-reconstructed geometric details and lifts them back into 3D space. For each pixel \(p\) in a training view, an importance score is defined:

\[\Omega_P(p) = \delta \cdot \Omega_D(p) + (1-\delta) \cdot \Omega_T(p)\]

where \(\delta = 0.03\) weights the depth and transmittance terms. The depth score \(\Omega_D(p)\) quantifies the normalized discrepancy between rendered depth and pseudo-ground-truth monocular depth \(\hat{D}_g(p)\) produced by Depth Anything. The accumulated transmittance score \(\Omega_T(p) = \sum_{i \in \mathcal{N}} \alpha_i T_i\) highlights pixels where rays penetrate through incomplete surfaces. Rendered depth \(D(p)\) is computed by enforcing a high opacity threshold \(\alpha'_i = 0.99\) to avoid floater bias:

\[D(p) = \sum_{i \in \mathcal{N}} \alpha'_i \prod_{j=1}^{i-1}(1-\alpha'_j) (W \|\mu_i - o\|_2 + t)\]

Using \(\Omega_P\) as a sampling distribution, approximately 10% of pixels per view are selected, back-projected into 3D coordinates using known camera intrinsics/extrinsics, and initialized with pixel colors to form the detail-aware Gaussian set \(\Theta_2\).

3. Cross-View Point Sampling: Anti-Overfitting Multi-View Ray Intersection Filtering

Gaussians spawned directly from single-view back-projection are vulnerable to single-view overfitting, which manifests as needle-like distortions or hollow cracks in unseen angles. Grounded in object-order rendering principles, cross-view point sampling (CVPS) aggregates the projection footprint of each Gaussian \(\theta_i \in \Theta_2\) across the entire collection of rendered training viewpoints \(\mathcal{I}\):

\[\Omega_G(\theta_i) = \sum_{I \in \mathcal{I}} \sum_{p \in I} \Pi(\theta_i, p) \alpha_i T_i\]

where \(\Pi(\theta_i, p) = 1\) if the projection of \(\theta_i\) covers pixel \(p\), and 0 otherwise. Rather than applying a hard Top-K selection—which would systematically favor oversized Gaussians and prune compact primitives carrying high-frequency boundary details—CVPS normalizes \(\Omega_G\) into a global categorical probability distribution:

\[\mathcal{P}(\theta_i) = \frac{\Omega_G(\theta_i)}{\sum_{j \in \mathcal{N}} \Omega_G(\theta_j)}\]

By stochastically retaining roughly 40% of the candidate primitives according to \(\mathcal{P}(\theta_i)\), CVPS filters out view-inconsistent single-view floaters while preserving critical multi-view consensus, producing the final re-initialized Gaussian set \(\Theta\).

Loss & Training

The pipeline runs through three consecutive optimization rounds: coarse-level pre-optimization, fine-level pre-optimization, and final optimization, totaling 6,000 iterations via the Adam optimizer. - Photometric appearance loss \(\mathcal{L}_A\) combines \(L_1\) color error with D-SSIM across all stages. - Geometric regularization adapts hierarchically: coarse pre-optimization applies hard depth regularization \(\mathcal{L}_C\) to aggressively flatten off-surface floaters into the approximate surface plane; fine-level pre-optimization and final optimization switch to soft depth regularization \(\mathcal{L}_F(p)\), preserving subtle surface variations without eroding intricate object contours.

Key Experimental Results

Main Results

Quantitative evaluations were performed on the standard forward-facing LLFF benchmark across challenging 2-view, 3-view, and 4-view settings. Methods marked with "Sparsity Discrepancy (!)" exploit dense SfM point clouds derived from all available views in the dataset; methods marked with "(%)" strictly adhere to few-shot training inputs without leakage.

Method Sparsity Discrepancy PSNR (2/3/4-view) ↑ SSIM (2/3/4-view) ↑ LPIPS (2/3/4-view) ↓ AVGE (2/3/4-view) ↓
RegNeRF - 16.55 / 19.41 / 21.49 0.468 / 0.627 / 0.713 0.417 / 0.306 / 0.257 0.207 / 0.149 / 0.105
FreeNeRF - 17.07 / 19.97 / 21.80 0.513 / 0.652 / 0.713 0.376 / 0.280 / 0.259 0.187 / 0.134 / 0.105
SparseNeRF - 17.74 / 20.33 / 21.90 0.513 / 0.657 / 0.720 0.386 / 0.302 / 0.260 0.171 / 0.127 / 0.101
3DGS ! 15.45 / 17.25 / 18.81 0.522 / 0.574 / 0.669 0.383 / 0.300 / 0.257 0.196 / 0.155 / 0.124
FSGS ! 18.45 / 20.31 / 21.11 0.617 / 0.652 / 0.703 0.325 / 0.288 / 0.220 0.137 / 0.104 / 0.098
DNGaussian ! 17.72 / 19.12 / 20.76 0.537 / 0.591 / 0.717 0.357 / 0.294 / 0.236 0.160 / 0.132 / 0.102
CoR-GS ! 18.73 / 20.45 / 21.70 0.636 / 0.712 / 0.769 0.247 / 0.196 / 0.169 0.125 / 0.101 / 0.082
3DGS (random) % 12.83 / 14.99 / 17.31 0.311 / 0.483 / 0.584 0.470 / 0.362 / 0.297 0.273 / 0.202 / 0.153
FSGS (random) % 15.58 / 18.28 / 20.30 0.461 / 0.572 / 0.673 0.463 / 0.312 / 0.273 0.217 / 0.147 / 0.110
DNGaussian (random) % 16.47 / 18.87 / 20.86 0.450 / 0.583 / 0.690 0.408 / 0.313 / 0.266 0.192 / 0.141 / 0.107
CoR-GS (random) % 14.96 / 16.32 / 18.03 0.363 / 0.452 / 0.561 0.429 / 0.354 / 0.295 0.225 / 0.188 / 0.151
FewViewGS % - / 18.96 / - - / 0.585 / - - / 0.307 / - - / 0.135 / -
ReInGS (Ours) % 19.39 / 20.88 / 22.24 0.599 / 0.687 / 0.748 0.296 / 0.256 / 0.239 0.132 / 0.108 / 0.092

On the object-centric DTU dataset under the 3-view setting, ReInGS demonstrates consistent superiority: - Vanilla 3DGS yields only 12.77 dB PSNR; FSGS drops to 12.93 dB without extended points. - DNGaussian and FewViewGS reach 18.85 dB and 19.13 dB, respectively. - ReInGS achieves 21.30 dB PSNR, 0.832 SSIM, 0.170 LPIPS, and 0.081 AVGE, outperforming the second-best unassisted method by over 2.17 dB in PSNR.

Ablation Study

On the LLFF dataset under the 3-view setting, an ablation study benchmarked against a reproduced unassisted DNGaussian baseline reveals the explicit contributions of EHPS, IVPS, and CVPS:

Configuration EHPS IVPS CVPS PSNR ↑ LPIPS ↓ SSIM ↑ Key Takeaway
Baseline (DNGaussian) - - - 18.87 0.313 0.583 Baseline sparse random point initialization
+ EHPS ✓ - - 19.38 0.348 0.589 Global spatial coverage anchors topology (+0.51 dB)
+ EHPS + IVPS ✓ ✓ - 19.79 0.260 0.645 Injects local geometric details (LPIPS drops sharply to 0.260)
+ EHPS + CVPS ✓ - ✓ 20.05 0.314 0.611 Multi-view consistency filtering boosts fidelity (+1.18 dB)
Standalone VDPS - ✓ ✓ 19.66 0.285 0.638 Without coarse macro scaffold, detail sampling yields limited gains
Full Model (ReInGS) ✓ ✓ ✓ 20.88 0.256 0.687 Synergistic combination achieves top scores across all metrics (+2.01 dB)

Key Findings

  • Complementary Roles of Intra-View and Cross-View Sampling: IVPS drives substantial perceptual quality gains (LPIPS improved from 0.348 to 0.260, SSIM from 0.589 to 0.645), confirming that residual-guided ray back-projection accurately targets missing surface geometry. Concurrently, CVPS drives PSNR gains (19.38 to 20.05 dB), validating that multi-view projection filtering suppresses single-view overfitting.
  • Interdependence Between Coarse and Fine Stages: Executing the fine-level VDPS stage alone without EHPS (Row 5) yields 19.66 dB PSNR. Integrating EHPS provides the essential global spatial scaffold, unlocking an additional 1.22 dB boost to reach 20.88 dB, underscoring the necessity of coarse-to-fine coupling.
  • Superiority Over Data-Leaked Counterparts: In the extreme 2-view setting on LLFF, ReInGS attains 19.39 dB PSNR, outperforming even CoR-GS with data leakage (18.73 dB), conclusively showing that principled re-initialization eliminates the need for extended-view stereo points.

Highlights & Insights

  • Exposing a Fundamental Benchmark Vulnerability: Systematically identifies and formalizes the "sparsity discrepancy" issue in few-shot 3DGS, providing the community with a rigorous, leakage-free benchmark standard.
  • Staged Geometric Regularization: Couples coarse pre-optimization with hard depth regularization to suppress floating artifacts, then transitions to soft depth regularization in fine stages to preserve subtle depth variations and surface contours.
  • Bridging Image-Order and Object-Order Graphics Principles: Translates classic rendering concepts—image-order ray querying and object-order projection accumulation—into actionable point-sampling mechanisms for discrete radiance field initialization.

Limitations & Future Work

  • Hollows and Cracks Under Severe Extrapolation: Because Gaussians cannot perfectly tile all pixels under drastic camera viewpoint alterations, unobserved gaps between adjacent primitives may appear as holes or cracked surfaces.
  • Computational Overhead: The multi-phase point generation and pre-optimization steps elevate training time and peak GPU memory compared to single-pass random point seeding, calling for lighter adaptive pruning in future iterations.
  • vs FSGS / CoR-GS: While prior works report strong metrics, they depend heavily on dense SfM points extracted from all available dataset views, failing when restricted to true few-shot inputs. ReInGS eliminates this dependency and achieves higher accuracy purely from training views.
  • vs DNGaussian: DNGaussian employs sparse random points without volume expansion or structured detail refinement, frequently leading to localized floater clumps. ReInGS builds an expressive coarse-to-fine hierarchy that couples macro space expansion with micro feature recovery.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates the sparsity discrepancy dilemma and designs an elegant coarse-to-fine re-initialization pipeline inspired by image/object-order rendering.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Extensively analyzes with-leakage and without-leakage settings across LLFF and DTU benchmarks, accompanied by clear visual and quantitative ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear exposition, cohesive logic, rigorous mathematical formulation, and faithful adherence to graphics principles.
  • Value: ⭐⭐⭐⭐⭐ Fixes a persistent evaluation loophole in few-shot 3DGS and establishes a robust baseline for unassisted sparse-view neural rendering.