InceptionGS: Generative Bootstrapping for Large-Scale Gaussian Splatting under Unstructured View Sampling¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: 3D Vision
Keywords: Large-scale scene reconstruction, 3D Gaussian Splatting, unstructured view sampling, generative bootstrapping, geometry-conditioned diffusion
TL;DR¶
To tackle unstructured view sampling and severe view scarcity in large-scale scene captures, InceptionGS adapts a latent diffusion model on dynamically rasterized G-buffer geometry from an initial Gaussian splatting, and refines the 3DGS field via photometric blending and Langevin importance sampling, reducing average FID by 32% without requiring additional real-world captures.
Background & Motivation¶
Creating high-fidelity, photorealistic 3D digital assets of large-scale real-world environments is fundamental for virtual reality, film production, and interactive simulation. In these applications, users expect seamless, immersive navigation across arbitrary viewing angles and scales. However, acquiring multi-view imagery for complex outdoor landmarks inherently suffers from unstructured view sampling. Unlike compact, tabletop, or bounded indoor objects where camera trajectories can be arranged in near-isotropic hemispherical distributions, large-scale outdoor captures demand convoluted camera trajectories shaped by terrain constraints, varying distances, and physical obstacles. As a result, while most open areas receive redundant coverage, critical corners, occluded faΓ§ades, and elevated angles inevitably experience severe view scarcity.
Existing reconstruction-based and generation-based novel view synthesis (NVS) paradigms struggle when faced with such unstructured sampling. On the one hand, reconstruction-based approaches such as 3D Gaussian Splatting (3DGS) and its surface-regularized variants excel given dense observations, but their performance collapses drastically once camera poses deviate into under-observed regions, yielding conspicuous floaters, blurring, and needle-like stretching artifacts. Laborious physical re-capture is often impractical or prohibited in real-world scenarios. On the other hand, recent generative NVS methods leverage video or multi-view diffusion priors to hallucinate unseen regions. However, generic diffusion models struggle to generalize to intricate, unique large-scale structures (such as the sprawling segments of the Great Wall or ornate architectural cornices) and typically condition generation on low-quality artifact-ridden RGB renderings or PlΓΌcker ray embeddings, which easily compromise long-term 3D consistency and precise structural control.
Reconstruction provides a verifiable scene-specific geometric foundation, whereas generative models offer realistic high-frequency textural statistics. Rather than treating reconstruction and generation in isolation, the key opportunity lies in mutual bootstrapping: using reconstructed geometry to customize generative priors, which in turn repair deficient regions in the reconstructive field. Core idea: mutually bootstrap Gaussian splatting and diffusion generation by adapting a latent diffusion model on dynamically extracted G-buffer geometry (normals, opacity, and multi-resolution hash features), followed by view-space Langevin importance sampling and reliable photometric blending to softly inject generative supervision into the 3DGS field.
Method¶
Overall Architecture¶
The InceptionGS framework operates across two interconnected stages: Stage I performs "Generative Adaptation", where an initial 3DGS field (anchored by PGSR planar regularization) optimizes geometric surfaces while dynamically rasterizing geometry buffers (normal maps, opacity masks, and learnable hash feature maps) to fine-tune a pre-trained latent diffusion model (LDM) via ControlNet, capturing the internal scene-specific geometry-appearance correspondence. Stage II performs "Gaussian Bootstrapping", which identifies artifact-prone target viewpoints, constructs candidate virtual viewpoints via pose clustering and interpolation, blends diffusion-generated appearances with reliable warped photometric clues, and dynamically injects generative supervision into the 3DGS field through Langevin Monte Carlo (LMC) importance sampling.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multi-view image inputs<br/>Unstructured sparse sampling"] --> B["Stage 1: Geometry-aware generative adaptation<br/>Dynamic G-buffer rasterization + ControlNet co-optimization"]
B --> C["Stage 2: Candidate virtual viewpoint generation<br/>Pose cluster retrieval + Quaternion SLERP interpolation"]
C --> D["Stage 3: Photometric blending<br/>Depth-based image warping + Confidence mask composition"]
D --> E["Stage 4: View-space importance sampling<br/>LMC posterior energy guidance + Soft 3DGS refinement"]
E --> F["High-fidelity refined 3DGS scene"]
Key Designs¶
1. Geometry-aware generative adaptation: Learning scene-specific appearance conditioned on dynamic G-buffers
Conditioning generative models on low-quality RGB renderings contaminated by artifacts inevitably leads to error propagation and spatial inconsistency. In contrast, scene geometry exhibits strict 3D consistency and is substantially more robust under planar regularization. InceptionGS builds upon PGSR, flattening Gaussian ellipsoids into local planes regularized by single-view normal alignment, NCC photo-consistency, and multi-view geometric reprojection constraints. During optimization, three geometry representations are dynamically rasterized on-the-fly into a G-buffer \(R = (N, O, F) \in \mathbb{R}^{H \times W \times (3 + 1 + Z)}\): planar surface normals \(N\), foreground opacity mask \(O\), and a learnable surface-aware multi-resolution hash feature map \(F\). Randomly sampled local image patches \(\hat{I}\) and their corresponding geometry patches \(\hat{R}\) are used to fine-tune ControlNet parameters \(\theta\) of a pre-trained LDM, exploiting internal self-similarity via shared convolutional kernels: $$ \mathcal{L}{D} = \mathbb{E}) - \epsilon|_2^2\right] $$ Full-resolution synthesis at inference time is achieved by fusing overlapping diffusion paths using MultiDiffusion, ensuring that the generated image }\left[|\epsilon_{\theta}(z_t, t, \hat{R\(I^* = \mathcal{D}(z_0^*)\) strictly respects the underlying reconstructed geometry while restoring vivid high-frequency appearance.
2. Candidate virtual viewpoint generation: Trajectory interpolation anchored by pose proximity
Directly optimizing an isolated problematic target viewpoint risks localized overfitting and breaks spatial continuity in 3D space. For any designated artifact-prone target viewpoint \(j^\star\), InceptionGS selects the top-\(K\) nearest training camera views \(V_{\text{src}}^{(j^\star)}\) based on a joint metric evaluating Euclidean position and optical axis orientation. Between target view \(j^\star\) and each selected reference view \(i \in V_{\text{src}}^{(j^\star)}\), camera centers are linearly interpolated while rotations are interpolated via spherical linear interpolation (SLERP): $$ V_{\text{vrt}} = \left{\text{interp}(j^\star, i, l) \mid i \in V_{\text{src}}^{(j^\star)}, l \in [L]\right} $$ This constructs \(|V_{\text{vrt}}| = LK\) smoothly distributed candidate virtual viewpoints along intermediate viewing rays, providing a stable spatial scaffolding for long-range generative refinement.
3. Photometric blending: Fusing trustworthy warped observations with generative completions
While conditional diffusion synthesis completes missing semantic structures, latent VAE decoders often introduce subtle domain shifts and high-frequency textural blur compared to real camera captures. To retain genuine photographic detail, InceptionGS performs depth-guided image-based rendering. For each virtual view \(j \in V_{\text{vrt}}\), nearby real training views are warped to view \(j\) using camera matrices and reconstructed PGSR depth maps to obtain warped images \(\{I_{i \to j}\}\). Blending weights \(W_{i \to j}\) measure visibility and apply a conservative depth threshold \(\delta\) to filter out occluded or unreliable regions, producing an initial photometric composite \(\widetilde{I}_j = \sum_{i} W_{i \to j} I_{i \to j}\). A binary visibility mask \(M\) identifies valid projection regions, and missing or occluded areas are seamlessly filled with the generative output \(I^*_j\): $$ I^{}_{j} \leftarrow M \circ \widetilde{I}_j + (\mathbf{1} - M) \circ I^{}_{j} $$ This hybrid supervision anchors the high-frequency textural fidelity to real ground-truth captures wherever visible, while restricting generative hallucination strictly to genuine blind spots.
4. View-space importance sampling: Langevin Monte Carlo exploration for balanced supervision
Uniformly injecting generative loss across all candidate virtual viewpoints risks over-exposing the 3DGS field to synthetic supervision, degrading well-reconstructed regions that already conform to photometric evidence. InceptionGS introduces an adaptive view-space importance sampling scheme based on Langevin Monte Carlo (LMC). The log-likelihood of posterior distribution \(P\) is modeled by the \(L_1\) discrepancy between the current 3DGS rendering \(C_q\) and the blended target image \(I^*_q\), i.e., \(P = \exp(\|C_q - I^*_q\|_1)\). The camera pose vector \(v \in \mathbb{R}^7\) (translation and quaternion) is updated by combining posterior gradients with stochastic noise \(\eta \sim \mathcal{N}(0, 1)\): $$ \hat{v}{\tau+1} = v\tau + a \nabla_{v_\tau} \log P + b \eta $$ The continuous update is mapped to the nearest pre-cached discrete viewpoint \(q = \arg\min_{j \in V_{\text{vrt}}} \|\hat{v}_{\tau+1} - v_j\|_2\). This focuses optimization steps on virtual viewpoints exhibiting the largest reconstruction-generation divergence, enabling stable, soft regularization without disturbing well-converged scene regions.
Loss & Training¶
In Stage I (Generative Adaptation), the 3DGS model and ControlNet are jointly trained using the color reconstruction loss \(\mathcal{L}_C = (1-\lambda_1)\mathcal{L}_1 + \lambda_1 \mathcal{L}_{\text{D-SSIM}}\), the PGSR geometric regularization loss \(\mathcal{L}_G = \lambda_{SV}\mathcal{L}_{SV} + \lambda_{MVC}\mathcal{L}_{MVC} + \lambda_{MVG}\mathcal{L}_{MVG}\), and the diffusion denoising loss \(\mathcal{L}_D\).
In Stage II (Bootstrapping), the selected virtual view \(j\) is supervised by an adaptive structural-perceptual loss: $$ \mathcal{L}P^{(j)} = \begin{cases} \lambda_2 \mathcal{L}(C_j, I^}j), & \text{if } \mathcal{L}(C_j, I^}_j) \le \tau \ \lambda_1 \mathcal{L}_1(C_j, I^j) + \lambda_2 \mathcal{L}(C_j, I^}_j), & \text{otherwise} \end{cases} $$ When structural similarity falls below threshold \(\tau\), the pipeline switches from rigid pixel-level losses to deep perceptual LPIPS supervision, mitigating blur caused by sub-pixel misalignments. The bootstrapping process supports iterative refinement across different target view defects.
Key Experimental Results¶
Main Results¶
Quantitative evaluations are conducted on seven massive outdoor scenes from the GigaNVS benchmark and public outdoor/indoor scenes from MipNeRF360. Unstructured view sampling is simulated by clustering camera poses into 6 groups via farthest point sampling, holding out 90% of views in a selected cluster as the sparse target test set, and holding out 5% across remaining clusters to measure standard view synthesis.
| Dataset | Metric | Ours (Full) | Best Diffusion Baseline (Difix3D+) | Geometry Reconstruction Baseline (PGSR) | Gain (vs Best Diffusion) |
|---|---|---|---|---|---|
| GigaNVS (Large-scale scenes) | PSNR β | 17.24 | 15.86 | 16.54 | +1.38 dB |
| GigaNVS | LPIPS β | 0.360 | 0.418 | 0.438 | -13.9% (better) |
| GigaNVS | FID β | 64.42 | 71.39 | 91.27 | -9.8% (better) |
| MipNeRF360 (Standard 360 scenes) | PSNR β | 20.72 | 19.08 | 18.92 | +1.64 dB |
| MipNeRF360 | LPIPS β | 0.297 | 0.321 | 0.315 | -7.5% (better) |
| MipNeRF360 | FID β | 62.78 | 66.73 | 83.59 | -5.9% (better) |
Across generative alternatives, video-based NVS baselines (e.g., MVSplat360 at 168.62 FID, AC3D at 177.49 FID, SEVA at 135.48 FID) suffer from severe hallucinations and drift on large-scale architectures, whereas InceptionGS maintains consistent perceptual quality and fidelity.
Ablation Study¶
The ablation analysis examines the individual contributions of Photometric Blending (PB), View-adaptive Sampling (VS), and generative Fine-Tuning (FT) on GigaNVS and MipNeRF360:
| Config | GigaNVS PSNR β | GigaNVS LPIPS β | GigaNVS FID β | MipNeRF360 PSNR β | MipNeRF360 FID β | Note |
|---|---|---|---|---|---|---|
| Ours (Full) | 17.24 | 0.360 | 64.42 | 20.72 | 62.78 | full model |
| w/o PB | 16.97 | 0.376 | 73.35 | 20.48 | 64.89 | Missing real photographic high-frequency details |
| w/o VS | 16.88 | 0.394 | 72.89 | 20.28 | 66.59 | Uniform virtual view injection causes negative conflict |
| w/o FT | 14.97 | 0.610 | 176.40 | 17.09 | 164.67 | Generic prior hallucinates severely without scene adaptation |
In severe scarcity stress tests (reducing visible views from 4 to 2 on MipNeRF360 Room), PSNR drops marginally from 26.82 to 26.60, and FID increases slightly from 50.89 to 53.35. Varying candidate virtual view counts by 0.5Γ or 1.5Γ yields remarkably stable performance (PSNR fluctuates within 16.24β16.30), verifying the robustness of the soft sampling strategy.
Key Findings¶
- Scene-specific geometric fine-tuning is the single most critical driver of quality: Removing FT causes FID on GigaNVS to surge from 64.42 to 176.40. Off-the-shelf diffusion priors lack structural understanding of unique real-world architectural patterns; on-the-fly G-buffer conditioning is essential to constrain generation.
- Iterative bootstrapping enables non-destructive progressive repair: On the TW-Pavilion scene, executing a second bootstrapping round targeting the same viewpoint (\(j^{\star(2)} = j^{\star(1)}\)) maintains stable metrics (PSNR 14.53 vs 14.66), whereas targeting a distinct defective viewpoint (\(j^{\star(2)} \ne j^{\star(1)}\)) boosts PSNR to 16.28 and drops FID from 143.93 to 53.79, without compromising previously refined areas.
Highlights & Insights¶
- Decoupling generative priors via geometric manifold conditions: Rather than attempting 2D image restoration over degraded RGB renderings filled with floaters, conditioning diffusion on surface normals, opacity, and hash features provides an invariant 3D geometric anchor that eliminates cross-view appearance inconsistency.
- Energy-guided Langevin importance mining: Using posterior rendering discrepancies to drive Langevin sampling dynamically shifts optimization capacity to the most informative viewpoints, avoiding arbitrary heuristic sampling schedules.
- Self-contained digital asset maintenance cycle: The method establishes an operational loop: inspect rendering artifacts, synthesize targeted virtual views via geometric diffusion, blend photometric evidence, and locally update the 3DGS field without requiring physical recollecting missions.
Limitations & Future Work¶
- Training latency and memory requirements: As acknowledged in the paper, Stage I generative adaptation requires approximately 30 minutes with a peak VRAM of 30 GB, while Stage II bootstrapping requires around 25 minutes with a peak VRAM under 15 GB, limiting real-time deployment.
- Dependency on base geometric initialization: In scenarios where extreme view occlusions leave regions with zero intersecting rays, PGSR cannot form even a coarse planar geometry, leaving the conditional diffusion model without geometric guidance.
- Future directions: Integrating test-time training (TTT) or feed-forward geometry conditioning modules to eliminate per-scene diffusion fine-tuning overhead.
Related Work & Insights¶
- vs PGSR / 3DGS: Pure reconstruction methods rely exclusively on observed photometric consistency; under severe unstructured view sampling, unobserved regions inevitably suffer from severe tearing and floaters. InceptionGS introduces generative statistics to fill perceptual voids.
- vs Difix3D+ / ViewCrafter / AC3D: Existing generative NVS baselines synthesize novel views conditioned on camera trajectories or noisy RGB inputs, frequently producing structural distortion and temporal drift on large landmarks. InceptionGS achieves strict 3D consistency by binding diffusion models to 3DGS G-buffer geometry.
Rating¶
- Novelty: βββββ Elegant synergy between dynamic 3DGS geometric surface extraction and conditional diffusion adaptation.
- Experimental Thoroughness: βββββ Rigorous clustered holdout evaluation on GigaNVS alongside extensive ablation and stress tests.
- Writing Quality: βββββ Clear problem formulation, detailed methodological exposition, and cohesive mathematical definitions.
- Value: βββββ Provides a practical, high-quality blueprint for digitizing and maintaining large-scale 3D virtual environments.