Geometric Probing for Isotropic Optimization Manifold in Sparse-View 3D Gaussian Splatting¶
Conference: ECCV 2026
Paper: ECCV Official Page
PDF: ECCV PDF
Code: https://github.com/zyl123456aB/StableGS
Area: 3D Vision
Keywords: 3D Gaussian Splatting, Sparse-View Reconstruction, Isotropic Optimization Manifold, Geometric Probing, Rendering Stability
TL;DR¶
Addressing the severe directional anisotropy of the 3DGS loss surface under sparse views which triggers floaters and splat collapse, StableGS steers optimization toward an isotropic manifold via Hessian-informed Attribute-Space Probing and Viewpoint-Space Probing without external priors, reducing anisotropy by 74% and advancing SOTA across LLFF, DTU, and Mip-NeRF360.
Background & Motivation¶
Reconstructing high-fidelity 3D radiance fields from sparse views (typically 3 to 9 images) is essential for autonomous driving, robotics, and virtual reality. However, standard 3D Gaussian Splatting (3DGS) degrades catastrophically when supervision is scarce, giving rise to persistent floaters, splat collapse or explosion, and view-dependent opacity tearing. Existing remedies predominantly address symptoms: they inject external monocular depth estimators, apply heuristic geometric regularizations, or randomly drop Gaussians for self-ensembling. Yet, such methods remain hostage to error-prone depth priors and fail to explain the intrinsic mechanism causing the differentiable splatting pipeline to disintegrate under sparse supervision.
Analyzing the problem through the lens of optimization manifold curvature reveals that the root failure stems from extreme "directional anisotropy" on the loss surface. Because sparse camera views leave vast regions of parameter space unconstrained, the splatting operator's sensitivities along different parameter directions diverge by orders of magnitude: infinitesimal positional shifts along the viewing ray trigger depth rank flips, extreme aspect ratios induce projection ill-conditioning, and opacity values overfit view-dependent visibility. This landscape—featuring steep ridges along fragile axes coexisting with flat valleys elsewhere—steers standard gradient descent into sharp, narrow minima that satisfy training images but collapse on novel viewpoints. Generic flat-minima algorithms like Sharpness-Aware Minimization (SAM) enforce uniform parameter-space smoothing, unable to account for the structured, non-linear geometric mechanics and multi-view projective constraints intrinsic to 3DGS.
To overcome this fundamental limitation without external depth priors, this paper introduces Stable 3DGS (StableGS). Core idea: by diagnosing curvature through Hessian trace estimation, Attribute-Space Probing (ASP) perturbs Gaussians along fragile geometric axes to penalize worst-case rendering loss, while Viewpoint-Space Probing (VSP) enforces cross-view consistency across occlusion boundaries and parallax-critical poses, actively reshaping the optimization manifold from directional anisotropy into an isotropic, stable basin.
Method¶
Overall Architecture¶
StableGS augments the standard 3DGS pipeline with geometric manifold diagnostics and regularized probing. Training begins with a 1,000-iteration standard reconstruction warm-up. Subsequently, the manifold probing loop executes: a matrix-free Hutchinson trace estimator periodically evaluates average curvature across each attribute dimension to scale probe magnitudes inversely with sensitivity; Attribute-Space Probing (ASP) simultaneously perturbs positions (with depth bias), rotation quaternions, log-scales, and logit opacities, rasterizing a single perturbed configuration to compute an adversarial reconstruction loss; concurrently, Viewpoint-Space Probing (VSP) constructs composite novel cameras targeting depth edges, geodesic trajectory interpolations, and parallax offsets, enforcing rendering consistency between base and perturbed models. The entire pipeline avoids explicit Hessian inversion, incurring zero test-time overhead while reducing the manifold condition number by 74%.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Sparse Multi-View Images and Camera Poses"] --> B["Standard 3DGS Field Initialization & Warm-up"]
B --> C["Hessian Curvature Estimator<br/>Hutchinson trace estimation of attribute curvature"]
C --> D["Attribute-Space Probing ASP<br/>Depth-biased position / ill-conditioned rotation-scale / opacity"]
C --> E["Viewpoint-Space Probing VSP<br/>Occlusion boundary normal / parallax translation / pose interpolation"]
D --> F["Perturbed Gaussian Rasterization & Loss Evaluation"]
E --> F
F --> G["Total Objective Optimization<br/>L_recon + λ_ASP·L_ASP + λ_VSP·L_VSP"]
G --> H["Isotropically Stable Manifold for Robust Novel View Synthesis"]
Key Designs¶
1. Optimization Manifold Diagnosis: Characterizing the Three Geometric Fragility Axes
The fundamental obstacle in sparse-view 3DGS is the ill-conditioned loss Hessian \(\mathcal{A}(\boldsymbol{\Theta}) = \lambda_{\max}(\mathbf{H}) / \lambda_{\min}(\mathbf{H})\). Second-order sensitivity analysis isolates three primary fragility axes: first, Depth Swap, where minute position perturbations along the viewing ray \(v_{\text{cam}} = \mu_i - t_m\) flip the depth sorting order in \(\mathcal{N}(p)\), with ray-aligned gradients exceeding lateral gradients by over \(100\times\); second, Projection Ill-conditioning, where elongated or obliquely oriented Gaussians cause 2D projected covariance condition numbers \(\kappa(\boldsymbol{\Sigma}'_i) = \lambda_{\max}(\boldsymbol{\Sigma}'_i) / \lambda_{\min}(\boldsymbol{\Sigma}'_i)\) to exceed 100, causing splats to collapse into needles or explode across the screen under minor rotations; and third, Opacity Instability, where logit opacity \(\alpha_i\) acts as a degenerate visibility switch to memorize training views, corrupting transmittance products on unseen angles.
2. Attribute-Space Probing (ASP): Hessian-Adaptive Fragility Suppression
To flatten these sharp loss ridges, ASP applies curvature-informed perturbations. Utilizing the Hutchinson trace estimator, it computes the attribute curvature scale \(\tau_\phi = \sqrt{\text{tr}(\mathbf{H}_\phi)/d_\phi}\) and assigns an adaptive probe magnitude \(\rho_\phi = c_\phi / (\tau_\phi + \epsilon)\), probing aggressively in flat directions while scaling down in steep zones. For positions, gradients are orthogonally projected onto the camera viewing ray to isolate depth and lateral components, applying a depth-weighted perturbation: $\(\boldsymbol{\epsilon}_{\mu, i} = \rho_\mu \cdot \frac{\gamma \cdot \mathbf{g}_{\text{depth}, i} + (1-\gamma) \cdot \mathbf{g}_{\text{lateral}, i}}{\|\gamma \cdot \mathbf{g}_{\text{depth}, i} + (1-\gamma) \cdot \mathbf{g}_{\text{lateral}, i}\|}\)$ with \(\gamma = 0.7\) explicitly stressing occlusion boundaries. For rotation, perturbations in the \(\mathfrak{so}(3)\) Lie algebra tangent space are scaled by the 2D footprint condition number \(\kappa(\boldsymbol{\Sigma}'_i)\); scales are perturbed in log-space to ensure positivity, and opacity is probed in logit space along the gradient sign. All perturbations are consolidated into a single global configuration \(\tilde{\boldsymbol{\Theta}}\), requiring just one extra rasterization pass to minimize worst-case rendering loss \(\mathcal{L}_{\text{ASP}}(\boldsymbol{\Theta})\).
3. Viewpoint-Space Probing (VSP): Cross-View Consistency Under Geometrically Critical Poses
Restricting probes to training views allows the model to remain trapped in sharp, view-specific basins. VSP synthesizes geometrically critical target poses to diagnose out-of-distribution stability. For each training camera \(P_m\), a single composite novel view \(P_k = \{R_k, t_k\}\) is constructed by uniting three geometric perturbations: rotating around depth-edge normals \(\exp(\theta \cdot \mathbf{n}_{\text{edge}, m})\) to challenge occlusion contours; interpolating/extrapolating poses along geodesic paths using Slerp; and injecting lateral translations \(t_{\text{parallax}} = \delta \cdot t_{\perp, m}\) perpendicular to the optical axis to maximize near-field parallax. By penalizing the discrepancy between base model renders \(\hat{\mathbf{I}}_k(\boldsymbol{\Theta})\) and perturbed model renders \(\hat{\mathbf{I}}_k(\tilde{\boldsymbol{\Theta}})\) via \(\mathcal{L}_{\text{VSP}}(\boldsymbol{\Theta})\), VSP prohibits view-specific memorization and enforces true 3D spatial coherence.
Loss & Training¶
The complete optimization objective combines three terms: $\(\mathcal{L}_{\text{total}}(\boldsymbol{\Theta}) = \mathcal{L}_{\text{recon}}(\boldsymbol{\Theta}) + \lambda_{\text{ASP}} \mathcal{L}_{\text{ASP}}(\boldsymbol{\Theta}) + \lambda_{\text{VSP}} \mathcal{L}_{\text{VSP}}(\boldsymbol{\Theta})\)$ where \(\mathcal{L}_{\text{recon}}\) combines photometric \(\ell_1\) and D-SSIM losses on training views. \(\mathcal{L}_{\text{ASP}}\) penalizes reconstruction error under perturbed parameters \(\tilde{\boldsymbol{\Theta}}\), and \(\mathcal{L}_{\text{VSP}}\) enforces \(\ell_1\) and SSIM agreement across \(K\) composite probe viewpoints. Standard hyperparameters selected on the LLFF validation set are fixed across all benchmarks: \(\lambda_{\text{ASP}} = 1.0\), \(\lambda_{\text{VSP}} = 0.5\), and global probing scale \(c_\phi = 0.01\). The Hutchinson estimator employs \(B = 10\) Gaussian vectors, refreshed every 100 iterations alongside densification and pruning.
Key Experimental Results¶
Main Results¶
Quantitative evaluations across LLFF (3/6/9 views), DTU (3 views), and Mip-NeRF360 (12/24 views) demonstrate consistent superiority over leading NeRF and 3DGS baselines:
| Dataset / Setting | Metric | 3DGS (Baseline) | DropGaussian (Prev. SOTA) | StableGS (Ours) | Gain vs Prev. SOTA |
|---|---|---|---|---|---|
| LLFF (3-view) | PSNR ↑ / SSIM ↑ / LPIPS ↓ | 16.46 / 0.440 / 0.401 | 20.76 / 0.713 / 0.200 | 21.19 / 0.726 / 0.208 | +0.43 dB / +0.013 / - |
| LLFF (6-view) | PSNR ↑ / SSIM ↑ / LPIPS ↓ | 21.09 / 0.699 / 0.229 | 24.74 / 0.837 / 0.117 | 25.05 / 0.841 / 0.127 | +0.31 dB / +0.004 / - |
| LLFF (9-view) | PSNR ↑ / SSIM ↑ / LPIPS ↓ | 23.21 / 0.785 / 0.176 | 26.21 / 0.874 / 0.088 | 26.47 / 0.885 / 0.097 | +0.26 dB / +0.011 / - |
| DTU (3-view) | PSNR ↑ / SSIM ↑ / LPIPS ↓ | 14.74 / 0.672 / 0.249 | 20.22 / 0.830 / 0.150* | 20.86 / 0.848 / 0.130 | +0.64 dB / +0.018 / -0.020 |
| Mip-NeRF360 (12-view) | PSNR ↑ / SSIM ↑ / LPIPS ↓ | 18.52 / 0.523 / 0.415 | 19.74 / 0.577 / 0.364 | 19.98 / 0.580 / 0.368 | +0.24 dB / +0.003 / - |
| Mip-NeRF360 (24-view) | PSNR ↑ / SSIM ↑ / LPIPS ↓ | 22.80 / 0.708 / 0.276 | 24.13 / 0.762 / 0.225 | 24.38 / 0.756 / 0.233 | +0.25 dB / -0.006 / - |
*Note: Baseline on DTU corresponds to DropoutGS; StableGS improves upon it by +0.64 dB PSNR with significantly sharper perceptual quality (lower LPIPS).
Ablation Study¶
Systematic ablations isolate the contribution of each fragility axis, probe sampling mode, and cross-dataset robustness:
| Experiment Category | Specific Configuration | LLFF 3-view PSNR (dB) | DTU 3-view PSNR (dB) | Mip-NeRF360 12-view PSNR (dB) | Anisotropy Index \(\mathcal{A}\) Reduction |
|---|---|---|---|---|---|
| Baseline | Vanilla 3DGS | 16.46 | 14.74 | 18.52 | Baseline (0%) |
| ASP Axis Breakdown | Position probing only | 18.92 | 18.35 | 19.31 | - |
| ASP Axis Breakdown | Opacity probing only | 17.85 | 16.98 | 19.02 | - |
| ASP Axis Breakdown | Rotation/Scale probing only | 17.52 | 16.42 | 18.89 | - |
| VSP Probe Modes | Occlusion Boundary Rotation only | 19.87 | - | - | - |
| VSP Probe Modes | Interpolation / Extrapolation only | 19.62 | - | - | - |
| VSP Probe Modes | Parallax Translation only | 19.43 | - | - | - |
| Full Model | StableGS (Full ASP + VSP) | 21.19 | 20.86 | 19.98 | 74% Reduction |
Key Findings¶
- Depth sorting is the primary failure mode: Position probing alone yields a +2.46 dB surge on 3-view LLFF, capturing 84% of the standalone ASP gain and proving that ray-direction depth instability is the main cause of floater artifacts.
- Fragility axes exhibit structured synergy: Standalone gains across the three axes sum to +4.91 dB while combined ASP provides +2.92 dB, demonstrating that the probes successfully target overlapping aspects of the splatting operator's failure modes.
- Superior training efficiency: On an RTX 5090 GPU, StableGS reaches 21.16 dB at the 9.7-minute mark—matching vanilla 3DGS wall-clock time—already beating all fully converged baselines. Completing training in 13.1 minutes (+1.3× time) yields +4.73 dB over baseline, incurring zero extra cost at rendering time.
Highlights & Insights¶
- Intrinsic manifold regularization without external priors: Departing from heuristic monocular depth priors that introduce systemic scale errors, StableGS shows that reforming the loss curvature directly yields superior and more physically plausible geometry.
- Efficient global probe bundling: Consolidating 5 attribute perturbations into a single global configuration allows the entire manifold exploration step to cost only one extra rasterization pass per iteration.
- Broad transferability to differentiable rendering: The depth-vs-lateral gradient decomposition and \(\mathfrak{so}(3)\) condition-adaptive perturbation principles apply directly to other differentiable representations, including 2D-GS, Mip-Splatting, and differentiable neural surface meshes.
Limitations & Future Work¶
- Lack of generative priors in extreme extrapolation: Under extreme camera viewpoints outside the training convex hull with zero visual overlap, the model cannot hallucinate unseen high-frequency textures without external generative priors.
- Stochastic variance in Hessian trace estimation: Hutchinson estimators can exhibit localized variance in scenes with millions of Gaussians; momentum-smoothed curvature tracking represents an appealing avenue for stabilization.
- Integration with surface-oriented primitives: Extending isotropic manifold constraints to Gaussian normal alignment and compact mesh extraction (such as 2D-GS and SuGaR) could unlock seamless conversion to high-fidelity CAD/mesh assets.
Related Work & Insights¶
- vs DNGaussian / FSGS: While DNGaussian and FSGS rely on monocular depth estimators to provide pseudo-ground truth depth, they suffer in non-Lambertian or textureless regions; StableGS operates without external supervision, outperforming DNGaussian by nearly 2 dB on DTU.
- vs DropGaussian / DropoutGS: Dropout methods randomly discard Gaussians to prevent overfitting through isotropic parameter dropping; StableGS targets the specific geometric failure axes of the splatting operator using second-order curvature diagnostics, achieving higher optimization efficiency.
- vs SAM (Sharpness-Aware Minimization): SAM optimizes for isotropic flat minima in generic parameter space; StableGS proves that generic flattening cannot resolve splatting instabilities, pioneerng physically grounded geometric and viewpoint probing.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First work to diagnose sparse 3DGS breakdown from the perspective of optimization manifold anisotropy, proposing targeted geometric probes.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarks across LLFF, DTU, and Mip-NeRF360, featuring per-axis ablations, curvature diagnostics, efficiency analysis, and condition number evaluations.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulations, clear geometric intuition, and cohesive structure from analysis to methodology.
- Value: ⭐⭐⭐⭐⭐ Provides a principled, prior-free framework for few-shot 3D reconstruction, presenting valuable insights for differentiable rendering optimization.