Skip to content

Robust 3DGS-based SLAM via Adaptive Kernel Smoothing

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/xju-zsh/Robust-3DGS-based-SLAM-via-Adaptive-Kernel-Smoothing.git
Area: 3D Vision
Keywords: 3D Gaussian Splatting, Dense Visual SLAM, Adaptive Kernel Smoothing, Tracking Robustness, Local KNN Correction

TL;DR

Challenging the assumption that rendering sharpness governs tracking accuracy in 3DGS-SLAM, this paper introduces Corrective Blurry KNN (CB-KNN) adaptive kernel smoothing into differentiable rasterization to regularize the non-convex optimization landscape, reducing trajectory tracking errors by 19% to 35% without degrading global mapping fidelity.

Background & Motivation

With the advent of 3D Gaussian Splatting (3DGS) in neural radiance field rendering, differentiable rasterization over explicit Gaussian primitives has rapidly emerged as a leading paradigm for dense visual Simultaneous Localization and Mapping (SLAM). Compared to implicit coordinate-based representations such as NeRF, 3DGS offers high rendering frame rates and explicit geometric flexibility, allowing visual SLAM systems to execute dense photometric and geometric camera tracking directly by synthesizing color and depth images. However, existing 3DGS-SLAM frameworks predominantly operate under an underlying dogma: maximizing visual fidelity and achieving razor-sharp rendering is assumed to be the prerequisite for accurate camera pose estimation.

This conventional assumption breaks down severely in the presence of sensor noise, aggressive camera motions, and viewpoint changes. Because high-dimensional 3D Gaussian attributes (positions, anisotropic covariance scales, rotation quaternions, and view-independent or spherical harmonic colors) are jointly optimized via back-propagation on noisy observational frames, peripheral regions of Gaussians inevitably harbor considerable parameter noise and structural outliers. When passed through standard differentiable rasterization, these flawed parameters induce high-frequency rendering artifacts, color jitter, and geometric discontinuities. Under optimization theory, these rendering artifacts transform the photometric alignment landscape into a severely non-convex surface riddled with sharp local minima, easily trapping gradient-based pose optimizers and causing catastrophic tracking drift. Traditional post-processing strategies, such as feature re-weighting or robust M-estimators, are tailored to external sensor noise and fail to mitigate structural noise intrinsic to the map representation itself.

To resolve this core tension, this paper rethinks the objective and proposes a counter-intuitive smoothing perspective: for camera pose tracking, an intentionally smoother, slightly blurred rendering is vastly more resilient to Gaussian parameter noise than an overly sharp one. Theoretically grounded in Graduated Non-Convexity (GNC), allowing each Gaussian to influence a wider, more continuous footprint flattens spurious high-frequency noise spikes and expands the basin of attraction toward the global minimum. Core idea: introduce Corrective Blurry KNN (CB-KNN) adaptive kernel smoothing directly into the rasterization stage as a transient, plug-and-play regularizer that dynamically rectifies spatial positions and color weights of local neighboring Gaussians, thereby smoothing the optimization landscape for robust tracking while strictly preserving canonical map fidelity.

Method

Overall Architecture

The proposed system integrates adaptive kernel smoothing into a full-featured 3DGS-SLAM pipeline taking synchronized RGB-D streams as input. To balance tracking robustness and real-time efficiency, the architecture establishes a dual-track scheme: non-keyframes rely on a constant-velocity motion model and standard rasterization over the canonical Gaussian map for low-overhead coarse tracking; keyframes, which anchor long-term mapping consistency, undergo the full Corrective Blurry KNN (CB-KNN) regularization at the CUDA rasterization level. The system adaptively selects the neighborhood size based on camera motion and local Gaussian density, transiently shifts 2D projected centers toward local cluster centroids, and reweights RGB values according to spatial influence. This regularized forward pass produces smoothed color, depth, and silhouette images that guide pose refinement, noise-resistant co-visibility keyframe selection, and floater-free map densification.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Input: Synchronized RGB-D Frame Stream"] --> Split{"Keyframe vs Non-Keyframe Dispatch"}
    Split -->|Non-Keyframe| FastTrack["Constant-Velocity Motion Prior<br/>Fast Tracking on Canonical Gaussian Map"]
    Split -->|Keyframe Stage| KernelAdapt["Dynamic Kernel Size Adaptation<br/>Adjust K via Motion Magnitude & Spatial Density"]
    KernelAdapt --> DualSmooth["Dual-Dimensional Local Kernel Smoothing<br/>2D Projected Centroid Pull & Photometric Reweighting"]
    DualSmooth --> RegularizedRend["Regularized Differentiable Rasterization<br/>Synthesize Smoothed Color/Depth/Silhouette Modalities"]
    RegularizedRend --> OptMap["Regularization-Driven Map Management & Dual Optimization<br/>Robust Mask Loss Tracking & Noise-Resistant Densification"]
    FastTrack --> Out["Output: Accurate Camera Trajectory & Compact 3DGS Map"]
    OptMap --> Out

Key Designs

1. Dynamic Kernel Size Adaptation: balancing smoothing strength and geometric detail via motion and density

A static smoothing kernel cannot generalize across varying operational conditions: an excessively large neighborhood over-smooths intricate geometric edges during stationary observations, whereas an insufficient neighborhood fails to bridge gaps during violent camera shakes or within sparse boundary areas. To resolve this trade-off, the framework dynamically calculates the neighborhood size \(K\) by coupling the baseline scale with the inter-frame motion magnitude \(\gamma \in [0, 1]\) and the local Gaussian density \(\rho\) evaluated across an \(8 \times 8\) pixel grid:

\[K = K_0 \cdot \max\left(1, \beta \cdot \frac{\gamma}{\rho + \epsilon}\right)\]

where the baseline neighborhood \(K_0\) is configured to 5 for clean synthetic sequences and 8 for noisy real-world benchmarks, \(\beta=0.3\) is the scaling factor, and \(\epsilon=10^{-6}\) ensures numerical stability. When violent camera rotations or sparse geometric regions are encountered, the system automatically expands \(K\) to enforce stronger regularization, preventing gradient explosion caused by inaccurate initial poses; conversely, in dense, slow-moving views, \(K\) gracefully shrinks back to \(K_0\) to preserve fine details.

2. Dual-Dimensional Local Kernel Smoothing: rectifying spatial centroids and color consensus to eliminate local minima

For each pixel \(p=(u,v)\), the CUDA rasterizer identifies its \(K\) nearest contributing Gaussians \(\mathcal{G}_p = \{g_{p1}, \dots, g_{pK}\}\). Permanently modifying global Gaussian parameters would corrupt the canonical map; therefore, CB-KNN executes corrections exclusively on transient parameters during the rasterization pass, tackling both spatial geometry and color modalities:

In the spatial domain, the 3D Gaussian centers \(\mu_{pj}\) are projected onto the image plane via camera extrinsics \(E_t\) and intrinsics \(K_{cam}\) to compute the local 2D centroid \(C_p = \frac{1}{K} \sum_{g_{pj} \in \mathcal{G}_p} \pi(\mu_{pj}, E_t, K_{cam})\). Each primitive's projected center is then pulled slightly toward this centroid:

\[\pi(\mu'_{pk}) = \pi(\mu_{pk}) + \alpha \cdot \frac{C_p - \pi(\mu_{pk})}{\|C_p - \pi(\mu_{pk})\| + \epsilon}\]

with the displacement strength \(\alpha \in [0.1, 0.3]\). This centripetal shift enhances 2D spatial overlap between adjacent Gaussians, bridging artifactual tears and depth discontinuities. Concurrently, in the color domain, normalized contribution weights \(\omega_{pj} = \frac{f_{pj}(p)}{\sum_{g_{pi} \in \mathcal{G}_p} f_{pi}(p)}\) are assigned according to the radial opacity attenuation \(f_{pj}(p)\), defining the smoothed color as \(c'_{pk} = \sum_{g_{pj} \in \mathcal{G}_p} \omega_{pj} c_{pj}\). Using this transient set \(\mathcal{G}'_p\), the renderer synthesizes smoothed color \(C(p)\), expected depth \(D(p)\), and cumulative silhouette \(S(p)\) via standard alpha compositing.

3. Regularization-Driven Map Management & Dual Optimization: robust masking and floater-free densification

The smoothed image modalities directly govern gradient optimization and map growth. For camera tracking, the synthesized silhouette map \(S(p)\) provides a clean confidence mask, confining the optimization strictly to well-covered pixels (\(S(p) > 0.99\)):

\[\mathcal{L}_{t} = \sum_{p: S(p) > 0.99} \left( \|D(p) - D_{GT}(p)\|_1 + 0.6 \cdot \|C(p) - C_{GT}(p)\|_1 \right)\]

Because transient smoothing filters out isolated parameter spikes, the resulting loss landscape exhibits broader basins of attraction, allowing gradient descent to navigate stably toward the true camera pose. For map management, conventional systems often generate spurious "floater" Gaussians triggered by transient rendering errors. This issue is eliminated by constructing a binary densification mask guided by the smoothed silhouette and depth:

\[M(p) = \max\left(\mathbb{I}(S(p) < 0.5), \; \mathbb{I}(D_{GT}(p) < D(p)) \cdot \mathbb{I}\left(|D_{GT}(p) - D(p)| > \lambda \cdot \text{MDE}\right)\right)\]

where \(\lambda = 50\) scales the Median Depth Error (MDE). New Gaussians are spawned strictly in genuinely unmapped or newly occluded areas, completely suppressing floater proliferation while maintaining a compact, high-fidelity canonical map.

Loss & Training

The camera pose is optimized by minimizing the masked robust loss \(\mathcal{L}_{t}\). During keyframe mapping, the canonical parameters \(\mathcal{G}\) are jointly optimized over the current keyframe and top co-visible historical keyframes identified via back-projected regularized depth maps. Pruning is periodically applied to discard Gaussians with low opacity (\(\sigma_i < 0.05\)) or excessive spatial footprints. All optimization steps are driven end-to-end by analytic gradients derived from the regularized rasterization pipeline without requiring pre-trained neural networks.

Key Experimental Results

Main Results

The method was comprehensively benchmarked across the synthetic Replica dataset (8 scenes) and challenging real-world datasets TUM-RGBD (5 sequences) and ScanNet (6 sequences), evaluated primarily on Absolute Trajectory Error (ATE RMSE in cm).

Dataset Sequence / Summary SplaTAM [22] GS-SLAM [47] NICE-SLAM [59] Ours Gain (vs SplaTAM)
Replica (ATE RMSE โ†“ cm) Room0 0.31 0.48 0.97 0.25 -19.4%
Replica Room1 0.40 0.53 1.31 0.23 -42.5%
Replica Room2 0.29 0.33 1.07 0.25 -13.8%
Replica Office0 0.47 0.52 0.88 0.32 -31.9%
Replica Office1 0.27 0.41 1.00 0.21 -22.2%
Replica Office2 0.29 0.59 1.06 0.27 -6.9%
Replica Office3 0.32 0.46 1.10 0.29 -9.4%
Replica Office4 0.55 0.70 1.13 0.45 -18.2%
Replica Average 8 scenes mean 0.36 0.50 1.07 0.28 -22.2% error reduction
TUM-RGBD (ATE RMSE โ†“ cm) fr1/desk 3.35 3.32 4.26 1.88 -43.9%
TUM-RGBD fr1/desk2 6.56 โ€” 4.99 4.48 -31.7%
TUM-RGBD fr1/room 11.76 โ€” 34.49 7.94 -32.5%
TUM-RGBD fr2/xyz 1.36 1.36 31.73 1.25 -8.1%
TUM-RGBD fr3/office 5.16 6.62 3.87 2.78 -46.1%
TUM-RGBD Average 5 sequences mean 5.63 โ€” 15.87 3.67 -34.8% error reduction
ScanNet Average 6 scenes mean 11.88 โ€” 10.70 8.46 -28.8% error reduction

In terms of mapping and rendering quality, evaluation on Replica confirms that the method enhances average PSNR from 34.11 dB (SplaTAM) to 34.31 dB, achieves SSIM of 0.97, and maintains LPIPS at 0.09. This validates that transient kernel smoothing does not degrade scene fidelity, but instead cleans up peripheral floaters.

Ablation Study

The individual impacts of Position Smoothing and Color Smoothing were isolated on TUM-RGBD fr1/desk and ScanNet 0169.

Sequence Position Smoothing Color Smoothing Depth L1 (cm) โ†“ ATE RMSE (cm) โ†“ PSNR (dB) โ†‘ Note
TUM-RGBD fr1/desk โœ— โœ— 2.93 3.36 21.85 Baseline without smoothing
TUM-RGBD fr1/desk โœ— โœ“ 2.59 2.65 23.52 Color only: significant PSNR boost
TUM-RGBD fr1/desk โœ“ โœ— 2.30 2.28 22.61 Position only: primary driver for tracking
TUM-RGBD fr1/desk โœ“ โœ“ 2.03 1.86 24.41 Full CB-KNN: optimal tracking and mapping
ScanNet 0169 โœ— โœ— 6.55 12.13 18.79 Baseline without smoothing
ScanNet 0169 โœ— โœ“ 6.20 11.03 19.86 Color only
ScanNet 0169 โœ“ โœ— 5.91 10.12 19.23 Position only
ScanNet 0169 โœ“ โœ“ 5.83 9.82 20.32 Full CB-KNN: error compressed under 10 cm

Furthermore, generalizability testing on MonoGS demonstrates an average ATE reduction from 0.42 cm to 0.29 cm (-31.0%) on Replica and from 1.48 cm to 1.31 cm on TUM-RGBD. In runtime analysis on Replica Room0 (NVIDIA A40), although KNN search marginally raises keyframe processing time from 3.05 ms to 3.25 ms, the smoothed landscape cuts required tracking iterations by approximately 40%, shrinking per-frame tracking latency from 1.19 s to 0.94 s and driving overall system throughput from 0.49 Hz up to 1.41 Hz.

Key Findings

  • Position smoothing governs pose convergence: Spatial position shifting provides the dominant performance gain for tracking (ATE on fr1/desk drops from 3.36 cm to 2.28 cm), confirming that geometric tears and coordinate noise are the primary sources of sharp local minima.
  • Color smoothing stabilizes photometric alignment: Color reweighting yields the largest PSNR gain (+1.67 dB when added alone), smoothing out intensity spikes and providing reliable photometric gradients.
  • Iteration savings outpace computational overhead: While neighbor searching adds a negligible overhead in the keyframe forward pass, the well-conditioned convex basin accelerates convergence, delivering an overall 2.88x speedup in system throughput.

Highlights & Insights

  • Rethinking the "Clarity = Accuracy" dogma: Rather than equating visual sharpness with pose estimation accuracy, this work demonstrates that controlled blur acts as an effective regularizer against non-convex parameter noise in 3DGS.
  • Transient regularization leaves canonical maps intact: Confining the smoothing operation to the rasterization forward-backward pass guarantees robust optimization without corrupting the underlying explicit scene geometry.
  • Plug-and-play compatibility: Requiring no radical structural overhaul of the rasterizer, the lightweight CUDA KNN correction integrates seamlessly into diverse 3DGS frameworks including SplaTAM and MonoGS.

Limitations & Future Work

  • Author-admitted limitations: The current CB-KNN neighborhood search and centroid shifts operate in 2D projected screen space rather than 3D Euclidean space, which may introduce minor perspective artifacts across extreme depth discontinuities.
  • Empirical hyper-parameter dependence: The baseline neighborhood size \(K_0\) and motion scaling factor \(\beta\) are tuned empirically across datasets (e.g., \(K_0=5\) for clean synthetic scenes vs. \(K_0=8\) for noisy sensor setups), rather than being learned end-to-end.
  • Future directions: Integrating 3D anisotropic covariance smoothing and combining kernel regularization with event cameras or Hessian curvature-aware uncertainty modeling are promising avenues for future exploration.
  • vs SplaTAM [22]: SplaTAM relies on unrectified Gaussian rasterization, performing well in smooth, low-noise environments but drifting under aggressive motion; CB-KNN provides transient regularization, cutting ATE by 34.8% on TUM-RGBD while accelerating convergence.
  • vs GS-SLAM [47]: GS-SLAM skips unconstrained Gaussians through aggressive geometric thresholds, risking the loss of thin structures; this paper adopts adaptive continuous smoothing to retain structural continuity while eliminating gradient spikes.
  • vs ORB-SLAM3 [6]: Classical sparse visual SLAM applies robust loss functions against external measurement outliers, but cannot tackle noise originating from the neural scene representation itself; this paper bridges that gap directly within the rendering loop.

Rating

  • Novelty: โญโญโญโญโ˜† Challenges conventional fidelity-centric wisdom by introducing adaptive kernel smoothing via Graduated Non-Convexity.
  • Experimental Thoroughness: โญโญโญโญโญ Rigorous validation across Replica, TUM-RGBD, and ScanNet, supported by cross-framework tests, ablations, and runtime curves.
  • Writing Quality: โญโญโญโญโญ Well-structured narrative, disciplined mathematical formulation, and insightful empirical analysis.
  • Value: โญโญโญโญโญ A practical, elegant plug-and-play blueprint for deploying robust 3DGS-SLAM on mobile robotics under noisy real-world dynamics.