Skip to content

Uncertainty-Driven Gaussian Sphere Propagation for 3D Semantic Segmentation

Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/LENGYI1221/UGSP
Area: Autonomous Driving
Keywords: 3D Point Cloud Segmentation, Uncertainty Modeling, Spherical Harmonics, Anisotropic Propagation, Feature Refinement

TL;DR

To address the breakdown of deterministic predictions and isotropic feature aggregation across physical occlusions and sparse boundaries in 3D point clouds, this paper presents Uncertainty-driven Gaussian Sphere Propagation (UGSP), which constructs continuous topological paths between Monte Carlo Dropout-partitioned uncertainty regions and injects reliable anchor semantics into geometric voids via Spherical Harmonics-based anisotropic Gaussian spheres.

Background & Motivation

Semantic segmentation of 3D point clouds serves as the foundational perception module for autonomous driving and mobile robotics. In recent years, Transformer architectures—most notably Point Transformer v3 (PTv3)—have achieved remarkable state-of-the-art performance by capturing extensive long-range contextual dependencies across large receptive fields. Nevertheless, existing methods almost universally operate under a deterministic point-wise prediction paradigm, implicitly assuming that spatial features across the entire domain possess uniform reliability during inference. This assumption frequently breaks down in physical environments subjected to sensor beam divergence, sparse distant sampling, inter-object occlusions, and intricate boundary contours, creating "geometric voids" where discrete points fail to provide continuous geometric support.

Crucially, segmentation errors in point cloud networks are not randomly scattered; rather, they correlate strongly with predictive uncertainty in geometrically complex regions. High uncertainty precisely delineates geometric voids such as object boundaries, distant sparse regions, and occluded surfaces. Conversely, low-uncertainty regions with dense, regular point distributions provide highly reliable semantic cues. Conventional discrete aggregation mechanisms (such as 3D convolutions, local k-NN graph aggregation, or isotropic self-attention) treat points as isolated feature carriers and gather features with isotropic weights that decay uniformly with Euclidean distance. As a consequence, they lack both the geometric continuity required to bridge discrete physical gaps and the directional awareness necessary to selectively guide valid semantic context from certain anchors into ambiguous areas.

To overcome these structural limitations, this paper pivots from conventional discrete point processing to continuous geometric field reconstruction. Core idea: explicitly treat predictive uncertainty as a structural prior, decouple reliable anchors from high-uncertainty voids via region growing to construct continuous topological paths, sample anisotropic Gaussian spheres along these paths, and employ Spherical Harmonics to capture direction-aware geometric distributions in the frequency domain, steering reliable semantic context into geometric voids via cross-attention.

Method

Overall Architecture

The UGSP framework accepts an input point cloud with 3D coordinates and color information, aiming to construct a continuous geometric field that bridges semantic discontinuities in ambiguous regions. The entire pipeline proceeds through three successive stages: first, feature extraction via a backbone network paired with Monte Carlo Dropout quantifies predictive uncertainty, followed by hierarchical region growing to partition the scene into reliable low-uncertainty anchors and ambiguous high-uncertainty voids; second, topological pathways connecting region centroids are established and intermediate aggregation centers are sampled under adaptive density constraints; third, within each Gaussian sphere, Spherical Harmonics feature aggregation captures directional frequency components, and after path trajectory encoding and Transformer sequence modeling, cross-attention channels reliable contextual features back into high-uncertainty regions to be concatenated with original point features for final classification.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    IN["Input Point Cloud P & PTv3 Backbone"] --> URE["Uncertainty Evaluation & Region Growing<br/>MC-Dropout Estimation + Voxel Hierarchical Merging"]
    URE --> PC["Topological Path Construction<br/>Centroid Directed Linking + Density-Adaptive Sphere Sampling"]
    PC --> SH["Spherical Harmonics Anisotropic Aggregation<br/>Directional Vectors to Spherical Basis + Gaussian Radial Weighting"]
    SH --> REF["Trajectory Encoding & Cross-Attention Refinement<br/>Path Transformer + Region-Level Context Injection"]
    REF --> OUT["Point-Wise Feature Concatenation & Segmentation Head"]

Key Designs

1. Uncertainty Evaluation & Region Growing: Decoupling Reliable Anchors from Geometric Voids Conventional models treat feature reliability uniformly across space, ignoring local prediction confidence. UGSP activates Dropout layers during inference and performs \(K\) stochastic forward passes (Monte Carlo Dropout) to generate class probability distributions \(\{\hat{y}_i^{(k)}\}_{k=1}^K\) for each point. The point-wise uncertainty \(u_i\) is computed as the standard deviation across samples relative to the predictive mean \(\bar{y}_i\):

\[u_i = \sqrt{\frac{1}{K}\sum_{k=1}^K (\hat{y}_i^{(k)} - \bar{y}_i)^2}\]

To isolate continuous structural ambiguity rather than isolated noisy points, the scene is partitioned into a regular voxel grid \(g\), and each voxel receives an average uncertainty value \(u_v\). The top 40% voxels with the highest uncertainty are extracted, from which seed voxels are selected via Farthest Point Sampling (FPS). Agglomerative region merging is then performed based on a joint distance-uncertainty affinity cost between adjacent regions \(R_a\) and \(R_b\):

\[\mathcal{C}(R_a, R_b) = \min_{v_i \in R_a, v_j \in R_b} d(v_i, v_j) \cdot \left(1 + \alpha \cdot |\bar{u}_{R_a} - \bar{u}_{R_b}|\right)\]

where \(\alpha\) balances spatial proximity and uncertainty homogeneity. Merging proceeds iteratively until reaching the target granularity \(N_H\), forming high-uncertainty void regions \(\mathcal{R}_h\). The bottom 40% voxels generate low-uncertainty anchor regions \(\mathcal{R}_l\) through an identical procedure, while intermediate points are assigned to nearest regions via KNN, establishing a complete structural decomposition.

2. Topological Path Construction: Erecting Continuous Geometric Bridges across Discrete Voids To prevent feature aggregation from failing when physical points are sparse or absent in geometric voids, this module establishes structured pathways linking reliable anchors to high-uncertainty voids. For each high-uncertainty region \(R_h\) and its centroid \(C_{R_h} = \frac{1}{|R_h|}\sum_{p_i \in R_h} p_i\), the system identifies its \(m\) nearest low-uncertainty anchor centroids \(C_{R_l}\) based on Euclidean distance to form directed geometric pathways.

Along each path, \(N_s\) uniformly spaced intermediate points are sampled as aggregation centers \(p_s\). To ensure these centers do not fall into completely empty voids devoid of local features, a density-aware adaptive radius constraint is enforced: starting from a base radius \(r\), the spherical neighborhood expands iteratively whenever voxel occupancy falls below a threshold \(T_{\min}\). This mechanism guarantees that every sampled Gaussian sphere retains sufficient physical support even in sparse or occluded regions, preserving the geometric continuity of the contextual bridge.

3. Spherical Harmonics Anisotropic Aggregation: Encoding Directional Geometry in the Frequency Domain Traditional feature aggregation methods apply isotropic convolutional kernels or pooling operations (such as max or average pooling), collapsing directional spatial cues like surface normals and boundary orientations. For each aggregation center \(p_s\), the relative displacement of neighborhood points \(x_i\) is normalized and mapped to spherical coordinates \((\theta_i, \phi_i)\). These directions are projected onto spherical harmonic basis functions \(Y_l^m(\theta_i, \phi_i) \in \mathbb{R}^{(L+1)^2}\) up to degree \(L\) (\(l=0,\dots,L\), \(m=-l,\dots,l\)). Point features \(f_i\) are then lifted into the spherical harmonic frequency domain:

\[F_i^{(l,m)} = f_i \cdot Y_l^m(\theta_i, \phi_i)\]

Simultaneously, a lightweight neural network \(\psi(r_i)\) computes radial distance-dependent Gaussian weights \(w_i = \psi(r_i)\) from Euclidean distance \(r_i = \|x_i - p_s\|\). The aggregated feature representation at center \(p_s\) for degree-order \((l, m)\) is obtained as:

\[F_{SH} = \sum_{i} w_i \cdot f_i \cdot Y_l^m(\theta_i, \phi_i)\]

This formulation allows each Gaussian sphere to capture not only distance-based proximity but also fine-grained, directionally sensitive geometric distributions across multiple semantic categories.

4. Trajectory Encoding & Cross-Attention Refinement: Directional Context Injection with Point-Wise Detail Preservation To coordinate contextual information along the sequence of Gaussian spheres, the relative scalar displacement of the \(i\)-th sphere from the path origin \(C_{R_h}\) is defined as \(t_i = (p_s - C_{R_h}) \cdot \hat{d}_i\). An MLP projects \(t_i\) into a trajectory positional encoding \(PE = \text{MLP}(t_i)\), yielding the composite sphere token \(g_i = [F_{SH}; PE]\). A Transformer encoder processes the sequence \(\{g_0, \dots, g_{N_s}\}\) to capture macro-level geometric evolution along the entire path into sequence feature \(h\).

During semantic transfer, cross-attention guides reliable features from anchors into ambiguous voids: path representations \(h\) serve as Queries (\(Q\)), while low-uncertainty anchor features \(f_{R_l}\) serve as Keys (\(K\)) and Values (\(V\)), producing refined context feature \(\tilde{f}\). Crucially, to prevent coarse region-level updates from producing blocky boundary artifacts, the refined feature \(\tilde{f}\) is concatenated with the original point-wise feature \(f_i\) from the backbone, providing both global structural guidance and fine-grained local detail for the final shared classification head.

Loss & Training

The framework adopts Point Transformer v3 (PTv3) as the feature extraction backbone. During training, the backbone segmentation head and the final refined prediction head are jointly supervised using standard Cross-Entropy Loss in an end-to-end manner. At inference time, \(K=16\) Monte Carlo Dropout passes are executed concurrently via batch-level parallelization on a single GPU, providing stable uncertainty estimates while bypassing the latency bottlenecks of serial passes.

Key Experimental Results

Main Results

On standard indoor benchmarks (ScanNet V2 and S3DIS Area 5) and large-scale outdoor autonomous driving benchmarks (nuScenes, SemanticKITTI, and Waymo Open Dataset), UGSP consistently outperforms both deterministic baselines and prior uncertainty-aware methods.

Dataset Metric (mIoU %) Ours (UGSP) Baseline (PTv3) Prev. SOTA / Representative Gain (vs Baseline)
ScanNet V2 (val) mIoU 77.9 76.8 76.4 (Swin3D) +1.1
S3DIS (Area 5) mIoU 75.6 73.3 72.5 (Swin3D) +2.3
nuScenes (val) mIoU 82.2 80.2 78.9 (OA-CNNs / Swin3D) +2.0
SemanticKITTI (val) mIoU 73.8 72.3 72.1 (Swin3D) +1.5
Waymo Open (val) mIoU 72.1 71.3 73.5 (Swin3D) +0.8

Ablation Study

Ablation experiments on S3DIS Area 5 rigorously examine the specific contributions of uncertainty-driven region partitioning, Spherical Harmonics encoding, and path attention mechanisms:

Config Uncertainty Region SH Encoding Path Attention S3DIS Area 5 (mIoU %) Note
PTv3 Baseline — — — 73.3 Standard backbone
+ RAND-KNN ✗ (Random seeds) ✓ ✓ 72.8 Degrades performance without uncertainty guidance (-0.5)
+ UND-KNN ✗ (Seeds only) ✓ ✓ 74.1 Partial guidance; lacks uncertainty in region growing
Full Model (UGSP) ✓ ✓ ✓ 75.6 End-to-end uncertainty guidance (+2.3)
w/o SH (Max-Pooling) ✓ ✗ (Max-Pool) ✓ 72.8 Severe drop due to loss of directional geometry (-2.8)
w/o SH (Avg-Pooling) ✓ ✗ (Avg-Pool) ✓ 73.1 Conflates anisotropic features across angles (-2.5)
w/o Attention (w/o path-dir) ✓ ✓ ✗ (No path-dir) 73.3 Spherical harmonic coefficients lose spatial alignment
w/o Attention (Concatenate) ✓ ✓ ✗ (Naive Concat) 73.5 Insufficient capture of high-order inter-channel interactions

Computational efficiency and refinement alternative comparison on a single NVIDIA H100 GPU (S3DIS Area 5):

Method / Variant mIoU (%) Latency (ms) FPS GPU Mem. (GB) Key Insight & Attribution
PTv3 (Backbone) 73.3 105 9.5 5.8 Standard deterministic baseline
+ MLP Refine 74.4 121 8.3 11.4 Parameter-matched MLP yields only +1.1%, proving gain is not mere capacity expansion
+ TTA (16× Test-Time Aug) 73.7 2,194 0.5 8.8 Blind ensembling is computationally prohibitive and yields marginal gain
+ RAND-KNN Path 72.8 183 5.5 8.2 Identical cost but underperforms baseline, isolating causal role of uncertainty guidance
+ UGSP (Ours) 75.6 183 5.5 8.2 Parallelized MC inference delivers +2.3% mIoU at practical overhead

Hyperparameter sensitivity on S3DIS Area 5: - Spherical Harmonic Degree \(L\): \(L=2\) reaches 74.2% mIoU; \(L=3\) surges to 75.6%; \(L=4\) saturates at 75.64%. \(L=3\) is adopted as default. - Gaussian Spheres along Path \(N_s\): Sampling 3 spheres yields 74.2%; 5 spheres achieves the optimal 75.6%; increasing to 7 spheres slightly declines to 75.5% due to redundant feature mixing. - Monte Carlo Samples \(K\): \(K=1\) equals the 73.3% baseline; \(K=4\) reaches 74.2%; \(K=8\) reaches 74.7%; performance plateaus at \(K=16\) (75.6%) and \(K=20\) (75.6%).

Key Findings

  • Complementarity between Spherical Harmonics and Path Direction: Removing path-direction modeling causes the SH performance boost to vanish entirely (dropping to 73.3%). Spherical Harmonics provide angular frequency bases, but only an aligned geometric path provides the structural coordinate frame necessary to exploit them.
  • Causal Necessity of Uncertainty Guidance: RAND-KNN uses identical Gaussian sphere sampling and Transformer refinement yet falls below the baseline to 72.8%. Arbitrary cross-region pathways introduce noise; only directed propagation from low-uncertainty anchors to high-uncertainty voids drives accurate semantic recovery.
  • Practical Parallel Inference Trade-off: By executing 16 MC passes in parallel batches on a single GPU, UGSP incurs an incremental latency of only 78 ms, maintaining a practical 5.5 FPS on an H100 GPU.

Highlights & Insights

  • Native 3D Gaussian Anisotropic Modeling for Point Clouds: Unlike prior 3DGS semantic segmentation works that passively distill 2D features (from SAM or CLIP) onto 3D Gaussians, UGSP formulates anisotropic Gaussian spheres and Spherical Harmonics natively on 3D geometry without requiring external 2D foundation models.
  • Uncertainty as a Structural Scaffold: Rather than using Bayesian uncertainty solely as an output confidence score or auxiliary loss, this work translates uncertainty into a spatial roadmap that actively governs geometric pathway creation.
  • Dual Representation to Prevent Boundary Artifacts: Combining path-level macro context with point-level backbone features ensures that ambiguous voids acquire global structural consistency without sacrificing fine local boundary sharpness.

Limitations & Future Work

  • Inference Latency Overhead: Despite parallel batching, executing 16 MC Dropout forward passes increases latency by ~74% (105 ms to 183 ms), posing deployment challenges on resource-constrained automotive edge chips (such as NVIDIA DRIVE Orin). Future directions could explore single-pass evidential deep learning for uncertainty distillation.
  • Extrapolation Bounds in Extreme Scenarios: While adaptive sphere radii maintain continuity across sparse voids, large-scale physical blind spots that fall entirely outside sensor fields of view could lead to hallucinated interpolations.
  • Multimodal Sensor Expansion: The current framework relies solely on LiDAR and RGB-D point clouds; incorporating camera imagery into the geometric field reconstruction represents a promising avenue for multimodal autonomous driving perception.
  • vs Deterministic 3D Networks (PTv3, Swin3D): Standard deterministic architectures struggle at occluded boundaries due to isotropic aggregation. UGSP functions as a plug-and-play refinement module on top of PTv3, achieving +1.1% to +2.3% mIoU gains across indoor and outdoor benchmarks.
  • vs Uncertainty-Aware Methods (SalsaNet, SalsaNext): SalsaNext utilizes uncertainty primarily as an auxiliary confidence indicator or loss weighting term without explicitly guiding feature flow. UGSP innovates by transforming uncertainty into a topological spatial scaffold for feature propagation.
  • vs 3DGS Semantic Distillation (Feature3DGS, GaussianGrouping): Existing 3DGS semantic methods depend heavily on distilling 2D foundation models into dense radiance fields. UGSP introduces spherical harmonic Gaussian modeling natively into sparse 3D point cloud segmentation.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Elegant integration of Gaussian sphere modeling, Spherical Harmonics frequency bases, and Bayesian uncertainty topology for 3D point cloud void completion.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 5 indoor and outdoor benchmarks, in-depth ablation studies, and concrete H100 latency/throughput profiling against four alternatives.
  • Writing Quality: ⭐⭐⭐⭐⭐ Cohesive narrative, clear mathematical formulation, and well-justified design choices.
  • Value: ⭐⭐⭐⭐⭐ Offers both theoretical depth and practical insights for resolving long-distance sparsity and boundary ambiguities in autonomous driving perception.