Don’t Starve the Boundaries: Boundary-Constrained Label Propagation for Weakly Supervised 3D Segmentation¶
Conference: ECCV 2026
Paper: ECCV Official Link
Code: https://github.com/paul-swu/Bound3D
Area: 3D Vision
Keywords: Weakly Supervised Learning / 3D Semantic Segmentation / Label Propagation / Boundary Constraints / Pseudo-labels
TL;DR¶
Addressing the "boundary starvation" issue where limited supervision fails to provide reliable boundary guidance in weakly supervised 3D point cloud segmentation, this paper introduces Bound3D—a framework that projects multimodal 2D geometric discontinuities to identify 3D boundaries, generates high-purity pseudo-labels via manifold-constrained propagation and farthest angle sampling, and optimizes interior and boundary regions through position-aware divide-and-conquer supervision.
Background & Motivation¶
Point cloud 3D semantic segmentation serves as a fundamental perceptual pillar for embodied AI, autonomous driving, and large-scale 3D scene understanding. Nevertheless, collecting dense per-point semantic annotations on large-scale point clouds requires tremendous human labor and expense, severely restricting the expansion of training corpora and impeding the development of 3D foundation models. To alleviate the heavy annotation burden of fully supervised paradigms, weakly supervised 3D semantic segmentation (WS3DSS) has emerged as a promising alternative, typically relying on extremely sparse supervision signals such as 0.01% random point annotations or merely a single labeled point per object. Existing dominant paradigms primarily resort to perturbation consistency regularization or self-training based pseudo-label generation. However, when deployed in complex indoor scenes, they consistently suffer from severe boundary distortion and semantic degradation—including boundary adhesion between adjacent clutter objects, semantic bleeding between doors and surrounding walls, and thin structural components being engulfed by dominant structural categories.
The root cause of these boundary failures is conventionally presumed to be an insufficiency of implicit boundary supervision. However, a minimal intervention experiment on standard consistency-based frameworks reveals a counterintuitive reality: when ground-truth boundary regions are explicitly identified and assigned a higher consistency loss weight, both the overall segmentation accuracy (mIoU) and the boundary-specific accuracy (B-mIoU) decline sharply rather than improve. Conversely, down-weighting the boundary consistency penalty yields superior boundary fidelity. This phenomenon indicates that indiscriminate consistency constraints blend semantic distinctions across interfaces, actively starving boundary regions of discriminative signals. Concurrently, conventional pseudo-label self-training approaches evaluate the network's own maximum softmax probabilities and impose a uniform confidence threshold (e.g., 0.90) to filter candidates. Because neural networks inherently produce lower confidence scores near ambiguous object interfaces even when predictions are correct, this uniform cutoff systematically favors interior points and filters out boundary labels, exacerbating boundary starvation. Forcing lower global thresholds introduces catastrophic confirmation bias, where erroneous boundary predictions are recirculated as seeds in iterative self-training rounds.
To resolve this dilemma, this work sidesteps both recursive reliance on the network's own noisy predictions and the costly multi-view 2D-3D back-projection of heavy foundation models such as SAM. Instead, it turns to native multimodal 2D-3D geometric discontinuities and local manifold smoothness. Core idea: decouple pseudo-label generation from the segmentation backbone, illuminate 3D boundaries via multi-view depth and normal discontinuities, perform boundary-constrained label propagation using manifold density verification and farthest angle sampling, and adaptively supervise interior and boundary points via a position-aware divide-and-conquer strategy.
Method¶
Overall Architecture¶
The Bound3D pipeline consists of three core components: the Boundary Illumination Module (BIM), the Boundary-constrained Pseudo Label Generator (BPLG), and the Position-Aware Divide-and-Conquer (PADC) supervision scheme. Given an unannotated point cloud accompanied by sparse seed annotations and calibrated multi-view RGB-D images alongside surface normal maps, BIM first extracts 2D geometric edges by filtering Canny gradients with depth jumps and normal orientation discontinuities, subsequently back-projecting them into 3D space to construct a coarse pseudo-boundary set. Next, BPLG takes the sparse seeds and propagates labels along verified manifold surfaces that do not cross boundary interfaces, applying Farthest Angle Sampling to maximize directional coverage. A Boundary Arrival Rate (BAR) metric dynamically halts iterations when boundary coverage saturates, while restricting new seeds exclusively to safe interior core regions to eliminate error accumulation. Finally, PADC routes pseudo-labels into an asymmetric dual-branch consistency network: a weak perturbation branch with a stop-gradient auxiliary head calibrates core interior representations, while a strong perturbation branch focuses on boundary points with distance-decaying loss weights. The overall data and control flow is illustrated below:
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Data<br/>Sparse Point Cloud + Multi-view RGB-D/Normals"] --> B["Boundary Illumination Module (BIM)<br/>Depth-filtered Canny + Normal Discontinuity"]
B --> C["Boundary-Constrained Label Propagation (BCLP)<br/>Manifold Density Check + Farthest Angle Sampling"]
C --> D["Boundary Arrival Rate Iterator (BAR)<br/>Dynamic Triggering + Core-zone Safe Seeds"]
D --> E["Position-Aware Divide-and-Conquer (PADC)<br/>Weak View Core Calibration + Strong View Distance-weighted CE"]
E --> F["Output Prediction<br/>Inference on Raw 3D Point Cloud Only"]
Key Designs¶
1. Boundary Illumination Module (BIM): Multimodal Geometric Cues Inducing High-Fidelity 3D Boundaries Detecting boundaries directly from sparse 3D point clouds is inherently ill-posed due to irregular spatial discretization and absence of explicit semantic contours. BIM circumvents this by harvesting native geometric boundaries from calibrated 2D observations. For each view \(j\), initial edges \(E_j\) are derived from RGB image \(I_j\) using the standard Canny operator. To suppress false edges induced by intra-object surface textures, a logarithmic depth gradient mask \(M_j^{\text{dep}}\) is constructed: $\(M_j^{\text{dep}} = \mathbf{1}\left[\sqrt{(\partial_u \log D_j)^2 + (\partial_v \log D_j)^2} > \tau_d\right] \cdot \text{valid}(D_j)\)$ yielding depth-filtered edges \(E'_j = E_j \cap M_j^{\text{dep}}\). Because depth measurements degrade at grazing angles and extended ranges, BIM complements depth edges with surface normal orientation discontinuities across a local radius \(r_n\): $\(M_j^{\text{norm}} = \mathbf{1}\left[\min_{u' \in \mathcal{N}_{r_n}(u)} \left| n_j(u)^\top n_j(u') \right| < \cos \theta \right]\)$ The unified 2D boundary pixels \(B_j = E'_j \cup M_j^{\text{norm}}\) are back-projected into 3D coordinate space via camera intrinsic \(K_j\), extrinsic matrix \(\mathcal{E}_j\), and depth map \(D_j\), assembling the scene-wide pseudo-boundary point set \(\Phi\) without invoking heavy foundation models.
2. Boundary-Constrained Label Propagation & Farthest Angle Sampling (BCLP & FAS): Manifold Proxy and Isotropic Expansion Physical objects in 3D point clouds are composed of locally smooth 2D surface manifolds embedded in \(\mathbb{R}^3\), where semantic categories remain constant across contiguous surfaces and exhibit abrupt transitions solely at physical boundaries \(\Phi\). BCLP enforces a "boundary-before-consistency" principle: candidate propagation paths connecting a labeled seed \(x_\ell\) to boundary points \(x_b \in \mathcal{B}(x_\ell)\) are evaluated along the Euclidean segment \(\text{Path}(x_\ell, x_b)\). To avoid intractable analytical manifold verification, BCLP introduces a local point cloud density proxy. Along the discretized path segment set \(\mathcal{G}\), every sampled point \(g\) must satisfy a minimal point count threshold \(\tau\) within radius \(r\): $\(\forall g \in \mathcal{G}: \quad |\mathcal{N}(g, r)| \ge \tau\)$ Paths meeting this density condition are certified to lie on the same continuous surface without jumping across spatial voids or bridging disparate objects. Unlabeled points within a search cylinder radius \(r_s\) along safe paths are harvested into set \(\mathcal{S}\) and assigned label \(y_{x_\ell}\). To bypass exhaustive path evaluation and mitigate spatial clustering, BPLG integrates Farthest Angle Sampling (FAS). FAS projects the concept of farthest point sampling into the directional angular domain, greedily selecting boundary endpoints that maximize minimal angular divergence relative to previously selected targets: $\(x_{k+1} = \arg\max_{x \in \mathcal{B}(x_\ell) \setminus S_k} \min_{x_m \in S_k} \theta\big(\mathbf{u}(x; x_\ell), \mathbf{u}(x_m; x_\ell)\big)\)$ This ensures that pseudo-labels expand outward isotropically and uniformly in all physical directions.
3. Boundary Arrival Rate Triggering & Dual-Threshold Iteration (BAR & BPLG): Anti-Bias Dynamic Pseudo-Labeling Conventional iterative self-training schemes operate on fixed epoch schedules and uniform confidence filtering, which leads to redundant pseudo-label accumulation in safe interiors while filtering out boundary labels. BPLG decomposes the point set into a boundary buffer \(\mathcal{R}_b\) (points within radius \(r_b\) of \(\Phi\)) and a core interior region \(\mathcal{R}_c\). The Boundary Arrival Rate (BAR) coverage metric measures boundary expansion: $\(\mathcal{C}_{\text{bar}} = \frac{\sum_{x \in \mathcal{P}^{(t)}} \mathbf{1}[x \in \mathcal{R}_b]}{|\mathcal{R}_b|}\)$ A new BCLP iteration is triggered only when the coverage increment satisfies \(\Delta \mathcal{C}_{\text{bar}} = \mathcal{C}_{\text{bar}}^{(t)} - \mathcal{C}_{\text{bar}}^{(t-1)} > \eta\). Stagnant boundary coverage halts iterations, preventing uninformative exploration of interior regions. To address confidence bias, PADC establishes region-specific thresholds: a relaxed threshold \(\tau_b\) for boundary points in \(\mathcal{R}_b\) where geometric constraints already guarantee structural fidelity, and a strict threshold \(\tau_c\) for core points (\(\tau_b < \tau_c\)). Furthermore, to eliminate catastrophic confirmation bias and error cascading, new seeds for subsequent iterations are selected exclusively from the high-confidence subset of core region \(\mathcal{R}_c\), strictly excluding boundary region \(\mathcal{R}_b\) from the seed bank.
4. Position-Aware Divide-and-Conquer (PADC): Weak Core Calibration and Strong Distance-Weighted Boundary Supervision Treating all points equally under uniform loss formulations induces severe feature smoothing at class transitions. PADC establishes an asymmetric dual-branch architecture. In the weak augmentation branch, a lightweight auxiliary head equipped with a stop-gradient operation fits high-confidence core pseudo-labels in \(\mathcal{R}_c\): $\(\mathcal{L}_{\text{w}} = \frac{1}{|\mathcal{R}_c|} \sum_{x \in \mathcal{R}_c} \text{CE}(O_{a_x}, \tilde{y}_x)\)$ serving as a clean, stable semantic anchor without letting noisy gradients distort the backbone representations. In the strong perturbation branch, the model confronts challenging boundary points in \(\mathcal{R}_b\). Because ambiguity scales with proximity to geometric transitions, each point receives a distance-aware weight based on its Euclidean distance \(d_x\) to the closest boundary in \(\Phi\): $\(w_x = 1 + \lambda_d \left( \frac{d_x}{r_d} \right)\)$ supervising the strong branch via weighted cross-entropy: $\(\mathcal{L}_{\text{s}} = \frac{\sum_{x \in \mathcal{R}_b} w_x \cdot \text{CE}(O_{s_x}, \tilde{y}_x)}{\sum_{x \in \mathcal{R}_b} w_x}\)$ To prevent the backbone from developing an overly narrow receptive field focused solely on edges, a supplementary loss \(\mathcal{L}_{\text{sup}}\) with a minor weight \(\gamma\) leaks a fraction of core supervision to the strong branch. Combined with ground-truth supervision \(\mathcal{L}_{\text{gt}}\) on labeled seeds and a Jensen-Shannon divergence consistency loss \(\mathcal{L}_{\text{consis}}\) across unlabeled points, the overall training objective is formulated as: $\(\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{gt}} + \mathcal{L}_{\text{w}} + \lambda_1 \mathcal{L}_{\text{s}} + \lambda_2 \mathcal{L}_{\text{consis}} + \gamma \mathcal{L}_{\text{sup}}\)$ thereby driving high semantic sharpness precisely at geometric boundaries while preserving interior cluster coherence.
Key Experimental Results¶
Main Results¶
Bound3D underwent extensive evaluation on S3DIS Area 5 and ScanNetV2 under standard sparse annotation protocols. The main results on S3DIS under both percentage-based and instance-based budgets are summarized below:
| Dataset / Budget | Method | Backbone | Venue & Year | mIoU (%) | B-mIoU (%) |
|---|---|---|---|---|---|
| S3DIS (0.01% GT) | SQN | SparseConv | ECCV 2022 | 45.3 | - |
| S3DIS (0.01% GT) | CPCM | MinkowskiNet | ICCV 2023 | 59.8 | - |
| S3DIS (0.01% GT) | PointMatch | PointNet++ | CG 2023 | 59.9 | - |
| S3DIS (0.01% GT) | AADNet | SparseConv | AAAI 2025 | 60.8 | - |
| S3DIS (0.01% GT) | Bound3D (Ours) | MinkowskiNet | ECCV 2026 | 61.2 | 54.5 |
| S3DIS (1 pt/obj) | OTOC (0.02% budget) | MinkowskiNet | CVPR 2021 | 50.1 | - |
| S3DIS (1 pt/obj) | DAT (0.02% budget) | MinkowskiNet | ECCV 2022 | 56.5 | - |
| S3DIS (1 pt/obj) | RAC-Net (0.02% budget) | MinkowskiNet | IJCV 2024 | 58.4 | - |
| S3DIS (1 pt/obj) | AO-PTv2 | Point Transformer v2 | CVPR 2024 | 62.7 | - |
| S3DIS (1 pt/obj) | Bound3D (Ours) | MinkowskiNet | ECCV 2026 | 61.9 | 55.1 |
| S3DIS (1 pt/obj) | Bound3D-PTv2 (Ours) | Point Transformer v2 | ECCV 2026 | 63.5 | 56.8 |
On the ScanNetV2 benchmark, Bound3D demonstrates consistent superiority: under 0.1% labels, Bound3D achieves 63.2% mIoU (+6.3% over SQN); under the 20 points/scene setup, it achieves 62.4% mIoU, outperforming MIL (+8.0%), DAT (+7.2%), and OTOC (+3.0%).
Ablation Study¶
Ablation experiments on S3DIS Area 5 (0.01% supervision) demonstrate the progressive contribution of each architectural component, alongside component-level leave-one-out evaluations:
| Stage / Component | Configuration Details | mIoU (%) | B-mIoU (%) | Note / Relative Impact |
|---|---|---|---|---|
| Pure Sparse Baseline | MinkowskiNet trained on sparse GT only | 49.3 | 44.4 | Lacks mechanism to leverage unlabeled points |
| Consistency Baseline | Baseline (perturbation consistency) | 55.8 | 49.0 | Effective unlabeled point utilization (+6.5% mIoU) |
| Single-round Propagation | Baseline + BCLP1 | 56.7 | 49.4 | Safe manifold label expansion (+0.9% mIoU) |
| Iterative Generator | Baseline + BCLPn (Full BPLG) | 60.7 | 52.9 | Dynamic BAR and dual thresholds (+4.0% mIoU) |
| Full Model | Bound3D (Baseline + BCLPn + PADC) | 61.2 | 54.5 | Divide-and-conquer supervision (+0.5% mIoU / +1.6% B-mIoU) |
| Leave-one-out Analysis | w/o FAS (replaced by random sampling) | 59.3 | 51.6 | Angular clustering restricts spatial diversity (-1.9% mIoU) |
| Leave-one-out Analysis | w/o Distinct Thresholds (uniform threshold) | 58.9 | 51.2 | Boundary confidence drop triggers boundary starvation (-2.3% mIoU) |
| Leave-one-out Analysis | w/o Distance Weighting (uniform boundary loss) | 60.0 | 52.8 | Equal penalty fails to sharpen near-boundary gradients (-1.2% mIoU) |
| Gradient Isolation Test | w/o stop-gradient in AuxHead | 60.3 | 53.0 | Core representation corrupted by boundary noise (-0.9% mIoU) |
Key Findings¶
- High Label Purity Outperforming SAM-based Paradigms: Under S3DIS 1 pt/obj, the initial offline propagation BCLP1 produces 93,565 pseudo-labels per scene with 94.4% overall accuracy and 92.7% boundary accuracy (B-Acc). In comparison, the SAM-projection method PP2S (CVPR 2024) generates 278,134 labels but attains only 77.1% accuracy (66.1% at boundaries). On ScanNetV2, BCLP1 yields 93.3% accuracy, surpassing SAM3D (81.8%).
- Balanced Performance Across Categories: Detailed per-class accuracy confirms high label purity across diverse structures: Floor (99.18%), Beam (96.13%), Ceiling (95.22%), Sofa (95.19%), Wall (93.15%), Table (92.02%), Chair (91.10%), Window (91.03%), Door (88.53%), and Board (82.71%). Only small, thin, and weakly observed objects such as Clutter (75.34%) show reduced precision.
- Criticality of Region-Specific Thresholds and Gradient Isolation: Ablation indicates that removing region-specific thresholds incurs a 2.3% drop in mIoU and a 3.3% drop in B-mIoU, verifying that uniform filtering is the primary cause of boundary starvation. Furthermore, omitting the stop-gradient operation in PADC causes boundary ambiguity to contaminate the backbone, resulting in an immediate 0.9% mIoU decline.
Highlights & Insights¶
- Geometric Discontinuities Over Heavy Vision Foundation Models: Rather than inheriting cross-modal projection errors and high inference latency from large foundation models (SAM/CLIP), Bound3D shows that classical Canny filtering combined with depth and surface normal orientation discontinuities reliably exposes true 3D geometric boundaries.
- Manifold-Constrained Propagation with Farthest Angle Sampling: Formulating point cloud surface patches as 2D manifolds and validating propagation segments via discrete local density checks prevents pseudo-labels from leaking across spatial boundaries or jumping between disparate objects. Farthest Angle Sampling guarantees isotropic coverage.
- Solving the Consistency Dilemma via Position-Aware Supervision: By unveiling the counterintuitive paradox where excessive boundary consistency degrades accuracy, PADC successfully decouples core semantic stabilization from boundary discrimination via stop-gradient auxiliary heads and distance-dependent weighting.
Limitations & Future Work¶
- Dependency on Calibrated Multi-view RGB-D and Normal Maps: The boundary illumination stage requires synchronized camera intrinsics/extrinsics, valid depth, and surface normal estimations. Performance may degrade in sparse-view outdoor LiDAR environments or low-texture lighting conditions.
- Suboptimal Resolution on Thin and Fine Objects: For complex clutter and thin geometries, 2D projections and discrete depth gradients may fail to resolve fine boundaries, explaining the lower accuracy on the clutter class (75.34%). Future extensions could incorporate continuous implicit distance fields or signed distance functions.
Related Work & Insights¶
- vs SAM-based Projection Methods (PP2S, SAM3D): SAM-based methods project dense 2D masks into 3D, introducing significant boundary occlusion noise and requiring heavy 2D foundation model inference. Bound3D uses lightweight geometric cues and manifold density checks, yielding higher pseudo-label purity (94.4% vs 77.1%) with much lower computational cost.
- vs Perturbation Consistency Frameworks (CPCM, DAT): Prior consistency models apply uniform perturbation losses across all points, which over-smooths semantic boundaries. Bound3D introduces PADC, separating core feature calibration from distance-weighted boundary sharpening.
- vs Boundary-Guided Segmentation (BAGE, WS3D): Fully supervised BAGE requires ground-truth boundary masks, while WS3D trains an extra boundary prediction network. Bound3D directly extracts boundaries from multi-view geometry without extra network parameters or dense supervision.
Rating¶
- Novelty: ⭐⭐⭐⭐ [Identifies boundary starvation in weakly supervised 3D segmentation and proposes a geometric manifold propagation and divide-and-conquer training framework]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across multiple datasets, baselines, per-class metrics, and component ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation driven by counterintuitive empirical observations, rigorous mathematical formulations, and structured presentation]
- Value: ⭐⭐⭐⭐ [Provides an effective, foundation-model-free blueprint for high-precision weakly supervised 3D scene understanding]