Skip to content

NegROI: Click-Centric Uncertainty-Guided Refinement with Scene-Conditioned Negative Prompts for Robust Interactive 3D Segmentation

Conference: ECCV 2026
Paper: ECCV Official
Code: To be released
Area: Autonomous Driving / 3D Vision
Keywords: Interactive 3D Segmentation, Point Clouds, Transformers, Negative Prompts, Uncertainty-Guided Refinement, Cross-Dataset Robustness

TL;DR

NegROI addresses boundary under-segmentation and persistent hard false positives in interactive point cloud segmentation by introducing click-centric multi-resolution ROI refinement alongside scene-conditioned negative prototype learning, significantly boosting click efficiency and outdoor LiDAR transfer robustness.

Background & Motivation

Interactive 3D point cloud segmentation aims to extract accurate object instance masks through minimal user-provided click feedback, serving as an indispensable tool for drastically cutting dense point-level labeling costs in autonomous driving and embodied scene understanding. Existing state-of-the-art interactive pipelines (such as InterObject3D, AGILE3D, and Easy3D) predominantly discretize continuous point clouds into regular voxel grids for computational scalability, using two-way transformer decoders to establish cross-attention between scene tokens and interactive click embeddings. However, throughout real-world interactive annotation cycles, this prevailing paradigm suffers from two persistent structural bottlenecks: boundary under-segmentation caused by coarse voxel resolution blurring fine geometric silhouettes under tight click budgets, and hard false positives (FPs) induced by structurally and visually confusing background distractors that trigger high-confidence misclassifications.

These failure modes are markedly exacerbated under distribution shifts, especially when transferring models trained on dense indoor RGB-D scans (e.g., ScanNet) to large-scale, sparse outdoor autonomous driving LiDAR scenes (e.g., KITTI-360). Indoor scans feature planar boundaries and dense contact surfaces, whereas outdoor vehicle LiDAR datasets exhibit non-uniform point distributions, severe sensor noise, and wide-ranging background clutter. Prevailing pipelines rely on heuristic fixed refinement windows or purely corrective negative clicks, lacking an explicit mechanism to represent and suppress recurring background distractors, which forces annotators to spend excessive corrective clicks on recurring false-positive regions.

This paper tackles these challenges through an orthogonal yet synergistic perspective: while coarse global representations are necessary for contextual awareness, fine-grained computation should be allocated selectively to ambiguous regions around the active interaction, and background distractors must be explicitly anchored via learned negative prototypes. Core idea: NegROI couples scene-conditioned negative prompts and boundary-aware hard negative mining to actively suppress structural false positives, with click-centric uncertainty-gated fine-grid ROI refinement to sharpen ambiguous boundaries, achieving high boundary fidelity and robust cross-domain generalization under sparse click budgets.

Method

Overall Architecture

The interactive inference pipeline of NegROI is built around three tightly coordinated stages: coarse-scale global encoding, scene-conditioned negative prototype modeling, and fine-grid click-centric local refinement. Given an input point cloud and the accumulated user clicks up to the current interaction step, the system first discretizes the scene onto a coarse voxel grid and extracts scene tokens via a sparse backbone encoder. After injecting relative positional encodings anchored at the initial click, a set of learnable negative queries cross-attends to the scene tokens to distill scene-conditioned negative prototypes. The augmented interaction tokens and scene tokens are fed into a two-way transformer decoder to yield coarse voxel logits. Subsequently, centered at the current click, an adaptive radius is predicted based on local geometric density, and a binary uncertainty gate filters out confident regions, retaining only ambiguous points to form a Region of Interest (ROI). The ROI points are re-voxelized at a finer resolution, processed by the shared encoder and an ROI decoder, and the resulting fine logits are max-aggregated and fused back into the coarse prediction via a residual update.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Point Cloud & User Click Tokens"] --> B["Scene-Conditioned Negative Prompts & Diversity Regularization<br/>Cross-attention extracts K background prototypes with cosine penalty"]
    B --> C["Boundary-Aware Hard Negative Mining Supervision<br/>Mines boundary high-confidence false positives and supervises prompt attention"]
    C --> D["Two-Way Transformer Decoder<br/>Foreground/background token difference yields coarse voxel logits"]
    D --> E["Uncertainty-Driven Adaptive ROI Selection<br/>Density-adaptive radius with binary uncertainty gating"]
    E --> F["Fine-Grid ROI Refinement & Residual Fusion<br/>Fine-scale re-voxelization + shared branch decoding & Max-aggregation"]
    F --> G["Final Refined Segmentation Mask"]

Key Designs

1. Scene-Conditioned Negative Prompts & Diversity Regularization: Explicitly Decoupling and Modeling Background Distractors

Standard interactive 3D segmentation architectures rely solely on generic foreground/background mask tokens and click tokens within the decoder, making them vulnerable to structural distractors that share geometric similarities with the target. NegROI introduces \(K\) learnable negative prompt queries \(\{q_k\}_{k=1}^K\) that cross-attend to the coarse scene tokens and relative position encodings prior to decoding, dynamically synthesizing \(K\) scene-conditioned negative prototype tokens that capture the holistic background context of the specific scene:

\[\hat{p}_k^{(c)} = \text{MLP}\left(\text{Attn}\left(q_k, H + P, H + P\right)\right) + e(\text{neg})\]

To prevent these \(K\) prototypes from collapsing into redundant, degenerate representations during optimization, a prompt diversity regularizer \(\mathcal{L}_{div}^{(c)}\) is applied at each step by penalizing the squared off-diagonal cosine similarity across all prototype pairs:

\[\mathcal{L}_{div}^{(c)} = \frac{1}{K(K - 1)} \sum_{i \neq j} \cos\left(p_i^{(c)}, p_j^{(c)}\right)^2\]

This diversity penalty forces the negative prototypes to scatter across the representation manifold, enabling them to simultaneously anchor heterogeneous background modes such as planar walls, floors, adjoining unrelated objects, and distant sparse sensor artifacts.

2. Boundary-Aware Hard Negative Mining Supervision: Guiding Negative Attention to Ambiguous Boundary Distractors

Relying solely on standard task segmentation loss does not guarantee that the negative prototypes attend to the most detrimental background regions. To directly enforce spatial selectivity, NegROI introduces a boundary-aware hard negative supervision objective. First, a set of background candidate voxels directly adjacent to ground-truth object boundaries is extracted via 6-neighborhood connectivity: \(\mathcal{B} = \{j \mid y_j = 0, \exists j' \in \mathcal{N}(j) \text{ s.t. } y_{j'} = 1\}\). Within \(\mathcal{B}\), the top-\(k\) voxels exhibiting the highest predicted coarse foreground probabilities \(p_j = \sigma(s_j^{(c)})\) are designated as the hard negative set \(\mathcal{H}\). A cross-entropy loss is then minimized between the average negative-prompt attention weights \(\bar{A}^{(c)}\) and a uniform distribution over \(\mathcal{H}\):

\[\mathcal{L}_{hn}^{(c)} = - \sum_{j=1}^V t_j \log\left(\bar{A}_j^{(c)} + \epsilon\right), \quad t_j = \frac{1}{|\mathcal{H}|}\mathbb{I}[j \in \mathcal{H}]\]

This explicit supervision drives the negative prototypes to actively target boundary-proximal, high-confidence false-positive candidates, allowing the decoder to sharply depress background activations along target contours.

3. Uncertainty-Driven Adaptive ROI Selection: Allocating High-Resolution Compute on Demand

Dense high-resolution voxelization across entire scenes is computationally prohibitive and memory-intensive. NegROI sidesteps this bottleneck through click-centric local refinement. To handle non-uniform point distributions, the refinement radius \(r_c\) is adaptively modulated rather than kept static. Starting from a decayed base radius \(r_0 \gamma^{c-1}\), local point density statistics (mean \(\mu\) and median \(m\) of \(k\)NN distances) and click features are passed through an MLP to predict a scale adjustment \(\Delta_c\), yielding a clamped radius \(r_c = \text{clip}(r_{base}^{(c)} \exp(\Delta_c))\). Furthermore, a coarse uncertainty gating mechanism is incorporated:

\[u_j = 1 - 2\left|\sigma(s_j^{(c)}) - 0.5\right| \in [0, 1]\]

Only points falling within the ROI sphere whose enclosing coarse voxel satisfies \(u_{v(i)} \ge \tau\) (default \(\tau = 0.20\)) are passed to the fine-grid refinement branch. This effectively concentrates fine-resolution computation on ambiguous boundary zones, bypassing already confident object interiors and open empty background.

4. Fine-Grid ROI Refinement & Residual Fusion: Cross-Resolution Boundary Recovery and Smooth Updating

The selected ROI points are re-voxelized at a finer voxel size \(v_f = v / \eta\) with scale factor \(\eta = 2\). To retain computational efficiency and avoid parameter bloat, the refinement branch completely shares weights with the coarse backbone encoder \(E\) and two-way decoder \(D\), generating fine-level prediction logits \(s_{f, u}^{(c)}\). When mapping fine predictions back to the coarse grid, fine voxels mapped to the same coarse voxel are pooled using max-aggregation:

\[\hat{s}_j^{(c)} = \max_{u \in \mathcal{M}(j)} s_{f, u}^{(c)}\]

The aggregated fine prediction is then integrated into the coarse prediction via a residual update: \(s_j^{(c)} \leftarrow (1 - \alpha) s_j^{(c)} + \alpha \hat{s}_j^{(c)}\) with default fusion weight \(\alpha = 0.7\). Max-aggregation preserves thin geometric structures and sharp protrusions that would otherwise be smoothed out by average pooling, while residual fusion ensures stable and non-destructive mask updating across successive clicks.

Loss & Training

During training, an oracle click simulation policy mimics user behavior: the first click is placed inside the ground-truth target mask, and subsequent clicks are iteratively assigned to the center of the largest error connected component (positive click for false-negative regions, negative click for false-positive regions), simulating up to \(C\) steps. The per-step objective combines the fused segmentation loss with the auxiliary prompt regularizers:

\[\mathcal{L} = \frac{1}{C}\sum_{c=1}^C \left[ \mathcal{L}_{bce}\left(s^{(c)}, y\right) + \mathcal{L}_{dice}\left(s^{(c)}, y\right) + \lambda_{hn}\mathcal{L}_{hn}^{(c)} + \lambda_{div}\mathcal{L}_{div}^{(c)} \right]\]

To ensure numerical stability in Distributed Data Parallel (DDP) training where hard geometric cropping can cause unused parameter warnings, a small \(L_2\) regularizer is imposed on the radius prediction offset: \(\mathcal{L}_{rad} = \lambda_{rad}\|\Delta_c\|_2^2\).

Key Experimental Results

Main Results

All models are evaluated under a unified protocol: trained exclusively on ScanNet40, and tested across ScanNet40 (in-domain), S3DIS (out-of-domain indoor), and KITTI-360 (out-of-domain outdoor LiDAR) across click budgets \(k \in \{1, 2, 3, 5, 10\}\) (IoU@k, %).

Dataset Method IoU@1 IoU@2 IoU@3 IoU@5 IoU@10
ScanNet40 (In-Domain) InterObject3D 40.8 55.9 63.9 67.6 77.6
AGILE3D 63.0 70.6 75.1 79.7 83.5
Easy3D 68.2 74.6 77.3 79.6 81.7
NegROI (Ours) 72.1 78.1 80.9 82.3 83.6
S3DIS (OOD Indoor) InterObject3D 38.5 54.0 62.5 72.4 79.9
AGILE3D 58.5 70.7 77.4 83.6 88.3
Point-SAM 38.8 n/a 67.1 72.2 80.6
Easy3D 65.7 76.0 80.8 84.9 87.8
NegROI (Ours) 66.9 77.8 82.2 85.4 87.8
KITTI-360 (OOD Outdoor) InterObject3D 2.0 5.1 8.5 72.4 83.6
AGILE3D 34.8 40.7 42.7 44.4 49.6
Point-SAM 44.0 n/a 67.1 72.2 80.8
Easy3D 46.3 58.7 66.7 76.2 83.6
NegROI (Ours) 47.5 62.6 71.4 80.1 87.7

Ablation Study

The contribution of each individual component is systematically ablated on ScanNet40 and S3DIS across click counts \(k \in \{1, 5, 10\}\):

Config Variant ScanNet40 IoU@1 ScanNet40 IoU@5 ScanNet40 IoU@10 S3DIS IoU@1 S3DIS IoU@5 S3DIS IoU@10 Note
Base (No ROI, No NEG Prompts) 67.0 78.8 80.8 62.5 80.8 83.8 Baseline two-way transformer
+ Scene-Cond. NEG Prompts 69.2 80.4 82.3 64.2 83.0 85.7 Explicit background prototype modeling
+ ROI Refinement (Fine-Grid) 70.9 81.6 83.0 65.8 84.1 86.9 Sharpens thin structural boundaries
+ Uncertainty-Driven ROI Selection 71.8 81.9 82.8 66.5 85.0 87.4 Prunes confident regions to focus compute
+ Boundary Hard Negatives (\(\mathcal{L}_{hn}\)) 71.2 82.1 83.1 65.6 85.2 87.1 Direct supervision on boundary confusers
+ Diversity (\(\mathcal{L}_{div}\)) (Full Model) 72.1 82.3 83.6 66.9 85.4 87.8 Prevents prompt collapse; optimal result

Key Findings

  • Substantial Gains Under Sparse Clicks: The performance advantage of NegROI is most prominent in the early interaction phase (1–3 clicks). On ScanNet40, IoU@1 reaches 72.1%, outperforming Easy3D by +3.9%. On KITTI-360, IoU@2 and IoU@3 increase by +3.9% (58.7% \(\rightarrow\) 62.6%) and +4.7% (66.7% \(\rightarrow\) 71.4%), confirming that early errors stem predominantly from background attachment and that negative prompts prevent large-scale false alarms.
  • Superior Outdoor Cross-Domain Transfer: When transferred directly to the challenging KITTI-360 outdoor LiDAR dataset without target-domain retraining, baselines like AGILE3D degrade severely due to domain shifts in sensor density (IoU@10 plateaus at 49.6%). In contrast, NegROI achieves 87.7% IoU@10 (+4.1% over Easy3D's 83.6%), demonstrating remarkable robustness across different spatial scales.
  • Complementarity of Dual Regularizers: The ablation study verifies that supervising boundary hard negatives (\(\mathcal{L}_{hn}\)) without prompt diversity (\(\mathcal{L}_{div}\)) risks prototype collapse. Enforcing diversity ensures each negative query specializes in distinct background distractors, leading to the best performance across all benchmarks.

Highlights & Insights

  • Active Distractor Modeling via Learned Prototypes: Traditional interactive segmentation passively relies on human users to supply negative clicks when false positives occur. NegROI proactively mines \(K\) scene-conditioned background prototypes from scene features via attention, shifting error correction from purely manual retries to model-driven suppression.
  • Decoupled Resolution via Uncertainty-Guided ROI: Global high-resolution voxelization is computationally intractable. NegROI leverages the physical locality of clicks alongside coarse uncertainty scores to focus high-resolution re-voxelization strictly where ambiguity persists, presenting an efficient paradigm for scalable 3D vision.
  • Max-Aggregation for Fine-to-Coarse Projection: Utilizing max-pooling instead of average-pooling during fine-to-coarse back-projection ensures that narrow boundaries and small protrusions detected on the fine grid are not diluted or washed out during discretization.

Limitations & Future Work

  • Non-Differentiable ROI Gating: The current hard thresholding for ROI point selection and radius prediction breaks end-to-end gradient flow; exploring continuous soft gating or differentiable sampling would be a promising direction.
  • Static Weights at Test Time: The model relies on static pre-trained parameters during test-time user interactions, leaving room for online test-time adaptation (TTA) that adapts feature weights dynamically along the interactive click trajectory.
  • Degradation in Extreme Far-Field LiDAR: In outdoor scenarios with ultra-sparse point returns (e.g., objects beyond 50 meters containing fewer than 10 points), local \(k\)NN density estimates exhibit high variance, occasionally affecting adaptive radius prediction.
  • vs Easy3D: While Easy3D establishes a strong, clean baseline using a two-way transformer decoder, it lacks specialized handling of boundary fuzziness and complex false positives. NegROI augments this baseline with negative prototype queries and multi-resolution ROI refinement, achieving consistent improvements across all interaction steps.
  • vs AGILE3D / InterObject3D: Earlier interactive models lack adaptive multi-scale reasoning, resulting in catastrophic failure when transferred to unobserved sparse outdoor domains (such as KITTI LiDAR). NegROI maintains superior generalizability owing to its localized refinement and robust background prototype representation.

Rating

  • Novelty: ⭐⭐⭐⭐☆ [Novel combination of scene-conditioned negative prototypes and uncertainty-guided click-centric ROI refinement]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive cross-dataset evaluations across indoor RGB-D and outdoor LiDAR, accompanied by detailed step-by-step ablations]
  • Writing Quality: ⭐⭐⭐⭐☆ [Clear formulation, well-motivated problem formulation, and solid empirical validation]
  • Value: ⭐⭐⭐⭐⭐ [Provides an efficient, robust solution for interactive 3D labeling, directly beneficial for large-scale autonomous driving point cloud annotation]