Online Segment 3D Gaussians via Launching Virtual Drones¶
Conference: ECCV 2026
Paper: ECCV Official
Area: 3D Vision
Keywords: 3D Gaussian Splatting, Interactive 3D Segmentation, Virtual Drones, Next-Best-View, Mask-shaped Frustum Filtering
TL;DR¶
SAGO reformulates interactive 3D Gaussian Splatting segmentation as an online Next-Best-View (NBV) planning task driven by virtual drones within a Markov process, completely eliminating the time-consuming per-scene setup stage and enabling clean 3D asset extraction with sub-second latency (0.4–0.9s).
Background & Motivation¶
3D Gaussian Splatting (3DGS) has rapidly emerged as a dominant representation for 3D novel view synthesis and explicit scene modeling, offering real-time rendering speeds alongside high visual fidelity. These attributes make 3DGS highly appealing for downstream domains such as embodied AI, robotic interaction, virtual reality, and immersive media. In these interactive workflows, interactive 3D segmentation serves as an indispensable prerequisite for scene editing, asset extraction, and spatial understanding. Users typically expect to provide simple 2D interaction cues—such as point clicks, bounding boxes, or text prompts—on a single reference viewpoint to instantly isolate the corresponding target object geometry directly from the underlying explicit 3D Gaussian field.
However, existing 3DGS segmentation methods encounter a major trade-off when deployed in real-time interactive settings. The vast majority of optimization-based approaches (such as SAGA, Click-Gaussian, and LangSplat) rely heavily on an onerous offline per-scene preparation stage: prior to interactive queries, they must render multi-view images across the entire training set, query 2D foundation models to generate 2D masks, and perform iterative contrastive learning or feature distillation into 3D Gaussians. This setup takes tens of minutes to hours per scene. Conversely, recent training-free online approaches (such as GaussianCut and iSegMan) attempt to circumvent offline training but shift massive computation onto the interactive phase: constructing dense Gaussian affinity graphs and solving global min-cuts can take 40–60 seconds, while epipolar-guided multi-view interaction propagation and Gaussian voting take several seconds to tens of seconds per interaction. Such delays severely degrade the interactive user experience.
The core tension lies in accurately capturing complete 360-degree geometric boundaries and pruning background Gaussians under sub-second latency without relying on scene-specific offline feature distillation. Inspired by active robotic vision where autonomous drones navigate unknown physical environments for high-fidelity reconstruction, the core idea of this paper is to launch virtual drones that navigate around the target in an online Next-Best-View (NBV) Markov decision process, dynamically selecting information-maximizing viewpoints to prune background Gaussians via center-based mask-shaped frustum filtering within sub-second latency.
Method¶
Overall Architecture¶
SAGO aims to achieve setup-free, instantaneous online 3D Gaussian segmentation. The input consists of a pretrained raw 3D Gaussian scene \(\mathcal{G} = \{G_i\}_{i=1}^N\) and user-provided 2D interaction prompts (e.g., clicks, bounding boxes, or text) on an initial view; the output is a partition of the scene into disjoint foreground Gaussians \(\mathcal{A}\) and background Gaussians \(\mathcal{B}\). The interactive segmentation procedure is modeled as a Markov process for virtual drones. Starting from the user-interaction viewpoint, two virtual drones are launched to navigate along clockwise and counter-clockwise trajectories, utilizing SAM2 predictions at dynamically planned viewpoints to iteratively filter out background Gaussians.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Raw 3DGS Scene<br/>+ User Interaction Prompt"] --> B["Dual Virtual Drone Modeling<br/>Depth Back-projection Centroid & Orbit Init"]
B --> C["Online NBV Planning & Self-Evaluation<br/>Coarse-to-Fine Yaw Search + Consistency Check"]
C --> D["Center-based Frustum Filtering & Update<br/>2D Mask Spatial Pruning + Centroid & Memory Update"]
D -->|Coverage < 180 deg| C
D -->|360 deg Full Coverage Achieved| E["Clean 3D Asset Extraction<br/>Real-time Scene Manipulation & Editing"]
Key Designs¶
1. Dual Virtual Drone Modeling: Establishing an Object-Centric Spherical Coordinate System
To overcome the inefficiency of previous methods that randomly traverse static training views or conduct dense view sweeping, SAGO introduces virtual drones operating within an object-centric spherical coordinate system. At time step \(t\), the system state is parameterized as a four-tuple \(\mathcal{S}_t = \{\mathcal{A}^{(t)}, \mathbf{v}_t, \mathcal{M}_t, \tau_t\}\), where \(\mathcal{A}^{(t)}\) is the current set of foreground Gaussians, \(\mathbf{v}_t = \{\mathbf{c}_t, r_t, \phi_t, \theta_t\}\) denotes the camera pose of the virtual drone, \(\mathcal{M}_t\) represents the SAM2 memory bank tracking target features, and \(\tau_t\) is the coverage indicator. SAGO launches two virtual drones simultaneously from the user-specified interaction viewpoint—one clockwise (\(\phi\) expanding from \(0^\circ\) to \(180^\circ\)) and one counter-clockwise (\(\phi\) expanding from \(0^\circ\) to \(-180^\circ\))—each exploring a \(180^\circ\) yaw span. At initialization (\(t=0\)), the observation radius \(r_t\) is fixed to the distance between the starting viewpoint and the object centroid \(\mathbf{c}_0\), maintaining a constant observation scale. The initial centroid \(\mathbf{c}_0\) is roughly estimated via depth back-projection of the initial 2D mask. As background Gaussians are iteratively pruned and the boundary sharpens, the centroid dynamically updates to the geometric mean of the refined foreground set.
2. Online NBV Planning & Self-Evaluation: Maximizing Background Pruning Gain with Drift Guardrails
To minimize the number of required view explorations, the framework adopts an Exploration-Evaluation (EE) strategy to determine the next-best-view pose by searching across yaw increments \(\Delta\phi \in \{90^\circ, 60^\circ, 30^\circ\}\) and pitch angles \(\theta \in \{\theta_0, 0^\circ, 30^\circ, 60^\circ\}\). The optimal candidate pose \(\mathbf{v}^*\) is formulated as an optimization objective that maximizes the number of pruned background Gaussians while satisfying semantic consistency with the initial mask:
where \(\Delta\mathcal{A}'\) represents the candidate change in the foreground set after applying SAM2 inference with memory bank \(\mathcal{M}_t\) and mask-shaped frustum filtering on the rendered candidate view. The self-evaluation consistency constraint against threshold \(\sigma\) acts as a crucial guardrail, preventing error accumulation along the Markov trajectory caused by tracking drift under drastic viewpoint transitions. The drone greedily attempts an aggressive \(90^\circ\) yaw step; if tracking consistency holds, full \(180^\circ\) coverage is reached in merely two exploration steps. If tracking degrades due to heavy occlusion, the planner gracefully falls back to \(60^\circ\) or \(30^\circ\) increments, balancing exploration speed and tracking reliability.
3. Center-based Mask-shaped Frustum Filtering & Update: Lightweight Geometric Pruning & Closed-Loop Update
Upon selecting and navigating to the next optimal viewpoint \(\mathbf{v}^*\), SAGO applies Mask-shaped Frustum Filtering (MFF) to transfer 2D mask semantics into 3D Gaussian pruning at minimal computational cost. Rather than relying on expensive 3D convex hull intersections or full 2D splat-mask overlaps, SAGO implements center-based MFF: the 3D center position \(\mathbf{x}_i \in \mathbb{R}^3\) of each Gaussian is projected onto the 2D image plane as \(\mathbf{x}_{i}^{2D} = \text{project}(\mathbf{x}_i)\). Any Gaussian whose projected center falls outside the predicted 2D mask is identified as background and immediately moved to background set \(\mathcal{B}\):
Along with foreground set shrinking, the new visual and mask features are appended to the SAM2 memory bank \(\mathcal{M}_{t+1} = \mathcal{M}_t + \mathcal{M}(\mathbf{v}^*, \mathcal{M}_t)\), and the object centroid updates dynamically to \(\mathbf{c}_{t+1} = \frac{1}{|\mathcal{A}^{(t+1)}|} \sum_{\mathbf{x} \in \mathcal{A}^{(t+1)}} \mathbf{x}\). Because center projection requires only lightweight matrix multiplication and 2D point-in-polygon tests, per-view pruning executes in milliseconds without back-propagation or graph partitioning.
Loss & Training¶
SAGO operates in a purely inference-only, optimization-free mode at test time, requiring no gradient computation or model parameter fine-tuning. View rendering is performed using standard 3DGS rasterization at \(512 \times 512\) resolution, directly matching the input dimensions of SAM2. On a single NVIDIA RTX 4090 GPU, the entire end-to-end process of dual drone exploration and iterative state updates executes in 0.4–0.9 seconds.
Key Experimental Results¶
Main Results¶
To comprehensively evaluate segmentation accuracy and latency, SAGO is tested across four standard 3D segmentation benchmarks: SPIn-NeRF, NVOS, LERF-Mask, and 3D-OVS, comparing against representative offline optimization-based and online optimization-free methods.
| Method | Type | SPIn-NeRF mIoU (%) | SPIn-NeRF mAcc (%) | NVOS mIoU (%) | NVOS mAcc (%) | Setup Time | Seg. Time |
|---|---|---|---|---|---|---|---|
| MVSeg | Offline | 90.4 | 98.8 | - | - | - | - |
| NVOS | Offline | - | - | 70.1 | 92.0 | - | - |
| ISRF | Offline | 71.5 | 95.5 | 83.8 | 96.4 | - | - |
| SA3D | Offline | 91.9 | 98.8 | 90.3 | 98.2 | 2–5 min | 15–30 s |
| LangSplat | Offline | 69.5 | 94.5 | 74.0 | 94.0 | ~2.5 h | - |
| SAGA | Offline | 93.4 | 99.2 | 92.6 | 98.6 | ~1 h | 10 ms |
| Flashsplat | Online | - | - | 91.8 | 98.6 | 30–60 s | 1–2 s |
| iSegMan | Online | 92.4 | 99.1 | 92.0 | 98.4 | 50–90 s | 4–6 s |
| GaussianCut | Online | 92.9 | 99.2 | 92.5 | 98.4 | None (N/A) | 40–60 s |
| SAGO (Ours) | Online | 92.5 | 99.3 | 92.7 | 98.7 | None (N/A) | 0.4–0.9 s |
On challenging multi-object datasets LERF-Mask and 3D-OVS:
| Dataset | Metric | SAGO (Ours) | Strong Baseline (SAGA) | Online / Interactive Baseline | Analysis |
|---|---|---|---|---|---|
| LERF-Mask | Average mIoU (%) | 91.0 | - | 89.1 (ClickGaussian) | Surpasses ClickGaussian by +1.9%; +5.4% gain on Teatime scene |
| 3D-OVS | Mean mIoU (%) | 96.2 | 96.0 | 86.8 (3D-OVS) | Surpasses offline SOTA SAGA by +0.2%; achieves best on 3 of 5 scenes |
Ablation Study¶
The ablation experiments examine the impact of the active NBV planning strategy compared to reusing static acquisition viewpoints, as well as center-based MFF versus splat-based intersection.
| Config | LERF-Mask mIoU (%) | NVOS mIoU (%) | SPIn-NeRF mIoU (%) | 3D-OVS mIoU (%) | Note |
|---|---|---|---|---|---|
| Full model | 91.0 | 92.7 | 92.5 | 96.2 | Active NBV planning + center-based MFF |
| w/o NBV strategy (relying only on original real views) | 36.7 (-54.3) | 89.1 (-3.6) | 90.1 (-2.4) | 91.6 (-4.6) | Severe catastrophic degradation under heavy occlusion |
| w/o Center-based MFF (using splat-based overlap) | 90.4 (-0.6) | 92.5 (-0.2) | 91.1 (-1.4) | 96.0 (-0.2) | Gaussian boundary opacity creates false retention and added latency |
Key Findings¶
- Active NBV planning is indispensable for resolving 3D occlusions: On the occluded and cluttered LERF-Mask dataset, omitting the NBV strategy and relying solely on existing training camera views drops mIoU precipitously from 91.0% to 36.7% (a -54.3% decrease). Fixed camera poses fail to expose occluded surfaces, whereas virtual drone planning actively uncovers unobserved angles.
- Center-based MFF yields both superior efficiency and discriminability: Splat-based MFF retains any Gaussian with an elliptical splat overlapping the mask, erroneously retaining fuzzy background Gaussians on object silhouettes. Center-based MFF provides cleaner boundaries and runs faster, achieving a 0.2%–1.4% mIoU lead across all benchmarks.
- Resolving the setup-free vs. low-latency trilemma: Prior setup-free work (GaussianCut) required 40–60 seconds per interaction. SAGO delivers higher accuracy while reducing latency to 0.4–0.9 seconds, representing a 50×+ acceleration that makes truly setup-free interactive 3DGS practical.
Highlights & Insights¶
- Bridging active robotic vision and explicit 3DGS fields: Transferring Next-Best-View (NBV) exploration from autonomous physical drone inspection into explicit 3D digital Gaussian spaces replaces passive offline feature optimization with active, goal-driven view selection.
- Aggressive exploration with conservative self-evaluation guardrails: Prioritizing wide \(90^\circ\) yaw exploration while integrating an mIoU consistency check enables full coverage in just two steps for standard geometries, falling back smoothly to finer angles under severe occlusion.
- Direct applicability to real-time 3D asset creation and editing: Because clean 3D assets are extracted in under 1 second without per-scene preprocessing, users can immediately perform translation, removal, and duplication on raw 3DGS scenes.
Limitations & Future Work¶
- Approximating Gaussians as dimensionless center points: Treating continuous 3D Gaussians as center points in center-based MFF can introduce subtle jagged artifacts along sharp instance boundaries. Integrating lightweight online post-processing (e.g., GaussianTrimmer) could smooth object contours.
- Dependence on initial viewpoint clarity: The spherical navigation trajectory relies on an initial centroid estimated via depth back-projection from the user's reference view. Severe occlusions or ambiguous prompts in the initial view can skew the centroid. Future work could investigate active initialization search by virtual drones.
Related Work & Insights¶
- vs SAGA: SAGA requires ~1 hour of offline feature distillation per scene. SAGO achieves competitive accuracy (92.7% vs. 92.6% on NVOS) without any preparatory training, operating instantly on raw 3DGS scenes.
- vs GaussianCut: GaussianCut is setup-free but constructs dense affinity graphs and solves global graph cuts during user interaction, resulting in 40–60s latency. SAGO replaces heavy graph optimization with virtual drone NBV exploration and center-based MFF, cutting latency to 0.4–0.9s (>50× speedup).
- vs iSegMan: iSegMan requires 50–90s of partial setup and 4–6s per interaction for epipolar-guided propagation and Gaussian voting. SAGO achieves a setup-free pipeline with sub-second turnaround and cleaner asset isolation.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates setup-free 3DGS interactive segmentation as an active virtual drone NBV planning Markov process.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Evaluated across four established 3D segmentation benchmarks with comprehensive ablation studies and latency benchmarks.
- Writing Quality: ⭐⭐⭐⭐⭐ Problem formulation is precise, architectural trade-offs are clearly articulated, and mathematical modeling is rigorous.
- Value: ⭐⭐⭐⭐⭐ Achieves a 50× speedup over existing setup-free frameworks, enabling practical sub-second 3D asset extraction and interactive scene manipulation.