Online 3D Instance Segmentation at task-oriented granularity with Unposed Monocular Video¶
Conference: ECCV 2026
Paper: CVF Open Access
PDF: ECCV 2026
Area: 3D Vision
Keywords: 3D instance segmentation, task-oriented granularity, unposed monocular video, dense SLAM, cross-view mask clustering
TL;DR¶
Addressing the online perception needs of embodied agents on unposed monocular video, this paper introduces a task-oriented 3D instance segmentation system that reuses feed-forward dense pixel correspondences from MASt3R-SLAM to build geometric association scores and intra-frame mutual exclusivity constraints, achieving real-time 3D instance disentanglement via confidence-priority clustering without requiring depth sensors or offline pose priors.
Background & Motivation¶
Embodied agents interacting with open physical environments must parse 3D spatial instance boundaries and semantic attributes of surrounding objects incrementally from live video streams. However, conventional 3D instance segmentation approaches predominantly assume the availability of offline-collected RGB-D sensor point clouds paired with precise camera trajectories obtained from external tracking systems or offline Structure-from-Motion (SfM). This setup is impractical in embodied deployments, where an agent cannot explore and scan an entire scene beforehand, and mobile platforms are often restricted to lightweight monocular cameras, creating a pressing need for online 3D instance perception directly from streaming unposed monocular video.
Recent visual foundation model (VFM) guided 3D instance segmentation paradigms broadly follow two routes, neither of which suits interactive embodied tasks. The first relies on bottom-up segmentation (e.g., SAM or superpoint clustering) to segment all visually distinguishable components prior to semantic recognition, inevitably producing severe over-segmentation and demanding heavy post-processing. The second depends on 2D or 3D segmentation backbones pretrained on closed-set categories, fixing the segmentation granularity at test time so that it cannot adapt to specific task instructions. In a query such as "Find the towel on the bathtub," bottom-up or fixed-granularity models frequently merge the towel into the bathtub as a single instance, causing the target object to disappear. Furthermore, when integrating modern feed-forward dense SLAM systems, the backend bundle adjustment and loop closures continuously re-optimize camera poses and per-pixel depth. These drifting optimization variables cannot serve as reliable references for incremental mask association, and the resulting noisy SLAM reconstructions cause existing 3D segmentation networks to suffer catastrophic performance drops.
To address these challenges, this paper shifts away from the conventional "segment before recognition" bottom-up paradigm toward a task-oriented segmentation approach guided by task-specific category sets. Core idea: reuse dense point-level correspondences from feed-forward SLAM that remain invariant to backend pose optimizations, combine them with intra-frame mutual exclusivity constraints, and perform confidence-descending online mask clustering to achieve real-time, task-adaptive 3D instance segmentation from unposed monocular video.
Method¶
Overall Architecture¶
The framework takes an unposed continuous monocular RGB video stream and an open-vocabulary category set \(\mathcal{C}\) (extracted from task instructions via an LLM) as inputs. MASt3R-SLAM incrementally maintains a keyframe pose graph, outputs dense pointmaps, and computes point-level pixel correspondences across frames. For each newly added keyframe, task-oriented open-vocabulary detection and promptable segmentation extract non-overlapping 2D instance masks. Guided by the sparse connectivity of the SLAM pose graph, candidate mask pairs are evaluated using a geometric association metric (GAM), a semantic similarity metric (SSM), and an intra-frame mutual exclusivity matrix (MEM). A priority-ordered online clustering algorithm merges cross-view masks, which are subsequently back-projected onto the reconstructed 3D point cloud to generate task-adaptive 3D instance and semantic maps in real time.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input<br/>Unposed monocular video stream + Task category set"] --> B["Task-Oriented 2D Mask Decoupling<br/>YOLO-World + SAM2 with small-mask retention"]
B --> C["Pose-Graph Guidance & Dense Correspondence Reuse<br/>MASt3R-SLAM keyframe edges and pixel mappings"]
C --> D["Feature Priors & Intra-Frame Mutual Exclusivity<br/>Class text embeddings + Exclusivity matrix MEM"]
D --> E["Priority-Ordered Clustering & 3D IoU Merging<br/>GAM-descending merges + Long-span 3D box clustering"]
E --> F["Output<br/>Task-adaptive 3D instance and semantic maps"]
Key Designs¶
1. Task-Oriented 2D Mask Decoupling: Top-down constraint of candidate categories and spatial granularity To overcome the fragmented over-segmentation and unintended merging typical of bottom-up segmentation, the system discards class-agnostic entity partitioning. Given a task-relevant category set \(\mathcal{C}\) inferred from natural language instructions (e.g., \(\{\text{towel}, \text{bathtub}\}\) for "Find the towel on the bathtub"), the open-vocabulary detector YOLO-World first generates coarse 2D bounding boxes. These boxes prompt SAM2 to generate high-fidelity boundary masks. To resolve hierarchical container-object overlaps (e.g., items resting on tables), a small-mask retention priority rule assigns conflicting pixels to finer instances, yielding a non-overlapping multi-instance mask representation \(M_i \in \mathbb{Z}^{H \times W}\) for each keyframe \(K_i\). This aligns perception directly with task requirements, isolating relevant objects while filtering out distracting scene clutter.
2. Pose-Graph Guidance and Dense Correspondence Reuse: Efficient geometric association independent of drifting backend variables In monocular dense SLAM systems, estimated depth maps and global camera poses are repeatedly deformed and updated during backend sliding-window BA and loop closure. Re-projecting 3D point clouds to compute spatial intersections would require updating past frame projections continuously and remains highly sensitive to reconstruction noise. In contrast, the bidirectional pixel correspondences \(\pi_{ij}: p_i \to p_j\) (along with valid match mask \(V_{ij}\)) established by MASt3R-SLAM rely solely on feed-forward feature matching and local non-linear optimization, remaining valid over time regardless of backend pose adjustments. Leveraging the sparsity of the incremental pose graph, the system computes pairwise geometric associations only for keyframe pairs \(\langle K_i, K_j \rangle\) connected by an edge. Using \(\pi_{ij}\) to project mask \(M_{id=n}^i\) onto the image plane of \(K_j\), the valid overlap ratio is defined as: $\(or(n, m) = \frac{|\pi_{ij}(M_{id=n}^i) \cap M_{id=m}^j|}{|(M_{id=n}^i, V_{ij})|}\)$ Because objects are often partially observed during camera exploration, a large size discrepancy (\(M_{id=n}^i \gg M_{id=m}^j\)) can cause matching noise to underestimate the one-way overlap ratio. The system therefore computes bidirectional overlap ratios symmetrically and takes the maximum as the geometric association metric: $\(GAM_{(n,m)} = \max\left(or(n, m), or(m, n)\right) \in [0, 1]\)$ This step runs in parallel on GPU, requiring only about 9 ms per keyframe pair with virtually zero extra overhead.
3. Feature Priors and Intra-Frame Mutual Exclusivity: Zero-overhead semantic verification and error-propagation prevention Geometric cues alone can cause spurious links near contact boundaries or under matching noise, while computing mask-pooled CLIP visual embeddings on cropped image patches across views is computationally prohibitive in real-time pipelines. Instead, the system pre-computes text embeddings for category set \(\mathcal{C}\) using CLIP once at initialization: \(F = \{f_c \mid c \in \mathcal{C}\}\). Each 2D mask directly inherits the text vector of its detected class, \(s_n^i = f_{c_n^i}\), simplifying the semantic similarity metric \(SSM_{(n, m)}\) to an \(O(1)\) cosine similarity lookup between unit vectors. More critically, to prevent matching noise and occasional 2D under-segmentation from triggering cascading merge errors across views, the system enforces an intra-frame mutual exclusivity metric (MEM). Grounded in the physical principle that distinct objects co-occurring within the same image frame cannot belong to the same 3D instance, masks with negligible overlap (\(\text{IoU}(\hat{M}_n^i, \hat{M}_m^i) < \epsilon\)) prior to explicit non-overlapping partitioning are flagged as mutually exclusive: $\(MEM_{(n, m)} = \mathbb{I}\left(\text{IoU}(\hat{M}_n^i, \hat{M}_m^i) < \epsilon\right)\)$ Stored in a dense array with \(O(1)\) lookup complexity, this matrix provides a hard non-merge constraint during online clustering.
4. Priority-Ordered Clustering and 3D Bounding-Box Merging: Long-term consistency without early noise contamination In multi-view incremental mask fusion, greedy sequential matching is vulnerable to ordering artifacts, where early noisy matches irrevocably contaminate clusters. The framework models cross-view mask linking as constrained online graph clustering. After threshold filtering, candidate mask pairs are sorted in descending order of their geometric matching score \(GAM\), prioritizing pairs with the highest confidence. When evaluating whether to merge two clusters or add an unassigned mask, the operation proceeds only if no mutually exclusive mask pair (according to MEM) exists between them; conflicting candidate edges are discarded. Furthermore, because the SLAM pose graph primarily connects spatiotemporal neighbors or detected loop closures, spatially proximate views separated by long time gaps may lack graph edges. Following the first clustering stage, the system estimates 3D bounding boxes for all clusters and performs a secondary merge in descending order of 3D box IoU, subject to the same MEM constraints. This guarantees global consistency across distant viewpoints while strictly preventing illegal entity mergers.
Loss & Training¶
The framework is fully training-free and zero-shot, requiring no fine-tuning or post-training on 3D annotated datasets. It utilizes off-the-shelf pre-trained weights for 2D detection and segmentation (YOLO-World-L and SAM2-Hiera-Large) alongside pre-trained MASt3R-SLAM for real-time dense reconstruction. A single set of hyperparameters is shared across all environments: mutual exclusivity threshold \(\epsilon = 0.2\), geometric association threshold \(\tau_{GAM} = 0.25\), semantic similarity threshold \(\tau_{SSM} = 0.85\), and inter-cluster 3D bounding box IoU threshold \(\tau_{IoU} = 0.1\). The method demonstrates high empirical stability across varying scenes.
Key Experimental Results¶
Main Results¶
Evaluation was carried out on the ScanNet200 validation set (312 scenes, 198 open-vocabulary categories) and the Replica dataset (48 categories). Metrics follow standard ScanNet evaluation, reporting Average Precision at 50% and 25% 3D mask IoU thresholds (AP50 and AP25) under both Open-Vocabulary and Class-Agnostic settings. For unposed monocular video inputs, reconstructed point clouds are aligned to ground-truth coordinates using a Sim(3) transformation via evo before transferring semantic and instance labels via nearest-neighbor vertex lookup.
| Dataset / Setting | Method | Pose & Depth Source | Online | Zero-shot | Segment Granularity | AP50 โ | AP25 โ | FPS โ |
|---|---|---|---|---|---|---|---|---|
| ScanNet200 Open-Vocabulary | Open-YOLO 3D | MASt3R-SLAM (Point Cloud + Pose) | ร | ร | Mask3D | 0.8 | 2.0 | - |
| ScanNet200 Open-Vocabulary | OnlineAnySeg | MASt3R-SLAM (Point Cloud + Pose) | โ | โ | CropFormer | 0.1 | 0.5 | 15 |
| ScanNet200 Open-Vocabulary | Ours | MASt3R-SLAM (Monocular Video Only) | โ | โ | YOLO-World + SAM | 7.4 | 19.3 | 7 |
| ScanNet200 Open-Vocabulary (Ref) | Open-YOLO 3D | Sensor Depth + GT Pose | ร | ร | Mask3D | 31.7 | 36.2 | - |
| ScanNet200 Open-Vocabulary (Ref) | EmbodiedSAM | Sensor Depth + GT Pose | โ | ร | ScanNet200 Pretrained | 19.2 | 23.9 | 10 |
| ScanNet200 Class-Agnostic | OnlineAnySeg | MASt3R-SLAM (Point Cloud + Pose) | โ | โ | CropFormer | 4.7 | 19.5 | 15 |
| ScanNet200 Class-Agnostic | Ours | MASt3R-SLAM (Monocular Video Only) | โ | โ | YOLO-World + SAM | 16.4 | 44.7 | 7 |
| ScanNet200 Class-Agnostic (Ref) | OnlineAnySeg | Sensor Depth + GT Pose | โ | โ | CropFormer | 36.1 | 53.5 | 15 |
| Replica Open-Vocabulary | PanSt3R | MUSt3R (Offline Feed-forward) | ร | ร | Offline PanSt3R | 20.9 | 34.3 | - |
| Replica Open-Vocabulary | Open-YOLO 3D | Sensor Depth + GT Pose | ร | ร | Mask3D | 28.6 | 34.8 | - |
| Replica Open-Vocabulary | Open3DIS | Sensor Depth + GT Pose | ร | ร | ISBNet | 24.5 | 28.2 | - |
| Replica Open-Vocabulary | OVIR-3D | Sensor Depth + GT Pose | ร | โ | Detic | 20.5 | 27.5 | - |
| Replica Open-Vocabulary | Ours | MASt3R-SLAM (Monocular Video Only) | โ | โ | YOLO-World + SAM | 23.4 | 37.3 | 11 |
Ablation Study¶
Ablations on the Replica dataset examine the individual contributions of the Geometric Association Metric (GAM), Semantic Similarity Metric (SSM), Mutual Exclusivity Metric (MEM), inter-cluster 3D IoU merging, and the priority-ordered clustering principle.
| Config | AP50 โ | AP25 โ | Note |
|---|---|---|---|
| Ours Final System | 23.4 | 37.3 | Full model with all metrics and priority clustering |
| w/o GAM | 17.3 | 31.5 | -6.1% AP50 without feed-forward correspondence guidance |
| w/o SSM | 21.7 | 33.1 | Minor drops due to cross-category false associations |
| w/o MEM | 8.2 | 12.0 | Severe drop (-15.2% AP50); noise causes uncontrolled instance merging |
| w/o IoU | 13.8 | 31.4 | -9.6% AP50 due to failure to connect long-temporal viewpoints |
| w/o priority-ordered | 20.3 | 33.8 | -3.1% AP50 when processing pairs in random order |
Key Findings¶
- Prior 3D instance segmentation collapses on SLAM point clouds: Feeding MASt3R-SLAM reconstructed geometry and poses into existing methods leads to catastrophic degradation: Open-YOLO 3D drops from 31.7 to 0.8 in AP50, while OnlineAnySeg reaches only 0.1. This stems from noise, localized warping, and geometric artifacts inherent in feed-forward monocular reconstruction, which break 3D convolutions and spatial projection assumptions. By grounding associations directly in point-level correspondences, the proposed system reaches 7.4 AP50 on the same SLAM output, proving robust against geometric degradation.
- Intra-frame mutual exclusivity (MEM) is the critical safeguard: Removing MEM causes performance to plummet from 23.4 to 8.2 in AP50 on Replica. SLAM point matching and 2D bounding boxes inevitably contain occasional noise; without MEM acting as a hard boundary, a single spurious link between adjacent objects triggers cascading merges across views, collapsing entire rooms into oversized composite clusters.
- Monocular online perception rivals offline and GT-pose systems: On Replica, the method achieves 37.3 AP25 from raw monocular video, outperforming all evaluated offline methods with ground-truth depth and poses (e.g., Open-YOLO 3D at 34.8, Open3DIS at 28.2). It also surpasses PanSt3R (23.4 vs. 20.9 AP50, 37.3 vs. 34.3 AP25), despite PanSt3R requiring offline global processing over all frames.
- Minimal computational footprint beyond SLAM: Running on an Intel i9-12900KS CPU and NVIDIA RTX 3090 GPU, the framework operates at 10.87 FPS on Replica office0 and 7.32 FPS on ScanNet scene0011_00, closely matching standalone MASt3R-SLAM speeds (11.23 FPS and 7.58 FPS). Per-keyframe processing takes 30.1 ms for YOLO-World, 132.4 ms for SAM2, 9.0 ms per keyframe pair for GAM calculation, and 32 ms for online mask merging, confirming that correspondence reuse incurs negligible overhead.
Highlights & Insights¶
- Correspondence reuse bypasses optimization drift: Exploiting feed-forward pixel correspondences that remain invariant during backend bundle adjustment sidesteps the instability of constantly shifting 3D coordinates, establishing a clean interface between dense SLAM and embodied visual perception.
- Task-oriented granularity eliminates bottom-up over-segmentation: Transforming language instructions into focused category lists guides 2D detectors and promptable segmentation models top-down, preventing small functional objects (e.g., towels, papers) from being absorbed into background furniture.
- Lightweight hard exclusivity halts error compounding: Leveraging the physical axiom that distinct intra-frame detections cannot belong to the same 3D instance provides an effective barrier against greedy clustering errors with negligible memory overhead.
Limitations & Future Work¶
- Static scene dependency: The framework relies on the rigid static world assumption of MASt3R-SLAM. In dynamic embodied environments with moving humans or manipulated objects, correspondences degrade, necessitating future integration with monocular 4D dynamic reconstruction models.
- One-way SLAM-to-perception coupling: Information currently flows unidirectionally from SLAM to the instance segmentation pipeline. Future work could close the loop by leveraging identified 3D instances to filter dynamic points, formulate object-level loop closure hypotheses, and optimize local geometry.
- Sensitivity to 2D detector recall on small objects: Open-vocabulary performance on ScanNet200 remains constrained by 2D detection recall on long-tail categories under fast camera pans or low lighting, where 2D misses cannot be recovered in 3D.
Related Work & Insights¶
- vs OnlineAnySeg (CVPR 2025): OnlineAnySeg relies on class-agnostic CropFormer proposals with multi-view averaged CLIP embeddings and spatial projection on sensor depth; on monocular SLAM reconstructions, its AP50 drops to 0.1, whereas this method maintains robust performance (7.4 AP50) via task-oriented detection and SLAM correspondence reuse.
- vs Open-YOLO 3D (2024): Open-YOLO 3D depends on pretrained Mask3D 3D proposals matched with 2D labels, which degrades severely on noisy SLAM reconstructions (AP50 falls from 31.7 to 0.8); the proposed pipeline bypasses 3D proposal networks, projecting directly from 2D foundation models to 3D via reliable 2D correspondences.
- vs PanSt3R (ICCV 2025): PanSt3R leverages feed-forward geometry (MUSt3R) but requires offline batch processing across all views; the proposed method runs in real-time (10.87 FPS) online while exceeding PanSt3R in both AP50 (23.4 vs. 20.9) and AP25 (37.3 vs. 34.3) on Replica.
Rating¶
- Novelty: โญโญโญโญโ (Combines feed-forward SLAM pixel correspondences with task-oriented 2D foundation models and mutual exclusivity clustering, directly targeting the core challenge of online unposed 3D perception)
- Experimental Thoroughness: โญโญโญโญโญ (Comprehensive validation across ScanNet200 and Replica benchmarks, insightful evaluation on real SLAM outputs, rigorous ablations, and detailed runtime breakdown)
- Writing Quality: โญโญโญโญโญ (Clear problem formulation, progressive narrative flow, detailed algorithm specification, and well-structured illustrations)
- Value: โญโญโญโญโ (Provides an effective, deployable paradigm for sensor-minimal embodied agents to perform online open-vocabulary 3D scene understanding from low-cost monocular cameras)