Skip to content

LEGO: Leveled Language Gaussian Splatting

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://pz0826.github.io/LEGO-Webpage/
Area: 3D Vision
Keywords: Gaussian Splatting, 3D Semantic Hierarchy, Open-Vocabulary Scene Understanding, 3D Scene Graph, Chain-of-Retrieval

TL;DR

Addressing the viewpoint dependency and semantic-scale decoupling in existing 3D language Gaussian splatting, LEGO self-adaptively re-grades multi-view SAM granularities into unified 3D-consistent structural levels, performs decoupled feature distillation, and constructs level-wise language scene graphs to empower LLMs for complex spatial reasoning.

Background & Motivation

Three-dimensional open-vocabulary scene understanding is rapidly evolving from coarse closed-set classification toward the open-ended decomposition of arbitrary real-world concepts. Crucially, real-world semantics are inherently hierarchical: when perceiving a scene, humans naturally interpret entities along recursive structural lineages from coarse to fine, such as "flowerpot → bouquet → bud → petal". A truly robust 3D open-vocabulary perception system must not merely assign isolated semantic tags to free-floating objects, but explicitly uncover and structure the hierarchical compositional logic embedded in 3D physical space.

Lifting Segment Anything Model (SAM) masks to build 3D hierarchies is intuitive, yet prior paradigms encounter two fundamental roadblocks. First, 2D granularity-based distillation directly projects SAM's three rigid granularities (whole, part, subpart) into 3D. However, SAM's granularities are strictly perspective-bound rather than intrinsic to the physical object: a flower appears as a complete bud from a distant viewpoint, but splits into individual petals in a close-up shot at the exact same granularity setting. Distilling such inconsistent cross-view labels induces severe granularity blurring. Second, 3D scale-based distillation groups Gaussian segments by absolute physical scale to guarantee viewpoint invariance. Yet, because semantic concepts exhibit substantial intra-class scale variance (e.g., flowers within the same bouquet vary drastically in size), a single global scale parameter inevitably causes semantic-scale decoupling—over-segmenting large flowers while leaving smaller ones intact.

To resolve this dilemma, this paper argues that the core representation must shift from view-dependent granularities and rigid physical scales to intrinsic structural levels—a compositional rank invariant to camera distance and absolute size. Core idea: treat volatile multi-view SAM masks as partial, uncalibrated observations of an underlying hierarchy, self-adaptively re-grade them into view-consistent structural levels via local co-visibility 3D scale clustering, optimize Gaussians in a decoupled subspace with dense axioms, and construct level-wise language scene graphs for context-aware spatial reasoning.

Method

Overall Architecture

LEGO initializes a 3D Gaussian field from geometric point clouds generated by MASt3R-SfM. In the preprocessing phase, 2D SAM masks from all viewpoints are pooled and lifted into 3D space to estimate their physical dimensions, followed by local co-visibility scale histogram peak detection to assign 3D-consistent structural levels. Guided by these leveled masks, LEGO establishes dense, conflict-free pixel-pair indicator constraints governed by hierarchical axioms, and trains Gaussian primitives via contrastive distillation over decoupled identity feature subspaces. Finally, optimal view selection aggregates pristine CLIP embeddings, and the resulting segments are organized into a level-wise language scene graph with containment and spatial proximity edges to support LLM-directed Chain-of-Retrieval.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multi-view RGB Images + Reconstructed Point Cloud"] --> B["3D Scale Estimation & Local Peak-based Level Re-grading<br/>Resolving viewpoint dependency and semantic-scale gap"]
    B --> C["Hierarchical Dense Indicator Construction<br/>Enforcing monotonicity, recursive inclusion, and inheritance"]
    C --> D["Decoupled Feature Space Contrastive Distillation<br/>Independent level optimization with unit hypersphere normalization"]
    D --> E["Recursive Top-Down Hierarchical Scene Clustering<br/>HDBSCAN clustering generating nested tree segments"]
    E --> F["View-Aware Language Grounding & Level-wise Scene Graph<br/>Optimal view CLIP binding and CoR spatial reasoning"]
    F --> G["Multi-level Open-Vocabulary Segmentation & Query Outputs"]

Key Designs

1. 3D Scale Estimation & Local Peak-based Level Re-grading: resolving cross-view granularity conflicts Conventional lifting approaches directly inherit SAM's view-dependent granularity tags, causing severe label collisions across viewpoints with different camera-to-object distances. LEGO tackles this by establishing a pixel-to-point mapping \(F: u \mapsto p\). For each 2D mask \(m_i\) in the global pool \(\mathcal{M}\), constituent pixels are unprojected to a 3D point set \(P_{m_i} = \{F(u) \mid u \in m_i\}\), and its physical spatial scale \(s_i\) is computed from the spatial standard deviation: \(s_i = 2\sqrt{\sum_{d \in \{x,y,z\}} \text{std}(P_{m_i, d})^2}\). To determine the relative compositional rank, LEGO gathers the spatially co-visible neighboring masks \(\mathcal{N}(m_i) \cup \{m_i\}\) and performs peak detection on their scale histogram, yielding \(L\) prominent peaks \(K_i = \{p_1, p_2, \dots, p_L\}\) ordered from coarse to fine. Mask \(m_i\) is then assigned to a discrete structural level based on scale proximity: $\(l_i = \arg\min_{l \in \{1,\dots,L\}} |s_i - p_l|\)$ This local consensus converts fragmented 2D observations into globally consistent 3D semantic levels, eliminating the manual tuning of global scale parameters.

2. Hierarchical Dense Indicator Construction: stabilizing supervision against mask sparsity Because the discovered 3D structural hierarchy contains significantly more levels than SAM's three nominal granularities, the masks assigned to any specific level are naturally sparse, causing optimization instability and structural artifacts. LEGO avoids explicit mask densification by formulating an equivalent, highly parallelizable pixel-pair indicator \(\mathbb{I}_k(i, j)\), indicating whether pixels \(i\) and \(j\) in view \(v\) belong to the same entity mask at level \(k\). To guarantee topological soundness, the indicator enforces three strict axioms: monotonicity (signals non-increase as hierarchy deepens; sub-parts split but never re-merge), recursive inclusion (identity at fine levels requires identity across all coarser parent levels), and structural inheritance (pixels lacking masks at the target level inherit labels from their nearest valid parent). This yields dense, conflict-free supervision for Gaussian learning.

3. Decoupled Feature Space Contrastive Distillation: preventing cross-level semantic entanglement To eliminate interference across hierarchical ranks, LEGO assigns each Gaussian primitive a composite identity feature \(f \in \mathbb{R}^{L \times d}\), decomposed into \(L\) mutually independent \(d\)-dimensional subspaces, routing level \(k\) directly to \(f_k \in \mathbb{R}^d\). On the rendered feature map \(F_{v,k}\), a scale-balancing weight \(w_k^{i,j} \propto (A_i^k \cdot A_j^k)^{-1}\) prevents massive background objects from dominating gradients. Combined with cosine contrastive loss, an auxiliary \(L_2\) penalty enforces intra-cluster tightness: \(\mathcal{L}_{pos} = \|F_{v,k}^i - F_{v,k}^j\|_2^2\). Furthermore, to prevent the volume rendering process from exploiting view-dependent blending shortcuts, both primitive features and rendered 2D features are constrained to a unit hypersphere: $\(\mathcal{L}_{3d} = \left|1 - \|\mathbf{f}_k\|_2\right|, \quad \mathcal{L}_{2d} = \left|1 - \|\mathbf{F}_{v,k}^i\|_2\right|\)$ After optimization, top-down recursive HDBSCAN clustering partitions Gaussians into strictly nested tree-structured segments.

4. View-Aware Language Grounding & Level-wise Scene Graph: powering LLM spatial chain-of-retrieval Naive multi-view feature averaging frequently incorporates occlusion artifacts and off-angle blur. LEGO introduces an Optimal View Selection (OVS) metric \(S(c, v)\) combining 3D visibility, 2D image coverage, and 2D SAM semantic alignment IoU. The top-\(\tau\) views maximizing \(S(c, v)\) are cropped to extract CLIP features, which are average-pooled into a pristine open-vocabulary representation \(E_c\). Furthermore, segments are organized into a level-wise language scene graph with vertical part-whole edges \(E_{hier}\) and horizontal bounding-sphere intersection edges \(E_{adj}\). For intricate compositional queries (e.g., "Find the handle of the pitcher beside the rolling pin"), an LLM executes Chain-of-Retrieval (CoR) along graph paths (rolling pin → pitcher → handle), utilizing coarse objects as spatial anchors to disambiguate and pinpoint tiny targets.

Loss & Training

The overall network optimization minimizes the composite distillation loss per viewpoint \(v\) and structural level \(k\): $\(\mathcal{L}_{total} = \mathcal{L}_{contrast} + \lambda_{pos} \mathcal{L}_{pos} + \lambda_{norm} (\mathcal{L}_{3d} + \mathcal{L}_{2d})\)$ The scale-balancing weight balances gradients across diverse object scales. Once the identity field converges, recursive tree clustering and optimal view CLIP grounding are performed without additional network re-training.

Key Experimental Results

Main Results

LEGO is evaluated across promptable segmentation (NVOS, SPIn-NeRF) and open-vocabulary understanding benchmarks (LERF-OVS, Mip-NeRF 360).

Dataset Metric LEGO (Ours) Strong Baseline (SAGA/LaGa) Second Best Gain
NVOS mIoU(%) 94.2 92.6 (SAGA) 92.2 (SA3D-GS) +1.6%
NVOS mAcc(%) 98.7 98.6 (SAGA) 98.6 (COB-GS) +0.1%
SPIn-NeRF mIoU(%) 94.2 93.4 (SAGA) 94.3 (OmniSeg3D) on par with SOTA
SPIn-NeRF mAcc(%) 99.3 99.2 (SAGA) 99.3 (OmniSeg3D) on par
LERF-OVS 3D Location mAcc(%) 88.4 80.3 (LaGa) 84.3 (LangSplat) +4.1% vs second-best
LERF-OVS 3D Segmentation mIoU(%) 68.4 64.0 (LaGa) 61.3 (Occam's LGS) +4.4%
Mip-NeRF 360 3D Location mAcc(%) 92.6 85.6 (LaGa) 88.7 (GAGS) +3.9%
Mip-NeRF 360 3D Segmentation mIoU(%) 73.0 63.7 (LaGa) 69.4 (LangSplatV2) +3.6%

On the fine-grained Chain-of-Retrieval (CoR) benchmark comprising 120 complex spatial and hierarchical compositional queries across four scenes, LEGO demonstrates commanding advantages over scene graph and Gaussian baselines:

Method Teatime mIoU(%) Kitchen mIoU(%) Bonsai mIoU(%) Counter mIoU(%) Overall mIoU(%)
BBQ 5.6 2.6 8.1 8.2 6.1
THGS 9.9 5.0 10.3 12.9 9.5
LaGa (Native) 5.8 5.2 12.0 10.4 8.3
LaGa w/ CoR 8.4 14.5 13.8 20.5 14.3
LEGO w/o CoR 25.2 10.0 30.2 25.6 22.7
LEGO (Full) 52.2 44.9 57.6 51.4 51.6

Ablation Study

Ablation on the complex indoor Room scene from the Mip-NeRF 360 dataset investigates the impact of each distillation loss component:

Config Re-weighting (W) pos (\(\mathcal{L}_{pos}\)) Norm (\(\mathcal{L}_{3d/2d}\)) mAcc(%) mIoU(%) Note
Full model ✓ ✓ ✓ 93.1 67.3 optimal performance
w/o Re-weighting ✗ ✓ ✓ 86.2 63.6 large entities dominate gradients (-3.7% mIoU)
w/o pos penalty ✗ ✗ ✓ 86.2 62.3 loose cluster boundaries (-1.3% mIoU)
w/o Normalization ✗ ✗ ✗ 82.8 60.7 view-dependent shortcut collapse (-1.6% mIoU)

Key Findings

  • Exceptional fine-grained sub-part isolation: LEGO cleanly isolates extreme sub-parts such as "corn" and "onion segments" in Ramen, achieving +11.2% mAcc and +11.9% mIoU boosts over previous state-of-the-art baselines that routinely collapse to the entire ramen bowl.
  • Vital role of scale-balancing weight: Omitting scale-balancing weights results in a massive 6.9% drop in location mAcc, proving that unweighted gradients are heavily distorted by large objects at the expense of fine structures.
  • Scene graph as an essential reasoning bridge: On CoR queries, standard flat CLIP matching achieves only 22.7% mIoU, whereas graph-guided chain reasoning attains 51.6% mIoU (+28.9%), verifying the necessity of coarse-to-fine relational anchoring.

Highlights & Insights

  • From 2D volatile granularity to 3D consensus levels: Reframing SAM's view-dependent 2D masks as uncalibrated continuous scale samples and clustering them via local scale histograms elegantly bypasses manual scale tuning and 3-level limitations.
  • Decoupled subspace feature distillation: Allocating orthogonal feature vectors to distinct structural levels prevents vertical semantic contamination, ensuring that fine-grained sub-parts retain sharp identity representations.
  • Hierarchical scene graphs bridging 3D fields and LLMs: Elevating dense continuous fields to structured scene graphs with vertical hierarchy and horizontal proximity equips LLMs with spatial anchoring tools for multi-step reasoning.

Limitations & Future Work

  • Reliance on initial geometry estimation: Accurate physical scale computation depends on reliable point clouds from MASt3R-SfM; geometric failures in textureless or specular regions can induce level re-grading errors.
  • Restriction to static scenes: The current pipeline assumes stationary environments, making it challenging to maintain hierarchical graphs when dynamic object state changes occur.
  • Future directions: Integrating closed-loop LLM feedback for interactive topology verification and leveraging level-wise scene graphs for embodied robotic manipulation planning.
  • vs SAGA / OmniSeg3D: Prior 3D interactive segmentation methods rely on manual scale thresholds or user clicks; LEGO achieves autonomous, top-down hierarchical decomposition.
  • vs LangSplat / LaGa: Existing language Gaussian models adhere rigidly to SAM's 3-level granularities, suffering from view-dependent level collisions; LEGO's 3D-consistent re-grading and decoupled subspaces guarantee view-invariant hierarchy.
  • vs BBQ / THGS: Previous scene graph approaches produce flat or shallow graphs; LEGO constructs arbitrarily deep, nested part-whole graphs that support fine-grained contextual navigation.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Elegant reformulation of multi-view 2D granularity into 3D structural levels via local co-visibility scale clustering.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluations across promptable segmentation, open-vocabulary benchmarks, and a new 120-query CoR evaluation.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear conceptual taxonomy, rigorous mathematical formulation, and well-structured arguments.
  • Value: ⭐⭐⭐⭐⭐ Provides a foundational milestone for fine-grained 3D scene representation, open-world spatial reasoning, and embodied AI.