Skip to content

LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding

Conference: ECCV 2026
Paper: ECCV Official
PDF: Official PDF
Area: 3D Vision
Keywords: 3D scene understanding, Sparse Voxel Rasterization, open-vocabulary, multimodal feature fusion, foundation model

TL;DR

Addressing spatial ambiguity and high preprocessing costs in 3D Gaussian Splatting, LESV establishes a deterministic confidence-gated fusion framework using explicit Sparse Voxel Rasterization and dense AM-RADIO foundation model features, slashing preprocessing overhead while achieving SOTA open-vocabulary 3D scene understanding.

Background & Motivation

Open-vocabulary 3D scene understanding provides fundamental spatial grounding for next-generation robotics, embodied navigation, and augmented reality. With the rapid evolution from implicit Neural Radiance Fields (NeRFs) to explicit 3D Gaussian Splatting (3DGS), pioneering approaches (e.g., LERF, LangSplat, Dr. Splat) have attempted to lift 2D vision-language features (such as CLIP and SAM) into 3D primitives. However, early distillation pipelines relying on 2D view synthesis optimization require hours of iterative optimization per scene. While recent direct feature registration methods shorten feature lifting to minutes via inverse volume rendering, they expose critical bottlenecks stemming from the unstructured, anisotropic physical nature of 3DGS.

Current methodologies suffer from two fundamental limitations: first, spatial ambiguity induced by unstructured geometry. 3DGS represents scenes via heavily overlapping, unconstrained Gaussian ellipsoids. Mapping 2D high-dimensional features back onto 3D Gaussians relies on probabilistic alpha-blending and view-dependent global depth sorting, causing severe semantic bleeding and ray-like spillovers across object boundaries that hinder point-level 3D interaction. Second, multi-level semantic ambiguity across spatial hierarchies. A 3D coordinate naturally embodies hierarchical semantics (e.g., a spatial point on a bear's nose belongs simultaneously to "nose", "head", and "bear"). Prevailing solutions resort to training multi-level independent feature fields via hierarchical SAM masks, driving per-scene preprocessing time to several hours and washing out fine-grained local textures through mask-level pooling.

To overcome these dual bottlenecks, this paper introduces a fundamental paradigm shift in both 3D geometric representation and 2D vision-language feature encoding. On the geometric side, it discards overlapping Gaussians in favor of Sparse Voxel Rasterization (SVRaster), augmented by monocular priors and multi-resolution TSDF depth confidence gating to ensure deterministic, bleed-free 3D projection. On the semantic side, it bypasses SAM mask pooling and multi-field hierarchies by harnessing the emergent dense language alignment of the agglomerative vision foundation model AM-RADIO. Core idea: LESV leverages explicit Sparse Voxel Rasterization with geometric confidence gating to enable deterministic 3D feature projection, while capitalizing on AM-RADIO's dense patch tokens to resolve multi-level semantics in a single pass without mask-level preprocessing.

Method

Overall Architecture

The overall architecture of LESV comprises four cohesive stages: structured geometric optimization, high-resolution dense language feature extraction, geometric confidence gating, and deterministic batch volume fusion. Given multi-view RGB images, the pipeline first trains a geometrically sound and watertight SVRaster field regularized by monocular patch-depth and surface-normal constraints. Concurrently, native language-aligned dense patch tokens are extracted from an agglomerative foundation model (AM-RADIO) via a self-correcting sliding window mechanism. Next, a global reference mesh extracted from the multi-level TSDF grid is projected to generate view-consistent depth confidence maps. Finally, dense 2D features are deterministically aggregated into disjoint 3D sparse voxels in isolated spatial batches with \(O(1)\) memory complexity, supporting zero-shot 3D object retrieval and point cloud semantic segmentation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multi-view RGB input"] --> B["Sparse Voxel Geometric Regularization<br/>Patch-depth & normal monocular prior constraints"]
    A --> C["Agglomerative Foundation Dense Semantic Alignment<br/>AM-RADIO patch language head projection"]
    B --> D["TSDF Depth Confidence Gating<br/>Multi-level TSDF mesh depth verification"]
    C --> E["Self-Correcting Sliding Window & Denoising<br/>Gaussian attenuation blending & hard-threshold cosine filtering"]
    D --> F["Deterministic Local Voxel Batch Fusion<br/>O(1) memory isolated spatial batch aggregation"]
    E --> F
    F --> G["Open-Vocabulary 3D Scene Representation<br/>3D object retrieval & zero-shot point cloud segmentation"]

Key Designs

1. Sparse Voxel Geometric Regularization: Incorporating monocular priors to eliminate geometric hollows and artifacts

Vanilla SVRaster optimization driven solely by photometric rendering loss frequently overfits surface colors at the expense of underlying structural integrity, creating fragmented, hollow geometry that degrades subsequent feature projection. To resolve this, LESV introduces monocular geometric constraints into the voxel optimization pipeline. To circumvent global scale and shift ambiguities between monocular estimates and scene coordinates, the model enforces a patch-wise depth consistency loss \(\mathcal{L}_{\text{patch}}\), which standardizes depth within local patches by mean and standard deviation:

\[\mathcal{L}_{\text{patch}} = \frac{1}{|\mathcal{P}|} \sum_{p \in \mathcal{P}} \left\| \frac{\mathbf{D}_{\text{ren}}(p) - \mu_{\text{ren}}}{\sigma_{\text{ren}}} - \frac{\mathbf{D}_{\text{prior}}(p) - \mu_{\text{prior}}}{\sigma_{\text{prior}}} \right\|_1\]

In addition, an analytic surface normal loss \(\mathcal{L}_{\text{norm}}\) penalizes normal discrepancies at grazing angles and object boundaries. These dual continuous geometric constraints ensure that the sparse voxels form topologically accurate, closed physical surfaces.

2. Multi-Resolution TSDF Depth Confidence Gating: Suppressing multi-view rendering jitter and semantic bleeding

Even after monocular regularization, rendered depths from SVRaster can exhibit view-dependent noise that corrupts exponential projection kernels and causes cross-surface semantic leakage. LESV exploits the explicit voxel structure to extract a view-independent surface prior directly from the Truncated Signed Distance Function (TSDF) field within the grid. To resolve the trade-off where coarse voxels erase high-frequency contours and fine voxels introduce holes in sparse regions, the method fuses TSDF fields across multiple resolution levels. Re-projecting the resulting watertight mesh yields a reference depth \(\mathbf{D}_{\text{mesh},k}\), which is compared against the rendered depth \(\mathbf{D}_{\text{ren},k}\) to form a continuous geometric confidence weight \(w_{\text{conf},k}\):

\[w_{\text{conf},k} = \exp\left( - \frac{|\mathbf{D}_{\text{mesh},k} - \mathbf{D}_{\text{ren},k}|^2}{2\sigma_{\text{conf}}^2} \right)\]

Modulating the spatial Euclidean projection weight \(w_{s,i,k}\) with this confidence yields the robust aggregation weight \(w_{i,k} = w_{s,i,k} \cdot w_{\text{conf},k}\), dynamically filtering out erroneous multi-view projections at occlusion boundaries.

3. Agglomerative Foundation Dense Semantic Alignment: Single-pass multi-scale semantic decoupling without mask pooling

Conventional methods rely on multi-tier SAM masks for feature pooling, discarding fine-grained local textures and multiplying preprocessing time. LESV adopts the AM-RADIO foundation model, which pre-distills DINOv2, SAM, and SigLIP into a unified backbone. Crucially, the spatial patch tokens of AM-RADIO exhibit strong emergent alignment with language space even without dense text supervision. Rather than relying on the global CLS token, LESV directly feeds dense spatial patch tokens \(\mathbf{T}_{\text{patch}}\) through the language MLP head \(h_{\text{lang}}(\cdot)\) to produce high-resolution language-aligned feature maps:

\[\mathbf{F}_{\text{lang}} = h_{\text{lang}}(\mathbf{T}_{\text{patch}})\]

Because each patch token encapsulates both fine-grained DINOv2 structural details and high-level SigLIP semantics, the resulting 3D feature field responds seamlessly to both macro-level concepts (e.g., "bear") and fine-grained sub-parts (e.g., "bear nose") in a single pass.

4. Self-Correcting Sliding Window & Self-Correlating Aggregation: Overcoming resolution ceilings and boundary seams

AM-RADIO is constrained by a native input resolution of \(512 \times 512\), where naively downsampling high-resolution images obscures small items, while learned deep upsamplers (e.g., AnyUp) distort the strict metric cosine space required for zero-shot retrieval. LESV employs a self-correcting sliding window strategy, cropping high-resolution inputs into overlapping patches and recombining them using smooth 2D spatial Gaussian kernels centered at patch anchors. To eliminate high-frequency noise introduced during window merging, LESV incorporates Self-Correction Recursive Attention (SCRA) and hard-thresholded cosine similarity filters, zeroing out negatively correlated token pairs to preserve sharp object boundaries prior to 3D lifting.

5. Deterministic Local Spatial Projection & Scalable Batch Fusion: Achieving bounded \(O(1)\) memory scaling

In 3DGS feature registration (e.g., Dr. Splat), computing view-dependent alpha-blending weights requires global depth sorting of all Gaussians along rays, leading to extreme memory consumption exceeding 256GB on dense scenes. In SVRaster, each voxel \(v_i\) possesses an explicit, fixed 3D coordinate whose spatial proximity to physical surfaces is strictly local and view-independent. This disjoint property allows the scene to be partitioned into arbitrary spatial batches, executing multi-view feature projection with constant \(O(1)\) memory footprint and bounding peak memory under ~25GB regardless of scene size.

Loss & Training

The geometric training objective combines RGB rendering loss with monocular geometric regularizers: \(\mathcal{L}_{\text{geom}} = \mathcal{L}_{\text{rgb}} + \lambda_1 \mathcal{L}_{\text{patch}} + \lambda_2 \mathcal{L}_{\text{norm}}\). Once the geometric voxel field converges and the multi-level TSDF mesh is extracted, feature extraction and 3D lifting proceed purely through deterministic forward aggregation without iterative distillation. For each voxel \(v_i\), the aggregated feature \(\mathbf{F}_i\) is computed across visible views \(\Omega_i\) as:

\[\mathbf{F}_i = \frac{\sum_{k \in \Omega_i} w_{i,k} \mathbf{f}_{i,k}}{\sum_{k \in \Omega_i} w_{i,k} + \epsilon}\]

where \(\mathbf{f}_{i,k}\) is bilinearly sampled from the view feature map and \(w_{i,k}\) combines spatial proximity and TSDF confidence.

Key Experimental Results

Main Results

Open-vocabulary 3D scene understanding was comprehensively evaluated on ScanNet point cloud semantic segmentation and LERF 3D open-vocabulary object retrieval benchmarks.

Table 1: Open-vocabulary 3D semantic segmentation on the ScanNet dataset (mIoU / mAcc %)

Method Backbone 19 Classes mIoU 19 Classes mAcc 15 Classes mIoU 15 Classes mAcc 10 Classes mIoU 10 Classes mAcc
LangSplat 3DGS 3.78 9.11 5.35 13.20 8.40 22.06
OpenGaussian 3DGS 24.73 41.54 30.13 48.25 38.29 55.19
LaGa 3DGS 32.50 49.10 35.50 53.50 42.60 63.20
Dr. Splat 3DGS 31.66 48.64 37.59 56.54 44.87 64.27
LESV (Ours) SVRaster 53.22 70.41 54.78 73.62 65.25 82.95

Table 2: Open-vocabulary 3D object retrieval on the LERF dataset (mIoU % / Acc@25 %)

Method Protocol / Setting ramen figurines teatime waldo_kitchen Average mIoU Average Acc@25
LangSplat 3DGS + Distillation 5.96 / 9.86 6.05 / 8.93 18.89 / 23.73 12.96 / 18.18 10.97 15.18
OpenGaussian 3DGS + Contrastive 31.01 / 42.25 39.29 / 55.36 60.44 / 76.27 22.70 / 31.82 38.36 51.43
Dr. Splat 3DGS + Direct Reg. 37.49 / 69.01 61.73 / 83.93 59.45 / 72.88 52.00 / 81.81 52.67 76.91
LaGa (single-level) 3DGS + Contrastive 31.60 / 46.48 49.01 / 83.93 58.73 / 89.83 38.80 / 63.64 44.54 70.97
LaGa† (multi-level) 3DGS + Hierarchical 47.95 / 70.42 55.87 / 80.35 69.53 / 89.83 61.92 / 90.09 58.82 82.88
LESV (Ours) SVRaster + Deterministic 53.34 / 81.69 55.87 / 87.50 71.40 / 89.83 43.84 / 81.81 56.11 85.21

Ablation Study

The ablation experiments decouple the contributions of the underlying geometry and feature extraction backbones.

Table 3: Geometric regularization and confidence gating ablation (fixed SAM+CLIP features)

Configuration LERF 3D mIoU LERF Acc@25 ScanNet mIoU ScanNet mAcc Note
Dr. Splat (3DGS baseline) 52.76 76.91 31.66 48.64 Standard inverse volume rendering baseline
SVRaster (vanilla) 50.80 70.63 30.62 48.90 Lacks monocular constraints; suffers from geometric hollows
+ Monocular Supervision (\(\mathcal{L}_{\text{patch}}+\mathcal{L}_{\text{norm}}\)) 53.32 71.56 31.37 49.78 Refines continuous surface topology
+ TSDF Depth Confidence (Full Geometry) 53.37 72.91 32.06 50.67 Suppresses multi-view noise; outperforms 3DGS on identical features

Table 4: Disentangling geometric representation vs. feature encoder

Method Feature Backbone LERF 3D mIoU LERF Acc@25 ScanNet mIoU ScanNet mAcc Note
Dr. Splat CLIP 52.76 76.91 31.66 48.64 Original Dr. Splat baseline
Dr. Splat AM-RADIO 53.42 70.60 37.57 47.80 Forced mask pooling due to memory limits negates dense benefits
LESV (Ours) CLIP 53.37 72.91 32.06 50.67 Demonstrates geometric superiority of sparse voxels
LESV (Ours) AM-RADIO 56.11 85.21 53.22 70.41 Dense patch tokens preserved without mask pooling, unlocking huge gains

Table 5: Computational efficiency and memory consumption (~300 images per scene)

Method Preprocessing Time Geometry Training Feature Lifting Average Peak RAM (GB)
LangSplat ~120 mins ~20 mins ~60 mins N/A
LaGa (multi-level) ~120 mins ~20 mins ~70 mins N/A
Dr. Splat - (precomputed codebook) ~20 mins ~10 mins 131.8 (>256 GB unoptimized)
LESV (Ours) ~14 mins ~15 mins ~3 mins 22.5

Key Findings

  • On the ScanNet 19-class point cloud segmentation benchmark, LESV sets a new state-of-the-art with 53.22% mIoU, achieving a massive +21.56% absolute improvement over the strongest 3DGS baseline Dr. Splat (31.66%), proving the decisive advantage of disjoint voxels over overlapping Gaussians for 3D interactions.
  • Feature disentanglement experiments reveal that upgrading Dr. Splat to AM-RADIO yields modest gains (+5.91% mIoU on ScanNet) because 3DGS cannot scale to dense per-pixel features and forces mask pooling. Conversely, LESV natively registers dense patch tokens into voxels, leaping to 53.22% mIoU.
  • LESV slashes data preprocessing from ~120 minutes down to ~14 minutes (an 8x reduction) by eliminating multi-level SAM mask extraction, performs feature lifting in just ~3 minutes, and bounds peak RAM to ~22.5 GB.

Highlights & Insights

  • Explicit representation eliminates probabilistic blur: Points out that the core limitation of 3DGS feature lifting is spatial ambiguity caused by overlapping ellipsoids, solving it at the root by using discrete, non-overlapping sparse voxels that completely eliminate semantic bleeding.
  • Capitalizing on emergent dense foundation models: Bypasses the costly industry standard of "hierarchical SAM masks + multi-tier feature fields" by harnessing AM-RADIO's dense patch tokens, resolving sub-part to global semantic hierarchies in a single unified pass.
  • Dual geometric safeguards: Combines monocular patch-depth and normal supervision during optimization with view-independent multi-resolution TSDF mesh confidence gating during lifting, robustly preventing multi-view noise injection.
  • Bounded \(O(1)\) memory scaling: Eliminating global depth sorting allows isolated spatial batch processing, offering an elegant engineering blueprint for scaling open-vocabulary representations to unbounded large-scale environments.

Limitations & Future Work

  • Dependence on monocular estimator quality: The geometric fidelity of SVRaster relies on the accuracy of monocular depth and normal estimators; severe specular reflections or extreme lighting variations may induce local surface distortions.
  • Physical resolution bound of voxel grids: Although hierarchical octrees are employed, the spatial granularity of extremely miniature objects remains bounded by the minimal physical voxel size, unlike continuous implicit coordinate representations.
  • Future directions: Integrating continuous neural implicit residual interpolation within voxels to bridge discrete voxel efficiency with continuous geometric fidelity; extending the formulation to temporal sequences for dynamic open-vocabulary 4D understanding.
  • vs LangSplat / LeGaussians: LangSplat follows a "render-and-compare" distillation pipeline requiring scene-specific autoencoders and hours of optimization, suffering in direct 3D retrieval (10.97% mIoU); LESV employs deterministic forward lifting without compression autoencoders, achieving 56.11% mIoU with a 20x speedup.
  • vs Dr. Splat: Dr. Splat registers features via inverse volume ray-tracing on 3DGS, which introduces Gaussian overlap bleeding and memory bloat (>256 GB); LESV uses explicit SVRaster to bound RAM at ~25 GB while improving ScanNet segmentation mIoU by +21.56%.
  • vs LaGa: LaGa relies on hierarchical SAM masks and multi-scale contrastive learning requiring 190 minutes per scene; LESV leverages emergent dense alignment from AM-RADIO to achieve superior accuracy in under 30 minutes total time.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Addresses the core geometric ambiguity of 3DGS by combining explicit sparse voxels with emergent dense foundation model features]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Thorough validation across ScanNet segmentation, LERF 3D retrieval, 2D localization, architectural ablations, and runtime/memory profiling]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Insightful motivation, rigorous mathematical formulation, and exceptionally clear narrative progression]
  • Value: ⭐⭐⭐⭐⭐ [Provides a practical, highly scalable, and accurate baseline for open-vocabulary 3D scene understanding and embodied robotics]