Skip to content

Triangle Splatting SLAM

Conference: ECCV 2026
Paper: ECCV Official Page (Poster 5458)
Project Page: https://nmjfry.github.io/triangle-splatting-slam/
Area: 3D Vision
Keywords: RGB-D SLAM, triangle splatting, differentiable rendering, Delaunay triangulation, online mesh reconstruction

TL;DR

The first dense RGB-D SLAM system that adopts a differentiable triangle soup as its sole underlying map representation, utilizing analytic camera pose Jacobians, equilateral geometric regularization, and online restricted Delaunay triangulation to deliver superior 3D geometric reconstruction and real-time mesh editing compared to 3D/2D Gaussian Splatting baselines.

Background & Motivation

Embodied AI and robotics fundamentally rely on dense explicit 3D maps constructed incrementally as a camera traverses an unknown environment. Over the past few years, dense visual SLAM has advanced from sparse landmark tracking to dense radiance field reconstruction driven by differentiable rendering. Nonetheless, the choice of 3D spatial representation remains a compromise: coordinate-based neural implicit fields (NeRFs) suffer from computationally prohibitive ray-marching and MLP queries, precluding interactive online operations; 3D Gaussian Splatting (3DGS) remarkably boosts novel-view synthesis frame rates and visual fidelity, yet anisotropic Gaussian ellipsoids remain inherently volumetric and disconnected primitives without surface topology, severely impeding seamless integration with standard physics engines, collision simulation, and mesh-based geometric editing.

Triangle meshes have been the indisputable standard representation across computer graphics and game engines for decades, featuring native GPU rasterization efficiency, adaptive geometric resolution, and explicit topological connectivity. However, this rigid topological connectivity turns into a formidable liability during incremental reconstruction. Dynamically maintaining manifold topological consistency while continuously integrating noisy streaming sensor measurements is notoriously difficult. Prior volumetric TSDF systems resort to full-grid Marching Cubes regeneration that incurs severe memory overhead and fixed resolutions, whereas point/surfel or Gaussian SLAM frameworks extract surfaces purely via indirect post-processing that receives no direct geometric supervision during the tracking and mapping loop, breaking consistency between rendered depth and the final extracted mesh.

Recent offline rendering breakthroughs revealed that an unstructured "triangle soup" can be continuously optimized via differentiable rasterization and subsequently meshed via Delaunay triangulation across posed views. Could one discard intermediate neural fields or Gaussian proxies altogether and adopt differentiable triangles as the native map primitive for SLAM? The core idea is: use differentiable triangles as the sole underlying map representation for dense visual SLAM, deriving analytic camera pose Jacobians for fast tracking, imposing equilateral and surface normal regularizations against geometric degeneration, and performing online restricted Delaunay triangulation to achieve lightweight, post-processing-free continuous mesh extraction and interactive scene editing.

Method

Overall Architecture

Triangle Splatting SLAM adopts a sequential tracking and mapping architecture. The pipeline consumes an incoming stream of RGB-D frames. The tracking frontend freezes the current map primitives and optimizes the camera pose parameters on the SE(3) manifold until convergence using combined photometric and geometric depth losses. Keyframes are determined adaptively based on the triangle co-visibility intersection-over-union (IoU) with respect to the prior keyframe. In the backend mapping stage, new equilateral triangles are spawned at keyframe depth points with radii adapted to local point density, and a local keyframe replay window (comprising maximum co-visibility frames, random historical frames, and the current keyframe) is jointly optimized across visual appearance and surface geometry. Unstable primitives are dynamically pruned according to opacity and screen footprint, while under-resolved blurred regions undergo 1-to-4 midpoint edge subdivision. When an explicit mesh is requested, restricted Delaunay triangulation connects vertices on-the-fly, producing a topologically coherent mesh ready for online deformation and collision testing.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Incoming RGB-D Frame Stream"] --> B["Analytic Pose Jacobian Tracking<br/>Photometric and Depth Optimization"]
    B --> C{"Co-Visibility Verification<br/>Triangle IoU vs. Prior Keyframe"}
    C -->|Regular Frame| B
    C -->|New Keyframe| D["Normal and Adaptive Size Initialization<br/>Equilateral Triangles from Depth Points"]
    D --> E["Multi-Constraint Joint Mapping<br/>Color + Depth + Normal + Equilateral Loss"]
    E --> F["Dynamic Pruning and Midpoint Subdivision<br/>Filter Opacity and Screen Coverage"]
    F --> G["Restricted Delaunay Triangulation<br/>Online Connected Mesh Extraction and Editing"]

Key Designs

1. Vertex-shared triangle parameterization and smooth rasterization: preventing topological discontinuities and enabling full-vertex gradient flow

The scene geometry is parameterized as a set of \(N\) vertices \(\mathcal{V}\), where each vertex stores world coordinates, RGB color, and opacity \(v_i = (x_i, y_i, z_i, c_i, o_i)\). A triangular face \(F_m = \{v_i, v_j, v_k\}\) defines the topological connectivity via vertex index triplets, and interior pixel colors are linearly interpolated via barycentric coordinates \(c_{F_m} = \lambda_i c_i + \lambda_j c_j + \lambda_k c_k\). Unlike classic Triangle Splatting where each face holds a single independent color, vertex sharing ensures smooth attribute transitions during dynamic re-indexing. Furthermore, while offline Mesh Splatting sets face opacity to the minimum vertex opacity \(\min(o_i, o_j, o_k)\)—which creates a discontinuous gradient path directed exclusively to the lowest-opacity vertex—this method adopts the arithmetic mean \(o_{F_m} = \frac{1}{3}(o_i + o_j + o_k)\), ensuring that all three vertices receive back-propagated gradients simultaneously. During rasterization, a 2D signed distance field \(\phi(p) = \max_{i \in \{1,2,3\}} L_i(p)\) and an incenter-referenced window function \(I(p) = \text{ReLU}\left(\frac{\phi(p)}{\phi(s)}\right)^\sigma\) impart smooth differentiable boundaries to projected faces, accumulating radiance along depth-sorted rays: $\(C(p) = \sum_{n=1}^N c_{F_n} o_{F_n} I_n(p) \prod_{i=1}^{n-1} (1 - o_{F_i} I_i(p))\)$

2. Analytic camera pose Jacobians on the SE(3) manifold: eliminating automatic differentiation overhead for efficient tracking

Camera pose estimation typically demands around 80 gradient descent iterations to achieve tight convergence. Relying on deep learning autograd graphs introduces heavy GPU memory footprint and latency that bottlenecks interactive operation. To resolve this, the authors derive exact analytical camera pose Jacobians and evaluate them directly inside the backward CUDA kernel. For Lie algebra perturbations \(\tau \in \mathfrak{se}(3)\), the derivative of camera-space vertex coordinates \(v_C\) with respect to the camera Lie algebra parameters follows the standard manifold formulation: $\(\frac{\mathcal{D} v_C}{\mathcal{D} \mathbf{T}_{CW}} = \begin{bmatrix} \mathbf{I}_3 & -[v_C]_\times \end{bmatrix}\)$ Chaining this expression with the image-space perspective projection derivative \(\frac{\partial v_I}{\partial v_C}\) and the loss gradient \(\frac{\partial L}{\partial v_I}\) allows closed-form gradient computation in a single fused CUDA pass, drastically reducing tracking latency to hundreds of milliseconds.

3. Normal guidance and equilateral regularization: suppressing geometric drift and degenerate sliver triangles

Online triangle optimization faces two critical failure modes: normal ambiguity under unconstrained view angles and catastrophic elongation into needle-like degenerate triangles under photometric gradient pulling. To enforce geometric consistency, sensor normal maps \(\bar{N} = \frac{\partial_x p_d \times \partial_y p_d}{\|\partial_x p_d \times \partial_y p_d\|}\) are computed directly via finite differences on back-projected keyframe depth maps. Mapping is supervised via a normal cosine distance loss \(E_{norm} = \sum_{p \in \mathcal{P}} (1 - N(V, \mathbf{T}_{CW})_p \cdot \bar{N}_p)\). Simultaneously, an equilateral triangle loss \(E_{equi}\) penalizes internal angles \(\theta_{m,1}, \theta_{m,2}, \theta_{m,3}\) that deviate from \(60^\circ\) (\(\cos 60^\circ = 0.5\)): $\(E_{equi} = \frac{1}{|\mathcal{F}|} \sum_{F_m \in \mathcal{F}} \frac{1}{3} \sum_{k=1}^3 (\cos \theta_{m,k} - 0.5)^2\)$ Coupled with photometric error \(E_{pho}\) and depth error \(E_{dep}\), these geometric constraints ensure that individual triangle primitives adhere tightly to physical object surfaces without folding or self-intersecting.

4. Adaptive window-based densification, pruning, and on-the-fly restricted Delaunay meshing

At keyframe creation, triangles are spawned around back-projected depth points with their circumscribed sphere radius \(r_i\) matched to the distance of nearest neighbors. During mapping, the system replays a selection of 7 keyframes (4 highest co-visibility, 2 random past keyframes, and the latest keyframe), each iterated for 30 steps. Triangles whose mean opacity falls below \(\epsilon_o\) or whose projected screen area exceeds \(\epsilon_a\) are pruned. Conversely, triangles whose screen coverage exceeds blur threshold \(\theta_{blur} \cdot H \cdot W\) are subdivided into 4 smaller equilateral children via edge midpoint interpolation. When downstream interaction or collision analysis is required, Restricted Delaunay Triangulation connects the optimized triangle soup into a continuous manifold mesh in seconds—bypassing voxel grids and Marching Cubes entirely—while directly facilitating interactive mesh deformation and physics simulation.

Loss & Training

The tracking loss balances photometric fidelity and depth error: $\(E_{track} = E_{pho} + \lambda_{dep} E_{dep}\)$ where \(E_{pho} = (1 - \lambda_{ssim}) \|I(V, \mathbf{T}_{CW}) - \bar{I}\|_1 + \lambda_{ssim} L_{\text{D-SSIM}}(I(V, \mathbf{T}_{CW}), \bar{I})\). The mapping stage minimizes a combined loss over the keyframe replay window: $\(E_{map} = E_{pho} + \lambda_{dep} E_{dep} + \lambda_{norm} E_{norm} + \lambda_{equi} E_{equi}\)$ Jointly updating keyframe poses, vertex positions, colors, and opacities guarantees global map consistency across viewpoints.

Key Experimental Results

Main Results

Quantitative 3D reconstruction quality (Chamfer Distance in cm evaluated with 1M samples, \(\downarrow\)) and mesh generation wall-clock time (seconds, \(\downarrow\)) on the Replica dataset:

Method / Representation Extraction Mode r0 r1 r2 o0 o1 o2 o3 o4 Avg. Chamfer (cm) ↓ Avg. Time (s) ↓
MonoGS TSDF Fusion 3.76 4.39 4.63 2.93 4.78 4.18 4.36 3.26 4.03 -
MonoGS-2D* TSDF Fusion 1.93 1.16 1.54 1.56 0.70 1.49 1.37 1.15 1.36 -
Ours TSDF Fusion 1.00 0.77 1.01 0.66 0.56 1.12 1.69 0.78 0.95 33.44
Ours Delaunay (pruned) 1.16 1.01 1.09 0.77 0.69 1.37 2.00 1.02 1.14 15.66
Ours Delaunay (unpruned) 1.99 1.35 1.41 1.23 1.18 1.55 2.38 1.28 1.55 11.18

Absolute Trajectory Error (ATE RMSE in cm, \(\downarrow\)) on the TUM-RGBD benchmark:

Input Loop Closure Method fr1/desk fr2/xyz fr3/office Avg. ATE (cm) ↓
RGB-D w/o iMAP 4.90 2.00 5.80 4.23
RGB-D w/o NICE-SLAM 4.26 6.19 3.87 4.77
RGB-D w/o Co-SLAM 2.40 1.70 2.40 2.17
RGB-D w/o Point-SLAM 4.34 1.31 3.48 3.04
RGB-D w/o MonoGS 1.50 1.44 1.49 1.47
RGB-D w/o MonoGS-2D* 1.58 1.20 1.83 1.54
RGB-D w/o Ours 1.77 1.12 1.83 1.57
RGB-D w/ BAD-SLAM 1.70 1.10 1.70 1.50
RGB-D w/ ORB-SLAM2 1.60 0.40 1.00 1.00

System execution latency, memory footprint, and primitive count across TUM and Replica sequences:

Metric fr1/desk fr2/xyz fr3/office Replica office1
Frame Rate (FPS) ↑ 0.82 2.33 1.29 0.55
Total Per-frame Latency (ms) ↓ 1225 429 775 1809
Frontend Tracking Latency (ms) 584 352 494 892
Backend Mapping Latency (ms) 642 77 282 918
Map Primitive Count (Triangles) 29.0k 24.0k 34.4k 152.2k
Model Storage Size (MB) ↓ 5.0 4.4 7.6 16.4
Peak GPU Memory Usage (GB) ↓ 0.47 0.47 0.53 1.25

Key Findings

  • Significant geometric accuracy gains over Gaussian baselines: across all 8 Replica scenes, the proposed method substantially outperforms MonoGS (4.03 cm) and MonoGS-2D* (1.36 cm), achieving a 0.95 cm Chamfer distance when coupled with TSDF fusion and 1.14 cm via direct pruned Delaunay extraction.
  • Favorable meshing efficiency vs. accuracy trade-off: pruned Delaunay meshing completes in just 15.66 seconds—less than half the 33.44 seconds required by TSDF fusion—while yielding competitive geometric precision. Unpruned Delaunay is even faster (11.18 s), introducing only minor artifacts in unobserved peripheral regions.
  • Robust tracking with minimal memory footprint: the system attains a 1.57 cm average ATE on TUM-RGBD without loop closures, matching SOTA Gaussian SLAM systems while maintaining an exceptionally compact map size (24k–152k triangles, 4.4–16.4 MB model size, and under 1.25 GB peak VRAM).

Highlights & Insights

  • Unified primitive representation: this work establishes the first dense SLAM system relying directly on a differentiable triangle soup, harmonizing real-time photorealistic novel-view synthesis with explicit mesh extraction without neural implicit intermediaries.
  • Analytical SE(3) Jacobian engineering: moving the camera manifold derivatives into fused CUDA kernels eliminates the computational burden of PyTorch autograd graphs, demonstrating that triangle rasterization can achieve competitive tracking speeds.
  • Effective geometric stabilization: combining vertex opacity averaging, equilateral angle constraints, and sensor normal supervision prevents the triangle soup from collapsing into degenerate slivers during gradient-based optimization.

Limitations & Future Work

  • Imperfect topological guarantees: Delaunay extraction in poorly constrained, sparsely viewed regions can still produce non-manifold boundaries or self-intersections; incorporating explicit topological manifold constraints during optimization remains an open challenge.
  • Sub-real-time throughput: frame rates currently range between 0.55 and 2.3 FPS (per-frame latencies of 429–1809 ms), falling short of 30 FPS hard real-time; multi-process parallelization, conjugate gradient solvers, and sparse bundle adjustment are identified as future optimizations.
  • Reliance on RGB-D sensors: the tracking frontend and normal supervision assume calibrated depth input; extending the pipeline to pure monocular video will require integrating strong feed-forward geometric priors such as MASt3R.
  • vs Gaussian Splatting SLAM (MonoGS / SplaTAM): 3D Gaussian SLAM relies on volumetric ellipsoids requiring density thresholding and Marching Cubes for surface recovery, which creates floating artifacts and prevents live topological deformation; Triangle Splatting SLAM maintains explicit polygonal surfaces natively compatible with standard graphics assets.
  • vs 2D Gaussian Splatting (MonoGS-2D / 4DTAM): while 2D surfels or flattened Gaussians constrain normals along flat disks, they remain detached primitives lacking connectivity; triangles directly connect into shared-vertex meshes suitable for rigid/non-rigid physics and collision detection.
  • vs Mesh Splatting (Held et al., CVPR 2026): Mesh Splatting was engineered for offline reconstruction with min-opacity routing and fixed shared topology; this work adapts the formulation to streaming incremental SLAM via average-opacity routing, adaptive Delaunay meshing, and equilateral regularization.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering adoption of differentiable triangles as the foundational primitive for dense visual SLAM.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous evaluation across trajectory accuracy, depth error, Chamfer distance, primitive scale, and runtime trade-offs on TUM and Replica.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical derivations, coherent system design narrative, and transparent ablation reporting.
  • Value: ⭐⭐⭐⭐⭐ Bridges the long-standing divide between differentiable radiance field SLAM and explicit mesh-based simulation and robotics pipelines.