Graph-GSReg: Leveraging 3D Scene Graphs for Gaussian Splatting Registration¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://lee-jaewon.github.io/Graph-GSReg/
Area: 3D Vision
Keywords: Gaussian Splatting, Scene Graph Registration, Scene Merging, Maximum Clique, Test-Time Optimization
TL;DR¶
Lifts unstructured 3D Gaussian Splatting scenes into object-level 3D scene graphs, solves robust rigid alignment via TRIMs distance consistency and maximum clique search without pretraining, and seamlessly eliminates occlusion artifacts and floaters through self-supervised test-time optimization.
Background & Motivation¶
In large-scale 3D mapping, robot navigation, and virtual reality, 3D Gaussian Splatting (3DGS) has rapidly emerged as a preferred map representation due to its real-time rendering speed and photorealistic visual fidelity. However, capturing entire real-world environments in a single continuous trajectory is practically infeasible due to physical occlusions, sensor range limits, and dynamic obstacles. Moreover, scaling end-to-end optimization of massive Gaussian primitives across large areas leads to prohibitive memory consumption and numerical instability. Dividing environments into independently reconstructed local submaps and subsequently registering and merging them into a globally consistent coordinate frame is thus indispensable for long-term large-scale mapping.
Despite its importance, existing 3DGS registration and merging paradigms face severe operational bottlenecks. On one hand, deep learning-based point cloud registration frameworks such as GaussReg depend on feature extraction networks trained on massive 3DGS datasets, limiting generalizability across unseen domains; furthermore, GaussReg merges overlapping scenes using heuristic distance thresholding relative to scene centers, which often mistakenly discards critical regions when scenes vary in scale or geometry, creating severe hollow regions. On the other hand, foundation model-assisted image matching frameworks like PhotoReg rely on 2D models such as DUSt3R across rendered image pairs followed by costly photometric optimization lasting hundreds of seconds; coarse failures under weak textures or large viewpoint variations lead directly to optimization divergence, while simply concatenating raw primitives introduces prominent floaters and redundant occlusions.
The core tension stems from the unstructured nature of low-level 3DGS primitives—carrying only position, covariance, opacity, and spherical harmonic colors—which lack high-level semantic distinctiveness and structural context for reliable matching under partial overlap. To overcome this limitation, this paper lifts low-level Gaussian primitives into 3D scene graphs with object centroids and topological context embeddings, transforming the task into a robust graph compatibility problem resolved via maximum clique search, complemented by self-supervised test-time optimization. Core idea: Lift low-level Gaussian primitives into object-level 3D scene graphs enriched with topological neighborhood embeddings, formulate rigid registration via pairwise distance consistency and maximum clique search on a compatibility graph, and seamlessly merge scenes through self-supervised test-time optimization guided by original view re-renderings without supervised training.
Method¶
Overall Architecture¶
Graph-GSReg accepts two independently reconstructed 3DGS submaps \(S^A=(G^A, C^A)\) and \(S^B=(G^B, C^B)\), consisting of Gaussian primitive sets \(G\) and camera poses \(C\). The framework first renders images from a uniformly sampled subset of viewpoints, detects 2D object masks via SAM, extracts CLIP visual embeddings, and projects the masks into 3D space to cluster Gaussians into consistent object nodes across frames. Candidate node correspondences are then established using enriched semantic-topological embeddings, and a compatibility graph is constructed using translation and rotation invariant measurements (TRIMs) to enforce pairwise rigidity. A maximum clique search extracts the largest mutually consistent inlier set, from which an initial rigid pose \(T_{\text{graph}} \in SE(3)\) is determined via SVD and refined by ICP on Gaussian centroids. Finally, the aligned scenes are voxelized at 0.01m resolution to prune physical duplicates, followed by a fast self-supervised test-time optimization that minimizes photometric \(L_1\) discrepancy against original re-renderings to eliminate floating artifacts and occlusions.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Independent 3DGS Scene Pair (S^A, S^B)"] --> B["Cross-Frame Object Association & Contextual Scene Graph Construction<br/>Render sampled views + SAM/CLIP extraction + 3D GIoU association + 3-hop local histogram"]
B --> C["TRIMs-based Compatibility Filtering & Maximum Clique Registration<br/>Semantic cosine similarity + pairwise distance ratio graph + maximum clique inliers + SVD/ICP"]
C --> D["Voxelized Pruning & Self-Supervised Test-Time Optimization Merging<br/>SE(3) aligned union + 0.01m spatial voxelization + L1 photometric refinement vs original views"]
D --> E["Output: Seamless Unified Merged 3DGS Scene"]
Key Designs¶
1. Cross-Frame Object Association & Contextual Scene Graph Construction: Abstracting Primitives into Semantic Nodes
To overcome the lack of semantic awareness and correspondence ambiguity in raw Gaussian primitives, this module lifts low-level primitives into an object-level 3D scene graph using the renderable nature of 3DGS. The pipeline uniformly samples a subset of camera poses \(C^X_{\text{sampled}}\) and renders corresponding RGB images. For each frame, SAM extracts object masks with confidence scores above \(c_{\text{mask}}\), and CLIP encodes the visual features of masked regions. Using camera intrinsics \(K\) and extrinsics \((R_j^X, T_j^X)\), each 3D Gaussian mean \(\boldsymbol{\mu}_i\) is projected to pixel coordinates \(\mathbf{p}_{i,j}^X\). Gaussians falling within mask \(m_{j,k}^X\) in front of the camera (\(z_{i,j}^X > 0\)) form a 3D cluster \(\mathcal{A}_{j,k}^X\), whose mean position defines the 3D centroid \(\mathbf{x}_{j,k}^X\) of the object node: $\(\mathbf{x}_{j,k}^X = \frac{1}{|\mathcal{A}_{j,k}^X|} \sum_{\boldsymbol{\mu}_i \in \mathcal{A}_{j,k}^X} \boldsymbol{\mu}_i\)$ To achieve globally consistent nodes across sequential views, an incremental association is performed: when the sum of CLIP cosine similarity and 3D Generalized Intersection over Union (GIoU) between a newly detected Gaussian cluster and an existing node exceeds threshold \(\tau_{\text{asso}}\), they are merged into the same entity by combining point sets, downsampling, and averaging normalized CLIP embeddings. To further resolve visual ambiguities among objects sharing repetitive appearances (e.g., identical chairs), a 3-hop local neighborhood CLIP histogram \(f_{\text{hist}}(v)\) is constructed over spatially proximate nodes and concatenated with normalized CLIP appearance: \(f(v) = [f_{\text{CLIP}}(v), f_{\text{hist}}(v)]\), capturing both local appearance and structural topological context.
2. TRIMs-based Compatibility Filtering & Maximum Clique Registration: Global Geometric Consensus
To eliminate spurious matches generated from appearance similarities, this module enforces rigid geometric invariance within a graph-theoretic compatibility formulation. Initial candidate matches \(\mathcal{V}_{\text{comp}} = \{(q', p') \mid s(q', p') \ge \tau_{\text{node}}\}\) are generated based on node embedding cosine similarities. Because genuine correspondences under rigid motion strictly preserve Euclidean distances between object pairs across scenes, the framework applies translation and rotation invariant measurements (TRIMs) to construct edge set \(\mathcal{E}_{\text{comp}}\) in the compatibility graph \(G_{\text{comp}}\): $\(((q'_1, p'_1), (q'_2, p'_2)) \in \mathcal{E}_{\text{comp}} \iff 1 - \tau_{\text{TRIMs}} < \frac{\|q'_1 - q'_2\|}{\|p'_1 - p'_2\|} < 1 + \tau_{\text{TRIMs}}\)$ In \(G_{\text{comp}}\), mutually compatible true inliers form a dense, fully connected clique, whereas outlier matches violate pairwise distance constraints with the majority of valid pairs and fail to join large cliques. Solving for the Maximum Clique on \(G_{\text{comp}}\) extracts the largest mutually consistent correspondence set while systematically discarding false positives. A closed-form SVD transformation yields initial rigid pose \(T_{\text{graph}} \in SE(3)\) (or Umeyama algorithm when scale estimation is required), followed by fine-grained ICP alignment over Gaussian primitive centroids to produce final transformation \(T_{AB}\).
3. Voxelized Pruning & Self-Supervised Test-Time Optimization Merging: Artifact-Free Fusion via Source Re-renderings
To resolve visual floaters, occlusions, and redundant densities caused by subtle rigid alignment imperfections, this design avoids destructive geometric thresholding and instead leverages the original submaps as physical ground-truth rendering references. The aligned Gaussian sets \(G^A\) and \(G^{B'} = T_{AB}(G^B)\) are first combined into a naive union \(G^A \cup G^{B'}\). A spatial voxel grid with a resolution of 0.01m is applied to merge overlapping Gaussians within each voxel cell, pruning spatial redundancies and providing a lightweight, stable initialization \(G^{\text{merged}}\). Because true ground-truth images of the merged environment do not exist, the key insight is that original submaps \(S^A\) and \(S^B\) remain renderable from their training camera poses \(C^A\) and \(C^{B'}\). During test-time optimization, merged scene \(G^{\text{merged}}\) is rendered from original viewpoints and optimized to match source renderings \(I_j^X\) by minimizing the photometric \(L_1\) loss: $\(\arg\min_{G^{\text{merged}}} \sum_{j} \left\| \textit{Render}(G^{\text{merged}} \mid R_j^X, T_j^X) - I_j^X \right\|_1, \quad X \in \{A, B'\}\)$ In approximately one minute of optimization, this self-supervised process refines Gaussian opacities, positions, and spherical harmonic colors in overlapping regions, naturally fading out spurious floaters and filling boundary gaps without discarding valid surface geometry.
Loss & Training¶
The registration pipeline is training-free and operates zero-shot using frozen SAM and CLIP models. The self-supervised test-time optimization runs per scene pair for roughly 60 seconds on a single NVIDIA RTX 4070 GPU, updating only the parameters of \(G^{\text{merged}}\) under the photometric \(L_1\) rendering objective across all source viewpoints.
Key Experimental Results¶
Main Results¶
The method is evaluated on the real-world indoor benchmark ScanNet-GSReg (82 partially overlapping scene pairs) and the synthetic multi-room benchmark uHumans2 (69 partially overlapping scene pairs). Evaluation metrics include Relative Rotation Error (RRE, in degrees), Relative Translation Error (RTE), Relative Scale Error (RSE), Absolute Translation Error (ATE, in meters), runtime in seconds, and rendering metrics (PSNR, SSIM, LPIPS).
Table 1: Quantitative comparison of Gaussian Splatting registration on ScanNet-GSReg
| Method | RRE (°) ↓ | RTE ↓ | RSE ↓ | Time (s) ↓ | Note |
|---|---|---|---|---|---|
| GaussReg (Coarse) | 3.403 | 0.061 | 0.034 | 3.700 | Supervised point cloud registration network |
| GaussReg (w/ Fine) | 2.827 | 0.042 | 0.032 | 4.800 | Rendered volumetric feature refinement |
| PhotoReg (Coarse) | 7.825 | 0.082 | - | 16.231 | DUSt3R image-pair initial alignment |
| PhotoReg (w/ Fine) | 7.366 | 0.072 | - | 585.216 | Iterative photometric optimization (~10 min) |
| Liu et al. (Skeleton) | 2.595 | 0.045 | 0.013 | 5.041 | Skeleton alignment and adaptive features |
| Graph-GSReg (Coarse) | 3.247 | 0.039 | 0.013 | 2.784 | 3D scene graph maximum clique matching |
| Graph-GSReg (w/ Fine) | 1.970 | 0.025 | - | 3.658 | Scene graph + Gaussian centroid ICP |
Table 2: Quantitative comparison of Gaussian Splatting registration on uHumans2
| Method | RRE (°) ↓ | RTE ↓ | ATE (m) ↓ | Time (s) ↓ | Failure Mode & Characteristics |
|---|---|---|---|---|---|
| TEASER++ & ICP | 9.099 | 0.071 | 1.669 | 4.570 | Traditional geometric FPFH features fail under noise |
| MAC (FPFH) | 108.089 | 0.652 | 17.118 | 6.359 | Feature mismatching causes massive rotational divergence |
| MAC (FCGF) | 60.021 | 0.335 | 8.862 | 5.462 | Deep geometric features degrade in low-overlap environments |
| PhotoReg (Coarse) | 15.895 | 0.154 | 6.916 | 17.814 | Image-level foundation model fails across wide baselines |
| PhotoReg (w/ Fine) | 9.959 | 0.077 | 2.754 | 342.754 | Trapped in local minima during fine optimization |
| Graph-GSReg (Coarse) | 1.748 | 0.010 | 0.243 | 1.805 | Semantic graph structure provides robust coarse pose |
| Graph-GSReg (w/ Fine) | 0.711 | 0.006 | 0.192 | 3.782 | Sub-degree rotation and sub-20cm translation accuracy |
Table 3: Quantitative comparison of Gaussian Splatting merging quality under fixed registrations
| Registration Source | Merging Strategy | ScanNet PSNR ↑ | ScanNet SSIM ↑ | ScanNet LPIPS ↓ | uHumans2 PSNR ↑ | uHumans2 SSIM ↑ | uHumans2 LPIPS ↓ |
|---|---|---|---|---|---|---|---|
| Oracle (Single Scene) | Upper Bound (No Merge) | 22.2913 | 0.8566 | 0.3347 | 32.0999 | 0.9218 | 0.1416 |
| Ground Truth Pose | PhotoReg (Direct Union) | 20.1588 | 0.8117 | 0.3844 | 20.9777 | 0.7081 | 0.3316 |
| GaussReg (Distance Cut) | 20.9512 | 0.8306 | 0.3551 | 23.3410 | 0.7368 | 0.2871 | |
| Graph-GSReg (Ours) | 21.8954 | 0.8465 | 0.3484 | 25.4032 | 0.7608 | 0.2801 | |
| Graph-GSReg Pose | PhotoReg (Direct Union) | 19.3298 | 0.7926 | 0.4026 | 20.7897 | 0.7048 | 0.3370 |
| GaussReg (Distance Cut) | 20.1133 | 0.8089 | 0.3715 | 23.1519 | 0.7321 | 0.2904 | |
| Graph-GSReg (Ours) | 21.5248 | 0.8304 | 0.3698 | 25.2425 | 0.7607 | 0.2868 |
Ablation Study¶
Ablations on ScanNet-GSReg dissect the contributions of contextual embeddings, geometric constraints, and voxelized fusion.
Table 4: Ablation study of 3D scene graph registration components on ScanNet-GSReg
| Method Variant | RRE (°) ↓ | RTE ↓ | RSE ↓ | Time (s) ↓ | Mechanism Insight |
|---|---|---|---|---|---|
| w/o CLIP Histogram | 4.865 | 0.052 | 0.019 | 2.755 | Missing topological context induces false matches among similar items |
| w/o TRIMs | 6.047 | 0.140 | 0.104 | 2.903 | Lacking pairwise rigidity check admits outliers, doubling RTE |
| w/o Maximum Clique | 6.575 | 0.151 | 0.120 | 2.801 | Inability to extract globally consistent inlier set collapses pose estimation |
| w/ FPFH Features | 3.190 | 0.039 | 0.013 | 3.645 | Marginally improves RRE but introduces point feature overhead |
| CLIP Histogram-5 (5-hop) | 3.867 | 0.045 | 0.015 | 4.341 | Overly distant hops incorporate irrelevant structural noise |
| Maximal Clique (CLIP-Clique) | 3.981 | 0.045 | 0.013 | 2.805 | Selecting single highest-scoring clique yields suboptimal node coverage |
| Ours (Full Model) | 3.247 | 0.039 | 0.013 | 2.784 | Optimal trade-off between structural robustness and computational cost |
Table 5: Ablation study and memory footprint comparison of 3DGS merging
| Setting | PSNR ↑ | SSIM ↑ | LPIPS ↓ | Storage / Size (MB) ↓ | Observation & Impact |
|---|---|---|---|---|---|
| GaussReg (Distance-based) | 20.1905 | 0.8106 | 0.3698 | 148.18 | Heuristic culling discards necessary Gaussians, creating hollows |
| Ours (w/o Voxelization) | 21.4538 | 0.8304 | 0.3675 | 238.10 | Redundant overlapping primitives inflate file size significantly |
| Ours (w/ Voxelization) | 21.5248 | 0.8304 | 0.3698 | 82.84 | Reduces storage by 65.2%, eliminates density conflicts, boosts PSNR |
Key Findings¶
- Geometric compatibility and maximum clique search form the core pillar of outlier rejection: Removing TRIMs or maximum clique search increases rotational error beyond 6° and triples translation error, confirming that high-level visual features alone cannot disambiguate repetitive indoor objects without rigid physical distance invariance.
- Training-free semantic graphs outperform supervised networks on out-of-distribution environments: On the complex multi-room synthetic dataset uHumans2, conventional geometric descriptors (FPFH, FCGF) fail dramatically (RRE of 60° to 108°), whereas Graph-GSReg achieves 0.711° RRE and 0.192m ATE, demonstrating superior zero-shot cross-domain robustness.
- Voxelized pre-culling provides dual benefits in storage and image fidelity: Introducing a 0.01m spatial voxel grid prior to test-time optimization drops merged model memory from 238.10MB to 82.84MB (surpassing GaussReg's 148.18MB) while slightly boosting PSNR by 0.07 dB by resolving photometric density contention.
Highlights & Insights¶
- Lifting continuous primitives to discrete semantic scene graphs: Rather than calculating geometric distances directly over unstructured, noisy Gaussian clouds, projecting Gaussians into mask-associated object nodes enables 3DGS to exploit mature, certifiable graph-theoretic matching algorithms.
- Self-supervised test-time optimization via source view re-rendering: Elegantly addresses the lack of ground-truth merged training data by using the original submaps as physical rendering oracles, resolving visual floaters and hollow boundaries in 60 seconds of test-time optimization.
- Transferable 3-hop contextual neighborhood embeddings: Augmenting local vision-language embeddings with topological neighborhood histograms effectively resolves semantic ambiguities in repetitive structural layouts, presenting a generalizable design for embodied spatial grounding and point cloud loop closure.
Limitations & Future Work¶
- Dependency on offline graph construction preprocessing: Although graph generation is a one-time operation that can be amortized across multiple pairwise registrations, re-rendering viewpoints and invoking 2D foundation models (SAM, CLIP) incurs upfront compute, posing challenges for real-time mobile SLAM backends.
- Performance degradation in textureless or object-sparse scenes: Environments lacking distinct objects (e.g., bare corridors or featureless tunnels) yield sparse or degenerate scene graph nodes, requiring complementary geometric primitives (such as supervoxels or structural wireframes) to maintain graph connectivity.
- Future directions: Integrating incremental graph maintenance into online Gaussian SLAM frameworks and extending self-supervised test-time optimization to multi-modal radiance field fusion (e.g., cross-registering NeRF and 3DGS submaps).
Related Work & Insights¶
- vs GaussReg (ECCV 2024): GaussReg relies on end-to-end supervised training on curated 3DGS datasets and prunes Gaussians via distance from scene centers, which causes hollow tears across varying scene scales; Graph-GSReg is fully training-free and combines voxelization with self-supervised test-time optimization for seamless boundaries.
- vs PhotoReg (arXiv 2024): PhotoReg relies on DUSt3R across image pairs and takes nearly 10 minutes of photometric optimization that frequently falls into local minima, while simply concatenating raw primitives; Graph-GSReg registers scenes in 3.6 seconds via scene graphs and eliminates floaters within 1 minute of self-supervised refinement.
- vs TEASER++ & MAC: Classical point cloud registration approaches fail in partially overlapping indoor spaces due to low-distinctiveness geometric descriptors; Graph-GSReg integrates multi-modal CLIP semantics with structural context to achieve robust global alignment.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ [Pioneering integration of object-level 3D scene graphs into 3D Gaussian Splatting registration with a self-supervised test-time optimization merging scheme.]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive quantitative registration and rendering ablations on both real-world ScanNet-GSReg and large synthetic uHumans2 benchmarks.]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear architecture narrative, rigorous yet restrained mathematical formulations, and thorough empirical analysis.]
- Value: ⭐⭐⭐⭐⭐ [Provides a practical, training-free engineering paradigm for large-scale multi-session robotic mapping and long-term 3DGS scene management.]