Skip to content

TopoGS: Planar Reconstruction via Topology-Aware 3D Gaussian Splatting

Conference: ECCV 2026
Paper: ECCV Official Link
Full Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-5769.txt
Area: 3D Vision
Keywords: 3D Gaussian Splatting, Planar Reconstruction, Topological Constraints, Indoor Scene Reconstruction, Multi-View Stereo

TL;DR

TopoGS couples 2D region adjacency graphs with 3D Gaussian Splatting via a tri-consistency optimization framework incorporating uncertainty-aware topological alignment, filtering view-dependent spurious adjacencies and producing sharp, topologically connected 3D planar models.

Background & Motivation

Recovering precise, editable 3D parametric representations from raw multi-view images is a fundamental pursuit across computer vision, computer graphics, robotics, and AR/VR. For human-engineered environments and architectural spaces, the ideal abstraction is a structured planar model—a mathematically concise and interconnected representation defined by parametric planes, their intersection line segments, and corner vertices. Historically, recovering such structured models has largely relied on decoupled, two-stage Multi-View Stereo (MVS) pipelines, which first reconstruct a dense point cloud or mesh and subsequently extract planes via RANSAC clustering and spatial partitioning. In typical indoor scenes featuring textureless walls, reflective surfaces, and variable illumination, MVS algorithms frequently generate noisy, outlier-ridden, or fragmented point clouds. Because this workflow is strictly decoupled and unidirectional, errors in the front-end geometry propagate catastrophically to the back-end plane fitting, while the rich photometric cues and multi-view consistency of the original images are entirely discarded.

Recently, 3D Gaussian Splatting (3DGS) combined with monocular geometric foundation models (such as Metric3D) has emerged as an attractive alternative, enabling direct, end-to-end recovery of planar primitives. Approaches like PlanarSplatting, PlanarGS, and PGS use normal regularization or planar Gaussian primitives to improve surface smoothness in weak-texture areas. Nonetheless, existing 3DGS-based methods continue to treat individual planes as isolated, independent geometric entities, completely lacking topological connectivity. In real-world environments, physical planes are governed by strict spatial dependencies, such as intersecting junctions between adjoining walls or contact seams between floors and walls. Without global topological awareness, optimization in textureless regions easily falls into scale-deviant local optima, yielding misaligned boundaries, floating floaters, and fractured plane fragments.

Lifting 2D image segmentations directly into 3D topology is fraught with ambiguity: perspective occlusions and depth discontinuities create numerous false-positive adjacencies where two regions overlap in 2D projections without any physical 3D contact. Blindly enforcing topological constraints would distort the underlying geometry. This paper's angle of attack is to anchor 3D Gaussian primitives to multi-view 2D planar segmentations and optimize geometry, appearance, and topology jointly within an iterative 3DGS loop, employing an uncertainty-aware weighting mechanism to naturally filter out view-dependent spurious adjacencies. Core idea: integrate 2D segmentation topological graphs into a 3DGS differentiable rendering loop under a photometric-geometric-topological tri-consistency framework, utilizing an uncertainty-aware weighting filter to snap genuine intersection lines together and assemble a globally connected, watertight-tending planar model.

Method

Overall Architecture

The objective of TopoGS is to recover a coherent scene structure represented by a connected planar model comprising a vertex set \(\mathcal{V} \subset \mathbb{R}^3\), an intersection edge set \(\mathcal{E}\), and a set of bounded planes \(\mathcal{P} = \{(\pi_k, \mathcal{E}_k)\}\) with plane parameters \(\pi_k = (\mathbf{n}_k, d_k)\). As illustrated below, the framework proceeds in three sequential stages: (1) Geometry & Topology Initialization, which segments 2D planar regions, fits initial 3D plane parameters, and constructs a candidate 2D adjacency graph; (2) Joint Optimization, which embeds plane-anchored Gaussians within 3DGS and optimizes appearance, geometry, and topology via a staged loss schedule; and (3) Structural Assembly, which prunes false-positive adjacencies, truncates 3D intersection lines using shared corners, and fuses multi-view boundaries into a structured mesh.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multi-View Images & Poses<br/>Metric3D Depth & Normal Estimation"] --> B["Hierarchical Geometry & Topology Initialization<br/>SAM Segmentation + Normal Mean-Shift + 2D Line Fitting"]
    B --> C["Plane-Anchored Gaussian Initialization & Tri-Consistency Optimization<br/>Grid Sampling Back-Projection + Staged Optimization Schedule"]
    C --> D["Uncertainty-Aware Topological Regularization<br/>Analytical 3D Intersection Projection + Distance-Variance Weighting"]
    D --> E["Multi-View Boundary Fusion & Structural Assembly<br/>Adjacency Verification + Line Intersection Truncation + Closed Mesh Assembly"]
    E --> F["Connected Structured Planar Model<br/>Parametric Planes, Bounding Edges & Corner Vertices"]

Key Designs

1. Hierarchical Geometry & Topology Initialization: Combining Semantic Masks with Normal Clustering To eliminate boundary ambiguities and jagged edges in textureless regions, initialization adopts a two-stage segmentation strategy. First, the Segment Anything Model (SAM) produces semantic masks that outline architectural boundaries such as wall-to-wall junctions. Second, because single semantic entities often span multiple distinct planes, Mean-shift clustering is executed within each mask based on predicted surface normal maps from Metric3D, subdividing masks into homogeneous planar regions \(s_i \in \mathcal{S}\). For each region, pixel coordinates are back-projected into 3D world space using predicted depth, and initial plane parameters \(\pi_i = (\mathbf{n}_i, d_i)\) are fitted using RANSAC. For topology, shared boundary pixels \(B_{ij}\) between adjacent regions \((s_i, s_j)\) are extracted and fitted with 2D straight lines. Pairs with severe non-linear residuals or insufficient contact length are discarded, yielding an initial 2D adjacency graph \(\mathcal{A}_{2D}\).

2. Plane-Anchored Gaussian Initialization & Tri-Consistency Optimization: Coarse-to-Fine Staged Guidance To bridge parametric planes with differentiable rendering, Gaussian primitives are anchored to 2D planar regions by uniformly sampling pixels with a stride of \((H/4, W/4)\) across each view and initializing their 3D means via back-projection. During training, in addition to standard photometric loss \(\mathcal{L}_{rgb}\) and scale regularization \(\mathcal{L}_{scale}\), a dedicated geometry loss \(\mathcal{L}_{geo}\) is enforced. This loss penalizes deviations between the rendered depth \(\hat{D}(p)\) and normal \(\hat{\mathbf{N}}(p)\) and the analytical plane depth \(D_{\pi}(p)\) and normal \(\mathbf{N}_{\pi}(p)\): $\(\mathcal{L}_{geo} = \sum_{p \in \mathcal{I}_k} \left( \lambda_1 |\hat{D}(p) - D_{\pi}(p)| + \lambda_2 \|\hat{\mathbf{N}}(p) - \mathbf{N}_{\pi}(p)\|_2 \right)\)$ To ensure stable convergence, optimization follows a two-stage schedule: Stage 1 (iterations 0–4k) applies \(\mathcal{L}_{rgb}\), \(\mathcal{L}_{scale}\), and \(\mathcal{L}_{geo}\) to establish solid 3D geometry; Stage 2 (iterations 4k–6k) introduces the topology loss \(\mathcal{L}_{topo}\) to snap adjacent planes into place. Furthermore, an adaptive plane control step executes every 2,000 iterations to merge co-planar overlapping regions and prune low-confidence primitives.

3. Uncertainty-Aware Topological Regularization: Exponential Decay Weighting for False-Positive Disambiguation To resolve false-positive 2D adjacencies caused by perspective occlusions (e.g., a foreground object visually touching a distant wall), the framework introduces an uncertainty-aware topology alignment mechanism. For any adjacent pair \((s_i, s_j) \in \mathcal{A}_{2D}\), their analytical 3D intersection line \(L_{ij}\) is defined by the intersection of plane equations \(\mathbf{n}_i \cdot \mathbf{X} + d_i = 0\) and \(\mathbf{n}_j \cdot \mathbf{X} + d_j = 0\). Projecting this 3D line onto the image plane yields a 2D line \(l_{ij}\). The mean Euclidean distance between all shared boundary pixels \(p \in B_{ij}\) and \(l_{ij}\) is computed as \(\mu_{ij} = \frac{1}{|B_{ij}|} \sum_{p \in B_{ij}} d(p, l_{ij})\). To prevent erroneous occlusion boundaries from corrupting genuine geometry, an exponential decay weight \(w_{ij}\) is applied: $\(w_{ij} = \exp\left( - \frac{\mu_{ij} \sigma_{ij}^2}{\tau} \right)\)$ where \(\sigma_{ij}^2\) denotes the variance of boundary pixel distances to \(l_{ij}\), and \(\tau\) is a temperature hyper-parameter set to 50.0. The overall topology loss is: $\(\mathcal{L}_{topo} = \sum_{(s_i, s_j) \in \mathcal{A}_{2D}} w_{ij} \cdot \mu_{ij}\)$ When regions are physically separated in 3D (high \(\mu_{ij}\)) or boundaries are irregular/curved (high \(\sigma_{ij}^2\)), \(w_{ij}\) decays rapidly to zero. This allows the optimization loop to naturally disregard spurious occlusions and only enforce structural alignment on genuine physical intersections.

4. Multi-View Boundary Fusion & Structural Assembly: Converting Infinite Intersections into Bounded Meshes After joint optimization, discrete primitives and verified relationships are assembled into a coherent, connected 3D polygonal model. First, by applying a strict distance threshold on converged \(\mu_{ij}\), spurious adjacencies are pruned to yield a verified 3D adjacency graph \(\mathcal{A}_{verified}\). Next, shared corners between 2D boundary segments \(B_{ij}\) are identified: when two segments share an image corner, their corresponding 3D intersection lines are treated as adjacent. Their 3D intersection point is analytically calculated to serve as a corner vertex, truncating infinite intersection lines \(L_{ij}\) into finite bounding segments. Finally, plane boundary segments observed across multiple views are aggregated into a unified global boundary and trimmed against the verified 3D intersection lines and corner vertices, eliminating inter-plane gaps and producing a cleanly connected structured model.

Loss & Training

The overall joint optimization loss function evolves dynamically across training phases: $\(\mathcal{L} = \begin{cases} \lambda_{rgb} \mathcal{L}_{rgb} + \lambda_{scale} \mathcal{L}_{scale} + \mathcal{L}_{geo}, & t < 4000 \\ \mathcal{L}_{\text{Stage1}} + \lambda_{topo} \mathcal{L}_{topo}, & 4000 \le t \le 6000 \end{cases}\)$ Here, the photometric loss balances \(L_1\) color error with D-SSIM structural similarity, and the scale regularization term \(\mathcal{L}_{scale}\) penalizes the maximum elongation axis of Gaussian ellipsoids to suppress floating artifacts in free space. In tandem with the loss progression, the adaptive plane control step executes every 2,000 iterations to evaluate the angular cosine difference and normal offsets of nearby planes, merging co-planar fragments into single entities and pruning sparse low-confidence primitives before Stage 2 concludes. This ensures that only compact and geometrically consistent planar primitives enter the final structural assembly phase.

Key Experimental Results

Main Results

Quantitative evaluations on the ScanNet++ dataset across 25 diverse indoor scenes compare TopoGS against two-stage pipelines (PlanarGS and Depth-Anything-V3 with AirPlanes RANSAC extraction) and direct planar reconstruction baselines (NeuralPlane and PlanarSplatting).

Method Geometry CD \(\downarrow\) Geometry F-score \(\uparrow\) Planar Fidelity \(\downarrow\) Planar Acc \(\downarrow\) Planar CD \(\downarrow\) Seg VOI \(\downarrow\) Seg RI \(\uparrow\) Seg SC \(\uparrow\)
PlanarGS + AirPlanes 14.02 23.77 24.74 19.25 22.00 3.400 0.926 0.379
DA3 + AirPlanes 8.49 74.40 20.35 12.85 16.60 3.041 0.919 0.431
NeuralPlane 7.56 53.83 14.50 11.83 13.17 2.872 0.944 0.464
PlanarSplatting 9.06 53.83 15.07 14.06 14.56 2.869 0.933 0.463
TopoGS (Ours) 6.01 70.18 12.34 9.96 11.15 2.933 0.946 0.463
TopoGS (+ Assembly) 6.09 69.80 12.96 10.91 11.94 2.867 0.946 0.472

Ablation Study

The impact of geometry loss and topology loss was quantitatively investigated on ScanNet++, reporting global geometry metrics, planar accuracy, and topological F-scores for vertices (V), edges (E), and planes (P):

Config Geometry CD \(\downarrow\) Geometry F-score \(\uparrow\) Planar Fidelity \(\downarrow\) Planar Acc \(\downarrow\) Planar CD \(\downarrow\) Topo V F-score \(\uparrow\) Topo E F-score \(\uparrow\) Topo P F-score \(\uparrow\) Note
w/o Topo 6.54 64.98 13.98 10.96 12.47 19.62 32.05 34.22 Removing topology loss degrades structural alignment; edge F-score drops by 3.13%
w/o Geom 7.56 62.32 16.50 10.66 13.58 12.72 21.28 33.38 Removing geometry loss causes severe depth drift and floating planes; vertex F-score plummets
Full loss (Full model) 6.01 70.18 12.34 9.96 11.15 21.52 35.18 36.55 Synergistic tri-consistency achieves optimal geometry and topological connectivity

Key Findings

  • Regularization role of topological constraints: Excluding topological constraints ("w/o Topo") causes clear intersection misalignment and raises Chamfer Distance from 6.01 to 6.54, while topological edge recall drops from 35.18% to 32.05%. This demonstrates that explicit topological constraints act as essential global regularizers, correcting scale and orientation drift in featureless regions.
  • Geometric loss prevents 2D topological overfitting: The "w/o Geom" variant displays deceptively low 2D projection error but exhibits severe 3D depth deviation, resulting in a poor Planar Fidelity of 16.50 and a catastrophic drop in vertex F-score to 12.72%. This confirms that 3D geometric loss provides the necessary anchor points to prevent 2D boundary constraints from distorting global geometry.
  • Superiority over state-of-the-art baselines: TopoGS significantly surpasses the leading direct reconstruction baseline NeuralPlane in Planar Fidelity (12.34 vs. 14.50) and Planar Accuracy (9.96 vs. 11.83), while resolving the error accumulation and boundary fragmentation problems of two-stage MVS pipelines.

Highlights & Insights

  • Exponential uncertainty weighting for automatic occlusion pruning: Weighting projected intersection residuals with an exponential function of distance mean and variance allows the optimization to filter out view-dependent perspective occlusions without manual thresholding, avoiding geometric distortion.
  • Bidirectional synergy between geometry and topology: Geometric constraints ground topological optimization in Euclidean 3D space, while topological closure provides long-range structural guidance that breaks scale ambiguities across textureless surfaces.
  • Skeletal intersection-to-polygon assembly: Deriving corner vertices from shared 2D segment endpoints and truncating analytical 3D intersection lines turns unorganized Gaussian clouds into clean, CAD-ready parametric planar meshes suitable for graphics rendering and physical simulation.

Limitations & Future Work

  • Non-watertight models in occluded regions: Due to camera trajectories and physical occlusions, unobserved corners or seams may lack sufficient multi-view visibility, resulting in small boundary gaps. Incorporating generative 3D shape priors could help hallucinate missing structures.
  • Sensitivity to 2D segmentation errors: When SAM masks or normal-based Mean-shift clustering fail on intricate details (e.g., complex moldings or small objects), redundant or noisy intersection lines can be introduced into the assembly phase.
  • Restriction to planar surfaces: The current framework assumes strictly planar surfaces and cannot natively parameterize curved walls, cylinders, or free-form furniture, which requires future extension to curved surface primitives.
  • vs PlanarSplatting (CVPR 2025): PlanarSplatting substitutes 3D Gaussians with planar discs to accelerate reconstruction, but it treats primitives independently, leaving jagged and unaligned boundary seams. TopoGS introduces explicit topological graph optimization, producing sharp, continuous intersection edges.
  • vs NeuralPlane (ICLR 2025): NeuralPlane optimizes planar primitives within an implicit neural field, where extracting topological connectivity and boundary corners requires cumbersome post-processing. TopoGS directly formulates analytical 3D intersection lines and corner vertices in 3DGS, yielding CAD-structured polygonal outputs.
  • vs AirPlanes (CVPR 2024): AirPlanes relies on an offline two-stage clustering and plane-fitting protocol that is vulnerable to noisy front-end point clouds. TopoGS operates end-to-end, enabling photometric and geometric multi-view cues to refine plane parameters jointly.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering integration of multi-view 2D topological graphs into 3D Gaussian Splatting for structured planar reconstruction.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 25 ScanNet++ scenes covering geometry, planar quality, segmentation, and topological recall metrics with detailed ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear narrative, rigorous mathematical formulations, and compelling ablation visualizations.
  • Value: ⭐⭐⭐⭐⭐ Highly valuable for indoor scene modeling, architectural CAD reconstruction, game asset generation, and robotic spatial reasoning.