MagnetGS-Mesh: High-Quality Multi-Object Mesh Reconstruction via Adaptive Surface Optimization¶
Conference: ECCV 2026
Paper: CVF Open Access
Code: https://github.com/MinsuPark0752/MagnetGS-Mesh
Area: 3D Vision
Keywords: 3D Gaussian Splatting, Mesh Reconstruction, Object Segmentation, Surface Optimization, Normal Consistency
TL;DR¶
Addressing surface holes and outlier scatter caused by object-wise 3DGS decomposition, MagnetGS-Mesh introduces Enhanced Occupancy Learning and Adaptive Local Normal Consistency (ALNC) Loss to magnetically relocate outlier Gaussians into sparse surface areas, enabling high-quality, lightweight, and artifact-free multi-object mesh extraction.
Background & Motivation¶
3D Gaussian Splatting (3DGS) has revolutionized novel view synthesis by modeling scenes as millions of anisotropic 3D Gaussians rendered via real-time differentiable rasterization at over 100 FPS. To convert these explicit point-based representations into polygonal meshes required by downstream editing, physical simulation, and rendering pipelines, recent works have incorporated auxiliary implicit fields, such as Mesh-in-the-Loop (MILo) occupancy fields or Gaussian Opacity Fields (GOF). However, existing 3DGS-to-mesh frameworks optimize entire scenes holistically, meshing all Gaussians indiscriminately without distinguishing foreground objects from background clutter. This inevitably yields bloated, heavy meshes crowded with redundant vertices in irrelevant regions, hindering per-object manipulation, relighting, and animation.
An intuitive solution is to decompose 3DGS representations into individual objects via 3D segmentation before extracting meshes. Nonetheless, decomposing scenes introduces severe geometric degradation that existing pipelines cannot accommodate: imperfect segmentation boundaries scatter numerous outlier Gaussians far from true surfaces, while the remaining surface Gaussians exhibit highly non-uniform spatial density, clustering heavily in textured regions while leaving high-curvature edges and thin structures sparsely populated. Because robust mesh extraction algorithms (such as tetrahedral meshing) fundamentally rely on uniform Gaussian coverage, this density disparity results in extensive surface holes, non-manifold artifacts, and broken topologies.
Existing surface regularization and normal consistency approaches uniformly enforce smoothness across fixed neighborhood windows, which either blurs sharp geometric features in dense regions or fails under noise in sparse regions. This paper approaches the problem from a novel perspective: rather than pruning outlier Gaussians as useless noise, they can be utilized as geometric material to fill surface voids under the guidance of local surface geometry. Core idea: through a geometry-aligned two-stage optimization procedure, Enhanced Occupancy Learning first sharpens ambiguous object boundaries into crisp binary transitions, and the Adaptive Local Normal Consistency (ALNC) Loss subsequently acts as a magnet to relocate outlier Gaussians into sparse surface regions, achieving high-fidelity, hole-free, and lightweight multi-object mesh reconstruction.
Method¶
Overall Architecture¶
MagnetGS-Mesh takes calibrated multi-view images as input and produces high-quality, independent object-wise triangle meshes along with a fully assembled scene mesh. The architecture operates through two geometry-aligned stages: Stage 1 performs Enhanced Occupancy Learning by augmenting MILo's hybrid representation with five synergistic geometric regularization terms, compressing ambiguous occupancy bands into sharp binary transitions to facilitate clean 3D object segmentation via SAGA and precise surface/outlier partitioning. Stage 2 executes ALNC-Guided Optimization: the first 30% of iterations pull outliers close to object surfaces via proximity constraints, while the remaining 70% of iterations apply the Adaptive Local Normal Consistency (ALNC) Loss with density-adaptive neighborhood scaling, driving outliers along coherent surface normals into sparse regions under photometric guidance. Finally, marching tetrahedra is applied over each object's continuous occupancy field to extract clean, artifact-free polygonal meshes.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multi-view Image Input"] --> B["Stage 1: Enhanced Occupancy Learning<br/>Fivefold Regularization for Crisp Binary Boundaries"]
B --> C["Stage 1: Object Segmentation & Surface Detection<br/>SAGA Decomposition + Dual Scoring"]
C --> D["Stage 2: Proximity-Driven Attachment<br/>First 30% Iterations: Coarse Surface Pulling"]
D --> E["Stage 2: Adaptive Local Normal Consistency Loss<br/>Density-Adaptive k + Magnet Relocation"]
E --> F["Mesh Extraction & Scene Assembly<br/>Marching Tetrahedra + Per-Object & Full Meshes"]
Key Designs¶
1. Enhanced Occupancy Learning: Sharpening Ambiguous Boundaries for Crisp Decomposition Vanilla MILo produces soft, ambiguous occupancy fields with values hovering between 0.3 and 0.7 around object boundaries, which confuses 3D segmentation and leaves scattered floaters. This design introduces five synergistic regularization losses during joint radiance field and occupancy field optimization. An entropy minimization loss \(\mathcal{L}_{\text{entropy}}\) forces occupancy values \(o_i\) toward binary extremes (0 for free space and 1 for object interior). A spatial smoothness term \(\mathcal{L}_{\text{smooth}}\) penalizes differences across 3-nearest neighbors to enforce local structural coherence. A consistency loss \(\mathcal{L}_{\text{consist}}\) dynamically aligns opacity \(\alpha_i\) with occupancy \(o_i\). A boundary sharpening term \(\mathcal{L}_{\text{boundary}}\) maximizes occupancy gradients near the isosurface (\(o_i \approx 0.5\)) while suppressing them elsewhere to create step-like geometric transitions. An exponential moving average (EMA) temporal loss \(\mathcal{L}_{\text{temporal}}\) stabilizes training dynamics. Consequently, 80% to 90% of Gaussians settle into decisive interior or exterior values, leaving a compact 10% to 20% boundary band that enables SAGA to segment objects with minimal spatial contamination.
2. Dual-Criterion Surface Detection: Accurate Separation of Surface Anchors and Free Outliers Within each segmented object cluster, residual floaters persist. This design establishes a dual-criterion scoring mechanism combining occupancy geometry and multi-view photometric consistency to partition boundary-layer Gaussians into reliable surface anchors \(\mathcal{S}\) and relocatable outliers \(\mathcal{O}\). The occupancy-based score measures deviation from the theoretical \(o=0.5\) isosurface: $\(s_i^{\text{occ}} = \exp\left(-\frac{(o_i - 0.5)^2}{2\sigma_{\text{occ}}^2}\right)\)$ Simultaneously, the photometric consistency score projects each Gaussian across all visible camera views \(\mathcal{V}\) and calculates the intersection-over-union (IoU) between the rendered mask \(R_v\) and the ground-truth segmentation mask \(M_v\), recording the fraction of views exceeding threshold \(\tau\): $\(s_i^{\text{photo}} = \frac{1}{|\mathcal{V}|} \sum_{v \in \mathcal{V}} \mathbb{I}(\text{IoU}(R_v, M_v) > \tau)\)$ The composite surface score is formed by \(s_i = w_{\text{occ}} s_i^{\text{occ}} + w_{\text{photo}} s_i^{\text{photo}}\). Gaussians scoring above threshold \(\theta\) are categorized as surface anchors \(\mathcal{S}\), whereas the remainder constitute outlier Gaussians \(\mathcal{O}\) slated for magnetic relocation.
3. Proximity-Driven Attachment: Coarse Spatial Pre-Alignment During the initial 30% of Stage 2 iterations, outlier Gaussians are scattered far from true surfaces, where computing higher-order surface normal guidance risks being trapped in local geometric minima. This design enforces an aggressive proximity loss that penalizes squared Euclidean distances to the nearest surface point: $\(\mathcal{L}_{\text{prox}} = \frac{1}{N_{\text{out}}} \sum_{i \in \mathcal{O}} \min_{j \in \mathcal{S}} \|\boldsymbol{\mu}_i - \boldsymbol{\mu}_j\|^2\)$ This objective acts as a coarse gravitational field, rapidly compressing the spatial gap between outliers and the object's hull and pulling floating Gaussians into the immediate vicinity of the true surface prior to fine-grained normal alignment.
4. Adaptive Local Normal Consistency Loss: Density-Adaptive Normal-Guided Relocation Serving as the primary technical contribution, the ALNC loss governs the remaining 70% of Stage 2 iterations, executing magnetic relocation of outliers into underpopulated surface voids. The method first computes the mean distance \(d_i\) from outlier \(i\) to its \(k\)-nearest surface points to estimate local surface density \(\rho_i = \frac{1}{d_i + \epsilon}\). Based on \(\rho_i\), an adaptive neighborhood size \(k_i\) is determined: $\(k_i = k_{\text{max}} - (k_{\text{max}} - k_{\text{min}}) \cdot \frac{\rho_i - \rho_{\text{min}}}{\rho_{\text{max}} - \rho_{\text{min}}}\)$ where \(k_{\text{min}}=3\) provides the theoretical minimum for stable plane fitting, and \(k_{\text{max}}=15\) defines the PCA robustness ceiling. Dense regions receive small \(k_i\) to preserve high-frequency details, while sparse regions receive large \(k_i\) to ensure noise-resistant normal estimation. Surface normals \(\mathbf{n}_j^{\text{surf}}\) are pre-computed via PCA (\(k=10\)) on surface anchors. When averaging neighbor normals, sign alignment \(s_j = \text{sgn}(\mathbf{n}_j^{\text{surf}} \cdot \mathbf{n}_1^{\text{surf}})\) prevents destructive cancellation, yielding a coherent averaged normal \(\bar{\mathbf{n}}_i\). The normalized direction from outlier \(i\) to its closest surface point \(\mathbf{d}_i\) is aligned with \(\bar{\mathbf{n}}_i\) via absolute cosine similarity \(a_i = |\mathbf{d}_i \cdot \bar{\mathbf{n}}_i|\), which eliminates sign ambiguity. The ALNC loss is formulated as: $\(\mathcal{L}_{\text{ALNC}} = \frac{1}{N_{\text{out}}} \sum_{i \in \mathcal{O}} (1 - a_i)\)$ Optimized alongside photometric loss \(\mathcal{L}_{\text{photo}}\) and attenuated proximity loss \(\mathcal{L}_{\text{prox}}\), outliers migrate along tangent-normal geometric pathways and consolidate seamlessly into sparse surface regions, establishing uniform Gaussian coverage across the object.
Loss & Training¶
The framework utilizes a two-stage training scheme on an NVIDIA RTX 4090 GPU. Stage 1 trains the 3DGS primitives jointly with the MILo occupancy field using decaying learning rates under the combined loss: $\(\mathcal{L}_{\text{Stage1}} = \mathcal{L}_{\text{MILo}} + \lambda_{\text{entropy}} \mathcal{L}_{\text{entropy}} + \lambda_{\text{smooth}} \mathcal{L}_{\text{smooth}} + \lambda_{\text{consist}} \mathcal{L}_{\text{consist}} + \lambda_{\text{boundary}} \mathcal{L}_{\text{boundary}} + \lambda_{\text{temporal}} \mathcal{L}_{\text{temporal}}\)$ Stage 2 employs Adam optimization with gradient clipping and cosine annealing learning rate schedules. Phase 1 (first 30% iterations) employs photometric rendering loss and high-weight \(\mathcal{L}_{\text{prox}}\). Phase 2 (remaining 70% iterations) executes joint fine-tuning: $\(\mathcal{L}_{\text{Phase2}} = \mathcal{L}_{\text{photo}} + \lambda_{\text{prox}} \mathcal{L}_{\text{prox}} + \lambda_{\text{ALNC}} \mathcal{L}_{\text{ALNC}}\)$ This progressive scheduling guarantees stable convergence, preventing distant outliers from destabilizing local geometric curvature.
Key Experimental Results¶
Main Results¶
The framework is quantitatively evaluated on Mip-NeRF 360 and LERF datasets by rendering extracted meshes from test camera viewpoints and comparing them against ground-truth images. Mesh complexity (vertex count) and storage footprints are also benchmarked. As reported in the main table, MagnetGS-Mesh achieves the highest PSNR and SSIM across both benchmarks, confirming superior visual fidelity. In terms of storage efficiency, it slashes mesh size to 0.3 M vertices and 12.4 MB on LERF. On DTU, it achieves 0.66 Chamfer distance without any cell culling post-processing. On Tanks & Temples, it delivers the highest mean F1 score of 0.51, outperforming MILo (0.49) and GOF (0.46).
| Dataset | Method | PSNR (dB) โ | SSIM โ | LPIPS โ | Vertices (Verts) โ | Storage (MB) โ |
|---|---|---|---|---|---|---|
| Mip-NeRF 360 | 2DGS | 15.36 | 0.4990 | 0.4750 | 9.3 M* | 670.3 MB* |
| PGSR | 16.43 | 0.5270 | 0.4060 | - | - | |
| SVRaster | 18.72 | 0.5490 | 0.4060 | - | - | |
| GOF | 20.78 | 0.5730 | 0.4650 | 1.8 M* | 1457.4 MB* | |
| MILo | 24.09 | 0.6890 | 0.3240 | 0.4 M* | 172.6 MB* | |
| Ours | 25.68 | 0.7660 | 0.3120 | 0.3 M* | 152.3 MB* | |
| LERF | 2DGS | 14.97 | 0.4726 | 0.4457 | 9.3 M | 518.6 MB |
| PGSR | 18.55 | 0.6520 | 0.3400 | - | - | |
| SVRaster | 18.71 | 0.6630 | 0.3360 | - | - | |
| GOF | 18.40 | 0.6425 | 0.3541 | 1.8 M | 71.8 MB | |
| MILo | 18.93 | 0.6820 | 0.2975 | 0.4 M | 17.1 MB | |
| Ours | 19.85 | 0.7516 | 0.3088 | 0.3 M | 12.4 MB |
*Note: Storage sizes on Mip-NeRF 360 are measured on the representative bicycle scene; LERF statistics represent dataset-wide averages.
Ablation Study¶
Ablations systematically evaluate Stage 1 regularization components, Stage 2 ALNC modules, and hyperparameter sensitivity regarding neighborhood scaling ranges.
| Component / Stage | Configuration | PSNR (dB) โ | SSIM โ | LPIPS โ | Size (MB) โ | Note |
|---|---|---|---|---|---|---|
| Stage 1 Losses (LERF) | Baseline (MILo) | 18.93 | 0.6820 | 0.2975 | 17.1 | Baseline with soft occupancy field |
| + \(\mathcal{L}_{\text{entropy}}\) | 19.05 | 0.6950 | 0.2988 | - | Entropy minimization suppresses ambiguity | |
| + \(\mathcal{L}_{\text{smooth}}\) | 19.18 | 0.7120 | 0.3002 | - | Spatial coherence across k-NN neighbors | |
| + \(\mathcal{L}_{\text{consist}}\) | 19.32 | 0.7250 | 0.3025 | - | Opacity-occupancy consistency alignment | |
| + \(\mathcal{L}_{\text{boundary}}\) | 19.48 | 0.7310 | 0.3042 | - | Boundary gradient sharpening (largest PSNR gain) | |
| + \(\mathcal{L}_{\text{temporal}}\) | 19.56 | 0.7380 | 0.3058 | 15.8 | EMA temporal stabilization | |
| + Seg. (SAGA) | 19.72 | 0.7450 | 0.3072 | 14.2 | Object-wise 3D scene decomposition | |
| Full (+ \(\mathcal{L}_{\text{ALNC}}\)) | 19.85 | 0.7516 | 0.3088 | 12.4 | Full model: 27.5% storage reduction | |
| Stage 2 ALNC Variants | Single Normal | 19.68 | 0.7430 | 0.3120 | 14.8 | Single neighbor normal, susceptible to point noise |
| Proximity Only | 19.70 | 0.7440 | 0.3100 | 13.9 | Lacks normal alignment, causing surface clutter | |
| Fixed \(k = 10\) | 19.76 | 0.7470 | 0.3100 | 13.2 | Uniform k fails to balance detail and smoothness | |
| Full ALNC | 19.85 | 0.7516 | 0.3088 | 12.4 | Adaptive k achieves optimal geometry guidance | |
| \(k\) Range Sensitivity | \(k_{\text{min}}=3, k_{\text{max}}=20\) | 19.74 | 0.7420 | 0.3160 | - | Overly large \(k_{\text{max}}\) causes over-smoothing |
| \(k_{\text{min}}=5, k_{\text{max}}=15\) | 19.76 | 0.7440 | 0.3140 | - | \(k_{\text{min}}=5\) blurs sharp geometric features | |
| \(k_{\text{min}}=3, k_{\text{max}}=10\) | 19.78 | 0.7460 | 0.3120 | - | Insufficient aggregation in sparse zones | |
| \(k_{\text{min}}=1, k_{\text{max}}=15\) | 19.81 | 0.7490 | 0.3150 | - | \(k_{\text{min}}=1\) lacks plane fitting stability | |
| \(k_{\text{min}}=3, k_{\text{max}}=15\) (Default) | 19.85 | 0.7516 | 0.3088 | 12.4 | Best balance between stability and detail |
Key Findings¶
- In Stage 1, boundary sharpening (\(\mathcal{L}_{\text{boundary}}\)) delivers the most pronounced PSNR boost (from 19.32 dB to 19.48 dB), while spatial smoothness (\(\mathcal{L}_{\text{smooth}}\)) drives the largest single-step SSIM gain (0.6950 to 0.7120), indicating that both binary crispness and spatial coherence are indispensable.
- In Stage 2, ablating normal averaging down to a single neighbor (Single Normal) degrades PSNR to 19.68 dB, highlighting that raw point cloud noise distorts single-point tangents and verifying the necessity of adaptive neighborhood aggregation.
- MagnetGS-Mesh reduces mesh storage by approximately 27% compared to baseline MILo across benchmarks (e.g., from 17.1 MB down to 12.4 MB on LERF) while improving visual rendering metrics. Consolidating floating Gaussians into sparse surface voids suppresses unnecessary vertices while fortifying true surface topology.
- On DTU, conventional methods rely on cell culling to prune floaters (e.g., MILo achieves 0.68 with culling). MagnetGS-Mesh achieves 0.66 with cell culling disabled. Enabling culling actually degrades performance to 0.71 because ALNC has already consolidated all floaters onto surfaces, causing post-processing culling to erroneously erase thin structures.
Highlights & Insights¶
- Turning Noise into Structure: Instead of naively pruning floaters via arbitrary distance thresholds, MagnetGS-Mesh innovatively repurposes outlier Gaussians as geometric material, magnetically pulling them into surface gaps under normal consistency guidance.
- Density-Adaptive Local Guidance: Moving beyond fixed-window normal regularization, ALNC dynamically scales its neighborhood size \(k \in [3, 15]\) according to local surface density, preserving sharp high-frequency edges in dense regions while filtering noise in sparse regions.
- Direct Multi-Object Mesh Extraction: The framework outputs clean, watertight, and topology-independent meshes for individual objects, providing ready-to-use 3D assets for physics simulations, independent scene editing, and interactive graphics engines.
Limitations & Future Work¶
- Generalization from Large-Scale to Micro-Scale Objects: The author notes that MagnetGS-Mesh is primarily designed for large-scale multi-object scenes. When applied to isolated objects with extremely fine, porous, or micro-fiber structures, local normal consistency assumptions may lead to slight over-smoothing across disconnected layers.
- Sensitivity to Upstream 3D Segmentation: The pipeline depends on initial object decomposition from models like SAGA or SAM2. Severe occlusions or ambiguous cross-view semantic boundaries can introduce erroneous grouping that misleads subsequent magnetic relocation.
- Future Directions: Exploring end-to-end differentiable object decomposition directly within the occupancy field without requiring pre-computed 2D/3D segmentation masks, and coupling reconstructed object meshes with physically-based inverse rendering engines.
Related Work & Insights¶
- vs MILo [14]: MILo pioneered mesh-in-the-loop occupancy training but suffers from bloated meshes and ambiguous boundary occupancy values. MagnetGS-Mesh sharpens these boundaries with fivefold regularization and introduces ALNC to reduce model size by 27% while eliminating boundary artifacts.
- vs 2DGS [18] & GOF [50]: 2DGS constrains Gaussians to planar disks, sacrificing representational flexibility in complex volumetric regions, while GOF yields noisy surfaces with inconsistent vertex densities. MagnetGS-Mesh maintains 3D anisotropic Gaussians, achieving higher F1 scores on Tanks & Temples (0.51 vs 0.30/0.46) with smooth, watertight surfaces.
- vs GeoGaussian [28] & Neuralangelo [29]: Prior normal consistency approaches apply uniform smoothness or rely heavily on monocular normal priors. MagnetGS-Mesh dynamically adapts neighborhood aggregation based on point density, effectively resolving the trade-off between detail preservation and noise resilience.
Rating¶
- Novelty: โญโญโญโญโญ The magnetic relocation concept and density-adaptive normal consistency loss offer an elegant paradigm shift for surface optimization.
- Experimental Thoroughness: โญโญโญโญโญ Rigorous evaluation across four major benchmarks (Mip-NeRF 360, LERF, DTU, and Tanks & Temples) with comprehensive ablations and parameter sensitivity studies.
- Writing Quality: โญโญโญโญโญ Clear exposition, intuitive physical analogies, and well-structured mathematical formulations.
- Value: โญโญโญโญโญ Substantially bridges the gap between explicit 3DGS rendering and lightweight, editable multi-object asset generation for computer graphics and robotics.