Skip to content

SMP-UWGS: Coupled Physics-Geometry Optimization for Scalable Multi-Partition Underwater 3D Reconstruction

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/vicliuuuu/SMP-UWGS
Area: 3D Vision
Keywords: Underwater 3D Reconstruction, 3D Gaussian Splatting, Physics-Aware Rendering, Multi-Partition Optimization, Attenuation and Backscatter

TL;DR

To tackle computational explosion and color degradation caused by the tight coupling of light attenuation and depth in large-scale marine environments, SMP-UWGS couples a scalable multi-partition Gaussian architecture with a differentiable physical rendering network (DPR-Net), enabling high-fidelity, scalable, and real-time underwater 3D reconstruction.

Background & Motivation

Underwater three-dimensional reconstruction is indispensable for marine ecological monitoring, underwater archaeological documentation, and autonomous underwater robotic intervention. However, optical imaging in aquatic environments faces severe physical degradation: wavelength-dependent absorption and severe backscattering from suspended particulate matter induce contrast attenuation, pervasive blue-green color casts, and distance-dependent haze artifacts. While neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) have revolutionized scene reconstruction, directly adapting them to regional-scale underwater environments reveals a fundamental dilemma.

On one side, physics-grounded approaches such as SeaThru-NeRF, UW-GS, and SeaSplat explicitly model spectral absorption and scattering to restore true colors, but their reliance on dense volumetric sampling or global joint optimization across the entire scene leads to prohibitive GPU memory consumption and training times when scaling up. On the other side, large-scale explicit representation frameworks like VastGaussian and Mega-NeRF achieve scalability through spatial decomposition and point-based pruning, yet they fundamentally treat the transmission medium as uniform air. By ignoring the depth-conditioned nature of underwater light transport, spatial partitioning fragments the attenuation field, resulting in severe color distortions, seam artifacts, and geometric collapse in deeper regions.

This impasse stems from an overlooked coupling: underwater radiative transport is inherently depth-conditioned and thus fundamentally coupled with scene geometry. Because attenuation and backscattering depend explicitly on camera-to-surface distance, spatial partitioning and physical light transport cannot be resolved as decoupled, sequential stages. The core angle of this work is to formulate spatial scalability and physical radiative transport as a single unified constrained learning problem. Core idea: reformulate regional-scale underwater reconstruction as a depth-aware radiative Gaussian learning problem, unifying a multi-partition Gaussian architecture with cross-region visibility and depth weighting for geometric continuity, and embedding a differentiable physical rendering network (DPR-Net) to decouple global optical priors from local attenuation-backscatter refinement under hybrid physical-statistical supervision.

Method

Overall Architecture

The input to SMP-UWGS comprises multi-view underwater images, calibrated camera poses, and an initial sparse point cloud from COLMAP, producing a geometrically consistent and medium-corrected global 3D Gaussian radiation field. The pipeline first projects camera trajectories and points onto the XZ horizontal plane for hierarchical spatial partitioning, establishing overlapping sub-regions with shared boundary cameras. To counteract gradient decay in deep water caused by severe signal attenuation, a depth-aware regional weighting strategy (DARWS) modulates backpropagation gradients per partition. In the physical modeling stage, a Differentiable Physical Rendering Network (DPR-Net) is embedded: the WaterParamPredict module regresses global optical priors from differentiable Gaussian depth maps, followed by the Dual-Branch Differential Refinement (DBDR) module to capture spatial heterogeneity in attenuation and backscatter. Finally, an end-to-end hybrid objective combining physical image formation, physics-aware dark channel prior (PA-DCP), and edge-preserving depth smoothness guides the entire training pipeline.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Underwater Images / Camera Poses / Sparse Points"] --> B["Hierarchical Multi-Partition & Boundary Enhancement<br/>XZ plane partitioning + Shared boundary cameras + Three-tier fusion"]
    B --> C["Depth-Aware Regional Weighting & Gradient Modulation<br/>Frustum intersection + Depth log-scaling + Adaptive gradient modulation"]
    C --> D["Two-Stage Differentiable Physical Rendering Network (DPR-Net)<br/>WaterParamPredict global prior + DBDR local refinement"]
    D --> E["Hybrid Physical-Statistical Loss & Channel-Adaptive Prior<br/>FiLM view decoupling + PA-DCP reweighting + Edge-aware depth smoothness"]
    E --> F["Output: Physically Consistent Global Gaussian Field & High-Fidelity Rendering"]

Key Designs

1. Hierarchical Multi-Partition & Boundary Enhancement: Guaranteeing Global Geometric Continuity and Scalability

Large underwater scenes feature extensive and irregular camera trajectories where standard 3D bounding-box grids lead to camera distribution imbalance and boundary primitive fragmentation. SMP-UWGS addresses this by projecting camera coordinates onto the XZ horizontal plane, sorting cameras by their X-axis coordinates into \(m\) primary partitions with shared boundary cameras, and then hierarchically dividing each partition into \(n\) sub-regions along the Z-axis to yield \(m \times n\) balanced overlapping sub-regions. To eliminate seams at partition interfaces, boundary-aware region enhancement (BARE) expands bounding boxes using adaptive expansion ratios (2.0 for edge partitions and 1.5 for internal regions). It coordinates a three-tier Gaussian fusion strategy consisting of a global Gaussian layer (maintaining global scene coherence), a regional Gaussian layer (local feature stitching via \(G_{\text{enhanced}} = G_{\text{global}} \cup G_{\text{focus}}\)), and a neighborhood Gaussian layer (direct cross-boundary continuity). Cross-region Gaussian stitching is strictly governed by a visibility coverage threshold \(\rho \ge \tau\) derived from camera frustum projections, enabling distributed parallel optimization while maintaining strict multi-view geometric consistency.

2. Depth-Aware Regional Weighting & Gradient Modulation: Compensating for Optimization Imbalance from Deep-Water Attenuation

Because underwater light decays exponentially through water columns, signals from deep regions are severely attenuated, generating weak reconstruction gradients during backpropagation. Consequently, standard optimizers overfit to high-SNR shallow regions while underfitting deep geometry. To rectify this disparity, SMP-UWGS formulates a regional weighting scheme integrating visibility statistics and median depth:

\[w_i = (\bar{r}_i + \beta) \cdot \tau_{\text{init}} \log(d'_i) \cdot \gamma_{\text{geo}}\]

where \(\bar{r}_i\) is the average visibility ratio from bounding-box projection analysis, \(\beta=0.01\) prevents vanishing weights, \(d'_i\) is the normalized regional depth computed from camera-space median depth \(\bar{d}_i = \text{median}\{z_c \mid \mathbf{p} \in \mathcal{P}_i\}\), and \(\gamma_{\text{geo}}\) applies geometric scaling factors (2.0 for corner partitions, 1.5 for edge partitions, and 1.0 internally). The total objective is minimized as a weighted sum \(\mathcal{L}_{\text{total}} = \sum_i w_i \mathcal{L}_i\). This mechanism adaptively boosts effective gradient magnitudes for deep water partitions without altering internal optimizer states, ensuring balanced convergence across all depth strata.

3. Two-Stage Differentiable Physical Rendering Network (DPR-Net): Disentangling Global Water Parameters from Local Heterogeneity

According to the revised underwater image formation model (UIFM), the observed underwater image is modeled as \(I(x,y) = J(x,y) e^{-\beta_d(x,y) z(x,y)} + B_\infty(1 - e^{-\beta_b(x,y) z(x,y)})\). Because direct attenuation \(\beta_d\), backscattering \(\beta_b\), and background illumination \(B_\infty\) enter through non-linear exponential terms, joint end-to-end optimization directly from scratch induces severe parameter oscillation and color degradation. DPR-Net resolves this through a two-stage estimation strategy. In the first stage, WaterParamPredict utilizes a U-Net architecture with SE-ResNet encoders and attention-gated decoders to infer a 9-dimensional global physical parameter vector directly from the differentiable depth map rendered by 3DGS, mapped through constrained activations:

\[\beta_d^c = \operatorname{softplus}(p_c), \quad \beta_b^c = \operatorname{softplus}(p_{c+3}), \quad B_\infty^c = \sigma(p_{c+6}), \quad c \in \{R,G,B\}\]

In the second stage, operating from this stable initialization, the Dual-Branch Differential Refinement (DBDR) module uses shared convolutional layers to predict spatially varying adjustments: \([\beta_d(x,y), \beta_b(x,y)] = \operatorname{ReLU}(\text{Scale} \cdot \sigma(W z(x,y)))\). Decoupling global baseline parameters from local spatially varying refinements mitigates the numerical sensitivity of exponential terms and prevents local minima during joint optimization.

4. Hybrid Physical-Statistical Loss & Channel-Adaptive Prior: Resolving Underwater Color Casts and Ambiguity

To eliminate severe blue-green color casts and visibility degradation in turbid water, SMP-UWGS integrates physical radiance constraints with statistical channel priors. Feature-wise linear modulation (FiLM) is combined with view embeddings to decouple view-dependent water artifacts from intrinsic scene appearance, optimized via a physical reconstruction loss \(\mathcal{L}_{\text{phys}}\) combining \(L_1\) and SSIM. Standard Dark Channel Prior (DCP) fails underwater by misidentifying severe red-channel absorption as heavy haze, creating intense blue-green artifacts upon correction. To prevent this, a physics-aware dark channel prior (PA-DCP) is introduced: after isolating the backscatter-compensated signal \(\hat{J} = \hat{I} - B\), channel contributions are reweighted using inverse attenuation coefficients \(w_c \propto (\beta_d^c + \epsilon)^{-1}\), defining the dark channel map as:

\[D_{PA}(\mathbf{x}) = \min_{c \in \{R,G,B\}} \left( \min_{\mathbf{y} \in \Omega(\mathbf{x})} w_c \hat{J}^c(\mathbf{y}) \right)\]

Color consistency is enforced through an asymmetric objective \(\mathcal{L}_{\text{color}} = \lambda_{\text{dcp}} \|\operatorname{ReLU}(D_{PA})\|_1 + \lambda_{\text{neg}} \|\operatorname{ReLU}(-D_{PA})\|_{\text{smooth}}\) with \(\lambda_{\text{neg}}=1000\), penalizing residual backscatter while strictly preventing negative over-compensation. Augmented with an edge-aware depth smoothness loss \(\mathcal{L}_{\text{depth}}\), background opacity regularization \(\mathcal{L}_{bg}\), and optical parameter regularizer \(\mathcal{L}_{\text{param}}\), the hybrid objective isolates true scene geometry and optical parameters under severe scattering.

Loss & Training

The framework adopts a phased training curriculum over 12,000 total iterations for large-scale partitioned scenes: 1. Iterations 0 to 3,000: Initialize 3DGS and activate WaterParamPredict under physical loss \(\mathcal{L}_{\text{phys}}\) to reliably establish global water parameters \(\beta_d, \beta_b, B_\infty\); 2. Iterations 3,000 to 8,000: Activate DBDR for spatially varying refinement, bringing in color loss \(\mathcal{L}_{\text{color}}\), depth smoothness \(\mathcal{L}_{\text{depth}}\), and background loss \(\mathcal{L}_{bg}\); 3. Iterations 8,000 to 12,000: Freeze optical parameter networks and refine Gaussian positions, rotations, scales, and opacity for seamless partition fusion.

The total objective is formulated as:

\[\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{phys}} + \mathcal{L}_{\text{color}} + \mathcal{L}_{\text{depth}} + \mathcal{L}_{bg} + \mathcal{L}_{\text{param}}\]

All input images are linearized using the inverse IEC 61966-2-1 transfer function prior to optimization, guaranteeing that optical parameters and radiative transport equations are solved in linear radiance space.

Key Experimental Results

Main Results

Quantitative evaluations across SeaThru-NeRF, BVI-Coral, UVEB, and the authors' real-world shipwreck benchmark OTNN demonstrate that SMP-UWGS outperforms baseline 3DGS, neural volume method SeaThru-NeRF, and specialized underwater splatting methods UW-GS and SeaSplat.

Dataset Scene Metric SeaThru-NeRF 3DGS UW-GS SeaSplat SMP-UWGS (Ours)
SeaThru-NeRF IUI3 Red Sea PSNR ↑ / SSIM ↑ / LPIPS ↓ 25.908 / 0.785 / 0.304 22.980 / 0.843 / 0.246 27.652 / 0.863 / 0.185 26.670 / 0.870 / 0.210 27.829 / 0.907 / 0.180
SeaThru-NeRF Curaçao PSNR ↑ / SSIM ↑ / LPIPS ↓ 30.193 / 0.873 / 0.210 28.313 / 0.873 / 0.221 30.171 / 0.851 / 0.172 30.300 / 0.900 / 0.190 31.551 / 0.937 / 0.154
SeaThru-NeRF J.G. Red Sea PSNR ↑ / SSIM ↑ / LPIPS ↓ 21.841 / 0.767 / 0.249 21.493 / 0.854 / 0.216 22.947 / 0.830 / 0.190 22.700 / 0.870 / 0.180 23.289 / 0.895 / 0.174
SeaThru-NeRF Panama PSNR ↑ / SSIM ↑ / LPIPS ↓ 27.846 / 0.834 / 0.224 29.200 / 0.893 / 0.152 30.190 / 0.911 / 0.160 28.760 / 0.900 / 0.150 31.642 / 0.925 / 0.137
BVI-Coral Lagoon PSNR ↑ / SSIM ↑ / LPIPS ↓ 23.542 / 0.759 / 0.286 22.351 / 0.739 / 0.285 23.210 / 0.754 / 0.188 26.250 / 0.779 / 0.289 25.953 / 0.816 / 0.221
BVI-Coral Seabed PSNR ↑ / SSIM ↑ / LPIPS ↓ 22.705 / 0.641 / 0.346 21.919 / 0.743 / 0.283 24.809 / 0.800 / 0.263 25.346 / 0.774 / 0.252 24.461 / 0.769 / 0.299
BVI-Coral Reef PSNR ↑ / SSIM ↑ / LPIPS ↓ 18.551 / 0.408 / 0.552 20.381 / 0.714 / 0.286 22.008 / 0.712 / 0.284 24.316 / 0.754 / 0.258 23.022 / 0.754 / 0.291
BVI-Coral Marine PSNR ↑ / SSIM ↑ / LPIPS ↓ 18.103 / 0.571 / 0.458 16.120 / 0.670 / 0.365 19.406 / 0.703 / 0.308 19.986 / 0.710 / 0.336 20.159 / 0.718 / 0.313
OTNN MingWreck2 PSNR ↑ / SSIM ↑ / LPIPS ↓ 22.186 / 0.791 / 0.254 20.089 / 0.756 / 0.287 23.719 / 0.797 / 0.229 23.828 / 0.732 / 0.267 23.706 / 0.801 / 0.237
OTNN BlueThistle PSNR ↑ / SSIM ↑ / LPIPS ↓ 23.934 / 0.796 / 0.265 21.964 / 0.772 / 0.299 24.008 / 0.802 / 0.352 22.012 / 0.785 / 0.263 24.530 / 0.813 / 0.293
UVEB BioConservation PSNR ↑ / SSIM ↑ / LPIPS ↓ 19.137 / 0.707 / 0.270 18.549 / 0.742 / 0.321 21.980 / 0.721 / 0.272 23.245 / 0.767 / 0.232 22.487 / 0.772 / 0.263
UVEB GeoExploration PSNR ↑ / SSIM ↑ / LPIPS ↓ 17.421 / 0.675 / 0.260 15.418 / 0.639 / 0.288 24.442 / 0.775 / 0.372 27.303 / 0.765 / 0.311 26.707 / 0.788 / 0.303

Computational efficiency benchmarks on MingWreck1 using an NVIDIA RTX 3090 GPU highlight substantial performance gains: - Training Time: SMP-UWGS requires only 13 minutes, achieving a 3.2× speedup over 3DGS (42 min), a 6.2× speedup over SeaSplat (81 min), and a 46× speedup over SeaThru-NeRF (10 hours); - VRAM Consumption: SMP-UWGS consumes only 10.9 GB, well below SeaSplat (16.0 GB) and SeaThru-NeRF (14.6 GB); - Rendering Speed: Delivers an unprecedented 434 FPS, providing fluid real-time visualization for large-scale underwater surveys.

Ablation Study

Component-wise ablations conducted on the Lagoon scene from BVI-Coral demonstrate the indispensability of each physical and partition-based module.

Config PSNR ↑ SSIM ↑ LPIPS ↓ Time Note
Baseline (3DGS) 22.351 0.739 0.285 43 min Uniform air assumption, causing color shift and geometric haze
Full Model w/o SMP 26.275 0.793 0.271 1 h 13 min Without spatial partitioning, training time increases 5.2×
Full Model w/o DPR-Net 23.002 0.726 0.297 10 min Lacks physical light modeling, dropping PSNR by 2.95 dB
Full Model w/o WaterParamPredict 24.803 0.770 0.240 17 min Unstable parameter initialization slows convergence
Full Model w/o DBDR 24.970 0.783 0.235 12 min Omits spatial heterogeneity in attenuation and scattering
Full Model w/o DARWS 25.230 0.796 0.230 14 min Without depth weighting, deep water gradients remain suppressed
Full Model w/o BARE 23.019 0.729 0.301 14 min Discarding boundary enhancement creates visible seams across partitions
Full Model 25.953 0.816 0.221 14 min Optimal balance between physical fidelity and scalable efficiency

Key Findings

  • Crucial Role of Physical Prior Initialization: WaterParamPredict contributes a +1.15 dB PSNR improvement while eliminating exponential gradient instability during early iterations, enabling rapid, non-oscillating convergence within 12k steps. Removing DPR-Net altogether results in an immediate 2.95 dB drop.
  • Boundary Preservation via BARE and Depth Balancing via DARWS: Disabling BARE severely degrades SSIM to 0.729 and LPIPS to 0.301, confirming that spatial partitioning without three-tier fusion fragments cross-region geometric consistency. Concurrently, DARWS balances the gradient contribution across shallow and deep water strata.
  • Breakthrough in Computational Efficiency: By projecting camera trajectories and partitioning optimization along the XZ plane, SMP-UWGS slashes large-scene training time to under 15 minutes, resolving the longstanding computational bottleneck of neural volumetric water modeling.

Highlights & Insights

  • Unified Physics-Geometry Co-Optimization: Recognizing that underwater radiative transport is inherently depth-conditioned, SMP-UWGS tightly couples spatial partitioning with exponential light modeling, avoiding post-hoc stitching artifacts.
  • Physics-Aware Dark Channel Prior (PA-DCP) with Inverted Attenuation Weights: Modulating RGB channels by \(w_c \propto (\beta_d^c + \epsilon)^{-1}\) rectifies the classical DCP failure where red-light absorption is mistaken for particulate haze, establishing an elegant prior for underwater dehazing and restoration.
  • Extensible Depth-Adaptive Gradient Modulation: By scaling partition losses by visibility and depth, backpropagation gradients in deep-water regions are amplified without invasive modifications to optimizer internals, offering a versatile paradigm for non-uniform medium reconstruction.

Limitations & Future Work

  • Uniform Ambient Illumination Assumption: The model assumes spatially uniform ambient background light, which holds under diffuse natural lighting but struggles with directional, co-moving active spotlights mounted on ROVs/AUVs;
  • Static Scene Constraint & Dynamic Marine Distractors: Following 3DGS, the framework assumes rigid scenes; dynamic entities such as schools of fish, undulating seaweed, or dense drifting marine snow particles may induce semi-transparent ghosting artifacts;
  • Future Directions: Integrating active non-uniform spotlight models and incorporating spatiotemporal 4D Gaussians or dynamic motion masks to filter out transient aquatic fauna and suspended particles.
  • vs SeaThru-NeRF: SeaThru-NeRF introduced UIFM into neural fields with accurate color restoration, but suffers from volumetric sampling latencies and heavy memory requirements (10 hours per scene). SMP-UWGS accelerates training by 46× and renders at 434 FPS using explicit Gaussian primitives and multi-partition parallelization.
  • vs UW-GS & SeaSplat: Prior underwater Gaussian splatting methods lack depth-adaptive gradient balancing and scalable spatial partitioning, risking memory exhaustion (16 GB) or deep-scene underfitting. SMP-UWGS reduces VRAM usage to 10.9 GB while achieving superior structural similarity and perceptual fidelity across diverse marine datasets.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Seamlessly couples underwater physical light transport with scalable multi-partition 3DGS, introducing PA-DCP and depth-adaptive gradient modulation.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluations across four diverse benchmarks, including quantitative synthesis, convergence profiling, and modular ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulations, clear problem framing, and logically cohesive methodology.
  • Value: ⭐⭐⭐⭐⭐ Resolves the scalability and efficiency bottleneck for large-scale underwater 3D reconstruction, with direct utility in marine robotics and ecological mapping.