Skip to content

Wat3R: Underwater 3D Geometry Learning without Underwater Annotations

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/LSXI7/Wat3R
Area: 3D Vision
Keywords: underwater 3D reconstruction / cross-domain semi-supervised learning / foundation geometry model / cross-view consistency / static foreground mask

TL;DR

Wat3R is the first semi-supervised feed-forward 3D geometry estimation framework for complex underwater environments, which synergizes physics-based synthetic underwater rendering with large-scale unlabeled real underwater video mining alongside a static-mask-guided cross-view geometric consistency loss, achieving high-accuracy feed-forward multi-view depth, pose, and point cloud reconstruction without requiring any real underwater 3D annotations.

Background & Motivation

Underwater visual 3D geometry estimation aims to reconstruct precise camera poses, dense depth maps, and metric 3D point clouds from multi-view underwater imagery in a single forward pass. This capability is pivotal for autonomous underwater vehicle (AUV) navigation, robotic obstacle avoidance, marine habitat mapping, and underwater archaeological preservation. However, distinct from terrestrial environments, water media severely absorb and scatter propagating light, causing rapid attenuation in red wavelengths, drastic contrast degradation, forward and backscattering haze, and view-dependent optical distortions. These physical bottlenecks make acquiring large-scale, high-fidelity 3D ground truth (such as dense point clouds and continuous 6-DoF trajectory poses) in open-water environments technically prohibitive and financially exorbitant, posing a persistent data barrier in marine visual perception.

Recently, feed-forward foundation geometry models, such as DUSt3R and VGGT, have reformed traditional multi-view geometry by replacing iterative, multi-stage Structure-from-Motion (SfM) pipelines with direct feed-forward neural regression. While these models exhibit strong geometric generalization across terrestrial datasets containing millions of annotated views, directly deploying them to underwater domains incurs severe performance degradation due to pronounced domain shifts. Dedicated underwater reconstruction techniques based on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting require per-scene optimization and calibrated camera poses. Furthermore, conventional two-stage cascaded approaches that pair underwater image enhancement (UIE) with off-the-shelf terrestrial 3D models inevitably introduce high-frequency artifacts or corrupt multi-view photometric consistency during color restoration.

The core tension lies in adapting the rich structural priors of terrestrial foundation geometry models to harsh underwater domains in the complete absence of dense underwater 3D annotations. The angle of attack in this work is to circumvent the reliance on manual 3D annotations by bridging physical degradation modeling and abundant in-the-wild video footage through mutual multi-view geometric redundancy. Core idea: build a Mean Teacher cross-domain semi-supervised feed-forward geometry learning framework that initializes geometric priors via physics-based underwater rendering on annotated terrestrial data, adapts to real underwater domains using filtered unlabeled videos with asymmetric perturbations, and introduces a static-mask-filtered cross-view consistency loss that leverages multi-view complementary observations to compensate for view-dependent optical attenuation and scattering.

Method

Overall Architecture

Wat3R builds upon an end-to-end feed-forward multi-task visual geometry architecture with VGGT as its foundational backbone. Given an input sequence of \(N\) underwater RGB frames \((I_i)_{i=1}^N\), the network predicts a unified set of geometric attributes \(y_i = (g_i, D_i, P_i)\) for each view \(i\) relative to the reference coordinate frame of the first camera in a single forward pass. Here, \(g_i \in \mathbb{R}^9\) parameterizes camera rotation (quaternion), translation, and field of view (FoV); \(D_i \in \mathbb{R}^{H \times W}\) denotes the dense depth map; and \(P_i \in \mathbb{R}^{3 \times H \times W}\) is the 3D point map expressed in the reference camera frame.

To achieve robust domain adaptation without underwater 3D ground truth, the framework adopts a Mean Teacher semi-supervised design. The Teacher network parameters are updated as the exponential moving average (EMA) of the Student network, generating stable pseudo-supervision under weak data augmentations. Concurrently, the Student network receives strong sequence-level perturbations and is optimized jointly via supervised losses on physically synthesized underwater data, per-view pseudo-label losses on unlabeled real videos, and cross-view reprojection consistency constraints.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input underwater multi-view frames (I_i)"] --> B["Physics-based underwater rendering and real video mining<br/>annotated terrestrial synthesis + unlabeled real video footage"]
    B --> C["Mean Teacher pseudo-label geometric self-supervision<br/>weak-augmented Teacher yields stable pseudo-targets, strong-augmented Student backpropagates"]
    C --> D["Static-mask-guided cross-view consistency loss<br/>reprojected depth verification, suppressing water background and dynamic motion"]
    D --> E["Feed-forward geometric outputs<br/>camera poses g_i / dense depth D_i / 3D point maps P_i"]

Key Designs

1. Physics-based underwater rendering and real video mining: establishing priors and domain generalization without real 3D annotations To resolve the cold-start problem stemming from the total lack of underwater 3D annotations, Wat3R simulates realistic underwater optical attenuation on existing annotated terrestrial multi-view datasets (the original training corpus of VGGT). Rather than synthesizing arbitrary 3D geometry, the image formation follows a revised physical water transport model: $\(I = J e^{-\beta^D z} + B^\infty (1 - e^{-\beta^B z})\)$ where \(J\) is the clean terrestrial image, \(I\) is the simulated underwater counterpart, and \(z\) denotes scene depth. Because raw terrestrial datasets frequently feature sparse or incomplete depth annotations that create rendering seam artifacts, Wat3R employs the monocular foundation model DA3MONO-LARGE to predict continuous, smooth depth fields specifically for driving the rendering process, while strictly retaining original camera poses and true 3D structures as regression targets. Physical attenuation coefficients \(\beta^D\) and backscattering parameters \(\beta^B\) are sampled under the physical constraint that red light attenuates faster than green and blue, modulated by smoothed Gaussian noise to emulate spatially non-uniform turbidity.

To bridge the synthetic-to-real gap, the authors curated over 10,000 raw underwater video clips from public scientific repositories and in-the-wild recordings. Following strict filtering criteria—pruning jump cuts, severe motion blur, water surface clips, and open-ocean views devoid of static structures—a clean benchmark of 5,504 continuous video sequences comprising 359,000 high-quality frames was established, providing a rich, diverse distribution of real underwater optical conditions for unsupervised adaptation.

2. Mean Teacher pseudo-label geometric self-supervision: stabilizing cross-domain transfer via asymmetric perturbations Unsupervised adaptation on underwater video sequences is vulnerable to representation collapse and depth hallucination triggered by pervasive scattering haze. Wat3R introduces a Mean Teacher architecture where Student parameters \(\theta_s\) are updated through gradient descent, and Teacher parameters \(\theta_t\) track an exponential moving average: $\(\theta_t^k = (1 - \lambda)\theta_t^{k-1} + \lambda \theta_s^k\)$ To overcome the limited viewpoint diversity of video recordings, an asymmetric sequence-level augmentation strategy is implemented. The Teacher branch processes 24 to 36 randomly sampled and shuffled frames under weak transformations to produce high-confidence pseudo-labels. In contrast, the Student branch samples 2 to 12 frames from the identical sequence, shuffles frame order, and applies 0°, 90°, 180°, or 270° discrete rotations to prevent pose collapse, supplemented by strong color jittering, grayscale conversion, and Gaussian blurring.

In the per-view consistency loss \(\mathcal{L}_{\text{per-view}}\), Teacher predictions \(\{g^t, D^t, P^t\}\) serve as pseudo ground truths. To maximize geometric consistency across prediction heads, the point map targets are derived directly by unprojecting Teacher depth and camera parameters, guiding the Student to output coherent geometry despite aggressive appearance perturbations.

3. Static-mask-guided cross-view consistency loss: exploiting multi-view reprojection complementarity against water degradation Severe attenuation and backscattering frequently obscure structural textures in specific viewpoints, while neighboring views retain clear local contours. Wat3R formulates a cross-view geometric consistency objective that forces the network to borrow geometric cues from complementary viewpoints. However, naive unconstrained reprojection across underwater frames introduces severe noise from uninformative water body background regions and dynamic marine life.

To isolate reliable regions, the method applies a two-cluster K-means algorithm to the Teacher depth map, automatically stripping away distant open water to obtain a coarse foreground mask \(M_i^{\text{fg}}(x)\). Next, pixels from view \(i\) are backprojected to 3D and reprojected into all other Teacher views \(j\). A multi-view visibility indicator verifies geometric agreement within a tolerance threshold \(\delta\): $\(V_{i \to j}(x) = \mathbf{1}\left(|D_j^t(u_{i \to j}(x)) - D_{i \to j}^t(x)| < \delta \land u_{i \to j}(x) \in \Omega_j\right)\)$ A pixel is retained in the static mask only if it maintains consistent reprojection depth across at least \(k = N - 2\) accompanying teacher views: $\(M_i^{\text{static}}(x) = \mathbf{1}\left(\sum_{j \neq i} V_{i \to j}(x) \ge k\right) M_i^{\text{fg}}(x)\)$ During Student training, for randomly sampled frame pairs \((i, j) \in S\), Teacher depth from view \(i\) is reprojected to frame \(j\) (\(D_{i \to j}^t\)) and compared against Student depth \(D_j^s\) under the static mask: $\(\mathcal{L}_{\text{cross-view}} = \frac{1}{|S|(|S|-1)} \sum_{i \in S} \sum_{j \in S, j \neq i} L_1\left(D_j^s, D_{i \to j}^t, M_j^{\text{static}}\right)\)$ This constraint forces the model to synthesize cohesive geometry in degraded regions while remaining impervious to dynamic moving objects and featureless open water.

Loss & Training

The overall training loss is a weighted formulation: $\(\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{supervised}} + \lambda_u \left(\mathcal{L}_{\text{per-view}} + \mathcal{L}_{\text{cross-view}}\right)\)$ Training is conducted across 4 NVIDIA RTX 4090 GPUs for 19,200 iterations with a 1,000-step linear warmup. The initial 6,400 steps train strictly on synthetic underwater labeled data, establishing foundational geometric capabilities. Subsequently, the unsupervised weight \(\lambda_u\) is ramped up linearly, reaching its maximum value of 0.5 at step 12,800. The batch sampling ratio of unlabeled real video frames to labeled synthetic frames is fixed at 1:3. The maximum learning rate is set to \(5 \times 10^{-6}\) for the ViT backbone and \(5 \times 10^{-5}\) for downstream regression heads, trained with bfloat16 mixed precision and gradient checkpointing to handle multi-view visual tokens efficiently.

Key Experimental Results

Main Results

Evaluation is performed across public underwater benchmarks (Sea-thru, FLSea-Stereo) and the authors' newly constructed Water3D benchmark (comprising 42 real underwater scenes with verified camera poses and dense depth).

Multi-view depth estimation performance is evaluated under a 10-view shuffled protocol and a 100-frame subsequence protocol, measured by absolute relative error (Rel), logarithmic error (log10), root mean square error (RMSE), and threshold accuracy (\(\delta_1 < 1.25\)):

Method Sea-thru: Rel ↓ Sea-thru: \(\delta_1\) ↑ Sea-thru: RMSE ↓ FLSea-Stereo: Rel ↓ FLSea-Stereo: \(\delta_1\) ↑ FLSea-Stereo: RMSE ↓
WaterSplatting (3DV'25) — — — 0.427 0.415 1.476
Fast3r (CVPR'25) 0.277 0.713 0.631 0.290 0.529 1.355
MapAnything (3DV'26) 0.216 0.800 0.454 0.146 0.841 0.798
\(\pi3\) (ICLR'26) 0.185 0.909 0.358 0.139 0.856 0.837
DA3 (ICLR'26) 0.187 0.892 0.333 0.141 0.851 0.872
VGGT (Baseline, CVPR'25) 0.190 0.891 0.380 0.137 0.849 0.760
VGGT + Semi-UIR (UIE Pipeline) 0.201 0.846 0.387 0.136 0.861 0.784
VGGT + PSPL (UIE Pipeline) 0.209 0.836 0.419 0.160 0.817 0.911
Wat3R (Ours) 0.167 0.946 0.290 0.119 0.885 0.720
Relative Improvement over VGGT +12.1% +6.2% +23.7% +13.1% +4.2% +5.3%

Multi-view 3D point cloud reconstruction comparison on Water3D (evaluated via Accuracy, Completeness, and Overall error after Sim(3) alignment):

Method Accuracy Mean ↓ Accuracy Median ↓ Completeness Mean ↓ Completeness Median ↓ Overall Mean ↓ Overall Median ↓
Fast3r (CVPR'25) 1.216 0.531 2.444 0.784 1.830 0.658
MapAnything (3DV'26) 0.643 0.284 0.666 0.277 0.655 0.281
\(\pi3\) (ICLR'26) 0.491 0.168 0.413 0.184 0.452 0.176
DA3 (ICLR'26) 0.679 0.187 0.528 0.160 0.604 0.174
VGGT (CVPR'25) 0.486 0.193 0.762 0.191 0.624 0.192
Wat3R (Direct Point Map) 0.444 0.148 0.409 0.165 0.427 0.157
Wat3R (Unprojected Depth+Pose) 0.446 0.162 0.366 0.143 0.406 0.153

Ablation Study

The progressive contribution of synthetic rendering, unlabeled real videos, strong sequence augmentations, cross-view consistency loss \(\mathcal{L}_{\text{cross-view}}\), and the static mask \(M^{\text{static}}\) are evaluated on multi-view depth estimation:

Config Syn Water Real Video Strong Aug. \(\mathcal{L}_{\text{cross-view}}\) \(M^{\text{static}}\) Sea-thru: Rel ↓ Sea-thru: \(\delta_1\) ↑ FLSea-Stereo: Rel ↓ FLSea-Stereo: \(\delta_1\) ↑
(1) Baseline (VGGT) — — — — — 0.190 0.891 0.137 0.849
(2) Supervised on Synthetic Only ✓ — — — — 0.173 0.920 0.135 0.860
(3) Add Real Video (Standard Aug.) ✓ ✓ — — — 0.181 0.906 0.165 0.793
(4) Add Strong Sequence Augmentation ✓ ✓ ✓ — — 0.172 0.936 0.126 0.871
(5) Add Full-frame Cross-view Loss ✓ ✓ ✓ ✓ — 0.167 0.949 0.126 0.869
(6) Masked Cross-view without Real Video ✓ — ✓ ✓ ✓ 0.174 0.929 0.130 0.871
(7) Wat3R Full Model ✓ ✓ ✓ ✓ ✓ 0.167 0.946 0.119 0.885

Key Findings

  • Degradation of Two-Stage Image Enhancement: Pre-processing images with underwater image enhancement (UIE) models like Semi-UIR or PSPL degrades depth estimation performance relative to raw inputs (Sea-thru Rel increases from 0.190 to 0.201 and 0.209). This confirms that visual enhancement models optimize perceived color contrast at the expense of multi-view geometric and photometric consistency; learning geometry directly under degradation is far superior.
  • Necessity of Strong Sequence-Level Perturbations: Comparing rows (2) and (3) in the ablation study, adding real underwater video without strong sequence-level augmentations degrades Rel error on FLSea-Stereo from 0.135 to 0.165, indicating severe representation collapse. Injecting temporal shuffling, rotations, and blur perturbations (row 4) resolves this failure mode, reducing error to 0.126.
  • Static Mask Purifies Challenging Water Bodies: In the relatively clear, static Sea-thru scenes, introducing the static mask yields identical Rel error (0.167). However, on FLSea-Stereo featuring complex currents and floating turbidity, the static mask reduces Rel error from 0.126 to 0.119, verifying that suppressing water background and dynamic regions is vital in real marine operational conditions.
  • Inter-Head Physical Agreement: The point cloud directly output by the point head and the point cloud derived from unprojecting predicted depth and camera poses show nearly identical metric performance on Water3D (Overall error 0.427 vs. 0.406), achieving up to a 52.0% completeness gain over VGGT. This demonstrates that multi-view consistency constraints enforce intrinsic physical coherence across distinct prediction heads without ground-truth supervisory labels.

Highlights & Insights

  • Zero-Annotation Cross-Domain Adaptation: Wat3R presents an elegant blueprint combining physics-guided terrestrial simulation with self-supervised real video adaptation, unlocking foundation 3D geometry models for harsh environments devoid of 3D annotations.
  • Static-Mask Reprojection Synergy: By leveraging multi-view geometric overlap filtered through K-means foreground clustering and temporal visibility verification, the model turns the physical weakness of view-dependent water attenuation into an explicit cross-view learning signal.
  • Water3D High-Standard Benchmark: Constructing 42 rigorously curated real underwater scenes equipped with verified camera poses and metric depth provides the community with a dependable benchmark, surpassing historical 4-to-5-scene evaluation protocols.

Limitations & Future Work

  • Sparse Geometric Cues in Featureless Open Water: When an underwater vehicle operates in completely open, deep blue waters or exceptionally turbid channels lacking static bottom structures, reliable matched features become virtually nonexistent. Under such conditions, the conservative static mask \(M^{\text{static}}\) filters out nearly all pixels, attenuating the cross-view consistency learning signal.
  • Exclusion of Dynamic Marine Entities: The current framework discards moving marine fauna and divers as transient noise through the static mask. It does not model non-rigid 4D dynamics or separate foreground motion fields, limiting geometric fidelity in biologically dense coral reef scenes.
  • Future Directions: Integrating differentiable underwater volumetric rendering to jointly estimate medium absorption coefficients alongside 3D geometry, and extending the feed-forward model into an online underwater visual SLAM system with real-time pose graph loop closure.
  • vs VGGT / DUSt3R: While VGGT and DUSt3R pioneers feed-forward uncalibrated 3D reconstruction, they are trained predominantly on clean terrestrial datasets and collapse under underwater optical scattering. Wat3R introduces a cross-domain Mean Teacher paradigm and physics-based degradation training that adapts VGGT to complex marine scenes without real underwater 3D labels.
  • vs SeaThru-NeRF / WaterSplatting: Volume rendering and 3D Gaussian Splatting approaches explicitly model light transport but require per-scene iterative optimization and pre-calibrated camera poses. In contrast, Wat3R provides fast feed-forward inference across unseen scenes without requiring known camera parameters.
  • vs Cascaded Underwater Image Enhancement (UIE): Traditional pipelines enhance images before geometric estimation, which frequently introduces high-frequency artifacts and alters multi-view lighting coherence. Wat3R demonstrates that end-to-end geometry learning directly on degraded inputs preserves structural integrity and outperforms cascaded pipelines.

Rating

  • Novelty: ⭐⭐⭐⭐☆ [Pioneering adaptation of feed-forward geometry foundation models to underwater environments without real 3D annotations using cross-view consistency]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluated extensively on monocular/multi-view depth, camera poses, and point clouds across five public benchmarks and a comprehensive new 42-scene dataset]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear mathematical formulation, well-structured pipeline diagrams, and insightful ablation analyses]
  • Value: ⭐⭐⭐⭐☆ [Provides an impactful practical framework for underwater robotic vision, ocean mapping, and neural reconstruction in adverse scattering media]