Skip to content

OmniPoint: Universal Monocular Metric Pointcloud from Any Camera

Conference: ECCV 2026
Paper: ECCV Official
Project Page: https://omnipoint.github.io/
Area: 3D Vision
Keywords: monocular metric point cloud, camera-agnostic model, ray-distance decoupling, bidirectional 3D augmentation, geometric prior injection

TL;DR

OmniPoint introduces the first unified monocular metric point cloud estimation framework that handles arbitrary camera geometries (pinhole, fisheye, 360° panoramic) by decoupling ray direction from radial distance, incorporating bidirectional 3D data augmentation, and dynamically injecting geometric priors via input-state embeddings and Gaussian smoothing.

Background & Motivation

Recovering metric 3D geometry from a single image is a cornerstone challenge in computer vision. Driven by large-scale 3D datasets and powerful vision foundation backbones such as DINOv2, modern monocular depth estimators have achieved remarkable zero-shot generalization across in-the-wild scenes. However, current methodologies remain heavily fragmented, locked into isolated networks by rigid camera assumptions: the vast majority of methods strictly presuppose perspective pinhole optics and regress planar depth \(z\) or its inverse disparity; a few specialized approaches target panoramic or fisheye lenses with bespoke unprojection formulations that cannot operate across domains. In practical applications such as autonomous driving and mobile robotics equipped with heterogeneous surround-view sensor suites, practitioners are forced to deploy multiple isolated models, incurring prohibitive computational overhead and impeding multi-sensor fusion.

The root cause of this fragmentation lies in the architectural rigidity of projection assumptions and the severe scarcity of non-pinhole 3D ground truth. Standard planar depth diverges to infinity as the field of view approaches 180° and is fundamentally undefined for points behind the camera plane (\(z < 0\)), completely failing in 360° panoramic environments. Conversely, directly predicting unprojected Cartesian coordinates \((x, y, z)\) forces the neural network to implicitly memorize non-linear projection distortions and physical scene distances within a shared latent space, triggering severe optimization conflicts. Furthermore, dynamically incorporating optional geometric cues such as sparse LiDAR points or camera intrinsics introduces destructive feature distribution shifts that destabilize transformer backbones.

OmniPoint tackles this problem by strictly factorizing camera-dependent projection optics from camera-agnostic physical scene geometry. The core idea is to represent universal monocular point cloud estimation as the pixel-wise product of unit ray direction and radial distance \(P = d \cdot r\), optimized via a ground-truth ray decoupled point loss, supported by bidirectional 3D perspective-omnidirectional data augmentation, and conditioned seamlessly using learnable input-state embeddings and vectorized Gaussian smoothing.

Method

Overall Architecture

OmniPoint takes a single RGB image \(\mathbf{I}\) captured by any camera model, along with optional geometric priors (camera intrinsics \(K\) and sparse depth \(\mathbf{S}\)), to predict a dense metric 3D point cloud \(\hat{\mathbf{P}}\) and an invalid sky mask. The model factorizes each pixel's 3D position into a unit ray direction vector \(\hat{\mathbf{r}}\), a scalar radial distance \(\hat{d}\), and a global metric scale factor \(s\). Built upon a DINOv2-ViT backbone paired with DPT prediction heads, the framework eliminates distribution shifts at the input via state embeddings and Gaussian smoothing, while the decoupled objective independently guides ray unprojection learning and invariant scene distance estimation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image & Optional Conditions<br/>RGB Image + (Intrinsics K / Sparse Depth S)"] --> B["Adaptive Geometric Prior Injection<br/>State embedding indication + Vectorized Gaussian smoothing"]
    B --> C["Bidirectional 3D Data Augmentation<br/>Perspective-to-Any synthesis + Any-to-Perspective self-training"]
    C --> D["Decoupled Ray and Distance Representation<br/>Unit direction vector r + Scalar radial physical distance d"]
    D --> E["Decoupled Point Learning Objective<br/>GT ray guidance for distance + Independent ray direction regression"]
    E --> F["Dense Metric Point Cloud Output<br/>P = s * d * r and invalid sky region mask"]

Key Designs

1. Adaptive Geometric Prior Injection: Resolving multi-modal feature shifts and robustly densifying irregular sparse measurements When toggling between unconditioned RGB, RGB with intrinsics, and RGB with sparse depth, the sudden shift in input feature statistics destabilizes the ViT backbone. OmniPoint resolves this architectural ambiguity by introducing learnable input-state embeddings (\(e_{\text{intrinsic}}\) and \(e_{\text{depth}}\)) added to the input token sequence, acting as boolean indicators that signal the active prior configuration. To ingest sparse, irregular, and noisy depth measurements \(\mathbf{S}\) without inducing gradient instability, the sparse depth is first normalized by its mean valid value to achieve scale invariance. OmniPoint then applies a fully vectorized Gaussian splatting procedure over an influence radius \(r\), computing spatial weights as: $\(w(\Delta u, \Delta v) = \exp\left(-\frac{\Delta u^2 + \Delta v^2}{2\sigma^2}\right)\)$ where kernel bandwidth \(\sigma\) scales proportionally with depth magnitude to preserve relative smoothness across varying depths. Parallel scatter-add accumulation yields a continuous smoothed depth map \(\mathbf{D}_{\text{smooth}}\), which is concatenated with a validity mask and projected into the feature space, allowing OmniPoint to seamlessly function as both a zero-shot monocular estimator and a high-precision depth completion system.

2. Bidirectional 3D Data Augmentation: Bridging the supervision gap between perspective truth and unlabeled omnidirectional domains Because high-quality 3D ground truth is extremely scarce for non-pinhole cameras, standard 2D image-space augmentations fail to simulate realistic optical distortions. OmniPoint bridges this gap in 3D geometric space using a complementary two-way strategy. In Perspective-to-Any (P2A) synthesis, large-scale labeled pinhole datasets are unprojected into dense 3D point clouds and reprojected onto virtual fisheye and equirectangular camera planes with occlusion-aware validity masking, generating rich, paired non-pinhole supervision. In Any-to-Perspective (A2P) self-training, virtual pinhole patches are sampled from unlabeled in-the-wild panoramas, annotated with pseudo-ground-truth depth using the strong perspective teacher, and projected back into the panoramic frame. Sampling only a single patch per image avoids multi-patch depth inconsistencies, creating a highly effective self-supervised cycle.

3. Decoupled Ray and Distance Representation: Disentangling lens projection distortion from physical scene geometry Standard planar depth \(z\) fails for fields of view exceeding 180° and cannot represent points behind the camera (\(z < 0\)), while directly regressing Cartesian coordinates \((x, y, z)\) forces the network to memorize incompatible unprojection mappings \(F: (u, v) \mapsto (x, y, z)\), resulting in optimization failure across diverse camera geometries. OmniPoint explicitly factorizes the 3D prediction into a unit ray direction \(\mathbf{r} \in \mathbb{R}^3\) and a positive radial scalar distance \(d \in \mathbb{R}^+\): $\(\mathbf{P}_i = d_i \cdot \mathbf{r}_i\)$ All sensor-specific optical distortions and field-of-view parameters are isolated within the unit ray \(\mathbf{r}\), leaving radial distance \(d\) as a strictly camera-agnostic representation of spatial depth. Whether viewed through a narrow pinhole or an ultra-wide fisheye lens, the structural distance to a physical surface remains invariant, resolving cross-camera optimization conflicts.

4. Decoupled Point Learning Objective: Preventing cross-contamination between ray direction and radial distance gradients Supervising 3D points with a naive L1 loss \(\sum_i \|\hat{d}_i \hat{\mathbf{r}}_i - d_i \mathbf{r}_i\|_1\) entangles error gradients: an error in ray prediction \(\hat{\mathbf{r}}_i\) penalizes an otherwise accurate distance prediction \(\hat{d}_i\), destabilizing training in peripheral, high-distortion image regions. OmniPoint decouples the learning objectives: $\(\mathcal{L}_{\text{ray}} = \sum_i \|\hat{\mathbf{r}}_i - \mathbf{r}_i\|_1, \quad \mathcal{L}_{\text{point}} = \sum_i \|s^* \cdot \hat{d}_i \cdot \mathbf{r}_i - d_i \cdot \mathbf{r}_i\|_1\)$ where \(s^*\) is the optimal global scale factor computed online via ROE alignment. By projecting the predicted distance \(\hat{d}_i\) strictly along the ground-truth ray \(\mathbf{r}_i\) in \(\mathcal{L}_{\text{point}}\), structural distance errors are completely isolated from ray inaccuracies, ensuring stable and robust optimization across all camera regimes.

Loss & Training

In addition to \(\mathcal{L}_{\text{ray}}\) and \(\mathcal{L}_{\text{point}}\), the overall loss incorporates a metric scale alignment objective \(\mathcal{L}_{\text{metric}} = \|\log(\hat{s}) - \text{stopgrad}(\log(s^*))\|_1\) to recover global metric scale, alongside surface normal consistency loss \(\mathcal{L}_{\text{normal}}\) and local structural loss \(\mathcal{L}_{\text{local}}\) to refine geometric details. The model is trained in three stages: (1) pretraining on large-scale perspective datasets using 72 A100 GPUs (batch size 576) for 10k steps; (2) fine-tuning on mixed synthetic and real fisheye/panoramic data for 5k steps; and (3) freezing the backbone to train the sky mask head using SegFormer pseudo-labels with \(\mathcal{L}_{\text{mask}}\).

Key Experimental Results

Main Results

On zero-shot benchmarks spanning small field-of-view perspective images (8 datasets: NYUv2, KITTI, ETH3D, iBims-1, GSO, Sintel, DIODE, HAMMER), large field-of-view fisheye imagery (KITTI360), and 360° panoramas (Stanford2D3D-S, PanoSUNCG), OmniPoint demonstrates comprehensive superiority over existing approaches.

Method Small FoV (S.FoV)
Depth Rel↓ / \(\delta_1\)↑
Small FoV (S.FoV)
Point Rel↓ / \(\delta_1\)↑
Large FoV (L.FoV)
Depth Rel↓ / \(\delta_1\)↑
Large FoV (L.FoV)
Point Rel↓ / \(\delta_1\)↑
360° Panoramic Images
Depth Rel↓ / \(\delta_1\)↑
360° Panoramic Images
Point Rel↓ / \(\delta_1\)↑
Depth Anything V2 6.35 / 95.0 — / — — / — — / — — / — — / —
Metric3D V2 6.06 / 95.3 — / — — / — — / — — / — — / —
UniDepth V2 4.23 / 96.6 6.06 / 95.1 7.81 / 93.0 18.4 / 79.6 19.3 / 69.5 102.8 / 2.89
Depth Pro 5.28 / 95.6 7.49 / 93.3 26.6 / 54.6 36.2 / 35.5 — / — — / —
MoGe V2 3.98 / 96.8 5.59 / 95.5 12.0 / 85.0 23.4 / 60.9 25.1 / 56.1 102.8 / 2.85
UniK3D 4.14 / 96.7 5.61 / 95.5 8.77 / 92.7 11.5 / 90.6 10.4 / 90.0 11.4 / 89.4
DA2 (Panorama-only) — / — — / — — / — — / — 6.65 / 94.7 7.30 / 94.0
OmniPoint (Ours) 4.04 / 96.8 5.55 / 95.6 6.37 / 94.0 6.66 / 93.8 5.79 / 95.1 5.90 / 95.1

In zero-shot metric depth estimation across six diverse benchmarks, OmniPoint consistently outperforms existing specialized metric depth estimators:

Config / Method NYUv2 KITTI ETH3D iBims-1 DIODE HAMMER Mean Metric (\(\delta_1\)↑)
ZoeDepth 91.9 85.4 33.7 67.2 29.3 3.23 51.8
Depth Pro 91.9 38.3 32.8 81.5 37.7 63.0 57.5
UniDepth V2 92.8 95.4 69.5 93.2 51.8 46.8 74.4
MoGe V2 96.1 62.9 90.8 83.0 66.4 65.6 77.5
UniK3D 94.4 93.6 83.7 92.8 73.0 58.3 82.6
OmniPoint (Unconditioned) 86.2 88.0 91.8 87.7 73.3 73.8 83.5
OmniPoint (+ Intrinsics K) 88.3 90.7 93.0 88.3 75.2 74.5 85.0
OmniPoint (+ Sparse Depth D) 98.5 98.4 94.1 99.0 98.4 99.3 98.0

Ablation Study

Ablation experiments validate the necessity of every architectural and optimization component:

Study Dimension Configuration Benchmark Scenario / Metric Quantitative Result Analytical Takeaway
Output Representation Ray + D (Ours) Fisheye Point Rel↓ / Pano Depth Rel↓ 6.37 / 5.79 Decoupling projection geometry avoids unprojection conflicts
Output Representation Direct XYZ Regression Fisheye Point Rel↓ / Pano Depth Rel↓ 6.74 / 6.28 Implicit projection memorization triggers severe distortion
Supervised Loss Decoupling P + GT Ray (Ours) Point Cloud Rel↓ / \(\delta_1\)↑ 5.55 / 95.6 Isolates structural distance errors from ray direction errors
Supervised Loss Decoupling Naive Point L1 (P) Point Cloud Rel↓ / \(\delta_1\)↑ 5.65 / 95.2 Direction inaccuracies inappropriately penalize distance accuracy
Supervised Loss Decoupling Disjoint (D + Ray) Point Cloud Rel↓ / \(\delta_1\)↑ 7.88 / 93.5 Lacks 3D spatial alignment constraint, causing structural drift
Bidirectional Augmentation P2A + A2P (Full) 360° Pano Depth Rel↓ / \(\delta_1\)↑ 5.79 / 95.1 Combines precise synthetic pairs with real-world scene diversity
Bidirectional Augmentation P2A Only 360° Pano Depth Rel↓ / \(\delta_1\)↑ 6.21 / 94.5 Provides largest single boost by synthesizing omnidirectional pairs
Bidirectional Augmentation A2P Only 360° Pano Depth Rel↓ / \(\delta_1\)↑ 8.21 / 90.9 Pseudo-labels alone drift without synthetic ground-truth anchors
Bidirectional Augmentation No Augmentation Baseline 360° Pano Depth Rel↓ / \(\delta_1\)↑ 8.45 / 89.8 Overfits heavily to scarce non-pinhole supervision
Geometric Prior Conditioning Full Conditioning Module Sparse Guided Depth Rel↓ / \(\delta_1\)↑ 3.21 / 98.2 State embeddings and Gaussian smoothing prevent feature shifts
Geometric Prior Conditioning w/o Gaussian Smoothing Sparse Guided Depth Rel↓ / \(\delta_1\)↑ 3.80 / 97.5 Raw discrete sparse points induce severe gradient instability
Geometric Prior Conditioning w/o State Embeddings Sparse Guided Depth Rel↓ / \(\delta_1\)↑ 3.38 / 97.7 Absence of explicit indicators induces ViT representation ambiguity

Key Findings

  • Radical performance breakthrough in wide-FoV and omnidirectional regimes: Conventional pinhole foundation models (e.g., MoGe V2, UniDepth V2) collapse on 360° panoramas (Point Rel exceeding 100, \(\delta_1\) below 3%), whereas OmniPoint attains 5.79 Depth Rel and 5.90 Point Rel, notably surpassing the dedicated panoramic model DA2 (6.65 Rel).
  • Seamless transformation into high-performance depth densification: Injecting sparse depth increases average zero-shot metric accuracy \(\delta_1\) from 83.5% to 98.0%, and simultaneously supplying intrinsics further reduces point error to 3.59 Rel, proving that conditional priors and monocular backbones reinforce each other without negative interference.

Highlights & Insights

  • Universal applicability of explicit ray-distance factorization: Factoring 3D geometry into unit directional vectors and scalar distances completely eliminates the over-smoothing artifacts typical of Spherical Harmonics parameterizations, providing a universal representation that accommodates any camera model without architectural redesign.
  • Gradient isolation via ground-truth ray substitution: Ground-truth ray projection (\(s^* \cdot \hat{d}_i \cdot \mathbf{r}_i\)) effectively severs harmful gradient coupling in multi-target physical prediction, offering a broadly applicable paradigm for training complex composite physical predictors.
  • State embeddings preserve multi-modal backbone stability: Learned binary indicator embeddings allow a single vision transformer backbone to toggle dynamically between unconditioned monocular estimation and sparse depth completion without catastrophic forgetting or feature divergence.

Limitations & Future Work

  • Specular reflections and transparent surfaces: High-dynamic-range panoramic indoor scenes with intense specular reflections can induce subtle ray refraction errors, as non-Lambertian surfaces are not yet explicitly modeled.
  • Non-uniform sparse prior distribution: While Gaussian smoothing stabilizes irregular LiDAR points, adaptive spatial bandwidth tuning across sparse-dense boundary transitions remains an open area for refinement.
  • Future directions: Extending this universal ray-distance factorization to monocular dynamic video sequences and 4D scene geometry perception across heterogeneous robotic camera rigs.
  • vs UniK3D: UniK3D models the camera ray field via Spherical Harmonics but suffers from over-smoothing in extreme peripheral distortions and lacks bidirectional non-pinhole data augmentation; OmniPoint reduces panoramic error by nearly 50% via explicit ray vectors and 3D data synthesis.
  • vs MoGe V2 / Depth Pro: Leading perspective metric depth estimators whose formulations collapse completely on fisheye and 360° imagery; OmniPoint matches their pinhole performance while fully extending coverage to arbitrary fields of view.
  • vs PromptDA / PriorDA: Existing conditional depth estimators are strictly restricted to perspective pinhole cameras and cannot operate without prompts; OmniPoint functions universally across all camera models and transitions between unconditioned and prior-conditioned modes with zero degradation.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Introduces the first unified formulation for camera-agnostic metric point cloud estimation with decoupled ray-distance objectives and bidirectional 3D augmentation.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Validated across 29 training datasets and 11 evaluation benchmarks, covering metric depth, relative geometry, ablation studies, and prior densification.
  • Writing Quality: ⭐⭐⭐⭐⭐ Highly coherent narrative, clear mathematical formulations, and thorough experimental analysis.
  • Value: ⭐⭐⭐⭐⭐ Provides a foundational, unified 3D vision backbone for autonomous robotics and embodied AI operating heterogeneous sensor suites.