Skip to content

CAM3R: Camera-Agnostic Model for 3D Reconstruction

Conference: ECCV 2026
Paper: ECCV Official
Full-Text Cache: ../paper_cache/ECCV2026/eccv-5512.txt
Area: 3D Vision
Keywords: 3D reconstruction / camera-agnostic model / fisheye & panorama / ray decoupling / global alignment

TL;DR

Addressing the severe degradation of recent 3D foundation models on non-pinhole cameras due to rectilinear optics assumptions, CAM3R decouples pairwise geometry into spherical harmonic ray direction fields and radial distances with relative poses, and introduces ray-aware global alignment to achieve robust, calibration-free 3D reconstruction across panoramic, fisheye, and pinhole views.

Background & Motivation

Recovering dense 3D geometry from uncalibrated and unposed 2D image collections remains a foundational challenge in computer vision. Recent geometric foundation models—such as DUSt3R, MASt3R, VGGT, and \(\pi^3\)—have revolutionized the field by directly regressing 3D pointmaps in a single feed-forward pass, circumventing fragile multi-stage Structure-from-Motion (SfM) pipelines. However, almost all existing foundation models implicitly rely on the standard pinhole camera assumption because perspective datasets dominate large-scale 3D training corpora. When evaluated on wide-angle imagery captured via non-rectilinear optics—such as fisheye lenses in robotics or \(360^\circ\) panoramic cameras in immersive sensing—these models suffer from severe geometric degradation, yielding heavily curved surfaces, distorted planar structures, and divergent camera trajectories.

A common industry compromise involves undistorting wide-angle images into rectilinear perspective projections before feeding them into downstream models. Yet, this rectification introduces extreme peripheral stretching or discards peripheral regions with high curvature, resulting in massive loss of the effective field of view. Furthermore, cascading multi-stage pipelines accumulates errors and defeats the elegance and robustness of feed-forward foundation models. While recent monocular approaches like UniK3D map pixels directly to continuous 3D rays without parametric camera models, they are restricted to single-view depth estimation and cannot extract cross-view parallax or contextual cues necessary for multi-view pose estimation and coherent dense 3D reconstruction.

The fundamental tension in formulating camera-agnostic multi-view geometry lies in simultaneously learning an unbiased representation of arbitrary non-linear lens optics while extracting cross-view correspondences across heterogeneous camera types to assemble a consistent 3D world. Core idea: strictly decouple two-view feed-forward regression into a spherical harmonic ray direction field capturing internal lens optics and radial distances with relative rigid poses capturing cross-view disparity, unified via ray-aware scene-graph pruning and alternating optimization in a distortion-free 3D ray space.

Method

Overall Architecture

The CAM3R pipeline consists of two complementary components: a pairwise two-view feed-forward network and a ray-aware global alignment framework. Given an unposed, uncalibrated image pair captured by arbitrary camera optics, the two-view network processes the views through two decoupled branches: a shared Ray Module (RM) regressing continuous per-pixel unit ray directions, and a Cross-view Module (CVM) with cross-attention decoders predicting per-pixel radial distances, confidence maps, and the relative rigid camera pose. The local pointmaps are computed via element-wise multiplication of ray directions and radial distances, and aligned into the reference frame using the regressed relative transformation. For multi-view reconstruction, a two-stage graph pruning filters out inconsistent edges, followed by alternating optimization over camera poses and per-view scale factors in a purely ray-consistent 3D coordinate frame.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Uncalibrated Image Pair<br/>I1 and I2"] --> B["Spherical Harmonic Ray Decoupled Representation<br/>Regress continuous unit ray fields"]
    A --> C["Cross-view Radial Distance & Relative Pose Regression<br/>Decouple metric distance and rigid transform"]
    B --> D["Asymmetric Angular & Scale-Anchored Supervision<br/>Quantile penalty & local scale normalization"]
    C --> D
    D --> E["Local Pointmap Fusion & Reference Alignment<br/>Xi,i = di · ri transformed via P2→1"]
    E --> F["Ray-Aware Graph Pruning & Alternating Global Alignment<br/>Two-stage pruning & scale-pose optimization"]
    F --> G["Globally Consistent Scene Pointclouds & Camera Poses"]

Key Designs

1. Spherical Harmonic Ray Decoupled Representation: parameterizing arbitrary lens optics beyond pinhole priors

Prior 3D foundation models directly regress Cartesian pointmaps \((X, Y, Z)\) in the reference frame, entangling camera projection, scene depth, and relative extrinsic poses into a single output. Under extreme non-linear radial distortions, this formulation fails completely. CAM3R resolves this by decoupling the imaging ray direction from spatial distance. For each image \(I_i\), a shared geometric vision encoder extracts class tokens that pass through Transformer Ray Encoder (T-Enc) layers to regress a compact set of Spherical Harmonics (SH) coefficients \(\mathbf{c}_{l,m}^i\) up to bandwidth degree \(L\). An inverse spherical harmonic transformation then reconstructs the normalized unit ray vector \(\mathbf{d}_i(\mathbf{u}) \in \mathbb{S}^2\) for any pixel coordinate \(\mathbf{u}=(u, v)\) mapped to spherical domain \(\psi(\mathbf{u}) = (\theta, \phi)\):

\[\mathbf{d}_i(\mathbf{u}) = \frac{\sum_{l=1}^L \sum_{m=-l}^l \mathbf{c}_{l,m}^i Y_l^m(\psi(\mathbf{u}))}{\left\|\sum_{l=1}^L \sum_{m=-l}^l \mathbf{c}_{l,m}^i Y_l^m(\psi(\mathbf{u}))\right\|_2}\]

Because spherical harmonic basis functions are continuous and smooth, they compactly fit complex non-linear ray geometries (from equidistant fisheye to equirectangular panorama) at arbitrary query resolutions, eliminating the parameter overfitting and field-of-view limits inherent in classical polynomial calibration models.

2. Cross-view Radial Distance & Relative Pose Regression: decoupling metric geometry from spatial extrinsics

Once ray directions are established, 3D reconstruction simplifies to determining the Euclidean distance from the optical center along each ray. Unlike conventional Z-depth along the principal axis—which diverges or becomes undefined in ultra-wide and panoramic lenses—radial distance \(r_i(\mathbf{u}) \in \mathbb{R}^+\) is mathematically well-defined for any lens geometry. The Cross-view Module employs a Siamese ViT encoder coupled with dual transformer decoders that interweave self-attention and cross-attention between views. Dense Prediction Transformer (DPT) heads then predict strictly positive radial distances \(r_i\) alongside confidence maps \(\sigma_i\). Dense local pointmaps are synthesized via direct element-wise product: \(X^{i,i}(\mathbf{u}) = \mathbf{d}_i(\mathbf{u}) \cdot r_i(\mathbf{u})\). To transform the second view's geometry into the reference coordinate system, an explicit relative pose network predicts the relative rotation \(R_{2\to 1} \in \mathrm{SO}(3)\) and unit translation direction \(\hat{\mathbf{t}}_{2\to 1} \in \mathbb{S}^2\). This cleanly remedies DUSt3R's architectural flaw of conflating the second camera's internal geometry with extrinsic translation.

3. Asymmetric Angular & Scale-Anchored Supervision: preventing inward collapse and resolving scale ambiguity

Standard symmetric regression objectives (\(L_1\) or \(L_2\)) tend to collapse toward the dominant distribution in training corpora—namely narrow-FoV pinhole captures—causing severe inward distortion when encountering wide-angle inputs. CAM3R introduces an asymmetric angular quantile loss \(\mathcal{L}_A\) that penalizes angular underestimation substantially more heavily than overestimation:

\[\mathcal{L}_A = \beta \mathcal{L}_{\mathrm{AA}}^{0.7}(\hat{\theta}, \theta^*) + (1-\beta) \mathcal{L}_{\mathrm{AA}}^{0.5}(\hat{\phi}, \phi^*)\]

To train local pointmaps independent of absolute scene scale, CAM3R normalizes predicted and ground-truth coordinates by their respective mean distances to the camera origin \(\eta_i\) and \(\overline{\eta}_i\):

\[\eta_i = \operatorname{mean}_{\mathbf{u}\in\Omega^i} \|X^{i,i}(\mathbf{u})\|_2, \quad \overline{\eta}_i = \operatorname{mean}_{\mathbf{u}\in\Omega^i} \|\overline{X}^{i,i}(\mathbf{u})\|_2\]
\[\mathcal{L}_{\mathrm{regr}} = \sum_{i \in \{1,2\}} \sum_{\mathbf{u}\in\Omega^i} \left\|\frac{1}{\eta_i} X^{i,i}(\mathbf{u}) - \frac{1}{\overline{\eta}_i} \overline{X}^{i,i}(\mathbf{u})\right\|_2^2\]

The relative pose network is supervised by the geodesic SO(3) distance \(\mathcal{L}_{\mathrm{rot}}\) and scale-anchored translation error \(\mathcal{L}_{\mathrm{trans}}\), where the ground-truth translation is scaled by the detached ratio \(\tilde{s}\) between predicted and true pointmap magnitudes, aligning translation scale strictly with regressed 3D geometry.

4. Ray-Aware Graph Pruning & Alternating Global Alignment: optimizing consistency in distortion-free 3D ray space

When scaling pairwise predictions to an unstructured multi-view collection, conventional bundle adjustment fails because non-linear distortions violate the assumption that pixel distances linearly correspond to 3D distances. CAM3R constructs an exhaustive scene graph and applies a rigorous two-stage pruning mechanism: (a) bidirectional pose consistency checks that discard edges where reciprocal rotation deviates beyond threshold \(\tau_{\mathrm{rot}}\) or translation directions invert; and (b) 3D Mutual Nearest Neighbor (MNN) overlap verification that prunes edges with less than 20% spatial overlap, rejecting visual doppelgangers. Surviving incident edges are fused into consensus ray fields \(D_i\) and median-aligned radial fields \(R_i\), freezing a per-camera local geometric prior \(\mathbf{x}_i(\mathbf{u}) = R_i(\mathbf{u}) D_i(\mathbf{u})\). Global alignment is then formulated over camera poses \(\{P_i\} \in \mathrm{SE}(3)\) and per-camera scales \(\{s_i\} \in \mathbb{R}^+\):

\[\min_{\{P_i, s_i\}} \sum_{(i,j)\in\mathcal{E}_{\mathrm{pruned}}} \sum_{\mathbf{u}} \sigma_{i,j}(\mathbf{u}) \left\| P_i(s_i \mathbf{x}_i(\mathbf{u})) - P_j(s_j \mathbf{x}_j(\mathbf{u})) \right\|_2^2\]

Optimization alternates between updating camera poses with fixed scales and updating scales with fixed poses. After several stabilizing cycles, poses and scales are jointly refined, eliminating trajectory drift while strictly preserving local geometric rigidity.

Loss & Training

The overall training objective combines all three components: $\(\mathcal{L}_{\mathrm{total}} = \lambda_A \mathcal{L}_A + \lambda_{\mathrm{regr}} \mathcal{L}_{\mathrm{regr}} + \lambda_{\mathrm{pose}} \mathcal{L}_{\mathrm{pose}}\)$ Training follows a two-stage curriculum: Phase 1 trains exclusively on homogeneous camera pairs (e.g., panorama-panorama, pinhole-pinhole) as CAM3R-homo to master intra-model multi-view geometry; Phase 2 introduces heterogeneous pairs (e.g., pinhole-panorama, fisheye-pinhole) to learn cross-modal geometric mappings. Optimization uses AdamW with an initial learning rate of \(5 \times 10^{-5}\), linear warmup, and cosine decay across four NVIDIA H200 GPUs with balanced dataset sampling.

Key Experimental Results

Main Results

Evaluation spans five challenging benchmark datasets encompassing panorama, fisheye, and perspective imagery: 2D3DS, MegaDepth, 360Loc, ADT, and zero-shot CO3Dv2.

Table 1: Two-view relative pose estimation accuracy (RRA@15° and RTA@15°, higher is better)

Model 2D3DS (Pano) RRA / RTA MegaDepth (Pinhole) RRA / RTA CO3Dv2 (Zero-shot) RRA / RTA 360Loc (Pano) RRA / RTA ADT (Fisheye) RRA / RTA
DUSt3R 10.6 / 6.0 95.6 / 80.8 94.7 / 43.1 0.0 / 0.0 91.0 / 63.6
MASt3R 18.3 / 9.3 69.7 / 56.4 98.4 / 33.4 39.8 / 5.3 96.6 / 63.5
Pow3R 7.5 / 6.0 96.2 / 74.2 95.8 / 38.3 0.0 / 0.0 96.6 / 79.2
VGGT 11.8 / 11.0 98.0 / 88.2 90.9 / 29.4 37.8 / 11.1 92.7 / 82.9
\(\pi^3\) 16.8 / 11.4 99.8 / 93.3 90.7 / 22.7 38.5 / 13.0 97.5 / 93.8
CAM3R-homo 65.4 / 56.8 97.2 / 92.6 96.1 / 66.5 58.3 / 54.7 98.2 / 93.4
CAM3R (Full) 97.7 / 94.3 96.8 / 94.2 97.5 / 88.2 96.0 / 91.0 99.0 / 95.0

Table 2: 3D pointcloud reconstruction quality and radial distance accuracy (Pointcloud Acc/Comp/CD: lower is better, NC: higher is better; Radial Dist. Abs Rel: lower is better, \(\delta_1\): higher is better)

Dataset & Modality Method Pointcloud Acc \(\downarrow\) Pointcloud Comp \(\downarrow\) Normal NC \(\uparrow\) Chamfer CD \(\downarrow\) Radial Abs Rel \(\downarrow\) Radial \(\delta_1\) \(\uparrow\)
2D3DS
(Pano-Pin-Fish)
VGGT
\(\pi^3\)
CAM3R
0.520
0.325
0.059
1.101
0.399
0.077
0.443
0.474
0.970
1.022
0.602
0.068
0.276
0.217
0.053
0.625
0.710
0.954
360Loc
(Pano-Pin-Fish)
VGGT
\(\pi^3\)
CAM3R
0.890
0.630
0.071
3.665
1.405
0.056
0.547
0.595
0.977
2.270
1.017
0.064
0.436
0.333
0.011
0.385
0.547
0.994
MegaDepth
(Pinhole-Pinhole)
VGGT
\(\pi^3\)
CAM3R
0.034
0.013
0.032
0.076
0.058
0.022
0.907
0.921
0.901
0.027
0.024
0.026
0.016
0.010
0.012
0.987
0.996
0.996
ADT
(Pinhole-Fisheye)
VGGT
\(\pi^3\)
CAM3R
0.260
0.166
0.016
0.162
0.084
0.013
0.653
0.763
0.982
0.211
0.125
0.014
0.067
0.040
0.018
0.547
0.737
0.966
CO3Dv2
(Zero-shot Cross-Modal)
VGGT
\(\pi^3\)
CAM3R
0.093
0.113
0.052
0.221
0.109
0.084
0.613
0.616
0.676
0.157
0.111
0.068
0.035
0.027
0.024
0.985
0.997
0.986

Ablation Study

Table 3: Architectural decoupling validation and model evolution (Two-view relative pose accuracy RRA@15° / RTA@15°)

Model Iteration 2D3DS RRA / RTA MegaDepth RRA / RTA CO3Dv2 (Zero-shot) RRA / RTA 360Loc RRA / RTA ADT RRA / RTA
Vanilla DUSt3R (Baseline) 10.6 / 6.0 95.6 / 80.8 94.7 / 43.1 0.0 / 0.0 91.0 / 63.6
Fine-tuned DUSt3R (Stage 1)
(Naive fine-tuning on distorted data)
17.8 / 10.9 94.2 / 72.5 94.3 / 52.7 13.0 / 9.2 87.5 / 65.4
CAM3R (Final - Stage 3)
(SH ray decoupling + pose network)
97.7 / 94.3 96.8 / 94.2 97.5 / 88.2 96.0 / 91.0 99.0 / 95.0

Key Findings

  • Data fine-tuning alone cannot resolve representation failure: Naively fine-tuning DUSt3R on distorted mixtures yields only marginal improvement on 2D3DS (RRA rising from 10.6% to 17.8%), demonstrating that model collapse on wide-angle optics stems from mathematical representation entanglement rather than dataset deficiency.
  • Plausible radial distance does not imply undistorted 3D geometry: On CO3Dv2, baseline \(\pi^3\) achieves \(\delta_1 = 0.997\) on per-pixel distance, yet its Chamfer Distance (0.111) is nearly double that of CAM3R (0.068). Entangled Cartesian pointmap regression introduces strong radial curvature, preserving local depth magnitudes while severely bending global planar and structural geometry.
  • Ray-aware optimization eliminates multi-view drift: Replacing CAM3R's ray-aware global alignment with DUSt3R's pinhole-based alignment drops multi-view trajectory accuracy (mAA@30) from 73.5% to 38.8% on 2D3DS and inflates ATE RMSE from 2.7 to 4.5 on 360Loc, validating the necessity of optimizing in ray-consistent 3D space.

Highlights & Insights

  • Continuous Spherical Harmonic Ray Parametrization: Approximating camera rays via compact SH expansion coefficients allows continuous querying across arbitrary resolutions without predefining polynomial lens distortion parameters.
  • Projection-Invariant Radial Distance: Shifting the prediction target from Cartesian Z-depth to ray-aligned radial distance decouples geometric depth from camera orientation and sensor field of view.
  • Robust Two-Stage Scene-Graph Pruning: Evaluating reciprocal SO(3) pose cycle consistency and 3D MNN overlap effectively eliminates erroneous graph connections caused by visual repetition or lack of overlap.

Limitations & Future Work

  • Dual ViT Backbone Memory Footprint: Operating independent ViT encoders for the Ray Module and Cross-view Module requires \(\sim 14.5\text{ GB}\) VRAM during pairwise forward passes; future iterations could explore cross-task feature distillation or a single unified encoder.
  • Pairwise Graph Inference Overhead: Full multi-view reconstruction requires \(\mathcal{O}(N^2)\) pairwise passes across the scene graph, suggesting exploration of multi-view attention transformers to process unstructured image collections in fewer passes.
  • Extreme Scale and Textureless Regions: In expansive outdoor vistas with large untextured skies, uncertainty estimation can produce local floaters where geometric correspondences are sparse.
  • vs DUSt3R / MASt3R: DUSt3R pioneers end-to-end pointmap regression but hardcodes pinhole geometry by regressing Cartesian coordinates directly in the first frame. CAM3R cleanly separates camera rays from radial distances and predicts explicit relative poses, achieving state-of-the-art results across omnidirectional and perspective imagery.
  • vs UniK3D: UniK3D demonstrates continuous ray regression for single-view depth estimation. CAM3R extends camera-agnostic modeling to multi-view geometry, introducing a dual-decoder cross-attention module, relative pose estimation, and global trajectory alignment.
  • vs \(\pi^3\) / VGGT: While permutation-equivariant foundation models handle dense perspective views, they lack explicit physical ray modeling and suffer severe geometric bending on fisheye and panoramic inputs. CAM3R establishes that physical ray decoupling is essential for universal 3D vision.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneering camera-agnostic formulation for multi-view 3D foundation models with elegant spherical harmonic ray decoupling]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Exhaustive multi-modal evaluations across panoramic, fisheye, and perspective domains with comprehensive pose, trajectory, and pointcloud metrics]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Rigorous mathematical formulation, seamless narrative flow, and precise alignment between architecture and design principles]
  • Value: ⭐⭐⭐⭐⭐ [Provides a crucial open-source foundation for uncalibrated 3D reconstruction in robotics, autonomous driving, and wide-angle consumer devices]