Skip to content

Pixel-wise Geo-registration of Drone and Satellite Images

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/dshatwell23/skyreg-dataset
Area: Autonomous Driving
Keywords: cross-view geo-registration, drone imagery, satellite remote sensing, feed-forward 3D reconstruction, dense geo-localization

TL;DR

Addressing the fundamental limitations of prior cross-view localization confined to image-level camera center estimation and planar homographies failing under severe parallax, SkyReg introduces a feed-forward 3D geometry-aware registration framework alongside the first pixel-wise geodetic benchmark SkyReg-Bench, establishing robust, high-accuracy per-pixel GPS coordinate prediction.

Background & Motivation

Cross-view geo-localization plays an indispensable role in unmanned aerial vehicle (UAV) infrastructure inspection, smart city monitoring, emergency disaster assessment, and defense geospatial intelligence. Nevertheless, the prevailing paradigm has long been framed as either image retrieval (matching a query view against a massive database of geo-referenced satellite tiles) or 3-DoF camera pose regression (predicting the query's location and orientation in the satellite image coordinate plane). Both approaches inherently reduce the task to estimating the camera's spatial origin. In practical UAV operations with oblique viewpoints and complex three-dimensional urban topographies, a tiny error in the estimated camera center compounds into substantial spatial displacement on distant elevated structures or ground targets; operational demands necessitate determining the exact geographic coordinate of every visible pixel rather than just where the camera is.

Concurrently, traditional remote sensing geo-registration methods rely heavily on planar homography assumptions or iterative digital elevation model (DEM) template matching, which are well-suited only for near-nadir satellite-to-satellite pairs with subtle perspective shifts. When confronted with low-altitude, perspective drone captures and orthorectified high-altitude satellite views, conventional 2D correspondence algorithms (such as SuperPoint+SuperGlue or RoMa-based homography fitting) fail dramatically under dramatic scale discrepancies, non-planar parallax, severe occlusions, and multi-modal radiometric differences. Meanwhile, emerging neural rendering registration pipelines necessitate multi-view image sequences, rendering them inapplicable to single-shot drone alignment. Compounding this challenge, the field has lacked standardized, large-scale benchmarks providing dense, per-pixel geodetic ground truth alongside 3D scene parameters.

This paper tackles these challenges by circumventing 2D planar matching altogether and infusing feed-forward multi-view 3D neural reconstruction into single-image cross-view alignment. Core idea: leverage a feed-forward 3D reconstruction backbone to explicitly reconstruct shared 3D point maps and relative camera transformations between drone and satellite views, reprojecting query pixels onto the reference image plane via a 3D-aware warping function to assign per-pixel GPS coordinates via bilinear interpolation on the reference geodetic grid.

Method

Overall Architecture

The SkyReg pipeline proceeds in two sequential, tightly integrated stages: feed-forward 3D scene reconstruction and pixel-wise cross-view geometric warping. The system takes two inputs: a georeferenced satellite reference image \(I_r\) with known camera intrinsics and poses defined in an Earth-Centered, Earth-Fixed (ECEF) world coordinate system, and an uncalibrated oblique drone query image \(I_q\) lacking prior pose metadata. First, the transformer-based feed-forward backbone (MASt3R) processes the image pair to jointly predict dense 3D point maps in respective camera frames, recovers camera focal lengths, and computes the relative transformation using Procrustes alignment, transforming the drone point map into the satellite reference coordinate frame. Next, the model reprojects these 3D points onto the reference image plane to establish a non-planar 3D warping function, and finally queries the satellite geodetic ellipsoid mapping to yield dense per-pixel GPS coordinates or synthesize orthorectified mosaic overlays.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image Pair<br/>Perspective Drone + Geo-referenced Satellite"] --> B["Feed-Forward 3D Reconstruction & Pose Alignment<br/>Predict dense point maps & relative transformation"]
    B --> C["Cross-View Geometric Warping & Spatial Reprojection<br/>Lift and project 3D points into reference plane"]
    C --> D["Pixel-Wise GPS Coordinate Assignment & Mosaicing<br/>WGS-84 coordinate mapping & layer overlay"]

Key Designs

1. Feed-Forward 3D Reconstruction & Pose Alignment: lifting cross-view reasoning beyond 2D planar assumptions Planar homography baselines inherently fail when matching low-altitude perspective drone images with satellite tiles due to building height variations and severe parallax. SkyReg employs MASt3R as its feed-forward 3D geometric backbone to perform dense feature matching and spatial elevation simultaneously. In a single forward pass, the network predicts the dense 3D point cloud \(X^{q,r} \in \mathbb{R}^{H \times W \times 3}\) and associated per-pixel confidence maps. For the uncalibrated drone camera, the model optimizes focal length \(f_q^*\) via reprojection residual minimization and computes the relative similarity transformation \(\hat{P}^{q \to r}\) via Procrustes alignment: $$ \hat{P}^{q \to r} = \arg\min_{P} \sum_{\mathbf{x}_q} \left| h^{-1}(P \, h(X^{q,q}(\mathbf{x}_q))) - X^{q,r}(\mathbf{x}_q) \right| $$ By operating explicitly in 3D Euclidean space, this design circumvents the severe geometric distortions and feature collapses that plague 2D matchers across extreme scale and viewpoint differences.

2. Cross-View Geometric Warping & Spatial Reprojection: closed-form non-planar coordinate transfer Once the query point cloud \(X^{q,r}(\mathbf{x}_q)\) is expressed in the reference camera frame, SkyReg maps pixel coordinates between the query and reference image planes without requiring external DEM data during inference. It projects 3D query points directly onto the reference image plane using the reference intrinsics \(K_r\): $$ \mathbf{x}'_q = \mathrm{warp}(\mathbf{x}_q) = h^{-1}(K_r X^{q,r}(\mathbf{x}_q)) $$ Unlike homography transformations that force all visible structures onto a single artificial plane, this warping formulation explicitly accounts for elevation relief, building facades, and depth discontinuities, preserving structural fidelity across vertical occlusions.

3. Pixel-Wise GPS Coordinate Assignment & Mosaicing: unified mapping onto the WGS-84 geodetic ellipsoid Following successful reprojection of each query pixel to reference coordinates \(\mathbf{x}'_q\), the framework computes real-world geodetic coordinates by evaluating the geodetic mapping function \(\mathcal{E}(\cdot)\) linked to the reference satellite tile: $$ \mathbf{X}_{\mathrm{WGS84}} = \mathcal{E}(\mathrm{warp}(\mathbf{x}_q)) = \mathcal{E}\left(h^{-1}(K_r X^{q,r}(\mathbf{x}_q))\right) $$ Because this operation executes in parallel across all valid image pixels, the network outputs both a dense latitude-longitude tensor and allows direct forward texture mapping of the drone query onto the satellite base map, producing an accurate, orthorectified mosaic overlay for downstream spatial analysis.

Loss & Training

The 3D point maps are supervised using a confidence-weighted regression objective adapted from DUSt3R. For each supervised pixel \(i \in \mathcal{D}^v\) in view \(v\), the objective minimizes the normalized Euclidean distance between predicted and ground-truth 3D coordinates while regularizing the predicted confidence \(C_i^{v,1}\): $$ \mathcal{L}{\mathrm{conf}} = \sum $$ The network is trained with AdamW using a base learning rate of } \sum_{i \in \mathcal{D}^v} C_i^{v,1} \left| \frac{1}{z} X_i^{v,1} - \frac{1}{z} \bar{X}_i^{\bar{v},1} \right| - \alpha C_i^{v,1\(1 \times 10^{-4}\), weight decay of 0.05, and batch size of 48 image pairs over 20 epochs on SkyReg-Train. To retain generic 3D priors while adapting to satellite-drone cross-view distributions, each epoch samples 20k pairs from Urban scenes, 20k pairs from Landmarks, and 20k pairs from general 3D indoor/outdoor datasets (ScanNet++, ARKitScenes, and CO3D).

Key Experimental Results

Main Results

Evaluation is conducted on the held-out SkyReg-Bench benchmark, evaluating zero-shot geographic transfer across the unseen Urban test set (Chicago, 13,191 query pairs) and the Suburban test set (WRIVA scenes, 950 query pairs). Baselines include feature matchers paired with homography estimation (SuperPoint+SuperGlue, RoMa, and RoMa Dense). Metrics report Mean Geodetic Error (GE, in meters) and registration Recall@\(\tau\) (%) at distance thresholds \(\tau \in \{20, 30, 40, 50\}\text{m}\).

Dataset Method GE (m) โ†“ Recall@20m (%) โ†‘ Recall@30m (%) โ†‘ Recall@40m (%) โ†‘ Recall@50m (%) โ†‘
Urban (CHI) SP+SG 209.40 1.40 1.94 2.34 2.82
Urban (CHI) RoMa 133.06 25.69 27.18 28.35 29.33
Urban (CHI) RoMa Dense 131.45 17.62 20.08 22.43 24.72
Urban (CHI) SkyReg (Ours) 44.88 78.78 80.03 80.53 80.94
Suburban SP+SG 223.69 36.42 37.15 37.47 37.47
Suburban RoMa 112.10 60.00 61.26 62.25 63.15
Suburban RoMa Dense 111.25 51.89 52.52 53.57 54.42
Suburban SkyReg (Ours) 21.34 74.94 93.57 97.78 98.52

Under the Median-GPS localization protocol, which compares against retrieval baselines (TransGeo, University-1652, EarthMatch), SkyReg achieves 36.40 m GE in Urban (vs. 104.70 m for RoMa and 160.98 m for TransGeo) and 21.45 m GE in Suburban with a 95.27% Recall@50m, substantially outperforming all competing baselines.

Ablation Study

The training data composition ablation assesses the generalization capability of SkyReg variants evaluated under the Median-GPS protocol on unseen Urban and Suburban test splits.

Training Config Urban GE (m) โ†“ Urban R@20m (%) โ†‘ Urban R@50m (%) โ†‘ Suburban GE (m) โ†“ Suburban R@20m (%) โ†‘ Suburban R@50m (%) โ†‘ Note
Landmark Subset + 3D 48.70 59.89 73.18 27.23 73.66 91.40 Diverse view poses but higher metric error
Urban Subset + 3D 50.38 58.47 69.37 24.79 72.30 93.07 Good urban priors but weaker out-of-domain transfer
Landmark + Urban (Full) 36.40 66.72 81.57 21.45 75.32 95.27 Combined multi-source training yields optimal accuracy

Key Findings

  • SkyReg reduces pixel-wise mean geodetic error by 88.18 m in Urban and 90.76 m in Suburban settings compared to RoMa, demonstrating the fatal breakdown of planar homography models under steep cross-view baselines.
  • In dense urban canyons with tall buildings, planar assumptions fail catastrophically (SP+SG achieves under 3% recall), while 3D geometric reasoning sustains over 80% recall across all distance thresholds.
  • Jointly training on orthorectified urban LiDAR pairs and diverse-viewpoint perspective landmark pairs provides complementary scale and pose distributions, yielding a 12โ€“14 m error reduction over single-subset variants.

Highlights & Insights

  • Paradigm shift from image-level retrieval to dense 3D registration: Rather than stopping at 3-DoF camera localization, the paper unifies cross-view geo-registration into a feed-forward 3D reconstruction and geometric warping task.
  • Rigorous benchmark with dense geodetic ground truth: SkyReg-Bench introduces ray-traced LiDAR meshes and MVS-calibrated depth to deliver high-quality, pixel-wise GPS labels across thousands of unseen test viewpoints.
  • Versatile bridge between vision geometry and geodetic reference: By coupling feed-forward relative point clouds directly with WGS-84 coordinate transforms, the method operates without requiring runtime DEMs or iterative optimization.

Limitations & Future Work

  • Sensitivity to severe temporal and illumination shifts: Training images rely primarily on clear weather conditions; cross-season appearance changes and deep shadows may degrade feed-forward depth estimation.
  • Dynamic scene elements: The current formulation assumes a static rigid world, leading to minor coordinate distortions on high-speed moving vehicles or water bodies.
  • End-to-end geodetic constraint integration: Geodetic reprojection currently acts as a decoupled post-processing step; future work could integrate gravity alignment and horizontal ground priors directly into the network architecture.
  • vs TransGeo & University-1652: Retrieval methods only return approximate image-level GPS estimates from discrete reference galleries, failing to capture intra-image spatial geometry; SkyReg predicts dense geodetic coordinates for every pixel.
  • vs SuperPoint+SuperGlue & RoMa: Homography-based matchers struggle with large parallax and height discontinuities, producing severe shearing and stretching; SkyReg explicitly models 3D elevation to handle complex urban geometry.
  • vs Neural Radiance Fields / 3DGS-based registration: Multi-view rendering approaches demand expensive image sequences and test-time optimization; SkyReg performs single-image feed-forward inference in a few hundred milliseconds.

Rating

  • Novelty: โญโญโญโญโญ Pioneering feed-forward 3D geometry modeling for single-view cross-view pixel-level geo-registration.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive benchmarks spanning thousands of unseen urban/suburban images, pixel-wise and image-level metrics, and rigorous ablations.
  • Writing Quality: โญโญโญโญโญ Rigorous mathematical modeling, clean task formulation, and coherent structural presentation.
  • Value: โญโญโญโญโญ Directly addresses crucial pain points in drone navigation, disaster response, and urban spatial analysis.