Sector-Level Cross-View Geo-Localization with Implicit Orientation via Azimuthal Scanning¶
Conference: ECCV 2026
Paper: ECCV Official
Project Page: https://1203ll.github.io/SectorGeo/
Area: Remote Sensing / Image Retrieval
Keywords: Cross-View Geo-Localization, Limited Field-of-View, Azimuthal Scanning, Sector-Level Matching, Implicit Orientation Estimation
TL;DR¶
Addressing the geometric mismatch and background noise between limited-FoV ground queries and omnidirectional aerial images, this paper introduces the Sector-Level Cross-View Geo-Localization (SLCVGL) task and the SectorGeo framework, which decouples aerial feature maps into overlapping azimuthal sectors to achieve fine-grained sector-level sub-region retrieval and implicit orientation estimation without continuous pose supervision.
Background & Motivation¶
Cross-View Geo-Localization (CVGL) is a foundational vision-based technique for GPS-denied autonomous navigation, unmanned aerial vehicle (UAV) positioning, and augmented reality. However, bridging the severe perspective disparity between ground and aerial views remains notoriously difficult due to extreme geometric distortions, perspective shifts, and dramatic visual appearance variations. Existing formulations generally follow three main paradigms: image-level retrieval, cross-view continuous pose regression, and object-level geo-localization.
Nevertheless, existing paradigms remain caught in an impasse between coarse retrieval granularity and prohibitive fine-grained supervision costs. Coarse image-level methods (e.g., Sample4Geo, PLGeo) scale efficiently across vast databases, but treat the entire overhead scene as an atomic entity, remaining oblivious to where within the aerial frame the camera sits or where it is looking. Conversely, continuous 3-DoF pose estimation formulates the task as coordinate and orientation regression under strict pairwise setups, suffering from optimization instability across severe domain gaps. Furthermore, object-level geo-localization relies heavily on labor-intensive bounding-box or segmentation annotations, rendering real-world scaling practically infeasible.
Returning to the underlying geometry, a limited Field-of-View (FoV) ground camera captures only a narrow horizontal perspective of its surroundings. When projected onto the omnidirectional aerial footprint, this horizontal visual cone naturally maps to an azimuthal wedge-shaped sector whose angular span equals the ground FoV and whose central azimuth corresponds directly to the camera's geographical heading. Core idea: reformulate fine-grained localization as Sector-Level Cross-View Geo-Localization (SLCVGL), using azimuthal scanning to decompose aerial features into overlapping candidate sectors, jointly retrieving the fine-grained aerial sub-region and implicitly inferring the ground heading via discrete sector matching without continuous pose regression or manual bounding-box supervision.
Method¶
Overall Architecture¶
SectorGeo operationalizes this geometric inclusion relationship through a parameter-shared dual-branch ConvNeXt-Base backbone. The ground branch extracts a global semantic feature vector \(\mathbf{f}^q \in \mathbb{R}^C\). The aerial branch retains the 2D spatial feature map \(\mathbf{F}^a \in \mathbb{R}^{C \times H \times W}\) and applies rotating binary sector masks centered at the image origin to decouple the feature map into an orientation-aware candidate sequence \(\mathbf{V}^a = \{\mathbf{v}_p\}_{p=1}^P\).
During feature matching, the architecture splits into two complementary paths: the Azimuthal Alignment Stream (AAS) computes lightweight cosine similarities between the ground query and candidate sectors to anchor the optimal sector index \(p^*\); the Relational Retrieval Stream (RRS) reconstructs heading-conditioned ground contexts via Azimuthal Context Aggregation (ACA) and models global spatial consensus across all sectors using Cross-Sector Relational Reasoning (CSRR) to produce a robust image retrieval score \(S_{ret}\). At inference, candidates are ranked first by \(S_{ret}\) at the image level, and AAS is evaluated on the top candidates to identify the exact sub-region and discrete orientation \(\hat{\theta}\).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Ground Query Iq (FoV=ฯ)<br/>and Omnidirectional Aerial Ia"] --> B["Dual-Branch Feature Extraction<br/>Shared ConvNeXt-Base Backbone"]
B --> C["Azimuthal Feature Decoupling<br/>Rotating Sector Masks and Masked Pooling"]
C --> D["Azimuthal Context Aggregation<br/>Direction-Adaptive Context Reconstruction"]
D --> E["Cross-Sector Relational Reasoning<br/>Self-Attention Graph Modeling Consensus"]
C --> F["Azimuthal Alignment Stream<br/>Cosine Geometric Sector Anchoring"]
E --> G["Relational Retrieval Output<br/>Global Image Similarity Sret"]
F --> H["Sector Localization & Heading Output<br/>Optimal Sector p* and Orientation ฮธ"]
G --> I["Coarse-to-Fine Joint Inference"]
H --> I
Key Designs¶
1. Azimuthal Feature Decoupling: Discretizing Continuous Overhead Space into Directional Sequences
Standard aerial representations encompass complete 360ยฐ surroundings, whereas a narrow ground image observes only a localized directional sub-cone. Direct global matching inevitably suffers from severe noise injected by unobserved overhead sectors. To address this spatial discrepancy, SectorGeo generates \(P = \lfloor 360^\circ / \Delta\theta \rfloor\) overlapping binary sector masks \(\{\mathbf{M}_p\}_{p=1}^P\) of angular width \(\phi\) rotated in steps of \(\Delta\theta\). Spatial features within each mask are aggregated using masked average pooling with \(L_2\) normalization: $\(\mathbf{v}_p = \left\| \frac{\sum_{h=1}^H \sum_{w=1}^W \mathbf{F}^a(h,w) \mathbf{M}_p(h,w)}{\sum_{h=1}^H \sum_{w=1}^W \mathbf{M}_p(h,w) + \epsilon} \right\|_2\)$ Normalizing by the mask area ensures scale invariance against boundary truncations, converting the dense 2D overhead representation into an orderly sequence of discrete geographic candidates.
2. Azimuthal Context Aggregation and Cross-Sector Relational Reasoning: Suppressing Local Visual Ambiguities
Evaluating isolated sector-level correspondences independently is susceptible to false positives because distinct geographical landmarks often exhibit similar local textures. The Relational Retrieval Stream overcomes this via a two-stage mechanism. First, Azimuthal Context Aggregation (ACA) calculates raw directional affinities \(\mathbf{A} = \text{LeakyReLU}(\mathbf{V}^a \mathbf{f}^q)\), which are normalized and passed through a smoothed Softmax to modulate the query vector across headings: \(\mathbf{C}_g = \text{Softmax}(\mathbf{A} / \|\mathbf{A}\|_2) (\mathbf{f}^q)^\top\). Next, Cross-Sector Relational Reasoning (CSRR) takes the fused node features \(\mathbf{E}_0 = (\mathbf{C}_g \odot \mathbf{V}^a)\mathbf{W}_t\) and runs \(T\) iterations of graph self-attention: $\(\mathbf{M}_t = \text{Softmax}\left((\mathbf{E}_{t-1}\mathbf{W}_Q)(\mathbf{E}_{t-1}\mathbf{W}_K)^\top\right), \quad \mathbf{E}_t = \text{ReLU}(\mathbf{M}_t \mathbf{E}_{t-1}\mathbf{W}_g)\)$ Averaging and projecting the final node representations produces the holistic retrieval score \(S_{ret}\). By reasoning over the entire 360ยฐ geometric configuration, CSRR verifies whether directional alignments conform to the global scene layout.
3. Ground-Truth Soft Labeling: Smoothing Orientation Discretization Ambiguities
In the Azimuthal Alignment Stream, matching responses are computed simply via \(\mathbf{S}_{align} = \mathbf{V}^a \mathbf{f}^q \in \mathbb{R}^P\). However, arbitrary ground headings \(\theta_{gt}\) rarely coincide perfectly with the center of a discrete sector. Training with hard one-hot targets severely penalizes adjacent sectors that partially cover the true view, creating noisy gradient updates. Ground-Truth Soft Labeling identifies the two nearest discrete sectors \(\{p_1, p_2\}\) and assigns weights \(w_1, w_2\) via linear angular interpolation (\(w_1 + w_2 = 1\)) to form a target distribution \(\mathbf{y} \in \mathbb{R}^P\). This continuous geometric prior delivers a smooth optimization landscape for robust discrete classification.
Loss & Training¶
The framework is optimized end-to-end under a balanced multi-task objective combining image retrieval and sector alignment: $\(\mathcal{L} = \beta \mathcal{L}_{loc} + (1 - \beta) \mathcal{L}_{img}\)$ The image-level retrieval loss \(\mathcal{L}_{img}\) employs InfoNCE with a learnable temperature \(\tau\): $\(\mathcal{L}_{img} = -\frac{1}{B} \sum_{i=1}^B \log \frac{\exp(S_{ret}(i, i) / \tau)}{\sum_{j=1}^B \exp(S_{ret}(i, j) / \tau)}\)$ The sector localization loss \(\mathcal{L}_{loc}\) adopts soft-target cross-entropy: $\(\mathcal{L}_{loc} = -\frac{1}{B} \sum_{i=1}^B \sum_{p=1}^P y_p^{(i)} \log \left(\text{Softmax}\left(\mathbf{S}_{align}^{(i)}\right)_p\right)\)$ The balancing factor \(\beta\) is set to 0.9 on CVUSA and 0.3 on CVACT, with \(T = 3\) CSRR iterations across all experiments.
Key Experimental Results¶
Main Results¶
Evaluations report standard Image-Level Recall (\(R@K\)) alongside a stricter joint metric, Sector-Level Recall (\(S\text{-}R@K\)), which requires the true aerial image to rank within the top \(K\) candidates and the predicted sector index to match the dominant ground-truth sector. The table below presents performance under challenging narrow FoVs (\(70^\circ\) and \(90^\circ\)) on repurposed CVUSA and CVACT benchmarks.
| Dataset | FoV | Method | R@1 (%) | R@5 (%) | S-R@1 (%) | S-R@5 (%) |
|---|---|---|---|---|---|---|
| CVUSA | \(70^\circ\) | GeoDTR+ | 4.78 | 10.24 | 3.67 | 5.84 |
| CVUSA | \(70^\circ\) | Sample4Geo | 35.33 | 62.83 | 34.20 | 60.73 |
| CVUSA | \(70^\circ\) | ConGeo | 40.49 | 64.13 | 39.74 | 62.66 |
| CVUSA | \(70^\circ\) | PLGeo | 44.56 | 69.35 | 28.69 | 44.17 |
| CVUSA | \(70^\circ\) | Ours (SectorGeo) | 63.79 | 84.62 | 59.92 | 78.47 |
| CVACT | \(70^\circ\) | GeoDTR+ | 4.24 | 9.24 | 3.24 | 5.24 |
| CVACT | \(70^\circ\) | Sample4Geo | 6.96 | 18.73 | 6.24 | 16.69 |
| CVACT | \(70^\circ\) | ConGeo | 8.31 | 20.05 | 7.60 | 18.25 |
| CVACT | \(70^\circ\) | PLGeo | 11.33 | 26.47 | 9.78 | 21.97 |
| CVACT | \(70^\circ\) | Ours (SectorGeo) | 41.50 | 64.27 | 39.81 | 61.02 |
| CVUSA | \(90^\circ\) | PLGeo | 53.94 | 77.96 | 40.42 | 58.34 |
| CVUSA | \(90^\circ\) | Ours (SectorGeo) | 72.44 | 89.87 | 66.41 | 81.43 |
| CVACT | \(90^\circ\) | PLGeo | 19.69 | 36.47 | 15.53 | 26.55 |
| CVACT | \(90^\circ\) | Ours (SectorGeo) | 49.30 | 71.05 | 40.75 | 56.74 |
Ablation Study¶
A comprehensive component ablation on CVUSA (\(\text{FoV}=180^\circ\)) illustrates the relative importance of relational reasoning and feature pooling strategies.
| Config | CSRR | ACA | Pooling | R@1 (%) | R@5 (%) | S-R@1 (%) | S-R@5 (%) |
|---|---|---|---|---|---|---|---|
| w/o CSRR | โ | โ | Avg | 77.74 | 92.23 | 75.74 | 89.82 |
| w/o ACA | โ | โ | Avg | 84.48 | 94.81 | 82.86 | 91.97 |
| Max Pooling | โ | โ | Max | 81.40 | 94.20 | 79.60 | 91.12 |
| Mixed Pooling | โ | โ | Mix | 81.28 | 94.44 | 79.78 | 91.75 |
| Full Model | โ | โ | Avg | 87.87 | 96.75 | 83.67 | 91.12 |
Key Findings¶
- Dominance in Narrow FoVs: In the challenging \(70^\circ\) setting, competing methods degrade severely (e.g., PLGeo attains only 11.33% R@1 on CVACT), whereas SectorGeo reaches 41.50% R@1 and 39.81% S-R@1, exceeding prior baselines by over 30 percentage points. Suppressing unobserved overhead background is paramount when the mutual field is tiny.
- CSRR Drives Global Consistency: Removing CSRR drops R@1 by 10.13% (from 87.87% to 77.74%), verifying that local directional matches require relational cross-sector graph propagation to rule out structural ambiguities.
- Average Pooling Outperforms Salient Peak Pooling: Masked average pooling (87.87% R@1) outperforms max pooling (81.40%) and mixed pooling (81.28%), indicating that directional matching benefits from holistic spatial distributions rather than sparse local activations.
- Step Size Granularity Trade-Off: Sampling at \(\Delta\theta = 50^\circ\) (\(P = 7\) sectors) provides the optimal balance. Increasing \(\Delta\theta\) to \(100^\circ\) inflates quantization errors, while over-dense sampling at \(40^\circ\) introduces feature redundancy and slight over-fitting.
Highlights & Insights¶
- Reformulation of Continuous Pose Estimation: Transforming intractable 3-DoF continuous pose regression into geometrically grounded discrete sector classification eliminates gradient instability and costly ground-truth pose requirements.
- Dual-Stream Reciprocal Synergy: RRS and AAS reinforce each other: discriminative global retrieval gradients backpropagate to refine directional representations, while AAS provides localized geometric anchors that bolster image retrieval consistency.
- Soft Labeling for Continuous Heading: Linear interpolation of discrete sector targets based on angular distance circumvents the rigid boundary penalties of one-hot classification, delivering a smooth optimization trajectory.
Limitations & Future Work¶
- Center-Point Invariance Assumption: Sector masks currently assume the camera viewpoint coincides with the geometric center of the aerial tile. Extreme camera offsets towards tile borders induce radial perspective distortions not fully captured by centered wedges.
- Known FoV Dependency: The scanning formulation requires the ground horizontal FoV \(\phi\) a priori, which may be uncalibrated in consumer mobile photos or zoom cameras.
- Absence of 3D Elevation Modeling: Dense urban high-rises cause strong vertical occlusions and parallax shifts that planar 2D sector masks do not explicitly account for.
Related Work & Insights¶
- vs PLGeo (ACM MM 2025): While PLGeo relies on patch-level shuffling to mitigate rotational variance, it remains confined to holistic image retrieval. In \(70^\circ\) CVACT, PLGeo achieves only 11.33% R@1, whereas SectorGeo achieves 41.50% R@1 while additionally outputting fine-grained sectors and orientations.
- vs SliceMatch (CVPR 2023): SliceMatch employs circular convolution for continuous pose regression, facing high optimization complexity and rigid pairwise training constraints; SectorGeo converts the problem into scalable discrete sector retrieval.
- vs ConGeo (ECCV 2024): ConGeo enhances ground-view invariance via contrastive regularization but does not explicitly filter unobserved overhead directions; SectorGeo directly suppresses background noise through azimuthal feature masking.
Rating¶
- Novelty: โญโญโญโญ [Clever reframing of cross-view fine-grained localization into discrete sector scanning without manual pose supervision]
- Experimental Thoroughness: โญโญโญโญโญ [Rigorous benchmarking across multiple restricted FoVs, custom S-R@K protocol, deep component ablations, and step-size sweeps]
- Writing Quality: โญโญโญโญโญ [Clear mathematical definitions, compelling geometric motivation, and coherent prose presentation]
- Value: โญโญโญโญ [Substantial practical utility for GPS-denied drone navigation and mobile robot positioning]