Trajectory-aware Cross-view Geo-Localization with Sequential Observations¶
Conference: ECCV 2026
Paper: CVF Open Access
Code: https://humblegamer.github.io/trajloc/
Area: Autonomous Driving
Keywords: Cross-view Geo-Localization, Sequential Observations, Trajectory Modulation, Route Descriptions, Multimodal Retrieval
TL;DR¶
To tackle the lack of linguistic route narratives and allocentric spatial layout awareness in cross-view geo-localization, this paper introduces SeqGeo-VL comprising ~39K video-text-satellite triplets and proposes TrajLoc, a unified framework leveraging two-stage curriculum learning and a lightweight TrajMod module that significantly enhances both video and text retrieval.
Background & Motivation¶
In dense urban environments characterized by towering skyscrapers, narrow street canyons, and heavy tree canopies, GNSS signals suffer severe multipath degradation or total blockage, making reliable global positioning unavailable for embodied AI systems such as autonomous vehicles and quadruped delivery robots. Cross-view geo-localization addresses this challenge by retrieving matching GPS-referenced overhead satellite imagery using ground-level observations, providing an infrastructure-free, low-storage global localization alternative. Early methodologies primarily rely on single-frame ground-level panoramic photos; however, in modern repetitive urban street layouts and uniform architectural facades, single-frame queries inevitably encounter catastrophic perceptual aliasing.
To alleviate visual ambiguities, recent research has shifted toward sequential observations, utilizing temporal continuity and viewpoint changes across trajectories. Nevertheless, existing video-based retrieval frameworks focus exclusively on dense visual pixels, ignoring high-level linguistic route descriptions. In collaborative human-robot scenarios such as pickup assistance or voice navigation, users naturally communicate routes using abstract spatiotemporal instructions (e.g., "head down the avenue, turn right past the brick plaza, and halt before the glass tower"). Furthermore, when vehicle cameras suffer occlusions or sensor degradation, linguistic narratives serve as the sole localization input. Concurrently, standard Vision-Language Models (VLMs) and CLIP variants act largely as bag-of-words estimators, lacking egocentric-to-allocentric spatial layout reasoning.
The paper attacks this dilemma by unifying video and linguistic route queries while grounding representations in physical trajectory geometry. Core idea: build a multimodal benchmark of aligned video-text-satellite triplets, and introduce TrajLoc, a unified cross-view retrieval framework that employs a two-stage curriculum with drift regularization alongside a lightweight TrajMod module to condition embeddings on trajectory geometry.
Method¶
Overall Architecture¶
TrajLoc is formulated to retrieve matching geo-tagged satellite patches given either sequential ground video clips or long-form linguistic route descriptions. The framework employs three separate encoders initialized from pretrained CLIP ViT-L/14: a video encoder, a text encoder with linearly extended positional embeddings, and a shared satellite encoder. To resolve optimization disparities arising from differing cross-modal alignment difficulties, TrajLoc utilizes a two-stage curriculum learning strategy. During downstream retrieval, all feature backbones remain frozen, and a lightweight Trajectory-conditioned Modulation module (TrajMod) takes heading angles and origin-destination bearings to generate affine scale and shift parameters, yielding spatially grounded query representations.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
InV["Sequential Ground Video<br/>N frames sampled along route"] --> EncV["Video Encoder<br/>ViT-L/14 + Mean Pooling"]
InT["Route Description<br/>Long-form narrative trajectory"] --> EncT["Text Encoder<br/>Linear positional interpolation"]
InS["Satellite Gallery<br/>Overhead geo-tagged tiles"] --> EncS["Satellite Encoder<br/>ViT-L/14 Shared Soft Anchor"]
EncV --> S1["Two-stage Curriculum<br/>Stage 1: Video-satellite contrastive alignment"]
EncS --> S1
S1 --> S2["Anchor Regularization<br/>Stage 2: Freeze video branch, align text with cosine constraint"]
EncT --> S2
Meta["Trajectory Metadata<br/>Heading sequence + OD bearing"] --> TM1["Trajectory Fourier Encoding<br/>Multi-scale harmonic mapping"]
TM1 --> TM2["Affine Feature Modulation<br/>TrajMod predicts scale and shift"]
S2 --> TM2
TM2 --> Out["Spatially Grounded Embedding<br/>Cosine similarity retrieval against satellite gallery"]
Key Designs¶
1. Two-stage Curriculum Learning with Drift Regularization: Bridging Asymmetric Multimodal Alignment Gaps
Directly aligning abstract, long-form route narratives with dense satellite imagery is substantially more difficult than matching ground videos to satellite views. Joint end-to-end co-training causes severe gradient conflicts that destabilize the shared satellite encoder. To prevent optimization degradation, TrajLoc introduces a two-stage curriculum learning pipeline. In Stage 1, the framework trains the video and satellite encoders using a symmetric InfoNCE contrastive objective, establishing an initial shared visual manifold between ground physical structures and overhead satellite layouts. In Stage 2, the video encoder is frozen, and the satellite encoder acts as a soft anchor to guide the alignment of the text encoder. To prevent catastrophic representation drift in the satellite encoder during text back-propagation, TrajLoc enforces a cosine-similarity regularization loss \(\mathcal{L}_{\text{reg}}\):
This explicit geometric constraint preserves the fine-grained visual correspondence established in Stage 1 while enabling the abstract text representations to smoothly integrate into the established satellite feature space.
2. Trajectory Geometry Fourier Encoding: Capturing Scale-Free Directional Priors
Standard foundation backbones fail to encode explicit anisotropic spatial orientations. Recognizing that consumer vehicles and smart robots inherently carry basic IMUs and electronic compasses, TrajLoc extracts two complementary directional cues: local turning angles at sequential waypoints \(\mathbf{h} = \{h_1, \dots, h_T\}\) and the global origin-to-destination (OD) bearing \(\theta_{\text{OD}}\), all referenced to True North. To allow neural networks to capture angular periodicity and high-frequency directional shifts, each angle \(\alpha \in \mathbf{h} \cup \{\theta_{\text{OD}}\}\) is transformed via a multi-scale Fourier basis:
Concatenating the harmonic features of sequential headings and OD bearing creates a compact, coordinate-free trajectory geometry embedding \(\mathbf{z}_{\text{traj}}\) that avoids brittle dependencies on metric-scale Cartesian coordinates or high-definition maps.
3. Affine Trajectory Modulation with Cross-Modal Regularization: Injecting Allocentric Spatial Layouts
To condition query representations on physical spatial geometry, TrajMod instantiates two modality-specific multi-layer perceptrons, \(\text{MLP}_v\) and \(\text{MLP}_t\), motivated by feature-wise affine modulation. For a video embedding \(\mathbf{f}_v\) (obtained via temporal mean pooling) or a route text embedding \(\mathbf{f}_t\) (handling an average of 129 tokens via extended positional encodings), TrajMod predicts dimension-matched scale vectors \(\boldsymbol{\gamma}_m\) and shift vectors \(\boldsymbol{\beta}_m\):
During TrajMod training, the three primary feature encoders remain completely frozen, training only the compact MLP layers. To prevent independent modulations from fitting modality-specific artifacts, TrajMod introduces an auxiliary text-to-video contrastive objective alongside the satellite contrastive losses. Because both modalities share the identical trajectory geometry \(\mathbf{z}_{\text{traj}}\), this cross-modal constraint enforces \(\mathbf{f}'_v\) and \(\mathbf{f}'_t\) to align tightly along shared layout invariants, purifying the representations toward physical topology.
Loss & Training¶
Both curriculum stages optimize a symmetric InfoNCE loss:
where \((q_i, k_i)\) denotes paired query-satellite instances within a batch of size \(N\), and \(\ell(\cdot, \cdot)\) denotes temperature-scaled cross-entropy over cosine similarities. Stage 2 optimizes \(\mathcal{L}_{\text{stage2}} = \mathcal{L}_{\text{ctr}}(T, S) + \lambda \mathcal{L}_{\text{reg}}\). The final TrajMod stage optimizes the combined contrastive loss across video-satellite, text-satellite, and text-video pairs. The architecture uses Adam with a base learning rate of \(1 \times 10^{-5}\) across 40 epochs for encoder adaptations and 20 epochs for TrajMod.
Key Experimental Results¶
Main Results¶
Evaluation is conducted on the SeqGeo-VL benchmark (80:20 split, 38,863 sequential triplets) across both cross-view video geo-localization and cross-view text geo-localization tasks using Recall@K (R@1, R@5, R@10, and top 1% recall R@1%).
Cross-view Video Geo-localization Comparison (Table 1):
| Method | Backbone | R@1 (%) | R@5 (%) | R@10 (%) | R@1% (%) |
|---|---|---|---|---|---|
| SeqGeo (WACV 2023) | VGG16 | 1.80 | 6.45 | 10.36 | 34.38 |
| SeqGeoโ (Reimplementation) | ViT-L/14 | 8.14 | 24.42 | 34.02 | 66.17 |
| GARet (ECCV 2024) | DeiT-m | 3.34 | 11.19 | 17.18 | 44.39 |
| FlexGeo (ISPRS 2026) | ConvNext-B | 3.51 | 14.22 | 20.77 | 48.53 |
| Qwen3-VL-Embedding* (LoRA) | Qwen3-VL-2B | 5.44 | 19.21 | 28.15 | 61.16 |
| TrajLoc (Ours) | ViT-L/14 | 12.09 | 35.77 | 47.68 | 80.21 |
Cross-view Text Geo-localization Comparison (Table 2):
| Method | Backbone | R@1 (%) | R@5 (%) | R@10 (%) | R@1% (%) |
|---|---|---|---|---|---|
| CLIP | ViT-B/16 | 0.53 | 1.58 | 3.33 | 15.94 |
| CLIP | ViT-L/14 | 0.80 | 3.32 | 6.01 | 22.99 |
| EVA2-CLIP | ViT-B/16 | 0.63 | 2.20 | 3.98 | 17.59 |
| EVA2-CLIP | ViT-L/14@336 | 0.89 | 3.78 | 6.70 | 26.99 |
| SigLIP | ViT-SO400M/14 | 0.80 | 2.92 | 5.24 | 21.86 |
| Perception Encoder | ViT-L/14@336 | 0.66 | 2.08 | 3.80 | 15.17 |
| Qwen3-VL-Embedding* (LoRA) | Qwen3-VL-2B | 0.93 | 3.38 | 5.55 | 20.91 |
| CrossText2Loc (ICCV 2025) | ViT-L/14@336 | 0.98 | 4.08 | 7.08 | 25.45 |
| TrajLoc (Ours) | ViT-L/14 | 2.52 | 9.82 | 16.11 | 45.48 |
Ablation Study¶
Systematic ablations analyze the contributions of TrajMod components, co-training synergies, and curriculum design (Table 3):
| Config / Ablation Step | Video R@1 (%) | Video R@10 (%) | Video R@1% (%) | Text R@1 (%) | Text R@10 (%) | Text R@1% (%) | Note |
|---|---|---|---|---|---|---|---|
| Full model (TrajLoc) | 12.09 | 47.68 | 80.21 | 2.52 | 16.11 | 45.48 | Full model |
| w/o Txt2Vid contrastive loss | 10.23 | 41.21 | 74.20 | 2.17 | 14.19 | 42.22 | Excludes cross-modal regularization |
| w/o TrajMod | 7.27 | 30.92 | 63.47 | 0.98 | 7.21 | 26.15 | Removes geometric modulation |
| Vid2Sat only (w/ TrajMod) | 9.69 | 40.30 | 73.60 | โ | โ | โ | Single modality training |
| Txt2Sat only (w/ TrajMod) | โ | โ | โ | 1.90 | 15.57 | 47.25 | Single modality training |
| Vid2Sat only (w/o TrajMod) | 6.87 | 31.60 | 64.09 | โ | โ | โ | Baseline unimodal model |
| Txt2Sat only (w/o TrajMod) | โ | โ | โ | 0.80 | 6.01 | 22.99 | Baseline unimodal model |
| w/o Curriculum learning | 5.01 | 27.25 | 61.81 | 0.86 | 6.52 | 24.40 | Direct uncurricled co-training |
Key Findings¶
- Crucial Impact of TrajMod: Removing TrajMod causes a sharp drop in Video R@1 from 12.09% to 7.27% (a ~40% relative decrease) and in Text R@1 from 2.52% to 0.98% (a ~61% relative drop). General pre-trained vision-language models struggle to comprehend viewpoint-to-satellite spatial transformations without explicit trajectory guidance.
- Mutual Synergy Amplification: Without TrajMod, co-training yields only marginal gains (Video R@1 improves modestly from 6.87% to 7.27%). With TrajMod present, co-training gains are dramatically amplified (Video R@1 leaps from 9.69% to 12.09%, and Text R@1 from 1.90% to 2.52%), demonstrating that TrajMod grounds disparate modalities onto a shared geometric anchor.
- Temporal Sequence Scaling: As the observation sequence expands from 1 to 6 frames, retrieval recall steadily increases across all models; moreover, the margin enabled by TrajMod widens from +2.6% R@1 at single-frame to +5.6% R@1 at 6 frames, showing that trajectory modulation effectively resolves feature smoothing in extended trajectories.
Highlights & Insights¶
- Explicit Modulation Beats Text Prompting: Feeding trajectory metadata into MLLMs via textual prompting produces a modest +1.34% boost in R@1%, whereas TrajMod's affine parameter modulation delivers a dramatic +19.33% gain, confirming that continuous geometric conditioning is far more expressive than discrete linguistic tokens for spatial reasoning.
- Minimalist Aggregation Outperforms Heavy Transformers: Employing simple mean pooling over video frames alongside TrajMod comfortably surpasses complex autoregressive transformer aggregators (such as GARet and SeqGeo), demonstrating that the core bottleneck in cross-view matching lies in geometric alignment rather than temporal attention stacking.
- Reusable Curriculum for Asymmetric Retrieval: The strategy of using a dense visual modality to warm up a shared aerial backbone followed by soft-anchor drift regularization offers a generalizable blueprint for other asymmetric multimodal retrieval tasks.
Limitations & Future Work¶
- Single Scale Overhead Imagery: The benchmark evaluates on fixed Level 20 satellite tiles and does not account for cross-season changes, severe illumination variances, or structural urban renovations.
- Strict Trajectory Continuity Requirement: TrajMod assumes uncorrupted heading and bearing readings; erratic sensor noise, severe wheel slippage, or IMU drift in consumer hardware may compromise modulation accuracy.
- Future Directions: Exploring self-supervised odometry estimation directly from ground imagery and incorporating road topology priors to bridge the gap from satellite patch retrieval to continuous metric coordinate regression.
Related Work & Insights¶
- vs SeqGeo / GARet / FlexGeo: Prior sequential works heavily engineered complex temporal attention adapters while remaining confined to video queries and ignoring compass orientations. TrajLoc proves that explicit trajectory conditioning allows a lightweight mean-pooling pipeline to substantially outperform complex architectures.
- vs CrossText2Loc: CrossText2Loc pioneered natural language cross-view matching on isolated, single-frame street views; TrajLoc scales natural language geo-localization to trajectory-level route narratives supported by the SeqGeo-VL benchmark.
Rating¶
- Novelty: โญโญโญโญโญ Pioneering sequential route text geo-localization and proposing TrajMod for explicit trajectory geometry conditioning.
- Experimental Thoroughness: โญโญโญโญโญ Complete evaluations across both video and text tasks, extensive ablations, sequence length analyses, and search radius simulations.
- Writing Quality: โญโญโญโญโญ Rigorous task formulation, clear architectural illustrations, and self-consistent empirical analysis.
- Value: โญโญโญโญโญ Establishes a foundational benchmark and methodology for GPS-denied robot localization, autonomous driving, and human-agent cooperative navigation.