SARA: Structure-Aware Riemannian-Guided Alignment for Drone Image-Text Retrieval¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Area: Remote Sensing
Keywords: Drone Image-Text Retrieval, Riemannian Manifold, Symmetric Positive Definite Matrix, Log-Euclidean Metric, Structure-Aware Alignment
TL;DR¶
SARA proposes a structure-aware Riemannian-guided alignment framework for drone image-text retrieval, capturing second-order spatial topologies via symmetric positive definite (SPD) covariance descriptors on the Riemannian manifold and enforcing marginal and conditional distribution alignment in the tangent space under the Log-Euclidean metric to overcome viewpoint distortion and representation-level geometric inconsistency.
Background & Motivation¶
Drone image-text retrieval (DITR) plays an indispensable role in low-altitude surveillance, disaster assessment, and urban monitoring. Unlike ground-level photography or satellite remote sensing, drone imagery is acquired from varying flight altitudes and oblique viewing angles, leading to substantial perspective transformations, drastic scale drift, and irregular spatial layouts of scene entities. Conventional retrieval architectures mostly adhere to flat Euclidean embedding spaces and rely on contrastive objectives or cross-attention matching (e.g., VSE++, CLIP, VCSR), presuming that inner products or cosine metrics in flat vector spaces suffice for multimodal semantic association.
However, spatial transformations induced by continuous viewpoint shifts in low-altitude aerial scenes are inherently nonlinear. While relative projective configurations undergo extreme distortion across camera perspectives, underlying structural dependencies and intrinsic topological adjacencies among visual entities remain invariant. Conventional Euclidean representations approximate these curved geometric relationships using linear distances, inevitably suffering from representation-level geometric inconsistency and cross-modal misalignment. Recent explorations into non-Euclidean geometry (such as hyperbolic embeddings in MERU and HyCoCLIP) primarily employ curved spaces as global containers for hierarchical semantic containment, failing to explicitly capture fine-grained second-order spatial correlations and local geometric topologies in aerial observations.
The key entry point of this paper is that Riemannian geometry provides an intrinsic structural inductive bias, where second-order feature statistics can faithfully model spatial structures invariant to viewpoint alterations, naturally complementing flat Euclidean semantic representations. Core idea: construct a Structure-Aware Riemannian-Guided Alignment (SARA) framework that captures second-order structural dependencies on the symmetric positive definite (SPD) Riemannian manifold, conducts joint marginal and conditional distribution alignment in the tangent space via the Log-Euclidean metric, and fuses manifold structural representations with Euclidean semantic embeddings for geometry-consistent retrieval.
Method¶
Overall Architecture¶
SARA employs a decoupled dual-branch architecture spanning semantic content modeling and Riemannian structural reasoning. Given an image-text pair with visual input \(I \in \mathbb{R}^{3 \times H \times W}\) and description \(T = \{w_1, \dots, w_L\}\), the semantic branch extracts first-order global vectors \(v_{\mathrm{sem}}\) and \(t_{\mathrm{sem}}\) using RestV2-Base and BERT encoders. Simultaneously, the structural branch leverages the Manifold Structural Feature Extraction (MSFE) module to convert visual feature maps and text token sequences into symmetric positive definite (SPD) matrices on the Riemannian manifold, refines them through hierarchical manifold layers, and maps them to the tangent space. Subsequently, the Log-Euclidean Structural Alignment (LESA) module enforces marginal distribution alignment (MDA) and conditional distribution alignment (CDA) in the tangent space. Finally, semantic and structural descriptors are fused and optimized via a second-stage contrastive loss.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Drone Images and Text Queries"] --> B["Semantic Branch: Global Semantic Representation"]
A --> C["Manifold Structural Feature Extraction (MSFE)<br/>Second-order Covariance & Manifold Layers"]
B --> D["Semantic Contrastive Alignment<br/>Bidirectional InfoNCE Loss"]
C --> E["Log-Euclidean Structural Alignment (LESA)<br/>Tangent Space MDA & CDA Alignment"]
B & E --> F["Semantic-Structural Fusion & Second-Stage Alignment<br/>Joint Projection & Symmetric Contrastive Optimization"]
D & E & F --> G["Output: Bidirectional Cross-Modal Retrieval"]
Key Designs¶
1. Manifold Structural Feature Extraction (MSFE): Capturing Second-Order Dependencies on the SPD Manifold To address the inability of first-order Euclidean embeddings to preserve nonlinear geometric layouts, MSFE extracts the intermediate feature map \(F_v \in \mathbb{R}^{C_v \times N}\) from the fourth stage of the visual backbone (\(C_v\) channels, \(N\) spatial tokens) and computes a second-order sample covariance descriptor: $\(C_I^{(0)} = \frac{1}{N}\sum_{i=1}^{N}(\mathbf{f}_v^{i} - \bar{\mathbf{f}}_v)(\mathbf{f}_v^{i} - \bar{\mathbf{f}}_v)^\top + \epsilon_I \mathbf{I}\)$ The perturbation \(\epsilon_I \mathbf{I}\) guarantees strict positive definiteness, embedding \(C_I^{(0)} \in \mathrm{Sym}_{C_v}^+\) on the Riemannian manifold to capture inter-channel correlations and spatial layouts. MSFE refines this descriptor through four Riemannian transformation layers: - BiMap Layer: Maps the matrix to a lower-dimensional manifold via \(C^{(k)} = W_k^\top C^{(k-1)} W_k\) using semi-orthogonal weights \(W_k\) on the Stiefel manifold, preserving manifold geometry; - ReEig Layer: Decomposes \(C^{(k)} = U \Sigma U^\top\) and rectifies eigenvalues via threshold \(\zeta > 0\): \(\tilde{C}^{(k)} = U \max(\Sigma, \zeta \mathbf{I}) U^\top\) to inject non-linearity while preventing singularity; - ReCov Layer: Amplifies off-diagonal covariance entries using coefficient \(\lambda\): \(\bar{C}^{(k)} = \tilde{C}^{(k)} + \lambda (\tilde{C}^{(k)} - \mathrm{diag}(\tilde{C}^{(k)}))\), enhancing inter-channel dependencies; - LogEig Layer: Projects the SPD matrix from the Riemannian manifold into its Euclidean tangent space via the Log-Euclidean metric: \(\hat{C}^{(k)} = \log(\bar{C}^{(k)}) = U \log(\Sigma) U^\top\). On the textual side, token embeddings from the final Transformer layer are processed by an MLP and transformed into a covariance matrix, which is then mapped to the visual manifold dimension using a shared semi-orthogonal matrix \(W_s\): \(C_T = \log(W_s \mathrm{Cov}(\mathrm{MLP}(F_t)) W_s^\top)\).
2. Log-Euclidean Structural Alignment (LESA): Tangent Space Macro-Micro Distribution Alignment Even after mapping SPD matrices to the tangent space, disparity in feature distributions across modalities may cause cross-modal misalignment. LESA introduces complementary distribution objectives: At the macro level, Marginal Distribution Alignment (MDA) forces the Fréchet means of image and text structural features across the batch to coincide in the tangent space: $\(\mathcal{L}_{\mathrm{MDA}} = \left\| \frac{1}{B}\sum_{i=1}^B C_I^{(i)} - \frac{1}{B}\sum_{i=1}^B C_T^{(i)} \right\|_F^2\)$ At the micro level, Conditional Distribution Alignment (CDA) minimizes the distance between matched positive image-text pairs: $\(\mathcal{L}_{\mathrm{CDA}} = \frac{1}{N_{\mathrm{pos}}}\sum_{i=1}^{N_{\mathrm{pos}}} \left\| C_I^{(i)} - C_T^{(i)} \right\|_F^2\)$ MDA centers both modalities in a shared manifold neighborhood, while CDA preserves fine-grained structural correspondence for paired instances.
3. Semantic-Structural Fusion & Second-Stage Alignment: Bimodal Space Complementarity and Theoretical Error Bounds While structural features encode geometric topology, discriminative cross-modal retrieval requires fine-grained category semantics. SARA fuses semantic vectors with vectorized tangent matrices via learnable linear projections: $\(\mathbf{z}_I = W_f [v_{\mathrm{sem}} \,\|\, \mathrm{vec}(C_I)], \quad \mathbf{z}_T = W_f [t_{\mathrm{sem}} \,\|\, \mathrm{vec}(C_T)]\)$ The normalized fused embeddings are optimized through a second-stage symmetric contrastive loss \(\mathcal{L}_{\mathrm{fusion}}\). Theoretically, the authors establish an upper bound on the expected alignment error. Assuming the similarity function is \(L\)-Lipschitz continuous with respect to the combined multimodal distance \(d_{\mathrm{comb}}(v, t)\), and adopting a margin-\(\gamma\) hinge surrogate loss \(\phi_\gamma(z) = \max(0, 1 - z/\gamma)\), the expected mismatch error is bounded by: $\(\mathcal{E} = \mathbb{E}_{(v,t)} [\mathbb{I}\{\mathrm{mismatch}(v,t)\}] \le \frac{L}{\gamma} \mathcal{L}_{\mathrm{SARA}}\)$ This theoretical derivation guarantees that minimizing SARA's objective directly compresses the upper bound of true cross-modal mismatch errors.
Loss & Training¶
The entire SARA framework is optimized end-to-end under the unified objective: $\(\mathcal{L}_{\mathrm{SARA}} = \mathcal{L}_{\mathrm{sem}} + \lambda_m \mathcal{L}_{\mathrm{MDA}} + \lambda_c \mathcal{L}_{\mathrm{CDA}} + \alpha \mathcal{L}_{\mathrm{fusion}}\)$ where \(\mathcal{L}_{\mathrm{sem}} = \frac{1}{2}(\mathcal{L}_{\mathrm{v2t}} + \mathcal{L}_{\mathrm{t2v}})\) represents the standard InfoNCE loss. Hyperparameters are set to \(\lambda_m = 0.3\), \(\lambda_c = 0.7\), \(\alpha = 0.7\), and temperature \(\tau = 0.07\). Backbones are RestV2-Base and BERT-base-uncased. Models are optimized using Adam / Riemannian Adam with an initial learning rate of \(2 \times 10^{-4}\) for 50 epochs on an NVIDIA RTX A6000 GPU (batch sizes 120 on ERA and 210 on UDV).
Key Experimental Results¶
Main Results¶
SARA is thoroughly evaluated against CNN-RNN, Transformer, VLP, and manifold baselines on ERA and UDV drone retrieval benchmarks:
| Dataset | Method | Category | Source | Image→Text R@1 | Image→Text R@5 | Image→Text R@10 | Text→Image R@1 | Text→Image R@5 | Text→Image R@10 | Mean Recall (mR) |
|---|---|---|---|---|---|---|---|---|---|---|
| ERA | CLIP | VLP | ICML21 | 11.31 | 31.92 | 43.91 | 12.73 | 37.33 | 51.52 | 31.45 |
| ERA | PCME | CNN-RNN | CVPR21 | 14.69 | 35.30 | 49.15 | 13.85 | 42.87 | 60.64 | 36.08 |
| ERA | VCSR | CNN-RNN | TGRS24 | 15.65 | 38.28 | 53.49 | 13.69 | 46.31 | 66.37 | 38.96 |
| ERA | MSA | Transformer | TGRS24 | 18.04 | 50.54 | 64.86 | 19.93 | 41.22 | 55.74 | 41.72 |
| ERA | HCCM | Transformer | ACM MM25 | 19.93 | 39.19 | 56.76 | 18.58 | 45.20 | 59.93 | 39.93 |
| ERA | MERU* | Hyperbolic | TGRS24 | 6.22 | 21.37 | 32.98 | 7.41 | 21.06 | 32.33 | 20.18 |
| ERA | HyCoCLIP* | Hyperbolic | ICLR25 | 7.54 | 20.57 | 33.21 | 8.11 | 20.85 | 34.56 | 22.31 |
| ERA | SARA (Ours) | Riemannian | ECCV26 | 23.98 | 48.64 | 63.17 | 21.55 | 56.82 | 76.01 | 48.36 |
| UDV | CLIP | VLP | ICML21 | 6.35 | 20.26 | 29.99 | 4.25 | 16.15 | 26.09 | 17.18 |
| UDV | VCSR | CNN-RNN | TGRS24 | 7.33 | 22.25 | 32.11 | 6.45 | 22.25 | 33.61 | 20.67 |
| UDV | HCCM | Transformer | ACM MM25 | 7.97 | 19.60 | 28.38 | 7.43 | 22.60 | 29.97 | 19.33 |
| UDV | MERU* | Hyperbolic | TGRS24 | 2.63 | 8.92 | 15.23 | 2.38 | 13.51 | 15.24 | 9.12 |
| UDV | HyCoCLIP* | Hyperbolic | ICLR25 | 3.12 | 10.45 | 17.02 | 3.38 | 11.02 | 17.55 | 10.42 |
| UDV | SARA (Ours) | Riemannian | ECCV26 | 8.09 | 24.55 | 34.68 | 6.88 | 23.41 | 35.35 | 22.16 |
Ablation Study¶
Ablation analysis on ERA and UDV evaluates the individual contributions of Visual Structural Features (VSF), Textual Structural Features (TSF), MDA, and CDA:
| Config | VSF | TSF | MDA | CDA | ERA I2T R@1 | ERA T2I R@1 | ERA mR | UDV I2T R@1 | UDV T2I R@1 | UDV mR | Note |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Baseline | \(\times\) | \(\times\) | \(\times\) | \(\times\) | 17.90 | 18.91 | 44.11 | 6.85 | 6.55 | 21.02 | Euclidean dual-branch baseline |
| m1 | \(\checkmark\) | \(\times\) | \(\times\) | \(\times\) | 20.06 | 19.93 | 45.12 | 7.21 | 6.67 | 21.35 | Visual second-order manifold features only |
| m2 | \(\times\) | \(\checkmark\) | \(\times\) | \(\times\) | 12.63 | 15.25 | 36.86 | 6.30 | 6.01 | 20.26 | Textual structural features only (severe degradation) |
| m3 | \(\checkmark\) | \(\checkmark\) | \(\times\) | \(\times\) | 18.71 | 22.63 | 45.43 | 7.38 | 6.79 | 21.57 | Dual-modal structural features without distribution alignment |
| m4 | \(\times\) | \(\times\) | \(\checkmark\) | \(\checkmark\) | 19.02 | 19.76 | 45.50 | 7.12 | 6.74 | 21.44 | MDA & CDA applied directly in flat Euclidean space |
| m5 | \(\checkmark\) | \(\checkmark\) | \(\checkmark\) | \(\times\) | 19.45 | 22.63 | 46.62 | 7.66 | 6.81 | 21.79 | Structural features + Global marginal distribution alignment |
| m6 | \(\checkmark\) | \(\checkmark\) | \(\times\) | \(\checkmark\) | 21.09 | 22.29 | 46.81 | 7.75 | 6.83 | 21.89 | Structural features + Fine-grained conditional alignment |
| SARA | \(\checkmark\) | \(\checkmark\) | \(\checkmark\) | \(\checkmark\) | 21.55 | 23.98 | 48.36 | 8.09 | 6.88 | 22.16 | Full model (Manifold modeling & two-level distribution alignment) |
Key Findings¶
- Visual geometry anchors cross-modal structural modeling: Adding Visual Structural Features alone (m1) raises ERA mR from 44.11 to 45.12. Conversely, applying Textual Structural Features alone (m2) causes a sharp drop to 36.86 mR. Textual token statistics are syntax-driven and do not natively represent 2D spatial layouts; imposing an unanchored SPD structure introduces spurious dependencies. Dual structural modeling (m3) only succeeds when visual geometry provides the reference anchor (45.43 mR).
- Synergy between Riemannian geometry and distribution alignment: Applying MDA and CDA directly on flat semantic vectors (m4) yields marginal improvement (45.50 vs 44.11). In contrast, applying them in the SPD tangent space yields +1.19 and +1.38 mR gains (m5, m6), converging to 48.36 mR in SARA. This indicates that distribution alignment requires non-Euclidean curvature-aware features to be truly effective.
- Strong generalization on remote sensing benchmarks: Evaluated on RSICD and RSITMD, SARA delivers mR scores of 29.18 and 40.68, outperforming UniAdapter (27.54 / 39.23) and SIPN (28.58 / 38.93), validating its adaptability to high-density geospatial scenes beyond drone footage.
Highlights & Insights¶
- Introducing second-order SPD manifold statistics to cross-modal retrieval: Departing from conventional flat cosine similarity matching, SARA encodes visual feature correlations as covariance SPD matrices, exploiting the intrinsic metric curvature of Riemannian manifolds to mitigate viewpoint-induced nonlinear transformations.
- Dual-level distribution alignment in Log-Euclidean tangent space: The combination of Fréchet-mean marginal alignment (macro-level domain centering) and positive-pair conditional alignment (micro-level instance matching) successfully bridges the distribution discrepancy between continuous visual manifolds and discrete text embeddings.
- Theoretically grounded generalization error bound: The objective is rigorously proven to minimize an upper bound on the expected cross-modal mismatch rate, providing formal mathematical justification beyond empirical heuristic tuning.
- Modular and plug-and-play manifold pipeline: The sequential pipeline of BiMap, ReEig, ReCov, and LogEig offers an elegant template for integrating Riemannian representation learning into other aerial tasks such as drone object detection and geo-localization.
Limitations & Future Work¶
- Fragility of text-side manifold priors: Ablation shows that directly computing covariance descriptors across text tokens damages representation quality, indicating that textual representations require external syntactic dependency trees or scene graphs to build meaningful structural manifolds.
- Fixed metric curvature without adaptive scaling: SARA relies on a static Log-Euclidean metric, which cannot dynamically adjust local curvature across differing flight altitudes (e.g., low-altitude vs. high-altitude aerial scenes). Future directions include exploring learnable Riemannian metrics.
- Residual confusion in dense multi-object scenes: Qualitative analysis indicates that in cluttered urban backgrounds with overlapping objects, subtle contextual ambiguities remain, requiring finer-grained entity-level grounding.
Related Work & Insights¶
- vs MERU / HyCoCLIP (Hyperbolic Retrieval): Hyperbolic methods leverage negative curvature to capture hierarchical semantic taxonomy, which is ill-suited for modeling 2D spatial grid dependencies and affine perspective changes; SARA utilizes SPD Riemannian manifolds and Log-Euclidean metrics specifically matched to spatial covariance statistics.
- vs VCSR / HCCM (Euclidean Drone Retrieval): Existing drone retrieval approaches design attention reweighting or hierarchical cross-granularity matching entirely within Euclidean vector spaces; SARA shows that moving beyond the flat space assumption yields superior structural robustness under complex viewpoint transformations.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First framework to systematically introduce symmetric positive definite (SPD) Riemannian manifolds and Log-Euclidean tangent alignment into drone image-text retrieval.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluations across drone datasets (ERA, UDV) and remote sensing benchmarks (RSICD, RSITMD), complemented by ablations, sensitivity analyses, and t-SNE manifold visualizations.
- Writing Quality: ⭐⭐⭐⭐⭐ Cohesive theoretical narrative, restrained mathematical formulation, and a complete generalization bound proof.
- Value: ⭐⭐⭐⭐⭐ Provides a mathematically sound and highly effective non-Euclidean perspective for solving cross-modal perspective distortion in aerial vision.