Skip to content

VIGA: View-Conditioned and Identity-Guided Adaptation for Aerial-Ground Person Re-Identification

Conference: ECCV 2026
Paper: ECCV Official
PDF: Conference PDF
Area: Human Understanding
Keywords: Aerial-Ground Person Re-ID, View-Conditioned Adaptation, Identity-Guided Sparsification, Low-Rank Adaptation, Soft Masking

TL;DR

To tackle extreme geometric deformations and overwhelming background clutter in aerial-ground person re-identification, VIGA introduces identity-guided soft masking to suppress background noise while preserving 2D spatial topology, and view-conditioned low-rank adaptation where lightweight hyper-networks dynamically modulate decoder feed-forward weights, achieving state-of-the-art cross-platform retrieval on LAGPeR and AG-ReID.v2.

Background & Motivation

With the extensive deployment of unmanned aerial vehicles (UAVs) in urban surveillance and search-and-rescue missions, Aerial-Ground Person Re-Identification (AG-ReID) has emerged as an indispensable vision task for cross-platform target association. Unlike conventional ground-to-ground person Re-ID where camera perspectives remain predominantly horizontal, AG-ReID demands identity matching between high-altitude top-down UAV viewpoints and ground-level surveillance cameras. This cross-platform configuration introduces severe distribution shifts: high-altitude drone captures feature expansive fields of view where pedestrians occupy only a negligible fraction of the image surrounded by extensive background clutter, while extreme top-down perspective foreshortening imposes severe anisotropic deformation and vertical scale compression on human bodies.

In the token activation space, Vision Transformer (ViT) architectures rely on self-attention mechanisms to redistribute token importance. However, the relative normalization of Softmax only attenuates irrelevant patches relatively; low-weight background activations persist through residual connections and accumulate across layers, progressively contaminating global identity representations. Existing dynamic token pruning methods physically discard low-importance tokens, but hard pruning destroys the spatial 2D grid topology and introduces sharp feature cutoffs around human silhouettes. In the transformation space, traditional shared-backbone architectures implicitly force feed-forward networks (FFNs) to approximate conflicting geometric mappings between upright silhouettes and top-down compressed views, yielding compromised representations that are suboptimal for both domains. Furthermore, existing decoupling or prompt-based techniques adapt representations with static backbone weights, leaving operator-level geometric conflicts unresolved.

To reconcile these coupled challenges, an effective framework must eliminate background interference in activation space without breaking topological continuity, while simultaneously providing instance-adaptive geometric adaptation in parameter space without duplicating network backbones. Core idea: jointly unify activation-level Identity-Guided Sparsification and parameter-level View-Conditioned Low-Rank Adaptation, where decoupled identity semantics steer soft background attenuation to preserve spatial continuity, and view tokens dynamically generate low-rank FFN weight residuals via hyper-networks to resolve cross-view geometric disparities.

Method

Overall Architecture

The VIGA framework comprises a three-stage end-to-end pipeline: first, a ViT encoder extracts multi-granularity visual tokens and decouples identity semantics from view bias using subtraction and orthogonality constraints; second, an Identity-Guided Sparsification (IGS) module computes semantic relevance between patch tokens and the global identity feature, applying deterministic soft attenuation to filter background clutter while preserving 2D spatial topology; third, a View-Conditioned Decoder performs two-way cross-attention between identity queries and purified local tokens, where View-Conditioned Low-Rank Adaptation (VC-LoRA) dynamically synthesizes low-rank weight increments via hyper-networks to modulate FFN layers according to viewing geometry.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    IN["Input Aerial & Ground Images<br/>(Top-down drone / Ground CCTV)"] --> ENC["ViT Encoder Feature Extraction<br/>cls token + view token + local patches"]
    ENC --> S1["Identity-View Disentanglement<br/>Subtraction de-biasing & orthogonal constraint"]
    S1 --> S2["Identity-Guided Sparsification (IGS)<br/>Semantic cosine scoring & soft masking"]
    S2 --> S3["View-Conditioned Decoder<br/>Bidirectional cross-attention interaction"]
    S3 --> S4["View-Conditioned LoRA (VC-LoRA)<br/>Hyper-networks synthesize dynamic FFN residuals"]
    S4 --> OUT["Discriminative View-Invariant Embeddings<br/>Cross-platform identity retrieval"]

Key Designs

1. Identity-View Disentanglement: Subtraction de-biasing with angular orthogonality constraints To prevent the global classification token (\(f_{cls}\)) from entangling identity cues with viewpoint biases, the backbone incorporates a dedicated view token (\(f_{view}\)). Across all transformer layers, the view token serves as a prototype of camera viewpoint characteristics and is subtracted from the global representation: $\(F_{id}^{(l)} = \text{LayerNorm}(f_{cls}^{(l)} - f_{view}^{(l)})\)$ To enforce complete semantic separation and prevent residual viewpoint leakage into the identity embedding, an explicit orthogonality constraint is imposed on the final representations, minimizing the squared cosine similarity between \(F_{id}\) and \(f_{view}\) in angular space.

2. Identity-Guided Sparsification (IGS): Soft attenuation preserving spatial topology without clutter accumulation To curb background noise propagation across residual streams without disrupting the regular 2D feature grid, IGS utilizes the decoupled identity feature \(F_{id}\) as a top-down semantic probe. The semantic relevance between \(F_{id}\) and each local patch token \(f_i\) is computed via cosine similarity: \(s_i = \frac{F_{id}^\top f_i}{\|F_{id}\| \|f_i\|}\). Rather than physically pruning low-scoring tokens, IGS selects the top-\(\rho\) proportion (\(\rho = 87.5\%\)) of relevant patches as \(\mathcal{I}_{top}\) and applies a deterministic attenuation factor \(\epsilon = 10^{-3}\) to the remaining tokens: $\(\mathcal{M}_i = \begin{cases} 1, & \text{if } i \in \mathcal{I}_{top} \\ \epsilon, & \text{otherwise} \end{cases}\)$ The refined local features are computed element-wise: \(\hat{f}_i = \mathcal{M}_i \cdot f_i\). This soft masking retains the complete sequence of \(L\) tokens, suppressing irrelevant background activations to near-zero magnitude while eliminating boundary discontinuities. Positional encodings are bypassed during attenuation and instead fed in parallel as spatial queries to the decoder, completely avoiding background noise re-amplification.

3. View-Conditioned Low-Rank Adaptation (VC-LoRA): Dynamic hyper-network modulation for geometric transformation To resolve geometric mapping conflicts within decoder FFN layers when handling disparate aerial and ground silhouettes, VC-LoRA injects input-dependent low-rank weight updates. Instead of keeping FFN weights static, the effective weight matrix is dynamically conditioned on the view token: $\(W(v) = W_0 + \alpha \Delta W, \quad \Delta W = A B\)$ where \(A \in \mathbb{R}^{d_{in} \times r}\) and \(B \in \mathbb{R}^{r \times d_{out}}\) are low-rank factor matrices (\(r = 16 \ll \min(d_{in}, d_{out})\)) generated on the fly by lightweight hyper-networks \(\Phi_A\) and \(\Phi_B\): $\(A = \Phi_A(f_{view}), \quad B = \Phi_B(f_{view})\)$ This hyper-network formulation allows decoder transformations to specialize for aerial versus ground geometries with negligible parameter overhead (<1%). Matrix \(B\) is initialized to zero following standard LoRA conventions, ensuring stable convergence from the pre-trained initialization while scaling factor \(\alpha\) balances the dynamic residual.

Loss & Training

The framework is optimized end-to-end under a multi-task objective: $\(\mathcal{L}_{total} = \mathcal{L}_{id} + \lambda_1 \mathcal{L}_{view} + \lambda_2 \mathcal{L}_{orth}\)$ 1. Identity Loss \(\mathcal{L}_{id}\): Combines Cross-Entropy classification loss and Hard-Mining Triplet Loss applied to the final aggregated identity embedding to enforce intra-class compactness and inter-class separation. 2. View Classification Loss \(\mathcal{L}_{view}\): Supervised Cross-Entropy loss forcing the view token \(f_{view}\) to accurately classify viewing platforms \(v_i \in \{\text{aerial}, \text{ground}\}\): $\(\mathcal{L}_{view} = -\frac{1}{B} \sum_{i=1}^B \log p(v_i \mid f_{view}^{(i)})\)$ 3. Orthogonality Loss \(\mathcal{L}_{orth}\): Penalizes inner-product alignment between identity and view tokens: $\(\mathcal{L}_{orth} = \left( \frac{\langle F_{id}, f_{view} \rangle}{\|F_{id}\|_2 \|f_{view}\|_2} \right)^2\)$ Experiments use an ImageNet-pretrained ViT-B/16 backbone with input resolution \(256 \times 128\). Training runs for 120 epochs using SGD (momentum 0.9, initial learning rate \(8 \times 10^{-3}\)) with a cosine decay schedule, using a batch size of 64 (16 identities, 4 images per identity).

Key Experimental Results

Main Results

Evaluations are conducted on two large-scale AG-ReID benchmarks: LAGPeR (21 cameras, 4,231 identities) and AG-ReID.v2 (UAV, CCTV, and wearable platforms, 1,605 identities). Standard Rank-1 accuracy (%) and mean Average Precision (mAP, %) are reported against competitive Transformer-based Re-ID baselines.

Table 1: Cross-platform performance comparison on the LAGPeR dataset (A: Aerial view, G: Ground view) | Method | Backbone | A โ†’ G Rank-1 | A โ†’ G mAP | G โ†’ A Rank-1 | G โ†’ A mAP | G โ†’ A+G Rank-1 | G โ†’ A+G mAP | |---|---|---|---|---|---|---|---| | ViT Baseline | ViT | 38.67 | 27.25 | 32.04 | 30.69 | 18.88 | 15.31 | | TransReID | ViT | 38.80 | 28.80 | 33.00 | 32.10 | 22.90 | 18.80 | | CLIP-ReID | CLIP-ViT | 24.40 | 17.60 | 21.30 | 20.80 | 12.30 | 10.20 | | MIP | ViT | 39.30 | 29.30 | 33.90 | 32.60 | 21.00 | 17.30 | | AG-ReID | ViT | 40.48 | 28.89 | 32.96 | 31.91 | 22.03 | 17.89 | | VDT | ViT | 40.15 | 28.97 | 33.55 | 31.98 | 19.50 | 16.45 | | SeCap | ViT | 41.79 | 30.37 | 35.26 | 33.42 | 24.39 | 19.24 | | VIGA (Ours) | ViT | 43.47 | 31.52 | 36.05 | 34.44 | 25.11 | 20.18 |

Table 2: Performance comparison on the AG-ReID.v2 dataset across Aerial (A), CCTV (C), and Wearable (W) platforms | Method | Backbone | A โ†’ C Rank-1 | A โ†’ C mAP | C โ†’ A Rank-1 | C โ†’ A mAP | A โ†’ W Rank-1 | A โ†’ W mAP | W โ†’ A Rank-1 | W โ†’ A mAP | |---|---|---|---|---|---|---|---|---|---| | BoT | ViT | 85.40 | 77.03 | 84.65 | 75.90 | 89.77 | 80.48 | 84.65 | 75.90 | | AG-ReID.v1 | ViT | 87.70 | 79.00 | 87.35 | 78.24 | 93.67 | 83.14 | 87.73 | 79.08 | | VDT | ViT | 86.46 | 79.13 | 86.14 | 78.12 | 90.00 | 82.21 | 85.26 | 78.52 | | AG-ReID.v2 | ViT | 88.77 | 80.72 | 87.86 | 78.51 | 93.62 | 84.85 | 88.61 | 80.11 | | SeCap | ViT | 88.12 | 80.84 | 88.24 | 79.99 | 91.44 | 84.01 | 87.56 | 80.15 | | VIGA (Ours) | ViT | 88.88 | 82.37 | 88.96 | 81.51 | 90.36 | 84.39 | 87.95 | 81.75 |

Ablation Study

Component-wise ablations validate the complementary impact of parameter-level VC-LoRA and activation-level IGS, as well as soft masking over hard pruning.

Table 3: Component ablation analysis on LAGPeR and AG-ReID.v2 (scores in %) | Setting | LAGPeR Aโ†’G R-1 / mAP | LAGPeR Gโ†’A R-1 / mAP | LAGPeR Gโ†’A+G R-1 / mAP | AG-ReID.v2 Aโ†’C R-1 / mAP | AG-ReID.v2 Cโ†’A R-1 / mAP | Note | |---|---|---|---|---|---|---| | Baseline | 41.79 / 30.37 | 35.26 / 33.42 | 24.39 / 19.24 | 88.12 / 80.84 | 88.24 / 79.99 | Disentangled ViT baseline | | + VC-LoRA | 42.78 / 31.00 | 35.36 / 34.00 | 24.39 / 19.90 | 88.12 / 82.01 | 88.18 / 81.20 | Adding view-conditioned FFN adaptation | | + VC-LoRA + IGS (Full) | 43.47 / 31.52 | 36.05 / 34.44 | 25.11 / 20.18 | 88.88 / 82.37 | 88.96 / 81.51 | Combining soft masking and dynamic FFN |

Table 4: Masking strategy comparison: Hard pruning vs. Soft attenuation (scores in %) | Masking Strategy | LAGPeR Aโ†’G R-1 / mAP | LAGPeR Gโ†’A R-1 / mAP | LAGPeR Gโ†’A+G R-1 / mAP | AG-ReID.v2 Aโ†’C R-1 / mAP | AG-ReID.v2 Cโ†’A R-1 / mAP | AG-ReID.v2 Wโ†’A R-1 / mAP | |---|---|---|---|---|---|---| | Hard Masking (Pruning) | 41.76 / 30.65 | 34.60 / 33.67 | 25.25 / 19.89 | 88.67 / 82.40 | 87.85 / 81.36 | 87.63 / 81.07 | | Soft Masking (VIGA) | 43.47 / 31.52 | 36.05 / 34.44 | 25.11 / 20.18 | 88.88 / 82.37 | 88.96 / 81.51 | 87.95 / 81.75 |

Key Findings

  • VC-LoRA enhances ranking consistency: Injecting VC-LoRA boosts mAP significantly across platforms (e.g., AG-ReID.v2 A โ†’ C mAP improves from 80.84% to 82.01%), demonstrating that view-adaptive weight adjustments establish a more consistent cross-view feature manifold than static transformations.
  • Soft attenuation outperforms hard pruning: Preserving the complete 2D token grid via soft attenuation improves LAGPeR A โ†’ G Rank-1 by +1.71% and G โ†’ A Rank-1 by +1.45% over hard token pruning, confirming the necessity of spatial continuity for fine-grained cross-view alignment.
  • Hyperparameter robustness: Performance peaks at preservation ratio \(\rho = 87.5\%\); discarding too many tokens (\(\rho = 50\%\)) degrades identity details, while retaining all tokens (\(\rho = 100\%\)) introduces clutter. Rank dimension \(r = 16\) achieves the optimal balance between geometric adaptation capacity and parameter efficiency, preventing overfitting.

Highlights & Insights

  • Dual-level calibration paradigm: Rather than relying solely on representation decoupling or full fine-tuning, VIGA harmonizes activation-level soft filtering with operator-level low-rank weight generation, tackling background clutter accumulation and geometric mapping conflicts concurrently with negligible parameter overhead.
  • Top-down semantic masking: Instead of using heuristic activation magnitudes that risk pruning low-contrast body parts, IGS leverages pure identity vectors as semantic probes to guide background suppression while preserving full 2D grid topology.
  • Transferable hyper-network design: The view-conditioned LoRA module offers a plug-and-play paradigm for cross-platform vision tasks, easily extending to visible-infrared person Re-ID or multi-altitude drone tracking where disparate viewpoints demand parameter-level adaptation.

Limitations & Future Work

  • Admitted limitation: The module relies on discrete viewpoint labels (aerial, ground, wearable) to drive the hyper-networks. In operational scenarios with continuous zoom or continuously varying gimbal tilt angles, coarse categorical labels cannot capture continuous perspective deformations.
  • Improvement directions: Future investigations could integrate continuous camera extrinsic parameters (e.g., pitch, roll, altitude) or monocular depth priors directly into hyper-network conditioning, and explore content-aware dynamic prediction of the preservation ratio \(\rho\).
  • vs VDT (CVPR 2024): While VDT uses subtraction-based feature disentanglement, its shared backbone weights remain entirely static; VIGA enhances disentanglement with angular orthogonality and introduces dynamic parameter-level adaptation via VC-LoRA.
  • vs SeCap (CVPR 2025): SeCap adjusts input distributions using static prompt learning; VIGA intervenes inside decoder FFN layers via hyper-network generated residuals and applies soft topological masking to suppress background noise.
  • vs DTST (ICME 2025): DTST enforces hard token removal that fractures spatial continuity and body contours; VIGA implements deterministic soft masking with parallel spatial queries, maintaining robust geometric reasoning across viewpoints.

Rating

  • Novelty: โญโญโญโญโ˜† [Elegant coupling of top-down soft masking with view-conditioned low-rank hyper-networks]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive cross-platform evaluations on LAGPeR and AG-ReID.v2 with comprehensive ablations and sensitivity studies]
  • Writing Quality: โญโญโญโญโญ [Clear structural organization, rigorous mathematical formulation, and solid motivation]
  • Value: โญโญโญโญโ˜† [Provides an efficient, generalizable framework for multi-platform UAV surveillance and aerial-ground vision systems]