Skip to content

VIGA: View-Conditioned and Identity-Guided Adaptation for Aerial-Ground Person Re-Identification

Conference: ECCV 2026
Paper: ECCV Original
PDF: Full Text PDF
Area: Remote Sensing / Human Understanding
Keywords: Aerial-ground person re-identification, view-conditioned adaptation, identity-guided sparsity, low-rank adaptation, soft mask attenuation

TL;DR

Addressing extreme view geometry deformation and heavy background clutter in cross-platform aerial-ground person re-identification, VIGA suppresses background noise while preserving spatial topology via identity-guided soft masking at the feature activation level, and dynamically generates low-rank weight residual modulations for decoder feed-forward layers via a view-conditioned hypernetwork at the parameter level, achieving state-of-the-art retrieval accuracy across LAGPeR and AG-ReID.v2.

Background & Motivation

With the extensive deployment of unmanned aerial vehicles (UAVs) in security patrol and smart surveillance, aerial-ground person re-identification (AG-ReID) has emerged as a fundamental technology for cross-platform visual analytics. Unlike conventional ground-based surveillance scenarios, query and gallery images in aerial-ground matching stem from radically different capture viewpointsโ€”wide-field overhead bird's-eye views versus horizontal terrestrial perspectives. This cross-platform discrepancy introduces dual severe distributional shifts: first, UAV overhead imagery captures large fields of view where pedestrian targets occupy only a tiny fraction of pixels amidst complex background clutter; second, steep bird's-eye angles induce severe anisotropic compression and foreshortening, distorting gait silhouettes and body contours that are typically clear from ground cameras.

At the feature activation level, Vision Transformer (ViT) backbones rely on self-attention for weight redistribution. However, Softmax relative normalization only attenuates background tokens numerically without zeroing them out, allowing low-weight background activations to accumulate and amplify through residual connections, contaminating global identity representations. Existing hard token pruning techniques prune background patches but disrupt 2D grid topology and spatial continuity, severing critical boundary cues. At the operator level, conventional models utilize a single shared backbone across both aerial and ground inputs, forcing feed-forward networks (FFNs) to strike suboptimal compromises between conflicting geometric mappings. Prior static prompt or decoupling strategies freeze backbones, failing to dynamically calibrate geometric shifts at the weight operator level.

Resolving these challenges requires topology-preserving background attenuation guided by clean identity semantics in activation space, coupled with dynamic, lightweight operator modulation in parameter space. Core Idea: Jointly design activation-level identity-guided sparsity and parameter-level view-conditioned low-rank adaptationโ€”the former leverages decoupled identity semantics as top-down queries to softly attenuate background clutter, while the latter uses view tokens to dynamically synthesize low-rank weight residuals for decoder FFNs, achieving end-to-end geometric alignment and clean feature learning.

Method

Overall Architecture

VIGA adopts a three-stage end-to-end framework: first, a ViT encoder extracts multi-granularity representations and orthogonalizes identity and view representations; second, the Identity-Guided Sparsity (IGS) module evaluates cosine correlation between local patch tokens and global identity semantics to apply soft mask attenuation; finally, purified patch tokens and identity queries are passed into a View-Conditioned Decoder where a view-conditioned hypernetwork dynamically synthesizes low-rank residual matrices to modulate FFN weights.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    IN["Input Aerial & Ground Images<br/>(UAV overhead / CCTV horizontal)"] --> ENC["ViT Encoder Representation<br/>cls token + view token + local patches"]
    ENC --> S1["Identity-View Decoupling<br/>Subtractive debiasing & orthogonal constraint"]
    S1 --> S2["Identity-Guided Sparsity (IGS)<br/>Cosine correlation scoring & soft mask attenuation"]
    S2 --> S3["View-Conditioned Decoder<br/>Bidirectional cross-attention fusion"]
    S3 --> S4["View-Conditioned LoRA (VC-LoRA)<br/>Hypernetwork dynamic FFN residual synthesis"]
    S4 --> OUT["Discriminative Re-ID Representation<br/>Cross-platform identity retrieval"]

Key Designs

1. Identity-View Decoupling: Subtractive Debiasing and Orthogonal Projection To disentangle camera perspective bias from global identity features in standard \(f_{cls}\), VIGA introduces a dedicated view token \(f_{view}\). Across Transformer layers, the view token acts as a viewpoint prototype, explicitly subtracting view bias from the global classification representation: $\(F_{id}^{(l)} = \text{LayerNorm}(f_{cls}^{(l)} - f_{view}^{(l)})\)$ To enforce complete separation between identity representations \(F_{id}\) and view features \(f_{view}\) in the latent manifold, an explicit orthogonality constraint penalizes squared cosine similarity at the final layer, ensuring extracted identity representations remain viewpoint-invariant.

2. Identity-Guided Sparsity (IGS): Soft Attenuation Preserving Spatial Topology To prevent background clutter from accumulating along residual streams without disrupting 2D spatial topology, IGS employs decoupled global identity features \(F_{id}\) as a top-down semantic probe. It computes cosine correlation scores \(s_i = \frac{F_{id}^\top f_i}{\|F_{id}\| \|f_i\|}\) for each local patch token \(f_i\). Top-scoring foreground patches within a retention ratio \(\rho = 87.5\%\) are preserved, while non-salient regions are scaled by an attenuation factor \(\epsilon = 10^{-3}\): $\(\mathcal{M}_i = \begin{cases} 1, & \text{if } i \in \mathcal{I}_{top} \\ \epsilon, & \text{otherwise} \end{cases}\)$ Modulated features \(\hat{f}_i = \mathcal{M}_i \cdot f_i\) retain complete grid coordinates, eliminating boundary artifacts caused by hard truncation while suppressing background activations.

3. View-Conditioned Low-Rank Adaptation (VC-LoRA): Hypernetwork-Modulated FFN Weights To resolve geometric mapping conflicts between aerial overhead compression and ground upright postures within shared FFN layers, VC-LoRA incorporates input-dependent low-rank weight branches into decoder feed-forward blocks: $\(W(v) = W_0 + \alpha \Delta W, \quad \Delta W = A B\)$ where \(A \in \mathbb{R}^{d_{in} \times r}\) and \(B \in \mathbb{R}^{r \times d_{out}}\) with rank \(r=16 \ll \min(d_{in}, d_{out})\) are dynamically predicted by lightweight hypernetworks \(\Phi_A\) and \(\Phi_B\) conditioned on view tokens \(f_{view}\): $\(A = \Phi_A(f_{view}), \quad B = \Phi_B(f_{view})\)$ Matrix \(B\) is initialized to zero following standard LoRA conventions, ensuring stable training convergence while introducing under 1% additional parameters.

Loss Function & Training Strategy

The architecture is trained end-to-end using a composite objective: $\(\mathcal{L}_{total} = \mathcal{L}_{id} + \lambda_1 \mathcal{L}_{view} + \lambda_2 \mathcal{L}_{orth}\)$ 1. Identity Loss \(\mathcal{L}_{id}\): Weighted combination of cross-entropy classification loss and hard-mining triplet loss applied to aggregated identity representations. 2. View Classification Loss \(\mathcal{L}_{view}\): Cross-entropy objective supervising view tokens \(f_{view}\) to discriminate between aerial and ground camera viewpoints. 3. Orthogonality Loss \(\mathcal{L}_{orth}\): Penalizes inner-product alignment between identity vectors and view tokens: $\(\mathcal{L}_{orth} = \left( \frac{\langle F_{id}, f_{view} \rangle}{\|F_{id}\|_2 \|f_{view}\|_2} \right)^2\)$ The model uses ViT-B/16 as the backbone with input resolution \(256 \times 128\), optimized by SGD with momentum 0.9 and initial learning rate \(8 \times 10^{-3}\) over 120 epochs using cosine annealing.

Key Experimental Results

Main Benchmarks

Evaluated on two large-scale cross-platform benchmarks: LAGPeR (21 cameras, 4,231 identities) and AG-ReID.v2 (UAV, CCTV, and wearable camera platforms with 1,605 identities), measuring Rank-1 accuracy (%) and mean Average Precision (mAP, %).

Table 1: Cross-platform retrieval performance on LAGPeR (A: Aerial view, G: Ground view) | Method | Backbone | A โ†’ G Rank-1 | A โ†’ G mAP | G โ†’ A Rank-1 | G โ†’ A mAP | G โ†’ A+G Rank-1 | G โ†’ A+G mAP | |---|---|---|---|---|---|---|---| | Baseline ViT | ViT | 38.67 | 27.25 | 32.04 | 30.69 | 18.88 | 15.31 | | TransReID | ViT | 38.80 | 28.80 | 33.00 | 32.10 | 22.90 | 18.80 | | CLIP-ReID | CLIP-ViT | 24.40 | 17.60 | 21.30 | 20.80 | 12.30 | 10.20 | | MIP | ViT | 39.30 | 29.30 | 33.90 | 32.60 | 21.00 | 17.30 | | AG-ReID | ViT | 40.48 | 28.89 | 32.96 | 31.91 | 22.03 | 17.89 | | VDT | ViT | 40.15 | 28.97 | 33.55 | 31.98 | 19.50 | 16.45 | | SeCap | ViT | 41.79 | 30.37 | 35.26 | 33.42 | 24.39 | 19.24 | | VIGA (Ours) | ViT | 43.47 | 31.52 | 36.05 | 34.44 | 25.11 | 20.18 |

Table 2: Cross-platform retrieval performance on AG-ReID.v2 (C: Ground CCTV, W: Wearable camera, A: Aerial UAV) | Method | Backbone | A โ†’ C Rank-1 | A โ†’ C mAP | C โ†’ A Rank-1 | C โ†’ A mAP | A โ†’ W Rank-1 | A โ†’ W mAP | W โ†’ A Rank-1 | W โ†’ A mAP | |---|---|---|---|---|---|---|---|---|---| | BoT | ViT | 85.40 | 77.03 | 84.65 | 75.90 | 89.77 | 80.48 | 84.65 | 75.90 | | AG-ReID.v1 | ViT | 87.70 | 79.00 | 87.35 | 78.24 | 93.67 | 83.14 | 87.73 | 79.08 | | VDT | ViT | 86.46 | 79.13 | 86.14 | 78.12 | 90.00 | 82.21 | 85.26 | 78.52 | | AG-ReID.v2 | ViT | 88.77 | 80.72 | 87.86 | 78.51 | 93.62 | 84.85 | 88.61 | 80.11 | | SeCap | ViT | 88.12 | 80.84 | 88.24 | 79.99 | 91.44 | 84.01 | 87.56 | 80.15 | | VIGA (Ours) | ViT | 88.88 | 82.37 | 88.96 | 81.51 | 90.36 | 84.39 | 87.95 | 81.75 |

Ablation Studies

Table 3: Component-wise ablation study on LAGPeR and AG-ReID.v2 (%) | Configuration | LAGPeR Aโ†’G R-1 / mAP | LAGPeR Gโ†’A R-1 / mAP | LAGPeR Gโ†’A+G R-1 / mAP | AG-ReID.v2 Aโ†’C R-1 / mAP | AG-ReID.v2 Cโ†’A R-1 / mAP | Description | |---|---|---|---|---|---|---| | Baseline | 41.79 / 30.37 | 35.26 / 33.42 | 24.39 / 19.24 | 88.12 / 80.84 | 88.24 / 79.99 | Disentangled ViT baseline | | + VC-LoRA | 42.78 / 31.00 | 35.36 / 34.00 | 24.39 / 19.90 | 88.12 / 82.01 | 88.18 / 81.20 | View-conditioned dynamic FFN modulation | | + VC-LoRA + IGS (Full) | 43.47 / 31.52 | 36.05 / 34.44 | 25.11 / 20.18 | 88.88 / 82.37 | 88.96 / 81.51 | Joint soft sparsity & dynamic adaptation |

Table 4: Comparison between soft masking and hard token pruning (%) | Masking Strategy | LAGPeR Aโ†’G R-1 / mAP | LAGPeR Gโ†’A R-1 / mAP | LAGPeR Gโ†’A+G R-1 / mAP | AG-ReID.v2 Aโ†’C R-1 / mAP | AG-ReID.v2 Cโ†’A R-1 / mAP | AG-ReID.v2 Wโ†’A R-1 / mAP | |---|---|---|---|---|---|---| | Hard Token Pruning | 41.76 / 30.65 | 34.60 / 33.67 | 25.25 / 19.89 | 88.67 / 82.40 | 87.85 / 81.36 | 87.63 / 81.07 | | Soft Masking (VIGA) | 43.47 / 31.52 | 36.05 / 34.44 | 25.11 / 20.18 | 88.88 / 82.37 | 88.96 / 81.51 | 87.95 / 81.75 |

Highlights & Insights

  • Dual-Level Calibration: Unifies activation-level background attenuation with operator-level parameter modulation, addressing both noise propagation and geometric misalignment at minimal parameter overhead.
  • Topology-Preserving Top-Down Guidance: Rather than relying purely on bottom-up attention magnitudes, IGS probes token relevance using decoupled identity vectors, preserving 2D grid continuity and boundary cues.
  • Generalizable Parameter-Efficient Design: Hypernetwork-driven low-rank adaptation provides an effective paradigm readily applicable to other cross-domain visual tasks such as visible-infrared matching or cross-altitude aerial tracking.

Limitations & Future Work

  • Discrete View Categorization: The hypernetwork relies on discrete view classification labels, which cannot model continuous pitch angle shifts during dynamic UAV maneuvers.
  • Static Sparsity Threshold: The retention ratio \(\rho\) is fixed at 87.5%; content-adaptive dynamic thresholds based on scene complexity represent an important avenue for future research.
  • vs VDT (CVPR 2024): VDT utilizes subtractive feature disentanglement but maintains static shared parameters; VIGA introduces VC-LoRA to dynamically calibrate operator weights according to viewpoint.
  • vs SeCap (CVPR 2025): SeCap applies input-level prompts for viewpoint adaptation; VIGA intervenes inside decoder feed-forward blocks via hypernetworks while applying soft spatial attenuation.
  • vs DTST (ICME 2025): DTST relies on hard token dropping which severs contour topology; VIGA soft masking retains full grid positions for superior robustness.

Rating

  • Novelty: โญโญโญโญโ˜† [Elegant combination of top-down soft masking and hypernetwork-driven dynamic LoRA]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive cross-platform evaluations on LAGPeR and AG-ReID.v2 with comprehensive ablations]
  • Writing Quality: โญโญโญโญโญ [Clear motivation, mathematically rigorous derivations, and coherent technical exposition]
  • Practical Value: โญโญโญโญโ˜† [Highly effective and efficient solution for UAV-assisted cross-platform surveillance]