Cross-Species Animal Re-Identification with Semantic Consistency Learning¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/Kemalau/ECCV-26-SCL
Area: Human Understanding
Keywords: Animal Re-Identification, Semantic Consistency Learning, Spectral Normalization, Reciprocal Neighborhood Modeling, Domain Generalization
TL;DR¶
To tackle embedding space fragmentation caused by drastic anatomical differences and complex ecological backgrounds in cross-species animal ReID, Semantic Consistency Learning (SCL) integrates Foreground–Background Decoupled Spectral Normalization (FDSNorm) for phase-preserving de-stylization and Cross-species Neighborhood Modeling (CNM) with dynamic memory for reciprocal relational alignment, establishing new state-of-the-art generalization across 11 benchmarks.
Background & Motivation¶
Individual animal re-identification (Animal ReID) aims to match individual animals across multiple camera encounters, viewpoints, and temporal spans, serving as an indispensable capability for biodiversity conservation, wildlife population monitoring, and ecological research. While person and vehicle ReID have achieved remarkable success in structured scenarios with consistent geometric layout assumptions (such as upright skeletal poses and shared part configurations), wildlife animal ReID operates in open-world, unstructured environments. Different animal species exhibit vast anatomical structural shifts, distinct surface patterns (e.g., spots, stripes, or featureless fur), and drastically heterogeneous natural habitats ranging from deserts and savannas to dense forests and marine ecosystems. Although species-specific single-species ReID models attain high accuracy on benchmark datasets, every newly monitored species necessitates collecting identity annotations and retraining dedicated models from scratch, which severely impedes practical deployment in broad-scale ecological surveillance.
Recent endeavors have explored multi-species animal ReID and large foundation models (such as MegaDescriptor, MiewID, and UniReID) to learn unified representation spaces across mixed species collections. Nevertheless, conventional pipelines directly borrow instance-level metric learning paradigms from person ReID (e.g., classification cross-entropy paired with batch hard triplet loss). When applied across heterogeneous species, these objectives induce significant representation collapse: embeddings cluster tightly around dominant species-specific visual signatures rather than learning transferable semantic structures, culminating in a highly fragmented embedding space where different species form topologically isolated islands. On the other hand, mainstream domain generalization (DG) strategies developed for person ReID—such as style normalization, feature disentanglement, and meta-learning—implicitly presuppose consistent underlying semantic topology across domains. When subjected to the severe anatomical deformities and morphological divergence inherent to cross-species transitions, standard global alignment methods fail to learn transferable representations and frequently suffer from negative transfer.
The fundamental tension in cross-species animal ReID is therefore twofold: the model must filter out spurious environmental background and style statistics without corrupting fine-grained foreground identity cues, while bridging inter-species morphological divides without collapsing identity-level discrimination within each species. The core idea is to propose Semantic Consistency Learning (SCL), which integrates Foreground–Background Decoupled Spectral Normalization (FDSNorm) to adaptively suppress environment-induced style variations while strictly preserving phase-dependent structural semantics, alongside Cross-species Neighborhood Modeling (CNM) that dynamically discovers reciprocal mutual neighbors from a teacher memory queue to weave isolated species manifolds into a cohesive, transferable relational topology under bounded margin constraints.
Method¶
Overall Architecture¶
The Semantic Consistency Learning (SCL) framework optimizes a Vision Transformer (ViT) backbone to simultaneously retain fine-grained intra-species individual identity discriminability and establish transferable cross-species topological continuity. Its pipeline comprises two synergistic mechanisms: in the spatial-spectral encoding stream, Foreground–Background Decoupled Spectral Normalization (FDSNorm) is strategically embedded into intermediate Transformer layers to decouple foreground and background feature maps, adaptively dampening style-bearing amplitude spectra while strictly preserving structural phase information; in the metric learning stream, an exponential moving average (EMA) teacher network maintains a dynamic feature memory queue to guide student representations via Cross-species Neighborhood Modeling (CNM), enforcing intra-species compactness while regularizing reciprocal cross-species neighborhoods under a bounded hinge objective.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multi-Species Animal Images<br/>(B×3×256×256)"] --> B["ViT Patch Embedding & Transformer Layers"]
B --> C["Foreground–Background Decoupled Spectral Normalization (FDSNorm)<br/>CLS-guided Masking + Spectral Decoupling + Phase-Preserving Reconstruction"]
C --> D["Deep Transformer Blocks & Representation Extraction"]
D --> E["Student Feature zs<br/>Teacher EMA Update & Dynamic Feature Queue"]
E --> F["Reciprocal Neighborhood Search & Memory Bank (CNM Retrieval)<br/>Bidirectional K1/K2 Validation Filtering Spurious Cross-Species Correlations"]
F --> G["Intra-Species Compactness & Bounded Cross-Species Constraint (CNM Loss)<br/>Same-Species Center Attraction + Bounded Relational Alignment"]
G --> H["Joint Multi-Task Loss Optimization<br/>L = λid Lid + λtri Ltri + LCNM"]
Key Designs¶
1. Foreground–Background Decoupled Spectral Normalization (FDSNorm): Adaptive Environment De-Stylization with Phase-Preserved Structural Semantics Camera trap images across diverse ecosystems exhibit severe domain distribution shifts induced by background habitats, weather conditions, and lighting variations. Visual frequency analysis reveals that Discrete Fourier Transform (DFT) amplitude spectra primarily capture domain styles (illumination, color temperature, background texture statistics), whereas phase spectra encode structural boundaries, shape contours, and spatial layouts. Standard frequency normalization indiscriminately suppresses global feature statistics, thereby inadvertently erasing subtle fine-grained identity patterns on animal bodies. FDSNorm overcomes this by introducing an explicit spatial decoupling mechanism into Transformer layers (\(l \in \{0, 4, 8\}\)). For token features \(F^{(l)}\), the CLS token \(c^{(l)}\) and patch tokens \(t_i^{(l)}\) compute cosine similarities \(s_i\), which are normalized and thresholded by a learnable quantile \(\tau = \text{Quantile}_{1-r}(\tilde{s})\) to generate a smooth foreground soft mask \(M^{(l)} \in [0, 1]^{1 \times H \times W}\). The spatial features and their instance-normalized counterparts \(\tilde{F}^{(l)}\) are decoupled into foreground and background branches: $\(F^{(l)}_{\text{fg}, *} = M^{(l)} \odot F^{(l)}_{*}, \quad F^{(l)}_{\text{bg}, *} = (1 - M^{(l)}) \odot F^{(l)}_{*} \quad (* \in \{\text{org}, \text{norm}\})\)$ Both branches undergo independent 2D Discrete Fourier Transforms into amplitude and phase components: \(F^{(l)}_{\text{fg}, *} = A^{(l)}_{\text{fg}, *} e^{j \Phi^{(l)}_{\text{fg}, *}}\) and \(F^{(l)}_{\text{bg}, *} = A^{(l)}_{\text{bg}, *} e^{j \Phi^{(l)}_{\text{bg}, *}}\). Independent learnable weights \(\alpha_{\text{fg}}\) and \(\alpha_{\text{bg}}\) modulate the amplitude spectra (\(\hat{A}_{\text{fg}} = \alpha_{\text{fg}} A_{\text{fg,norm}} + (1 - \alpha_{\text{fg}}) A_{\text{fg,org}}\)), while recombining with the unmodified original phases \(\Phi^{(l)}_{\text{fg,org}}\) and \(\Phi^{(l)}_{\text{bg,org}}\) through inverse DFT before spatial summation. Progressively deployed at shallow (\(l=0\), cleaning low-level illumination), middle (\(l=4\), aligning structural shifts), and deeper layers (\(l=8\), damping residual leakage), FDSNorm realizes structure-preserving style suppression.
2. Reciprocal Neighborhood Search & Memory Bank: Bidirectional Filtering of Spurious Cross-Species Correlations Mini-batch training provides insufficient sample diversity to uncover reliable cross-species semantic manifolds and often suffers from spurious visual correlations (e.g., two unrelated species photographed against similar savannah grasslands). To provide temporally smoothed, long-term semantic anchors, an EMA teacher network updates its parameters (\(\theta_t \leftarrow \mu \theta_t + (1 - \mu) \theta_s\)) and maintains a FIFO dynamic memory queue \(M = \{(f_i, sp_i, y_i)\}_{i=1}^{4096}\) storing normalized feature embeddings along with species and identity tags. For any student anchor feature \(z_s\), candidate neighbors are retrieved from the combined pool of batch teacher features and memory entries. For cross-species retrieval (\(sp_i \neq sp\)), SCL enforces a rigorous reciprocal nearest-neighbor rule: a retrieved top-\(K_1\) candidate \(B\) (with \(K_1=8\)) is retained if and only if anchor \(A\) simultaneously appears within candidate \(B\)'s own cross-species top-\(K_2\) nearest-neighbor list (with \(K_2=3 \le K_1\)). This bidirectional mutual verification filters out accidental, asymmetric visual similarities, ensuring that only semantically compatible pairs contribute to the cross-species relational center \(c_{\text{cross}}\).
3. Intra-Species Compactness & Bounded Cross-Species Constraint: Balancing Intra-Class Cohesion with Inter-Species Manifold Continuity The central challenge of cross-species representation learning lies in bridging disparate species distributions without eroding fine-grained individual identity boundaries. If cross-species features are pulled together unconditionally, individual discriminability breaks down into severe representation collapse. CNM addresses this trade-off via an asymmetric dual-objective loss formulation. For intra-species neighbors, the student feature is pulled toward the same-species center \(c^{(i)}_{\text{intra}}\) via cosine distance to foster intra-species manifold compactness: $\(L_{\text{intra}} = 1 - \frac{1}{B} \sum_{i=1}^B \text{sim}(z_i, c^{(i)}_{\text{intra}})\)$ For the cross-species reciprocal center \(c^{(i)}_{\text{cross}}\), CNM introduces a hinge-based bounded relational constraint: $\(L_{\text{cross}} = \frac{1}{B} \sum_{i=1}^B \max\left(0, m - \text{sim}(z_i, c^{(i)}_{\text{cross}})\right)\)$ Setting margin \(m=0.2\) creates a bounded attraction: once cross-species similarity reaches \(m\), the gradient instantly saturates to zero. This softly weaves disjoint species clusters into a connected global semantic topology while strictly preventing inter-identity feature collision.
Loss & Training¶
The complete learning objective unifies traditional ReID supervision with cross-species neighborhood modeling: $\(L = \lambda_{\text{id}} L_{\text{id}} + \lambda_{\text{tri}} L_{\text{tri}} + L_{\text{CNM}}\)$ where \(L_{\text{id}}\) is the cross-entropy identification loss, \(L_{\text{tri}}\) is the hard-mining triplet loss, and \(L_{\text{CNM}} = \lambda_{\text{intra}} L_{\text{intra}} + \lambda_{\text{cross}} L_{\text{cross}}\). The architecture employs a standard ViT-B/16 backbone pre-trained on ImageNet-1K with \(256 \times 256\) input resolution. Optimization uses SGD with an initial learning rate of 0.004 scheduled by cosine decay across 60 epochs on 4 NVIDIA RTX 4090 GPUs. Total batch size is 128 (8 identities \(\times\) 4 images per GPU \(\times\) 4 GPUs). Both FDSNorm and CNM adopt a 3-epoch warm-up schedule. At test time, only the student backbone's raw feature vectors are extracted for Euclidean/cosine distance evaluation, requiring zero test-time adaptation or auxiliary components.
Key Experimental Results¶
Main Results¶
Evaluation follows two rigorous open-set cross-species protocols where both identities and species are strictly disjoint between training and testing splits. Protocol-1 trains on the 67 seen species of Wildlife71 and evaluates across 10 unseen wildlife datasets as well as PetFace (13 domestic pet species). Protocol-2 trains on 9 public wildlife datasets and evaluates generalizability back onto Wildlife71. The table below presents Protocol-1 performance across representative unseen wildlife domains.
| Method | Venue | ELPephants Rank-1 / mAP |
SealID Rank-1 / mAP |
ATRW (Tiger) Rank-1 / mAP |
GZGC (Giraffe) Rank-1 / mAP |
SeaTurtleID Rank-1 / mAP |
WhaleSharkID Rank-1 / mAP |
|---|---|---|---|---|---|---|---|
| ViT-Base | ICCV 2021 | 31.9 / 7.9 | 79.4 / 24.8 | 96.2 / 58.9 | 21.0 / 25.3 | 46.7 / 7.3 | 37.4 / 6.9 |
| TransReID | ICCV 2021 | 32.3 / 8.1 | 78.6 / 23.3 | 96.5 / 59.1 | 20.0 / 25.6 | 51.2 / 8.1 | 42.2 / 7.8 |
| META | ECCV 2022 | 26.4 / 5.8 | 79.0 / 20.2 | 95.7 / 51.6 | 13.5 / 9.3 | 39.3 / 5.0 | 35.1 / 5.4 |
| CLIP-ReID | AAAI 2023 | 28.9 / 6.8 | 77.6 / 20.8 | 95.6 / 58.1 | 22.9 / 25.7 | 45.9 / 6.9 | 38.8 / 7.0 |
| UniReID | NeurIPS 2023 | 25.2 / 6.1 | 79.4 / 23.6 | 96.1 / 55.9 | 24.7 / 26.0 | 45.8 / 6.9 | 30.9 / 5.2 |
| AdaFreq | ECCV 2024 | 33.3 / 8.4 | 78.6 / 22.8 | 96.4 / 58.0 | 21.6 / 26.0 | 53.1 / 8.4 | 42.5 / 8.0 |
| ReNorm | ECCV 2024 | 20.7 / 5.0 | 78.6 / 23.6 | 96.3 / 56.1 | 23.5 / 24.1 | 60.4 / 9.3 | 32.9 / 5.4 |
| MegaDescriptor | WACV 2024 | 32.0 / 8.2 | 76.5 / 22.1 | 95.9 / 57.4 | 23.5 / 26.8 | 45.3 / 6.9 | 39.1 / 7.1 |
| MiewID | CoRR 2024 | 16.2 / 4.3 | 72.9 / 19.9 | 94.6 / 51.2 | 17.9 / 20.8 | 42.0 / 5.5 | 27.2 / 4.5 |
| SCL (Ours) | ECCV 2026 | 34.4 / 9.0 | 81.5 / 25.6 | 98.0 / 60.3 | 22.9 / 27.3 | 58.8 / 10.0 | 45.4 / 8.9 |
Note: Under Protocol-2 (training on 9 datasets, testing on Wildlife71), SCL attains 93.8% mAP, 78.5% mINP, and 97.6% Rank-1, markedly surpassing ViT-Base (90.1% mAP / 75.1% mINP), MegaDescriptor (87.3% mAP), and UniReID (84.4% mAP).
Ablation Study¶
The step-wise contributions of FDSNorm, intra-species compactness (\(L_{\text{intra}}\)), and cross-species relational modeling (\(L_{\text{cross}}\)) under Protocol-1 are detailed below:
| Config | FDSNorm | Intra | Cross | ELPephants Rank-1 / mAP |
SealID Rank-1 / mAP |
Wildlife71 Rank-1 / mAP |
SeaTurtleID Rank-1 / mAP |
WhaleSharkID Rank-1 / mAP |
|---|---|---|---|---|---|---|---|---|
| (a) Baseline (ViT) | – | – | – | 31.9 / 7.9 | 79.4 / 24.8 | 96.6 / 90.1 | 46.7 / 7.3 | 37.4 / 6.9 |
| (b) + FDSNorm | ✓ | – | – | 32.8 / 8.4 | 80.9 / 25.3 | 97.0 / 91.4 | 55.5 / 9.2 | 44.0 / 8.4 |
| (c) + CNM (Intra only) | – | ✓ | – | 33.1 / 8.5 | 80.8 / 24.8 | 97.2 / 92.3 | 56.7 / 8.9 | 44.3 / 8.5 |
| (d) + CNM (Intra + Cross) | – | ✓ | ✓ | 33.3 / 8.6 | 81.2 / 25.0 | 97.4 / 93.0 | 57.4 / 9.3 | 44.7 / 8.7 |
| (e) Full SCL | ✓ | ✓ | ✓ | 34.2 / 9.0 | 81.3 / 25.6 | 97.6 / 93.8 | 58.6 / 10.0 | 45.2 / 8.9 |
Additional structural and algorithmic ablation insights confirm: 1. Dynamic Memory Queue Necessity: Without memory (in-batch retrieval only), same-ID retrieval purity drops to 78.24% with 19.33% mAP. Incorporating the 4096-capacity memory queue boosts same-ID purity to 89.36% and mAP to 20.69% (Rank-1 reaches 58.75%), validating its stabilizing role. 2. Layer Allocation for FDSNorm: Inserting FDSNorm at sparse multi-depth layers \(l \in \{0, 4, 8\}\) delivers optimal performance (20.69% mAP), whereas dense placement (\(l \in \{0..11\}\), 19.37% mAP) degrades accuracy due to over-suppression of discriminative spatial details. Omitting teacher EMA updates degrades mAP to 20.14%. 3. Foreground-Only Robustness: When background cues are stripped via the MVANet segmentation model, ViT-Base mAP drops by 18.9% and MegaDescriptor drops by 17.5%, while SCL experiences a substantially smaller drop of 14.8% (79.0% vs. 93.8%), proving SCL relies far less on background context shortcuts.
Key Findings¶
- Manifold Connectivity Mitigates Species Fragmentation: As evidenced by the ratio of inter-ID to intra-ID distance across all 10 unseen datasets (e.g., 2.816 vs. 2.106 on ATRW, 1.379 vs. 1.263 on SealID against MegaDescriptor), SCL forms significantly more compact intra-identity groupings alongside cleaner inter-identity separation boundaries, validating that cross-species neighborhood regularization establishes continuous manifolds rather than isolated clusters.
- Phase Preservation Precludes Structural Distortion: Mainstream person ReID domain generalization methods (e.g., ReNorm, META) degrade substantially across unfamiliar species (e.g., dropping to 13.5% Rank-1 on GZGC giraffe). FDSNorm preserves exact phase angles while modulating amplitude, ensuring geometric body contours remain intact regardless of habitat variations.
Highlights & Insights¶
- Spatial-Spectral Dual-Decoupling Formulation: The marriage of CLS-driven spatial attention mask decoupling with 2D-DFT amplitude/phase decomposition yields an elegant mechanism that purges environmental style noise while protecting fine-grained morphological cues.
- Reciprocal Bounded Neighbor Alignment: Filtering cross-species candidates via mutual \(K_1/K_2\) nearest-neighbor verification combined with a small-margin hinge loss (\(m=0.2\)) provides a principled topological bridge across species without risking identity collapse.
- Training-Free Test-Time Inference: SCL requires no test-time batch statistics adaptation, auxiliary pose estimation, or handcrafted textual prompt engineering, executing direct standard distance matching with outstanding deployment efficiency.
Limitations & Future Work¶
- Segmentation Vulnerability Under Hostile Capture Conditions: Under severe water turbidity (marine turtles) or night-vision infrared noise (hyenas), CLS token attention maps may experience minor boundary fuzziness, occasionally perturbing edge amplitude statistics.
- Absence of Hierarchical Biological Taxonomy: Cross-species nearest-neighbor retrieval currently operates purely on geometric feature distance. Integrating phylogenetic hierarchical structures (order, family, genus, species) could provide structured evolutionary priors for long-range cross-taxa generalization.
Related Work & Insights¶
- vs UniReID: UniReID relies on CLIP vision-language prompting with whole-body textual descriptions. When deployed to face-centric domains like PetFace, its generic prompts become misaligned and degrade performance. SCL relies purely on intrinsic visual spectral and topological properties, achieving uniform robustness across full-body and facial datasets.
- vs MegaDescriptor / MiewID: These foundation models lean heavily on massive data scaling under standard single-species metric losses, failing to prevent embedding space fragmentation. SCL demonstrates that principled spectral and neighborhood modeling outmatches brute-force scaling on unseen species.
- vs Person ReID Domain Generalization (ReNorm / META): Person DG models hinge on rigid human skeletal layouts and suffer negative transfer under animal anatomical mutations. SCL resolves this by anchoring invariant phases and forming bounded cross-species neighborhoods.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates cross-species embedding fragmentation as a fundamental bottleneck, offering an original combination of decoupled spectral normalization and reciprocal neighbor alignment.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorously tested across 11 datasets spanning diverse wildlife and domestic animal species under two strict open-set protocols with extensive ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear theoretical exposition, mathematically sound formulations, and high-quality visualizations.
- Value: ⭐⭐⭐⭐⭐ Sets a new benchmark and provides a compelling architectural foundation for universal animal re-identification and ecological biodiversity surveillance.