CGCC: Towards Generalizable Clothes-Changing Person Re-Identification¶
Conference: ECCV 2026
Paper: ECCV 2026 Oral/Poster
Code: https://github.com/zhi-time/CGCC
Area: Human Understanding / Autonomous Driving
Keywords: Clothes-Changing Person Re-Identification, Generalization, Text-Guided Identity Refinement, Singular Value Decomposition, Orthogonal Projection
TL;DR¶
To tackle poor generalization in real-world clothes-changing person re-identification, this paper constructs CGCCโa large-scale synthetic dataset (4,101 IDs, 217k images) featuring global cultural/spatiotemporal diversity and structured semantic annotationsโand proposes the TGIR framework that projects visual features orthogonally away from an SVD-modeled identity-irrelevant semantic subspace.
Background & Motivation¶
Person Re-Identification (ReID) aims to match pedestrians across non-overlapping camera networks. Conventional methods rely heavily on the short-term assumption that individuals maintain identical clothing appearance. In practical long-term surveillance or cross-scene tracking, clothes-changing is frequent and inevitable, rendering appearance-centric texture features ineffective. While Clothes-Changing Person Re-Identification (CC-ReID) has emerged to address this, existing research remains constrained by two fundamental bottlenecks when deployed to unconstrained open-world scenarios.
First, existing CC-ReID datasets suffer from insufficient distribution diversity and lack fine-grained semantic supervision. Most benchmarks focus narrowly on solitary garment variations, neglecting how regional cultural differences profoundly alter dressing habits, while failing to integrate temporal seasons, illumination transitions, and adverse weather dynamics. Furthermore, relying purely on discrete ID and clothes labels deprives models of descriptive linguistic context needed to characterize subtle appearance transformations.
Second, existing semantic-guided CC-ReID approaches lack holistic modeling of identity-irrelevant interference. Prior works isolate clothes and background factors, applying ad-hoc adversarial suppression or instance-level penalties without unifying them into a structured "identity-irrelevant semantic space." Consequently, models struggle to delineate global distribution boundaries of non-identity noise. Core idea: by introducing a multi-factor synthetic dataset CGCC with fine-grained semantic annotations and proposing the TGIR framework that applies SVD-based spectrum-weighted orthogonal projection onto an identity-irrelevant semantic subspace, the approach forms a mutually reinforcing closed loop between data diversity expansion and subspace disentanglement.
Method¶
Overall Architecture¶
The TGIR framework approaches CC-ReID through semantic subspace disentanglement. Multimodal textual descriptions generated by MLLMs are partitioned into identity-relevant biometric texts (\(T_{bio}\)) and identity-irrelevant clothes (\(T_{cloth}\)) and environmental (\(T_{env}\)) texts. During training, two complementary modules refine visual representations: Biometric-Guided Prototype Alignment (BGPA) adaptively calibrates global identity prototypes against sample noise, while Spectrum-Aware Orthogonal Projection (SAOP) applies Singular Value Decomposition (SVD) on non-identity descriptions and enforces spectrum-weighted orthogonal constraints to strip away visual interference.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Image x & Disentangled Texts<br/>Tbio / Tcloth / Tenv"] --> B["Multimodal Feature Extraction<br/>Visual Encoder Ei + Text Encoder Et"]
B --> C["Biometric-Guided Prototype Alignment (BGPA)<br/>Adaptive Momentum Calibration via Consistency Score"]
B --> D["Spectrum-Aware Orthogonal Projection (SAOP)<br/>SVD on Non-Identity Texts for Subspace Basis"]
C --> E["Joint Objective Optimization<br/>LID + LTriplet + LITC + LBGPA + LSAOP"]
D --> E
E --> F["Inference Time Output<br/>Pure Visual Encoder Yields Refined Identity Feature"]
Key Designs¶
1. Multi-Factor CGCC Benchmark Construction: Breaking the Data Isolation of Single Clothing Variations
Acquiring extreme variation for identical pedestrians in real-world settings is physically prohibitive and constrained by privacy regulations. The authors establish a systematic generative pipeline grounded on MSMT17 seed images. First, SCHP torso integrity verification, portrait quality assessment networks, and an intra-class diversity criterion (>20 images per identity) filter out motion blur and severe occlusions. Next, ChatGPT compiles a two-dimensional prompt pool: one covering global cultural dressing across 6 continents and 66 countries, and the other detailing dynamic seasons, weather conditions, and functional wear, alongside an instruction recombination strategy for rare corner cases. Qwen-Image-Edit executes controllable synthesis, and Qwen3-VL performs quality audits to filter facial distortion before generating structured biometrics, clothes, and environmental texts, yielding 4,101 identities and 217,248 high-fidelity images.
2. Biometric-Guided Prototype Alignment (BGPA): Mitigating Semantic Drift via Adaptive Momentum Calibration
Averaging raw visual instances into identity prototypes leads to severe semantic drift when samples are dominated by background or garment biases. BGPA uses a lightweight Adapter MLP to compute a semantic consistency score \(\alpha_i\) between visual embedding \(f_v^i\) and biometric text embedding \(t_{bio}^i\): $$ \alpha_i = \sigma(\text{Adapter}([f_v^i; t_{bio}^i])) $$ Using \(\alpha_i\) as a confidence weight, the batch-level weighted feature center \(\bar{f}_k = \frac{\sum_{i \in \mathcal{B}_k} \alpha_i \cdot f_v^i}{\sum_{i \in \mathcal{B}_k} \alpha_i}\) and mean batch reliability \(\bar{\alpha}_k\) are computed. Global prototype \(P_k\) updates with an adaptive momentum step \(\mu_k = (1 - m) \cdot \bar{\alpha}_k\): $$ P_k^{(t)} = (1 - \mu_k) P_k^{(t-1)} + \mu_k \bar{f}_k $$ When samples faithfully align with biometric descriptors, the prototype smoothly incorporates new visual features; for corrupted or outlier instances (\(\bar{\alpha}_k \to 0\)), updates vanish, protecting global identity centers from non-identity contamination.
3. Spectrum-Aware Orthogonal Projection (SAOP): SVD-Driven Subspace Decoupling of Non-Identity Interference
Rather than treating clothing and environment as discrete adversarial targets, SAOP models non-identity factors as a continuous semantic subspace. Concatenating clothing text features \(t_{cloth}\) and environment text features \(t_{env}\) forms an identity-irrelevant matrix \(M_{irr} = [t_{cloth}; t_{env}] \in \mathbb{R}^{2N \times D}\). Performing Singular Value Decomposition yields: $$ M_{irr} \xrightarrow{\text{SVD}} U \Sigma V^T $$ where the right singular vectors \(V = \{v_1, v_2, \dots, v_D\}\) provide an orthogonal basis spanning the identity-irrelevant subspace, with singular values \(\lambda_1, \dots, \lambda_D\) quantifying directional energy. SAOP minimizes visual projections onto this subspace via a spectrum-weighted orthogonal loss: $$ \mathcal{L}{SAOP} = \sum |f_v^T v_j|^2 $$ This mathematical orthogonalization strips away non-rigid environmental and apparel distortions globally without attenuating rigid human skeletal and structural cues.}^D \frac{\lambda_j}{\sum_{d=1}^D \lambda_d
Loss & Training¶
The architecture is trained end-to-end under a composite loss integrating classical ReID supervision, vision-language alignment, and dual disentanglement constraints: $$ \mathcal{L}{total} = \mathcal{L}} + \mathcal{L{Triplet} + \mathcal{L}} + \mathcal{L{BGPA} + \mathcal{L} $$ where \(\mathcal{L}_{ID}\) and \(\mathcal{L}_{Triplet}\) are cross-entropy and triplet losses; \(\mathcal{L}_{ITC}\) aligns visual features with positive biometric descriptions; \(\mathcal{L}_{BGPA}\) forces visual features to cluster tightly around calibrated prototypes; and \(\mathcal{L}_{SAOP}\) enforces the spectrum-weighted orthogonal penalty. Models employ EVA02-CLIP-L/14 visual backbones and EVA02-CLIP-bigE-14 text encoders, trained for 60 epochs. Text encoders and SVD operations operate strictly during training; test inference runs purely on visual features with zero auxiliary computational overhead.
Key Experimental Results¶
Main Results¶
Evaluation assesses both zero-shot cross-dataset generalization (Source \(\to\) Target) and standard intra-dataset benchmarks.
Table 1 details cross-dataset retrieval performance when models trained on CGCC are transferred to unseen target datasets (LTCC, PRCC, DeepChange):
| Method | Source Dataset | LTCC (R1 / mAP) | PRCC (R1 / mAP) | DEEP (R1 / mAP) | Average (R1 / mAP) |
|---|---|---|---|---|---|
| TransReID | CGCC | 21.9% / 9.2% | 37.1% / 32.5% | 46.7% / 12.7% | 35.2% / 18.1% |
| MADE | CGCC | 38.3% / 19.6% | 60.3% / 55.9% | 59.6% / 20.9% | 52.7% / 32.1% |
| Differ | CGCC | 36.2% / 17.7% | 55.4% / 50.6% | 57.9% / 20.0% | 49.8% / 29.4% |
| TGIR (Ours) | CGCC | 40.3% / 19.1% | 62.0% / 56.9% | 62.1% / 22.4% | 54.8% / 32.8% |
Table 2 evaluates intra-dataset performance on the CGCC benchmark under both clothes-changing (CC) and same-clothes (SC) evaluation protocols:
| Method | Venue | CC Rank-1 | CC mAP | SC Rank-1 | SC mAP |
|---|---|---|---|---|---|
| TransReID | ICCV 2021 | 41.0% | 17.0% | 75.0% | 47.5% |
| CAL | CVPR 2022 | 32.3% | 12.0% | 65.3% | 35.9% |
| MADE | CVPR 2024 | 51.2% | 23.3% | 81.2% | 54.9% |
| Differ | CVPR 2025 | 50.7% | 23.3% | 83.7% | 60.1% |
| Eva02-CLIP | arXiv 2024 | 54.7% | 26.3% | 85.0% | 62.0% |
| TGIR (Ours) | ECCV 2026 | 57.3% | 28.1% | 86.0% | 64.3% |
Ablation Study¶
Table 3 reports ablation experiments on the CGCC \(\to\) PRCC cross-dataset transfer protocol:
| Index | Configuration | BGPA Prototype Alignment | SAOP Orthogonal Projection | Rank-1 (%) | mAP (%) |
|---|---|---|---|---|---|
| 1 | Baseline (Eva02-CLIP) | - | - | 54.5 | 50.6 |
| 2 | + BGPA | โ | - | 56.4 (+1.9) | 52.1 (+1.5) |
| 3 | + SAOP | - | โ | 59.0 (+4.5) | 55.3 (+4.7) |
| 4 | Full TGIR | โ | โ | 62.0 (+7.5) | 56.9 (+6.3) |
Key Findings¶
- Generalization Gains via CGCC: Compared to models trained on large-scale real data (LaST), MADE trained on CGCC achieves 59.6% Rank-1 on DeepChange, exceeding its LaST counterpart by 8.2%, highlighting that multi-factor environmental and cultural diversity in training data directly drives open-world transferability.
- Superiority of SAOP Subspace Decoupling: In ablation studies, adding SAOP boosts cross-domain Rank-1 by 4.5% over baseline, outperforming single-modality prototype calibration (+1.9%), substantiating that orthogonal subspace projection provides cleaner decoupling than unweighted feature suppression.
- Simultaneous Gains across CC and SC Settings: On CGCC intra-dataset testing, TGIR outperforms the nearest competitor on CC by 2.6% R1 / 1.8% mAP, while concurrently improving SC Rank-1 from 85.0% to 86.0%, proving that identity-irrelevant orthogonalization does not sacrifice intrinsic biometric discrimination.
Highlights & Insights¶
- Subspace-Level Disentanglement via SVD: Moving beyond heuristic negative-sample penalties, the method constructs a unified non-identity basis from multimodal text and applies spectrum-weighted orthogonal projections, providing an elegant and mathematically sound formulation for feature purification.
- Closed-Loop Synergy of Generative Diversity and Decoupled Supervision: The approach avoids naive visual data scaling by pairing structured prompt expansion (ChatGPT cultural/temporal taxonomy) with MLLM-driven visual editing (Qwen-Image-Edit) and automated semantic parsing (Qwen3-VL), tightly coupling dataset expressiveness with algorithm disentanglement.
- Zero-Cost Inference Deployment: Rich text embeddings and singular value decompositions serve exclusively as supervisory training guides, maintaining standard visual feed-forward inference without additional memory or latency overhead.
Limitations & Future Work¶
- Author-Admitted Limitations: Despite CGCC's extensive multi-factor variation, synthetic generation cannot encompass every corner-case in open-world surveillance; automated text captions may miss fine-grained visual cues under heavy occlusion or low resolution; and training-time SVD calculations incur extra GPU memory and computational costs.
- Broader Perspective: Constructing SVD bases on mini-batches may exhibit variance across training steps. Future work could maintain an online orthogonal memory bank or dictionary across batches to stabilize global interference bases. Furthermore, this orthogonal subspace projection paradigm holds promise for cross-weather autonomous driving tracking and open-vocabulary re-identification.
Related Work & Insights¶
- vs Differ (CVPR 2025): Differ leverages semantic cues via standard adversarial learning and instance-level feature separation, lacking global geometric orthogonality. TGIR explicitly models combined apparel and environment matrices via SVD, outperforming Differ by 5.0% average Rank-1 across unseen domains.
- vs MADE (TMM 2024): MADE relies on masked attribute embeddings over fixed discrete vocabularies. TGIR exploits continuous open-vocabulary descriptions and dynamic momentum prototype calibration, providing superior adaptability across unconstrained environments.
Rating¶
- Novelty: โญโญโญโญโญ Elegant mathematical formulation modeling non-identity factors as an SVD orthogonal subspace.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive cross-domain zero-shot evaluations across multiple target datasets alongside rigorous intra-domain benchmarks.
- Writing Quality: โญโญโญโญโญ Rigorous methodology presentation, logical narrative, and well-grounded motivation.
- Value: โญโญโญโญโญ The CGCC benchmark and TGIR framework significantly advance generalizable CC-ReID for practical long-term surveillance.