Skip to content

Virtual Category-Guided Continual Generalized Category Discovery

Conference: ECCV 2026
Paper: ECCV Official
PDF: EventHosts PDF
Code: https://github.com/Mrxjh105/VC-CGCD
Area: Self-Supervised Learning
Keywords: Continual Generalized Category Discovery, Virtual Category Learning, Expanded Neighborhood Contrastive Learning, Open-World Recognition, Pseudo-Label Disambiguation

TL;DR

To tackle ambiguous unlabeled samples that induce noisy pseudo-labels and familiar-class bias in non-rehearsal Continual Generalized Category Discovery (C-GCD), this paper introduces Virtual Category Learning to establish a safe optimization buffer and combines it with Expanded Neighborhood Contrastive Learning to reinforce novel-class discovery while suppressing catastrophic forgetting.

Background & Motivation

Visual recognition systems deployed in open-world environments—such as home robotics, autonomous driving, and persistent surveillance—must continually absorb newly emerging visual concepts without full model retraining, while retaining solid competence on previously learned categories. Generalized Category Discovery (GCD) relaxed the restrictive assumption of Novel Category Discovery (NCD) by considering unlabeled streams mixed with both known and novel classes. Continual Generalized Category Discovery (C-GCD) further pushes this setting into realistic non-rehearsal streaming regimes: after an offline initialization stage on labeled known categories, the model encounters a sequence of unlabeled sessions containing old and emerging classes without storing historical data, which renders traditional global one-shot clustering mechanisms inapplicable.

Under this continual setting without memory replay, a central failure mode stems from the pervasive ambiguity of streaming unlabeled data. Unlabeled sessions contain not only distinctive known instances and clearly clusterable novel samples, but also many confusing instances near decision boundaries or in overlapping semantic regions. Existing C-GCD pipelines rely on hard pseudo-labeling, clustering assignments, or nearest-neighbor consistency, forcing these ambiguous instances into either existing known classes or distinct novel partitions. Early incorrect assignments introduce stubborn label noise and trigger confirmation bias; because previous instances cannot be revisited, these mistakes compound across successive sessions, causing models to favor familiar known classes and underutilize challenging unlabeled samples that are critical for shaping novel decision boundaries.

The key insight is to rethink how ambiguous samples participate in continual optimization: rather than coercing every unlabeled instance into a concrete semantic category, the model requires a controlled buffering mechanism to accommodate undetermined semantics. Adapting ideas from semi-supervised learning into a continual discovery framework, this work introduces Virtual Category Learning (VCL) to route confusing samples to temporary virtual categories as a neutral gradient buffer, and pairs it with Expanded Neighborhood Contrastive Learning (ENCL) to counteract representation drift. Core idea: dynamically assign ambiguous unlabeled samples to temporary virtual categories generated via self-attention to build a neutral optimization buffer, while enlarging neighborhood relations via second-order neighbors and adaptive margins to prevent noisy pseudo-label injection and fully exploit hard unlabeled structural representations.

Method

Overall Architecture

The proposed VC-CGCD framework operates as a multi-branch architecture. Given an input image, weak and strong augmentations are generated and fed into teacher and student encoders respectively to extract feature representations. The architecture is driven by two complementary modules: first, the Virtual Category Learning (VCL) module identifies confusing samples via prediction competition and teacher-student consistency to form a Potential Category (PC) set, dynamically linking ambiguous samples to a Transformer-generated virtual weight vector and excluding conflicting candidate classes from the normalization denominator; second, the Expanded Neighborhood Contrastive Learning (ENCL) module reaches beyond direct nearest neighbors to incorporate neighbors-of-neighbors, reinforcing feature discriminability and separation between previously learned and novel emerging categories. Finally, the combined objective is trained end-to-end alongside the underlying C-GCD baseline loss.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Streaming Unlabeled Input Images<br/>Weak and Strong Augmented Views"] --> B["Feature Extraction & Prediction Competition<br/>Teacher-Student Encoders evaluate Top-2 gap & consistency"]
    B --> C["Potential Category Set Construction<br/>Dual criteria identify confusing samples and PC candidates"]
    C --> D["Virtual Category Learning<br/>Generate virtual weights and compute buffered VC loss"]
    D --> E["Expanded Neighborhood Contrastive Learning<br/>Retrieve second-order neighbors & apply margin contrast"]
    E --> F["Unified End-to-End Optimization<br/>Continuous representation update & label space expansion"]

Key Designs

1. Potential Category Set Construction: identifying confusing samples via dual uncertainty cues Under continuous streaming distribution shifts, standard top-1 prediction confidence becomes unreliable, and premature confidence thresholds introduce severe confirmation bias. To systematically isolate boundary samples, the framework combines prediction competition and temporal teacher-student consistency. The competition metric measures the relative margin between the top-two predicted probabilities: $\(D_{\text{top2}} = \frac{p_{\text{1st}} - p_{\text{2nd}}}{p_{\text{2nd}}}\)$ When \(D_{\text{top2}}\) falls below a calibrated threshold, strong competition between leading candidate classes is indicated. Concurrently, the predicted labels from the teacher network \(y'_t\) and student network \(y'_s\) are cross-compared, where discrepancy signals representation instability across temporal views. Categories surfaced by these two cues form the Potential Category (PC) set. Any instance whose PC set contains multiple plausible classes is tagged as a confusing sample (\(\delta_j = 1\)), avoiding premature filtering or incorrect hard pseudo-label assignments.

2. Virtual Category Learning: neutral optimization buffering via self-attention virtual weights Conventional discovery pipelines either aggressively pull ambiguous samples toward an uncertain pseudo-label center or discard them entirely, causing severe underutilization of hard samples. This work introduces a virtual weight vector \(w^v\) generated dynamically by a self-attention Transformer layer, extending the classifier dimension in parallel with existing categories. For identified confusing samples, the optimization directs representations toward the virtual category channel while explicitly excluding all candidate classes in the PC set from the cross-entropy denominator: $\(\ell_{\text{VC}} = \log \left( e^{l^v} + \sum_{i \notin \text{PC}}^K e^{l^i} \right) - l^v\)$ where \(l^v\) denotes the logit from projecting student features onto the virtual weight vector \(w^v\), and \(l^i\) is the logit for category \(i\). By excluding ambiguous candidates in PC, the sample applies no incorrect pulling force to plausible classes while retaining sharp repulsive gradients from confidently non-candidate negative categories. This provides a neutral, safe optimization path that regularizes feature compactness without corrupting true semantic boundaries.

3. Expanded Neighborhood Contrastive Learning: second-order neighbor expansion for enhanced class separation While virtual categories provide safe optimization buffers, continual streaming updates cause embedding spaces to drift over time. Relying solely on direct \(k\)-nearest neighbors can overlook semantic manifold structures, blurring boundaries between previously learned and novel emerging categories. ENCL expands the positive neighborhood pool to neighbors-of-neighbors. For each first-order neighbor \(q \in N(x_i)\), the framework retrieves its local \(m\)-nearest neighbors to construct an expanded neighborhood set \(E_M(x_i)\). The expanded neighborhood contrastive loss is formulated as: $\(\ell_{\text{EN}} = -\frac{1}{m} \sum_{q \in E_M(x_i)} \log \frac{\exp(f_i \cdot f_q / \tau)}{\sum_{j=1}^m \exp(f_i \cdot f_j / \tau)}\)$ In this construction, instances appearing repeatedly across multiple neighborhood expansions naturally receive higher implicit affinity weighting. Combining first-order neighborhood contrast \(\ell_N\) with expanded contrast \(\ell_{\text{EN}}\) via coefficient \(\beta\) enriches positive sample diversity and guards against feature collapse between emerging categories and historical classes.

Loss & Training

At each online session \(S_t\), the entire model is optimized over unlabeled mini-batches \(\mathcal{B}\) using the unified objective: $\(\mathcal{L} = \mathcal{L}_{\text{VC}} + \lambda_1 \mathcal{L}_{\text{EN}} + \lambda_2 \mathcal{L}_{\text{baseline}}\)$ where \(\mathcal{L}_{\text{VC}}\) is the average virtual category loss across confusing samples (\(\delta_j = 1\)), \(\mathcal{L}_{\text{EN}}\) is the expanded neighborhood contrastive loss, and \(\mathcal{L}_{\text{baseline}}\) is provided by an underlying decoupled C-GCD baseline (e.g., Happy). The loss coefficients are configured as \(\lambda_1 = 0.8\), \(\lambda_2 = 1.0\), with the virtual loss weight set to 1.0; neighborhood weighting is set to \(\beta = 0.1\), and temperature \(\tau = 0.1\). The offline stage is trained for 100 epochs, and each subsequent online session trains for 30 epochs with batch size 128 on a ViT-B/16 backbone initialized with DINO weights, strictly requiring zero rehearsal of past session data.

Key Experimental Results

Main Results

Clustering accuracy (ACC) is evaluated across all sessions and specifically at the final session (Session-5) on CIFAR-100, Tiny-ImageNet, and ImageNet-100 benchmarks. Evaluation covers All classes, previously learned Old classes, and newly introduced New classes.

Dataset Method Offline Init (All) Session-1 (All/Old/New) Session-5 Final (All/Old/New)
CIFAR-100 VanillaGCD 90.82 72.32 / 78.50 / 41.40 51.36 / 53.70 / 30.30
CIFAR-100 SimGCD 90.36 73.37 / 86.44 / 8.00 43.53 / 47.86 / 4.60
CIFAR-100 FRoST 90.36 76.87 / 79.58 / 63.30 48.03 / 48.17 / 46.80
CIFAR-100 GM 90.36 76.58 / 79.80 / 60.50 54.11 / 54.74 / 48.40
CIFAR-100 MetaGCD 90.82 76.12 / 83.60 / 38.70 55.78 / 58.47 / 31.60
CIFAR-100 Happy 90.36 80.40 / 85.26 / 56.10 59.99 / 60.96 / 51.30
CIFAR-100 Ours 90.64 83.22 / 83.56 / 81.50 61.28 / 62.70 / 48.50
Tiny-ImageNet VanillaGCD 84.20 55.93 / 58.92 / 41.00 45.94 / 48.06 / 26.90
Tiny-ImageNet FRoST 85.86 75.15 / 78.56 / 58.10 40.15 / 42.73 / 16.90
Tiny-ImageNet GM 85.86 76.42 / 82.40 / 46.50 46.90 / 50.62 / 13.40
Tiny-ImageNet Happy 85.26 76.67 / 81.72 / 51.40 51.99 / 76.60 / 27.38
Tiny-ImageNet Ours 85.15 77.15 / 80.88 / 58.50 53.25 / 75.98 / 30.52
ImageNet-100 Happy 96.16 91.03 / 95.16 / 70.40 74.98 / 92.40 / 57.56
ImageNet-100 Ours 96.21 91.10 / 95.52 / 69.00 76.72 / 91.88 / 61.55

Ablation Study

Ablation experiments isolate the contributions of Virtual Category Learning (VCL) and Expanded Neighborhood Contrastive Learning (ENCL) across all continual stages on CIFAR-100 and Tiny-ImageNet. The forgetting metric \(M_f\) (lower is better) and discovery metric \(M_d\) (higher is better) characterize continual retention and discovery dynamics.

Configuration CIFAR-100 (All / Old / New) Tiny-ImageNet (All / Old / New) Forgetting \(M_f\) (C100/Tiny) Discovery \(M_d\) (C100/Tiny)
Baseline (w/o VCL, w/o ENCL) 69.00 / 71.82 / 51.36 63.22 / 79.79 / 35.23 29.40 / 8.66 (Happy) 51.36 / 35.23 (Happy)
Only VCL (w/o ENCL) 70.04 / 80.06 / 50.45 63.13 / 79.92 / 35.14 - / - - / -
Only ENCL (w/o VCL) 69.57 / 80.96 / 51.10 64.22 / 78.93 / 40.22 - / - - / -
Full Model (VCL + ENCL) 71.27 / 78.64 / 56.61 64.40 / 78.96 / 40.56 27.94 / 9.17 56.61 / 40.56

Key Findings

  • Substantial boost in sustained discovery (\(M_d\)): On the discovery metric \(M_d\), the proposed method outperforms the previous SOTA (Happy) from 51.36 to 56.61 (+5.25%) on CIFAR-100, and from 35.23 to 40.56 (+5.33%) on Tiny-ImageNet. In an extended 10-session stress test (Table 3), our method maintains a 1.90% and 1.08% average lead over Happy, proving that mitigating premature pseudo-label commitments arrests long-term error accumulation.
  • Orthogonal contributions of VCL and ENCL: The ablation analysis demonstrates that enabling VCL predominantly bolsters retention on Old classes (Old accuracy rises from 71.82 to 80.06 on CIFAR-100) by shielding known prototypes from boundary noise. Conversely, ENCL sharpens cluster discriminability for New classes (New accuracy rises to 51.10 / 40.22). Combining both achieves the top overall performance across all categories.
  • Dynamic stabilization of confusing sample dynamics: Batch monitoring reveals that the proportion of confusing samples fluctuates early in Session 1 before settling into a steady range. In addition, the proportion of novel samples erroneously misclassified as old classes consistently decreases within each session, showing that the virtual category buffer prevents familiar-class confirmation bias.
  • Low compute overhead with FAISS: Assisted by GPU-accelerated FAISS for nearest-neighbor lookups, the full model adds modest computational overhead on CIFAR-100, shifting training time from 338s to 384s per epoch and memory from 17.9 GB to 18.8 GB.

Highlights & Insights

  • Repurposing prediction uncertainty as a constructive buffer: Instead of regarding low-confidence boundary predictions as toxic noise to be discarded, the method leverages virtual categories as neutral geometric anchors, retaining representation manifold continuity while severing corruptive pseudo-label gradients.
  • Dual-criteria capture of streaming representation drift: Coupling local classification competition (\(D_{\text{top2}}\)) with temporal cross-view inconsistency (teacher-student disagreement) ensures sensitive detection of both immediate boundary ambiguity and temporal feature drift.
  • Second-order neighborhood manifold regularization: In evolving feature spaces where direct nearest neighbors are fragile, second-order neighborhood expansion naturally assigns higher affinity weights to tightly connected manifold nodes, providing robust topological support for unsupervised clustering.

Limitations & Future Work

  • Single virtual prototype capacity: The framework currently allocates a single virtual category channel per batch. When a batch contains heterogeneous ambiguous samples spanning vastly different semantic spaces, a single virtual slot may face representation capacity limits; multi-virtual prototype slots warrant future investigation.
  • Dependence on upstream pretraining quality: The geometry of nearest neighbors relies on DINO initialization. In domains where initial representations exhibit poor feature topology (e.g., severe fine-grained long-tail distributions), neighborhood expansion benefits could diminish.
  • Extension to streaming multimodal category discovery: Extending virtual category buffering to vision-language models (VLMs) in streaming open-world scenarios could help disambiguate complex multimodal alignment streams.
  • vs Happy [Ma et al., NeurIPS 2024]: Happy mitigates old-class bias via loss reweighting and prediction calibration, yet still applies conventional pseudo-label supervision on ambiguous samples; this work fundamentally neutralizes ambiguous gradient directions via virtual categories, providing complementary gains.
  • vs FRoST [Roy et al., ECCV 2022] & GM [Zhang et al., NeurIPS 2022]: FRoST uses feature freezing and GM employs grow-and-merge model capacity expansion, both lacking robust defense against ambiguous boundary data in mixed sessions; the proposed method achieves superior discovery without expanding backbone capacity.
  • vs AMEND [Banerjee et al., WACV 2024]: While AMEND explored expanded neighborhoods in static single-stage GCD, this paper adapts and formulates neighborhood contrast specifically for non-rehearsal continual discovery, demonstrating its critical role in mitigating representation drift.

Rating

  • Novelty: ⭐⭐⭐⭐ [Introduces virtual category learning to non-rehearsal C-GCD with an elegant buffering formulation]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across 3 benchmarks, 5-stage and 10-stage streams, and detailed dynamic analyses]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Well-structured narrative, precise math, and clear motivation-to-design alignment]
  • Value: ⭐⭐⭐⭐☆ [Provides an effective blueprint for mitigating confirmation bias in open-world continual discovery]