Skip to content

Debiased Textual Prompt Tuning for Enhancing Unknown Class Discovery

Conference: ECCV 2026
Paper: ECCV Official
Area: Multimodal VLM
Keywords: Open-World Learning, Unknown Class Discovery, Textual Prompt Tuning, Optimal Transport, Label Propagation

TL;DR

Addressing known-class model bias and cross-modal graph sparsity in text-guided open-world semi-supervised learning, this paper proposes an optimal-transport-debiased prompt tuning scheme paired with an asymmetric bimodal graph label propagation framework, achieving a 20.7% relative accuracy improvement on unknown classes across large-scale benchmarks.

Background & Motivation

Open-world semi-supervised learning (OWSSL, also recognized as generalized category discovery) aims to simultaneously recognize known classes and discover novel unknown classes within unlabeled data, where labeled supervision is confined strictly to known categories. Traditional semi-supervised learning relies upon a closed-world assumption and lacks intrinsic mechanisms for novel category discovery. Recent advances have incorporated CLIP-style vision-language models to exploit class-specific textual descriptions as rich semantic anchors, tuning learnable textual prompts to align visual features with text semantics and mitigating the expressiveness limits of one-hot symbolic category representations.

However, existing text-guided OWSSL paradigms suffer from two acute bottlenecks when scaled to large label spaces. First, due to the complete absence of label supervision for novel categories during textual prompt tuning, models naturally develop an overwhelming inductive bias favoring known classes. This leads to severe misclassification of unknown samples into known categories and systematically suppresses the optimization of discriminative textual prompts for novel classes. Second, prior methods establish only fragile, one-to-one cross-modal pairings between an image and its own class description. Viewed through the lens of graph machine learning, this formulation creates severe intra-modal and inter-modal sparsity, obstructing the manifold propagation of visual cues and semantic knowledge across the broader sample distribution.

To overcome these structural limitations, this paper proposes the Textual Prompt Optimization and Bimodal Graph Enhancement framework (TPOBGE). The core idea is to enforce marginal uniform class constraints via entropy-regularized optimal transport to eliminate known-class bias during prompt tuning, while constructing an asymmetric bimodal proximity graph that propagates class confidence across image-image and image-prompt connections while avoiding noisy text-text modality gaps.

Method

Overall Architecture

TPOBGE operates through a decoupled two-stage pipeline: feature extraction with debiased textual prompt tuning, followed by bimodal graph construction and label propagation inference. Given a labeled set \(D_l\) and an unlabeled set \(D_u\), the pre-trained CLIP image and text encoders remain frozen throughout training to extract image embeddings and class description vectors. In the first stage, entropy-regularized optimal transport with uniform class marginal constraints is solved via the Sinkhorn algorithm to remove prediction bias toward known classes, producing hardened one-hot pseudo-labels that guide learnable textual prompt tuning. In the second stage, an asymmetric bimodal proximity graph is assembled over image and tuned prompt nodes, explicitly zeroing out text-text relations to evade modality gap artifacts, and running symmetric normalized label propagation to infer final category assignments for all unlabeled samples.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Data & Frozen CLIP Feature Extraction<br/>Labeled Set Dl + Unlabeled Set Du"] --> B["Optimal Transport Debiased Prompt Optimization"]
    B --> C["High-Confidence Hardened Pseudo-Label Filtering"]
    C --> D["Bimodal Proximity Graph Construction"]
    D --> E["Bimodal Manifold Label Propagation Inference"]
    E --> F["Output Unlabeled Sample Predictions"]

Key Designs

1. Optimal Transport Debiased Prompt Optimization: eliminating known-class prior bias via marginal distribution constraints

When unsupervised models learn prompts for novel categories, representations readily collapse into supervised known-class anchors. TPOBGE frames the assignment between unlabeled samples and textual descriptions as an entropy-regularized optimal transport problem. Given the cosine similarity matrix \(S\) between image features and textual descriptions, the assignment matrix \(\tilde{Q}\) is obtained by maximizing matching utility under Shannon entropy regularization \(H(Q)\):

\[\tilde{Q} = \arg\max_{Q \in \mathcal{U}} \langle Q, S \rangle + \tau_T H(Q)\]

The feasible constraint polytope \(\mathcal{U}\) enforces two explicit marginal conditions: sample-wise normalization ensures every sample allocates a fixed probability mass, while class-wise marginal constraints mechanically enforce an exact uniform prior \(\frac{1}{|C_u|}\) across all \(|C_u|\) known and unknown categories. This structural equality strips away the competitive advantage of known classes, ensuring unknown classes receive balanced gradient feedback and semantic representation space. The objective is efficiently solved within 20 Sinkhorn iterations.

2. High-Confidence Hardened Pseudo-Label Filtering: suppressing ambiguous noise to supervise learnable prompts

Continuous assignment probabilities frequently harbor low-confidence background noise and semantic ambiguity. To prevent these artifacts from degrading prompt learning, the optimal transport matrix is scaled by the total sample size \(\tilde{Q}^{norm} = (|D_u| + |D_l|) \cdot \tilde{Q}\) and filtered against a confidence threshold \(\alpha = 0.5\). Samples whose normalized category confidence exceeds \(\alpha\) are converted into strict one-hot pseudo-labels, while sub-threshold assignments preserve their continuous distribution. The hardened target distribution \(\hat{Q}\) subsequently supervises the learnable prompt vectors \(Prompt = \{u_j\}_{j=1}^{|C_u|}\) via cross-entropy loss, steering prompt embeddings toward discriminative visual features rather than noisy outliers.

3. Bimodal Proximity Graph Construction: resolving intra-modal and inter-modal graph sparsity

Conventional text-guided techniques evaluate only isolated image-description pairings, yielding an extremely sparse bipartite graph. TPOBGE integrates the unlabeled image visual nodes \(Vision = \{v_i\}_{i=1}^{|D_u|}\) and learnable prompt nodes \(Prompt = \{u_j\}_{j=1}^{|C_u|}\) into a unified topological graph. For each visual node \(v_i\), the top-\(k\) nearest visual neighbors within \(Vision\) and the top-\(k\) most similar semantic prompt neighbors within \(Prompt\) are retained (default \(k=5\)), setting remaining connections to zero. Crucially, recognizing that vision-language models exhibit a pronounced modality gap that distorts geometric distances purely between text vectors, the framework explicitly zeros out all prompt-to-prompt entries (\(P_{text} = 0\)), preventing unreliable intra-textual similarities from corrupting graph topology.

4. Bimodal Manifold Label Propagation Inference: diffusing semantic confidence across symmetric normalized hybrid graphs

The resulting proximity matrix \(\tilde{P}\) is symmetrized and normalized by its degree matrix \(D\) to formulate the graph diffusion operator \(\hat{P} = D^{-1/2}(\tilde{P} + \tilde{P}^T)D^{-1/2}\). A pseudo-label matrix \(Y \in \mathbb{R}^{(|D_u| + |C_u|) \times |C_u|}\) is initialized where visual nodes receive zero vectors and prompt nodes are assigned identity one-hot ground-truth anchors. Global label propagation iterates according to:

\[\hat{Y}^{(r+1)} = \beta \hat{P} \hat{Y}^{(r)} + (1 - \beta) Y\]

The propagation weight \(\beta = 0.5\) balances topological diffusion against initial anchor consistency. Through this iterative diffusion, sharp semantic category signals flow smoothly from prompt anchors across the dense visual manifold to unlabeled image nodes, and final predictions are directly obtained at inference via maximum a posteriori selection \(\hat{y}_i = \arg\max_c \hat{Y}(i, c)\).

Key Experimental Results

Main Results

On four large-scale benchmarks with up to 1,000 classesโ€”Hybrid (498 classes), ImageNet-500 (500 classes), ImageNet-1K (1,000 classes), and Webvision (1,000 classes)โ€”TPOBGE is evaluated against leading label-only GCD baselines and state-of-the-art text-guided OWSSL methods (all utilizing ViT-B/16 backbones). Metrics report clustering accuracy (ACC, %) across All classes (A), Known classes (K), and Unknown classes (U).

Methods Pretrain Hybrid (A/K/U) ImageNet-500 (A/K/U) ImageNet-1K (A/K/U) Webvision (A/K/U) Average (A/K/U)
GCD DINO 39.45 / 48.28 / 30.37 58.89 / 65.71 / 55.47 51.25 / 56.06 / 48.84 45.10 / 50.59 / 42.35 48.67 / 55.16 / 44.26
SimGCD DINO 49.35 / 55.73 / 42.78 45.60 / 66.56 / 35.12 34.34 / 60.84 / 21.09 31.80 / 56.67 / 19.36 40.27 / 59.95 / 29.59
LegoGCD DINO 48.30 / 52.64 / 43.85 46.49 / 69.81 / 34.82 34.20 / 61.13 / 20.74 31.60 / 56.85 / 18.97 40.15 / 60.11 / 29.60
ProtoGCD DINO 39.68 / 46.84 / 32.27 33.59 / 63.23 / 18.77 31.28 / 63.05 / 15.40 31.57 / 56.80 / 18.96 34.03 / 57.48 / 21.35
GCD CLIP 52.95 / 54.30 / 51.56 51.26 / 64.22 / 44.78 43.82 / 53.55 / 38.95 41.41 / 50.13 / 37.05 47.36 / 55.55 / 43.08
SimGCD CLIP 64.62 / 64.14 / 65.11 54.73 / 76.56 / 43.82 36.54 / 61.73 / 23.94 34.93 / 58.59 / 23.11 47.70 / 65.25 / 38.99
LegoGCD CLIP 63.45 / 64.40 / 62.46 54.82 / 72.75 / 45.85 38.00 / 63.17 / 25.41 36.00 / 58.50 / 24.75 48.06 / 64.70 / 39.61
ProtoGCD CLIP 48.19 / 50.39 / 45.93 34.58 / 65.15 / 19.30 32.22 / 62.53 / 17.07 30.15 / 57.29 / 16.58 36.28 / 58.84 / 24.72
TP-OWSSL CLIP 44.81 / 42.78 / 45.15 68.98 / 71.71 / 67.69 โ€” / โ€” / โ€” โ€” / โ€” / โ€” โ€” / โ€” / โ€”
CSC-OWSSL CLIP 62.91 / 68.35 / 60.19 56.74 / 72.13 / 49.04 39.03 / 63.05 / 27.02 34.94 / 58.22 / 23.29 48.41 / 65.44 / 39.89
SSR2-GCD CLIP 55.70 / 58.71 / 52.60 65.41 / 70.54 / 62.84 58.73 / 60.02 / 58.08 55.42 / 58.43 / 53.92 58.81 / 61.92 / 56.86
SpectralGCD CLIP 63.79 / 63.47 / 64.15 56.49 / 71.96 / 48.75 64.01 / 78.53 / 56.74 53.92 / 68.74 / 46.50 59.55 / 70.67 / 54.03
TPOBGE (Ours) CLIP 63.81 / 62.40 / 65.27 73.41 / 73.47 / 73.40 69.70 / 71.39 / 68.98 66.73 / 66.63 / 66.96 68.41 / 68.47 / 68.65

Ablation Study

The table below validates the individual mechanisms of debiased prompt optimization and bimodal graph enhancement, reporting unknown-to-known misclassification counts, graph edge density, and computational training efficiency.

Dataset Evaluation Dimension CSC-OWSSL (Prev. SOTA) TPOBGE (Ours) Gain & Observations
Hybrid Unknown misclassified as known 2.9k 1.5k 48.3% reduction in misclassified unknown samples
ImageNet-1K Unknown misclassified as known 15k 1.6k 89.3% reduction, preventing known-class representation collapse
Hybrid Graph edge count 15.1k 114k 7.5x increase in graph connectivity
ImageNet-1K Graph edge count 50k 375k 7.5x denser graph topology enabling manifold diffusion
Hybrid Training time / GPU memory 543 mins / 18.97 GB 3 mins / 0.95 GB 181x faster training, 95.0% memory reduction
ImageNet-1K Training time / GPU memory 7,303 mins / 18.93 GB 10 mins / 9.84 GB 730x training speedup (120+ hours down to 10 minutes)

Key Findings

  • Eradication of Severe Known-Class Bias: In the challenging 1,000-class ImageNet-1K setting, prior state-of-the-art CSC-OWSSL misclassifies 15,000 unknown samples into known categories, collapsing unknown accuracy down to 27.02%. TPOBGE slashes these misclassifications down to 1,600, driving unknown class accuracy up to 68.98% (+41.96% absolute gain).
  • Graph Densification Overcomes Modality Gaps: Expanding graph edges from 50k to 375k via mutual \(k\)-NN visual neighbors and image-to-prompt edges while isolating uninformative prompt-to-prompt edges successfully avoids cross-modal noise and empowers smooth label diffusion.
  • Drastic Computational Efficiency Boost: By operating entirely on frozen pre-extracted CLIP embeddings and restricting updates to prompt vectors and label propagation iterations, TPOBGE cuts full ImageNet-1K training time from over 120 hours to just 10 minutes on a single NVIDIA A6000 GPU.

Highlights & Insights

  • Optimal Transport Uniform Prior Formulation: Rather than tuning heuristic loss reweighting hyperparameters, the method imposes an exact mathematical column-sum constraint in optimal transport, guaranteeing equal opportunity for unknown classes during prompt learning.
  • Modality-Gap-Aware Asymmetric Topology: Acknowledging that text-to-text cosine similarities in multi-modal models often deviate from visual manifold geometry, setting text-text graph weights to zero eliminates negative transfer while retaining crisp anchor-to-image supervision.
  • Inference-Friendly Frozen Backbone Pipeline: Bypassing heavy end-to-end backpropagation over high-resolution vision transformers makes large-scale OWSSL viable in memory- and compute-constrained real-world environments.

Limitations & Future Work

  • Reliance on Multi-Modal Pre-training Generalization: Performance heavily leverages the general semantic coverage of pre-trained CLIP encoders; in specialized or niche domains (e.g., rare industrial defect inspection, complex SAR satellite imagery) lacking strong textual alignment, semantic priors may weaken.
  • Fixed Unknown Category Cardinality Assumption: Experiments assume that the total number of classes \(|C_u|\) is known or provided by external clustering estimation algorithms; exploring dynamic open-vocabulary streaming discovery where novel categories emerge continuously remains open.
  • Complex Long-Tailed and Open-Set Scenarios: Future extensions could generalize the standard optimal transport formulation to unbalanced optimal transport to handle extreme class imbalance and out-of-distribution noise.
  • vs GCD / SimGCD / LegoGCD / ProtoGCD: Standard category discovery methods operate solely on visual representations using self-distillation or prototype clustering, lacking semantic grounding and suffering heavy degradation on fine-grained unknown classes; TPOBGE exploits multi-modal language guidance with explicit debiasing.
  • vs TP-OWSSL / CSC-OWSSL: Prior text-guided OWSSL frameworks suffer from extreme known-class bias, sparse bipartite graphs, and prohibitive training times (tens to hundreds of hours); TPOBGE achieves an order-of-magnitude reduction in runtime and dramatically boosts unknown class discovery.

Rating

  • Novelty: โญโญโญโญโ˜† [Combines optimal transport marginal constraints with asymmetric bimodal label propagation in an elegant, effective formulation]
  • Experimental Thoroughness: โญโญโญโญโญ [Evaluated across four massive benchmarks with up to 1,000 classes, including misclassification counts, graph density, and efficiency benchmarks]
  • Writing Quality: โญโญโญโญโญ [Clear mathematical derivations, crisp problem formulation, and self-contained structure]
  • Value: โญโญโญโญโญ [Solves a critical failure mode in large-scale category discovery while achieving minute-level training speeds]