Debiased Textual Prompt Tuning for Enhancing Unknown Class Discovery¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Multimodal VLM
Keywords: Open-World Learning, Unknown Class Discovery, Textual Prompt Tuning, Optimal Transport, Label Propagation
TL;DR¶
Addressing known-class model bias and cross-modal graph sparsity in text-guided open-world semi-supervised learning, this paper proposes an optimal-transport-debiased prompt tuning scheme paired with an asymmetric bimodal graph label propagation framework, achieving a 20.7% relative accuracy improvement on unknown classes across large-scale benchmarks.
Background & Motivation¶
Open-world semi-supervised learning (OWSSL, also recognized as generalized category discovery) aims to simultaneously recognize known classes and discover novel unknown classes within unlabeled data, where labeled supervision is confined strictly to known categories. Traditional semi-supervised learning relies upon a closed-world assumption and lacks intrinsic mechanisms for novel category discovery. Recent advances have incorporated CLIP-style vision-language models to exploit class-specific textual descriptions as rich semantic anchors, tuning learnable textual prompts to align visual features with text semantics and mitigating the expressiveness limits of one-hot symbolic category representations.
However, existing text-guided OWSSL paradigms suffer from two acute bottlenecks when scaled to large label spaces. First, due to the complete absence of label supervision for novel categories during textual prompt tuning, models naturally develop an overwhelming inductive bias favoring known classes. This leads to severe misclassification of unknown samples into known categories and systematically suppresses the optimization of discriminative textual prompts for novel classes. Second, prior methods establish only fragile, one-to-one cross-modal pairings between an image and its own class description. Viewed through the lens of graph machine learning, this formulation creates severe intra-modal and inter-modal sparsity, obstructing the manifold propagation of visual cues and semantic knowledge across the broader sample distribution.
To overcome these structural limitations, this paper proposes the Textual Prompt Optimization and Bimodal Graph Enhancement framework (TPOBGE). The core idea is to enforce marginal uniform class constraints via entropy-regularized optimal transport to eliminate known-class bias during prompt tuning, while constructing an asymmetric bimodal proximity graph that propagates class confidence across image-image and image-prompt connections while avoiding noisy text-text modality gaps.
Method¶
Overall Architecture¶
TPOBGE operates through a decoupled two-stage pipeline: feature extraction with debiased textual prompt tuning, followed by bimodal graph construction and label propagation inference. Given a labeled set \(D_l\) and an unlabeled set \(D_u\), the pre-trained CLIP image and text encoders remain frozen throughout training to extract image embeddings and class description vectors. In the first stage, entropy-regularized optimal transport with uniform class marginal constraints is solved via the Sinkhorn algorithm to remove prediction bias toward known classes, producing hardened one-hot pseudo-labels that guide learnable textual prompt tuning. In the second stage, an asymmetric bimodal proximity graph is assembled over image and tuned prompt nodes, explicitly zeroing out text-text relations to evade modality gap artifacts, and running symmetric normalized label propagation to infer final category assignments for all unlabeled samples.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Data & Frozen CLIP Feature Extraction<br/>Labeled Set Dl + Unlabeled Set Du"] --> B["Optimal Transport Debiased Prompt Optimization"]
B --> C["High-Confidence Hardened Pseudo-Label Filtering"]
C --> D["Bimodal Proximity Graph Construction"]
D --> E["Bimodal Manifold Label Propagation Inference"]
E --> F["Output Unlabeled Sample Predictions"]
Key Designs¶
1. Optimal Transport Debiased Prompt Optimization: eliminating known-class prior bias via marginal distribution constraints
When unsupervised models learn prompts for novel categories, representations readily collapse into supervised known-class anchors. TPOBGE frames the assignment between unlabeled samples and textual descriptions as an entropy-regularized optimal transport problem. Given the cosine similarity matrix \(S\) between image features and textual descriptions, the assignment matrix \(\tilde{Q}\) is obtained by maximizing matching utility under Shannon entropy regularization \(H(Q)\):
The feasible constraint polytope \(\mathcal{U}\) enforces two explicit marginal conditions: sample-wise normalization ensures every sample allocates a fixed probability mass, while class-wise marginal constraints mechanically enforce an exact uniform prior \(\frac{1}{|C_u|}\) across all \(|C_u|\) known and unknown categories. This structural equality strips away the competitive advantage of known classes, ensuring unknown classes receive balanced gradient feedback and semantic representation space. The objective is efficiently solved within 20 Sinkhorn iterations.
2. High-Confidence Hardened Pseudo-Label Filtering: suppressing ambiguous noise to supervise learnable prompts
Continuous assignment probabilities frequently harbor low-confidence background noise and semantic ambiguity. To prevent these artifacts from degrading prompt learning, the optimal transport matrix is scaled by the total sample size \(\tilde{Q}^{norm} = (|D_u| + |D_l|) \cdot \tilde{Q}\) and filtered against a confidence threshold \(\alpha = 0.5\). Samples whose normalized category confidence exceeds \(\alpha\) are converted into strict one-hot pseudo-labels, while sub-threshold assignments preserve their continuous distribution. The hardened target distribution \(\hat{Q}\) subsequently supervises the learnable prompt vectors \(Prompt = \{u_j\}_{j=1}^{|C_u|}\) via cross-entropy loss, steering prompt embeddings toward discriminative visual features rather than noisy outliers.
3. Bimodal Proximity Graph Construction: resolving intra-modal and inter-modal graph sparsity
Conventional text-guided techniques evaluate only isolated image-description pairings, yielding an extremely sparse bipartite graph. TPOBGE integrates the unlabeled image visual nodes \(Vision = \{v_i\}_{i=1}^{|D_u|}\) and learnable prompt nodes \(Prompt = \{u_j\}_{j=1}^{|C_u|}\) into a unified topological graph. For each visual node \(v_i\), the top-\(k\) nearest visual neighbors within \(Vision\) and the top-\(k\) most similar semantic prompt neighbors within \(Prompt\) are retained (default \(k=5\)), setting remaining connections to zero. Crucially, recognizing that vision-language models exhibit a pronounced modality gap that distorts geometric distances purely between text vectors, the framework explicitly zeros out all prompt-to-prompt entries (\(P_{text} = 0\)), preventing unreliable intra-textual similarities from corrupting graph topology.
4. Bimodal Manifold Label Propagation Inference: diffusing semantic confidence across symmetric normalized hybrid graphs
The resulting proximity matrix \(\tilde{P}\) is symmetrized and normalized by its degree matrix \(D\) to formulate the graph diffusion operator \(\hat{P} = D^{-1/2}(\tilde{P} + \tilde{P}^T)D^{-1/2}\). A pseudo-label matrix \(Y \in \mathbb{R}^{(|D_u| + |C_u|) \times |C_u|}\) is initialized where visual nodes receive zero vectors and prompt nodes are assigned identity one-hot ground-truth anchors. Global label propagation iterates according to:
The propagation weight \(\beta = 0.5\) balances topological diffusion against initial anchor consistency. Through this iterative diffusion, sharp semantic category signals flow smoothly from prompt anchors across the dense visual manifold to unlabeled image nodes, and final predictions are directly obtained at inference via maximum a posteriori selection \(\hat{y}_i = \arg\max_c \hat{Y}(i, c)\).
Key Experimental Results¶
Main Results¶
On four large-scale benchmarks with up to 1,000 classesโHybrid (498 classes), ImageNet-500 (500 classes), ImageNet-1K (1,000 classes), and Webvision (1,000 classes)โTPOBGE is evaluated against leading label-only GCD baselines and state-of-the-art text-guided OWSSL methods (all utilizing ViT-B/16 backbones). Metrics report clustering accuracy (ACC, %) across All classes (A), Known classes (K), and Unknown classes (U).
| Methods | Pretrain | Hybrid (A/K/U) | ImageNet-500 (A/K/U) | ImageNet-1K (A/K/U) | Webvision (A/K/U) | Average (A/K/U) |
|---|---|---|---|---|---|---|
| GCD | DINO | 39.45 / 48.28 / 30.37 | 58.89 / 65.71 / 55.47 | 51.25 / 56.06 / 48.84 | 45.10 / 50.59 / 42.35 | 48.67 / 55.16 / 44.26 |
| SimGCD | DINO | 49.35 / 55.73 / 42.78 | 45.60 / 66.56 / 35.12 | 34.34 / 60.84 / 21.09 | 31.80 / 56.67 / 19.36 | 40.27 / 59.95 / 29.59 |
| LegoGCD | DINO | 48.30 / 52.64 / 43.85 | 46.49 / 69.81 / 34.82 | 34.20 / 61.13 / 20.74 | 31.60 / 56.85 / 18.97 | 40.15 / 60.11 / 29.60 |
| ProtoGCD | DINO | 39.68 / 46.84 / 32.27 | 33.59 / 63.23 / 18.77 | 31.28 / 63.05 / 15.40 | 31.57 / 56.80 / 18.96 | 34.03 / 57.48 / 21.35 |
| GCD | CLIP | 52.95 / 54.30 / 51.56 | 51.26 / 64.22 / 44.78 | 43.82 / 53.55 / 38.95 | 41.41 / 50.13 / 37.05 | 47.36 / 55.55 / 43.08 |
| SimGCD | CLIP | 64.62 / 64.14 / 65.11 | 54.73 / 76.56 / 43.82 | 36.54 / 61.73 / 23.94 | 34.93 / 58.59 / 23.11 | 47.70 / 65.25 / 38.99 |
| LegoGCD | CLIP | 63.45 / 64.40 / 62.46 | 54.82 / 72.75 / 45.85 | 38.00 / 63.17 / 25.41 | 36.00 / 58.50 / 24.75 | 48.06 / 64.70 / 39.61 |
| ProtoGCD | CLIP | 48.19 / 50.39 / 45.93 | 34.58 / 65.15 / 19.30 | 32.22 / 62.53 / 17.07 | 30.15 / 57.29 / 16.58 | 36.28 / 58.84 / 24.72 |
| TP-OWSSL | CLIP | 44.81 / 42.78 / 45.15 | 68.98 / 71.71 / 67.69 | โ / โ / โ | โ / โ / โ | โ / โ / โ |
| CSC-OWSSL | CLIP | 62.91 / 68.35 / 60.19 | 56.74 / 72.13 / 49.04 | 39.03 / 63.05 / 27.02 | 34.94 / 58.22 / 23.29 | 48.41 / 65.44 / 39.89 |
| SSR2-GCD | CLIP | 55.70 / 58.71 / 52.60 | 65.41 / 70.54 / 62.84 | 58.73 / 60.02 / 58.08 | 55.42 / 58.43 / 53.92 | 58.81 / 61.92 / 56.86 |
| SpectralGCD | CLIP | 63.79 / 63.47 / 64.15 | 56.49 / 71.96 / 48.75 | 64.01 / 78.53 / 56.74 | 53.92 / 68.74 / 46.50 | 59.55 / 70.67 / 54.03 |
| TPOBGE (Ours) | CLIP | 63.81 / 62.40 / 65.27 | 73.41 / 73.47 / 73.40 | 69.70 / 71.39 / 68.98 | 66.73 / 66.63 / 66.96 | 68.41 / 68.47 / 68.65 |
Ablation Study¶
The table below validates the individual mechanisms of debiased prompt optimization and bimodal graph enhancement, reporting unknown-to-known misclassification counts, graph edge density, and computational training efficiency.
| Dataset | Evaluation Dimension | CSC-OWSSL (Prev. SOTA) | TPOBGE (Ours) | Gain & Observations |
|---|---|---|---|---|
| Hybrid | Unknown misclassified as known | 2.9k | 1.5k | 48.3% reduction in misclassified unknown samples |
| ImageNet-1K | Unknown misclassified as known | 15k | 1.6k | 89.3% reduction, preventing known-class representation collapse |
| Hybrid | Graph edge count | 15.1k | 114k | 7.5x increase in graph connectivity |
| ImageNet-1K | Graph edge count | 50k | 375k | 7.5x denser graph topology enabling manifold diffusion |
| Hybrid | Training time / GPU memory | 543 mins / 18.97 GB | 3 mins / 0.95 GB | 181x faster training, 95.0% memory reduction |
| ImageNet-1K | Training time / GPU memory | 7,303 mins / 18.93 GB | 10 mins / 9.84 GB | 730x training speedup (120+ hours down to 10 minutes) |
Key Findings¶
- Eradication of Severe Known-Class Bias: In the challenging 1,000-class ImageNet-1K setting, prior state-of-the-art CSC-OWSSL misclassifies 15,000 unknown samples into known categories, collapsing unknown accuracy down to 27.02%. TPOBGE slashes these misclassifications down to 1,600, driving unknown class accuracy up to 68.98% (+41.96% absolute gain).
- Graph Densification Overcomes Modality Gaps: Expanding graph edges from 50k to 375k via mutual \(k\)-NN visual neighbors and image-to-prompt edges while isolating uninformative prompt-to-prompt edges successfully avoids cross-modal noise and empowers smooth label diffusion.
- Drastic Computational Efficiency Boost: By operating entirely on frozen pre-extracted CLIP embeddings and restricting updates to prompt vectors and label propagation iterations, TPOBGE cuts full ImageNet-1K training time from over 120 hours to just 10 minutes on a single NVIDIA A6000 GPU.
Highlights & Insights¶
- Optimal Transport Uniform Prior Formulation: Rather than tuning heuristic loss reweighting hyperparameters, the method imposes an exact mathematical column-sum constraint in optimal transport, guaranteeing equal opportunity for unknown classes during prompt learning.
- Modality-Gap-Aware Asymmetric Topology: Acknowledging that text-to-text cosine similarities in multi-modal models often deviate from visual manifold geometry, setting text-text graph weights to zero eliminates negative transfer while retaining crisp anchor-to-image supervision.
- Inference-Friendly Frozen Backbone Pipeline: Bypassing heavy end-to-end backpropagation over high-resolution vision transformers makes large-scale OWSSL viable in memory- and compute-constrained real-world environments.
Limitations & Future Work¶
- Reliance on Multi-Modal Pre-training Generalization: Performance heavily leverages the general semantic coverage of pre-trained CLIP encoders; in specialized or niche domains (e.g., rare industrial defect inspection, complex SAR satellite imagery) lacking strong textual alignment, semantic priors may weaken.
- Fixed Unknown Category Cardinality Assumption: Experiments assume that the total number of classes \(|C_u|\) is known or provided by external clustering estimation algorithms; exploring dynamic open-vocabulary streaming discovery where novel categories emerge continuously remains open.
- Complex Long-Tailed and Open-Set Scenarios: Future extensions could generalize the standard optimal transport formulation to unbalanced optimal transport to handle extreme class imbalance and out-of-distribution noise.
Related Work & Insights¶
- vs GCD / SimGCD / LegoGCD / ProtoGCD: Standard category discovery methods operate solely on visual representations using self-distillation or prototype clustering, lacking semantic grounding and suffering heavy degradation on fine-grained unknown classes; TPOBGE exploits multi-modal language guidance with explicit debiasing.
- vs TP-OWSSL / CSC-OWSSL: Prior text-guided OWSSL frameworks suffer from extreme known-class bias, sparse bipartite graphs, and prohibitive training times (tens to hundreds of hours); TPOBGE achieves an order-of-magnitude reduction in runtime and dramatically boosts unknown class discovery.
Rating¶
- Novelty: โญโญโญโญโ [Combines optimal transport marginal constraints with asymmetric bimodal label propagation in an elegant, effective formulation]
- Experimental Thoroughness: โญโญโญโญโญ [Evaluated across four massive benchmarks with up to 1,000 classes, including misclassification counts, graph density, and efficiency benchmarks]
- Writing Quality: โญโญโญโญโญ [Clear mathematical derivations, crisp problem formulation, and self-contained structure]
- Value: โญโญโญโญโญ [Solves a critical failure mode in large-scale category discovery while achieving minute-level training speeds]