TecoPrompt: Temporal-Conservative Prompt Learning for Vision-Language Models¶
Conference: ECCV 2026
Paper: ECCV Official Page
Area: Multimodal VLM
Keywords: Prompt Learning, Optimal Transport, Noisy Labels, Vision-Language Models, Temporal Consistency
TL;DR¶
TecoPrompt is a closed-loop robust prompt learning framework that leverages global optimal transport pseudo-labeling coupled with a K-epoch temporal stability window and EMA confidence gating to conservatively rewrite noisy labels, optimized via a tri-group loss for noise-resilient few-shot adaptation.
Background & Motivation¶
Vision-language pre-trained models (VL-PTMs) such as CLIP adapt seamlessly to downstream recognition tasks via prompt learning (e.g., CoOp, CoCoOp) by tuning a small set of continuous context tokens while keeping backbone encoders frozen. However, under few-shot supervision where training samples are scarce (e.g., 16-shot per class), prompt optimization is vulnerable to label noise. Even moderate label corruption can quickly misguide the context representations, leading to overfitting on erroneous targets and degrading the well-aligned multimodal representation space.
Existing defenses against label noise typically rely on sample filtering, sample grouping, or noise-robust loss functions designed for full-model fine-tuning. These strategies often downweight or discard suspicious samples entirely, which proves wasteful in few-shot regimes where every labeled instance is critical for learning discriminative category prototypes. Conversely, label correction presents a more proactive approach, but naive or aggressive pseudo-label rewriting easily alters clean samples into incorrect targets, amplifying confirmation bias due to early training fluctuations.
Motivated by this challenge, the authors investigate the learning dynamics of optimal transport (OT) pseudo-labeling under noisy supervision. Empirical observations reveal a stark temporal asymmetry: when recomputed across epochs with a globally coupled assignment, noisy samples frequently exhibit long consecutive runs of matching the true class, whereas clean samples rarely sustain prolonged incorrect assignments. Core idea: combine globally consistent optimal transport matching with multi-epoch temporal trajectory verification to trigger hard label rewriting only when candidates remain stable over a K-epoch window with high EMA confidence, closing the loop via a clean/mid/noisy tri-group objective.
Method¶
Overall Architecture¶
TecoPrompt introduces a closed-loop prompt optimization pipeline consisting of global candidate generation, trajectory stability verification, conservative hard rewriting, and tri-group supervised updating. Given frozen CLIP vision and text encoders, image embeddings and prompt-conditioned text prototypes are computed. An entropically regularized balanced optimal transport problem is solved via Sinkhorn iterations to produce global pseudo-label distributions. A temporal stability window of size \(K\) alongside EMA-smoothed confidence evaluates candidate reliability: only candidates showing persistent agreement and high confidence trigger hard label rewriting. Finally, samples are partitioned into clean, mid, and noisy subsets and supervised with cross-entropy, generalized cross-entropy, and mean absolute error losses, respectively.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Image & Text Inputs<br/>Frozen CLIP encoders + learnable prompt tokens"] --> B["Optimal Transport Global Candidate Generation<br/>Entropic Sinkhorn matching between prototypes and image features"]
B --> C["EMA Smoothing & Confidence Filtering<br/>Moving average across epochs to suppress transient score noise"]
C --> D["Temporal Consistency Hard Label Rewriting<br/>Joint gating via K-epoch stability window and confidence threshold"]
D --> E["Tri-Group Sample Partitioning & Tiered Loss Optimization<br/>Supervision via clean CE, mid GCE, and noisy MAE objectives"]
E -->|Gradient backpropagation updating prompt tokens| A
Key Designs¶
1. Optimal Transport Global Candidate Generation: moving beyond greedy instance-wise matching
Greedy per-instance classification is prone to local error propagation under heavy noise. TecoPrompt formulates pseudo-labeling as an optimal transport problem in the joint embedding space, treating class text prototypes as supplies and image features as demands. Given similarity matrix \(S = TI^{\top} \in \mathbb{R}^{C \times N}\) between text prototypes \(T\) and image representations \(I\), the cost matrix is defined via element-wise negative log-likelihood \(C = -\log S\). The balanced entropic optimal transport problem is formulated as: $\(\min_{Q \ge 0} \langle C, Q \rangle - \epsilon H(Q) \quad \text{s.t.} \quad Q \mathbf{1}_N = \frac{1}{C}\mathbf{1}_C, \quad Q^{\top}\mathbf{1}_C = \frac{1}{N}\mathbf{1}_N\)$ where \(H(Q) = -\sum_{c,i} Q_{c,i} \log Q_{c,i}\) denotes the entropy regularizer. Solved via Sinkhorn iterations, the optimal coupling \(Q^{\star}\) provides normalized soft assignments. The candidate label \(\hat{y}_i^{(t)} = \arg\max_c Q^{\star}_{c,i}\) and confidence \(c_i^{(t)} = \max_c Q^{\star}_{c,i}\) benefit from the global marginal constraints, avoiding degenerate clustering towards noisy classes.
2. EMA Smoothing & K-Epoch Temporal Consistency Gating: conservative hard rewriting against confirmation bias
Single-epoch OT assignments can fluctuate due to gradient updates. To stabilize confidence estimation, TecoPrompt tracks an exponential moving average (EMA) confidence \(\bar{c}_i^{(t)} = \beta \bar{c}_i^{(t-1)} + (1-\beta) c_i^{(t)}\) across epochs. For hard label rewriting, two rigorous criteria are enforced over a sliding window of size \(K\): the consistency event \(\mathcal{E}_K(i) = \{\hat{y}_i^{(t-K+1)} = \cdots = \hat{y}_i^{(t)} \neq \tilde{y}_i\}\) and the sustained confidence event \(\mathcal{D}_K(i) = \{c_i^{(t-K+1)} \ge \theta, \dots, c_i^{(t)} \ge \theta\}\). Hard rewriting \(\tilde{y}_i \leftarrow \hat{y}_i^{(t)}\) is executed if and only if candidate event \(\mathcal{A}_K(i) = \mathcal{E}_K(i) \cap \mathcal{D}_K(i)\) holds. Furthermore, epochs \(t < K\) are reserved as a warm-up phase where statistics accumulate without modifying labels, preventing irreversible mis-corrections during early optimization.
3. Tri-Group Sample Partitioning & Tiered Loss Optimization: aligning supervision strength with subset reliability
Treating all training instances with a uniform loss function causes either underfitting on clean samples or severe overfitting on corrupted ones. TecoPrompt partitions dataset \(\mathcal{D}\) into three disjoint groups: - Clean group \(\mathcal{D}_{\text{cln}}\): Samples where the OT pseudo-label matches the observed label and the EMA confidence satisfies \(\bar{c}_i^{(t)} \ge \tau_{\text{clean}}\). High reliability warrants standard cross-entropy loss \(\mathcal{L}_{\text{cln}} = \sum_{i \in \mathcal{D}_{\text{cln}}} -\log s_{i,\tilde{y}_i}\) to sharpen discriminative boundaries; - Mid group \(\mathcal{D}_{\text{mid}}\): Samples exhibiting moderate confidence (\(\tau_{\text{mid}} \le \bar{c}_i^{(t)} < \tau_{\text{clean}}\) with label agreement) or rewrite-eligible candidates \(\mathcal{A}_K(i)\) undergoing transition. These are supervised via generalized cross-entropy \(\mathcal{L}_{\text{mid}} = \sum_{i \in \mathcal{D}_{\text{mid}}} \frac{1 - s_{i,\tilde{y}_i}^q}{q}\) (with \(q=0.5\)), downweighting low-confidence signals while retaining informative gradients; - Noisy group \(\mathcal{D}_{\text{nos}}\): Remaining inconsistent or low-confidence instances (\(\mathcal{D} \setminus (\mathcal{D}_{\text{cln}} \cup \mathcal{D}_{\text{mid}})\)), supervised via symmetric mean absolute error (MAE) loss \(\mathcal{L}_{\text{nos}} = \sum_{i \in \mathcal{D}_{\text{nos}}} (1 - s_{i,\tilde{y}_i})\) to prevent corrupted labels from dominating the prompt gradient.
Loss & Training¶
The total prompt optimization loss combines the three subset objectives: $\(\mathcal{L} = \lambda_{\text{c}} \mathcal{L}_{\text{cln}} + \lambda_{\text{m}} \mathcal{L}_{\text{mid}} + \lambda_{\text{n}} \mathcal{L}_{\text{nos}}\)$ Training utilizes frozen CLIP with ResNet-50 visual backbone and Transformer text encoder. The prompt consists of 16 continuous context tokens with the class token appended at the end. Optimization runs for 200 epochs using SGD with an initial learning rate of 0.002 scheduled via cosine annealing.
Key Experimental Results¶
Main Results¶
Evaluation is conducted across 7 vision benchmarks under 16-shot symmetric (Sym) and asymmetric (Asym) noise settings (accuracy, %). Table 1 presents performance under challenging noise ratios:
| Dataset | Noise Type & Ratio | CoOp | GCE | JoAPR | NLPrompt | TecoPrompt (Ours) |
|---|---|---|---|---|---|---|
| Flowers102 | Sym 50.0% | 70.1 | 84.1 | 70.2 | 89.9 | 90.5 |
| Flowers102 | Sym 75.0% | 37.2 | 70.4 | 66.9 | 76.8 | 82.3 |
| Flowers102 | Asym 50.0% | 42.6 | 69.9 | 73.8 | 81.1 | 89.4 |
| Flowers102 | Asym 75.0% | 12.6 | 39.2 | 13.3 | 55.3 | 66.6 |
| DTD | Sym 50.0% | 34.4 | 50.7 | 53.0 | 55.2 | 58.1 |
| DTD | Asym 75.0% | 11.7 | 18.2 | 28.3 | 28.4 | 38.4 |
| EuroSAT | Sym 75.0% | 26.7 | 31.4 | 27.3 | 43.8 | 52.1 |
| OxfordPets | Asym 50.0% | 38.7 | 68.1 | 75.8 | 77.5 | 84.3 |
| OxfordPets | Asym 75.0% | 14.9 | 32.0 | 43.6 | 48.6 | 75.3 |
| StanfordCars | Asym 75.0% | 12.8 | 26.6 | 23.0 | 39.5 | 51.9 |
| UCF101 | Asym 75.0% | 13.2 | 36.4 | 47.5 | 49.3 | 58.1 |
| Caltech101 | Asym 75.0% | 20.3 | 62.1 | 81.9 | 81.1 | 85.7 |
On the real-world web-noise benchmark Food101N, TecoPrompt obtains 78.67% accuracy, outperforming CoOp (69.50%), GCE (71.32%), JoAPR (72.57%), and NLPrompt (76.46%).
Ablation Study¶
On OxfordPets with varying synthetic noise ratios (10% to 70%), ablation tests analyze the impact of EMA, rewriting, window size \(K\), and group loss schemes:
| Config | EMA | Rewrite | Window K | Group Loss Scheme | 10% Noise | 30% Noise | 50% Noise | 70% Noise | Avg Accuracy (%) |
|---|---|---|---|---|---|---|---|---|---|
| (a) Baseline (no EMA/rewrite) | โ | โ | โ | โ | 87.05 | 87.03 | 85.46 | 78.22 | 84.44 |
| (b) Rewrite only | โ | โ | โ | โ | 87.21 | 86.77 | 86.02 | 85.03 | 86.26 |
| (c) EMA only | โ | โ | โ | โ | 86.05 | 85.98 | 85.34 | 81.30 | 84.67 |
| (d) Altered loss combination | โ | โ | โ | CE-0.3-0.5 | 87.58 | 86.91 | 85.98 | 85.36 | 86.46 |
| (e) Suboptimal loss (no MAE) | โ | โ | โ | 0.7-0.5-0.3 | 86.70 | 85.95 | 84.35 | 58.62 | 78.91 |
| (f) Under-conservative window | โ | โ | 1 | Default tri-group | 82.23 | 80.85 | 78.12 | 80.15 | 80.59 |
| (g) Intermediate window | โ | โ | 4 | Default tri-group | 86.52 | 85.75 | 85.98 | 86.02 | 86.07 |
| (h) Full model (default) | โ | โ | 8 | Default tri-group | 87.52 | 87.19 | 86.40 | 86.72 | 86.96 |
| (i) Over-conservative window | โ | โ | 16 | Default tri-group | 87.70 | 87.61 | 85.87 | 84.74 | 86.48 |
Key Findings¶
- Hard label rewriting serves as the primary engine for noise resilience at high noise rates, while EMA provides stability against short-term oscillations. Without proper temporal constraints (\(K=1\)), aggressive rewriting degrades performance to 80.59% due to transient mis-corrections.
- The temporal window size \(K\) governs the trade-off between correction precision and coverage. Setting \(K=8\) achieves an optimal balance: on OxfordPets with 50% noise, 234 samples are rewritten, of which 222 correspond to true ground truth (94.87% correction precision), covering 75.00% of all 296 noisy labels, while only 2.56% of clean instances are erroneously altered.
- Erroneous rewrites are predominantly confined to fine-grained visual categories with strong textural and structural similarity (e.g., Bengal cat vs. Egyptian Mau, Leonberger vs. Keeshond).
Highlights & Insights¶
- Learning Dynamics in Robust Prompt Tuning: The paper identifies and harnesses the temporal asymmetry of optimal transport pseudo-labels during training, providing theoretical and empirical grounding for when pseudo-labels can be safely converted into hard supervisory targets.
- Differentiated Tri-Group Optimization: Rather than a binary clean/discard partition, the method creates a three-tier reliability buffer that pairs confidence tiers with matching loss functions (CE for clean, GCE for transitional, MAE for noisy), ensuring robust optimization without discarding supervision.
- Negligible Computational Overhead: On a single NVIDIA RTX 3080 Ti GPU, computing optimal transport and managing trajectory queues adds only 0.20 seconds per epoch over NLPrompt on Caltech101, making it highly practical for efficient downstream adaptation.
Limitations & Future Work¶
- Omission of High-Frequency Oscillating Instances: To prioritize correction precision, samples with frequent label flips (OT-forgettable) are excluded from rewriting, leaving some persistent noisy targets uncorrected throughout training.
- Static Window Hyperparameter: The window size \(K=8\) and EMA momentum \(\beta\) are globally fixed; adaptive window sizing conditioned on class-level difficulty or trajectory convergence speed could enhance performance on severely imbalanced datasets.
- Fine-Grained Semantic Ambiguity: Mis-corrections persist among visually indistinguishable categories due to frozen representation limits, suggesting potential benefits from integrating multi-scale or local region-aware prompt tuning.
Related Work & Insights¶
- vs CoOp / CoCoOp: Standard prompt learning assumes reliable supervision and suffers rapid degradation under label corruption; TecoPrompt introduces a closed-loop conservative correction paradigm that preserves few-shot prompt adaptability under severe noise.
- vs JoAPR: JoAPR relies on a multi-stage Gaussian Mixture Model (GMM) pipeline for sample partition and re-training; TecoPrompt provides an online, single-stage framework driven by optimal transport and temporal dynamics.
- vs NLPrompt: NLPrompt applies OT for soft target assignment and subset splitting without multi-epoch temporal verification; TecoPrompt demonstrates that temporal conservatism is vital to suppress confirmation bias, delivering substantial gains especially under 75% asymmetric noise.
Rating¶
- Novelty: โญโญโญโญ [Introduces temporal consistency and learning dynamics into optimal transport pseudo-labeling for robust prompt adaptation]
- Experimental Thoroughness: โญโญโญโญโญ [Evaluates 7 synthetic noise benchmarks across symmetric and asymmetric noise levels, 1 real-world noise dataset, and detailed shot-budget analyses]
- Writing Quality: โญโญโญโญโญ [Clear mathematical formulation, thorough motivation analysis with trajectory histograms, and transparent qualitative error cases]
- Value: โญโญโญโญโญ [Provides a practical, high-precision, low-overhead solution for adapting vision-language models under corrupted supervision]