Probe, Anchor, and Amend: Active Test-Time Adaptation of Vision-Language Models¶
Conference: ECCV 2026
Paper: ECCV 2026
Area: Multimodal VLM
Keywords: active test-time adaptation, vision-language models, prototype alignment, memory-driven refinement, binary verification
TL;DR¶
The paper presents PAA, a framework that models visual and textual prototypes as an evolving target-domain semantic state and updates it via a cyclic "Probe-Anchor-Amend" routine, turning sparse binary verification feedback into structured and robust test-time adaptation for vision-language models.
Background & Motivation¶
Large-scale pretrained vision-language models (VLMs) such as CLIP exhibit impressive zero-shot generalization across diverse computer vision benchmarks. However, in realistic deployment environments, testing streams frequently suffer from severe, non-stationary distribution shifts—ranging from sudden environmental variations and corruptions to style deviations and natural adversarial perturbations. Such shifts perturb the joint visual-textual embedding space and cause classification decision boundaries to drift drastically away from the pre-trained alignment. Conventional unsupervised test-time adaptation (TTA) approaches depend heavily on high-confidence pseudo-labels or heuristic entropy minimization objectives. Under complex distribution shifts, these methods easily fall prey to confirmation bias and noise accumulation, and crucially lack explicit guidance indicating where and how decision boundaries should move.
To overcome the lack of reliable supervision, active test-time adaptation (ATTA) has emerged as an appealing paradigm that queries a restricted annotation budget during online inference to steer model adaptation. Nevertheless, existing ATTA methods were primarily conceived for single-modality vision classifiers, treating queried annotations as ephemeral, local gradient penalties on isolated samples. They lack a systematic mechanism to consolidate discrete, sparse supervision into a structured update of the shared multimodal decision space. In VLMs, domain shifts simultaneously misalign visual representations and text embeddings. Applying naive local updates disrupts the global topology of the multimodal manifold, triggering catastrophic forgetting of core source semantics and causing boundary collapse across classes.
To resolve this fundamental tension, the paper argues that test-time adaptation should not operate via ad-hoc local corrections; instead, class-level visual prototype caches and learnable text representations must be maintained as a "continuously evolving state" of the target domain. The core idea is to establish a cyclic "Probe-Anchor-Amend" (PAA) adaptation routine: class prototypes actively steer the querying of informative center and boundary instances, lightweight binary verification regularizes the prototype-anchored decision geometry via three coordinated forces, and a memory bank of past hard cases distills temporal confidence discrepancies into delayed self-refinement.
Method¶
Overall Architecture¶
PAA operates on an incoming mini-batch stream by treating class-wise visual prototypes and adapted textual embeddings as dynamic anchors defining the decision manifold. For each incoming batch, initial predictions and uncertainties are produced with a single forward pass over the current prototype state. The pipeline then executes three tightly integrated steps within each adaptation cycle: 1. Probe (Prototype-Guided Querying): Under a dynamic budget balancing long-term global class counts and short-term batch proportions, visual prototypes guide the selection of representative center samples and ambiguous boundary samples, which are sent to an oracle for minimal binary verification (i.e., whether the top-1 prediction is correct). 2. Anchor (Prototype-Regularized Update): Using binary verification signals, a binary cross-entropy loss calibrates predictive confidence while a triad of geometric prototype regularizations—cross-modal alignment, intra-class compactness, and inter-class separation—jointly updates visual LayerNorm parameters and per-class text residual vectors. 3. Amend (Memory-Driven Self-Refinement): Test instances verified as false by the oracle are enqueued into a hard-case memory bank with decay counters. In subsequent adaptation cycles, when an error flips to an alternate class with lower predictive entropy, this temporal discrepancy is harvested as a delayed, high-confidence rectification signal for self-distillation.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Streaming Test Input & Initial Forward Pass<br/>Zero-shot visual LayerNorm & text residual inference"] --> B["Prototype-Guided Querying<br/>Global-local class budget; center & boundary sampling"]
B --> C["Prototype-Regularized Anchor Update<br/>Binary verification with cross-modal, intra & inter regularization"]
C --> D["Memory-Driven Self-Refinement<br/>Revisit error cache; entropy-drop label flips trigger distillation"]
D --> E["Evolving State & Parameter Update<br/>Update visual LayerNorm & class-wise text residuals"]
Key Designs¶
1. Prototype-Guided Querying: Balancing Core Semantic Retention with Boundary Exploration
When incorporating active querying into test-time adaptation, selecting samples purely via prediction entropy or random exploration easily leads to sampling bias toward dominant classes or vulnerability to transient noise bursts. To ensure balanced and informative exploration, PAA devises an adaptive allocation policy combining long-term global statistics with short-term batch distributions, alongside a dual-prototype geometric sampling strategy. The system tracks the cumulative empirical class distribution \(\mathbf{G} = [g_1, \dots, g_C]\) across the entire test stream alongside the local batch distribution \(\mathbf{L} = [l_1, \dots, l_C]\). Given the cumulative budget \(K\) and historical query counter \(\mathbf{S} = [s_1, \dots, s_C]\), the per-class query quota for class \(c\) is computed as: $\(k_c = \left\lfloor \max(K \cdot g_c - s_c, \, \kappa \cdot l_c) \right\rfloor\)$ Here, \(K \cdot g_c - s_c\) enforces long-term demographic balance to prevent class starvation, while \(\kappa \cdot l_c\) maintains agility toward immediate batch fluctuations. Using this allocated budget \(k_c\), PAA queries two complementary sets guided by the visual prototype \(\mathbf{p}_c^v\) (the normalized mean of the lowest-entropy cached visual features for class \(c\)). First, it selects the top-\(k_c\) central samples \(\mathcal{S}_c^{\mathrm{center}}\) with maximum cosine similarity \(\cos(\mathcal{E}_v(\mathbf{x}), \mathbf{p}_c^v)\), reinforcing well-established semantic clusters and preventing forgetting. Second, it computes the nearest-prototype distance \(d(\mathbf{x}) = \min_{c'} (1 - \cos(\mathcal{E}_v(\mathbf{x}), \mathbf{p}_{c'}^v))\) and picks the top-\(k_c\) boundary samples \(\mathcal{S}_c^{\mathrm{boundary}}\) with highest distance, uncovering ambiguous regions where the decision boundary requires shifting. The final query set \(\mathcal{S}\) merges both sets across all classes.
2. Prototype-Regularized Anchor Update: Triad Regularization for Stable Yet Plastic Manifolds
Querying full categorical labels in real-time streaming creates prohibitive human latency and interaction bottlenecks. PAA relies strictly on lightweight binary verification feedback \(v_{\mathbf{x}} = \mathcal{O}(\mathbf{x}, \hat{y}) \in \{0, 1\}\), where \(v_{\mathbf{x}}=1\) indicates a correct prediction and \(v_{\mathbf{x}}=0\) indicates an error. Model confidence is calibrated via binary cross-entropy \(\mathcal{L}_{\mathrm{BCE}}\). However, naively optimizing solely on scalar binary verification over few samples quickly overfits and distorts the shared embedding manifold. To ensure the decision space remains stable yet plastic, the Anchor step introduces three coordinated geometric forces regularizing the updates of visual LayerNorm parameters \(\theta^v_{\mathrm{LN}}\) and per-class text residuals \(\mathbf{r}_c^t\) (defining \(\mathbf{p}_c^t = \mathrm{norm}(\mathbf{z}_c^t + \mathbf{r}_c^t)\)): - Cross-Modal Alignment (\(\mathcal{L}_{\mathrm{align}}\)): Symmetrized Info-NCE loss enforces high mutual information between paired visual and textual prototypes \(\mathbf{p}_c^v\) and \(\mathbf{p}_c^t\), preventing cross-modal drift: $\(\mathcal{L}_{\mathrm{align}} = -\frac{1}{C} \sum_{c=1}^C \left( \log \frac{\exp(\mathbf{p}_c^{t\top} \mathbf{p}_c^v / \tau)}{\sum_{c'} \exp(\mathbf{p}_c^{t\top} \mathbf{p}_{c'}^v / \tau)} + \log \frac{\exp(\mathbf{p}_c^{t\top} \mathbf{p}_c^v / \tau)}{\sum_{c'} \exp(\mathbf{p}_{c'}^{t\top} \mathbf{p}_c^v / \tau)} \right)\)$ - Intra-Class Compactness (\(\mathcal{L}_{\mathrm{intra}}\)): Pulls verified positive instances (\(v_{\mathbf{x}}=1\)) closer to their corresponding visual class prototype \(\mathbf{p}_{\hat{y}}^v\), tightening target clusters. - Inter-Class Separation (\(\mathcal{L}_{\mathrm{inter}}\)): Since the learnable text residual \(\mathbf{r}_c^t\) models domain shift directions, it must not align with inter-class semantic difference vectors \((\mathbf{z}_{c'}^t - \mathbf{z}_c^t)\) from the source domain, which would collapse class discriminability. Thus, \(\mathcal{L}_{\mathrm{inter}}\) penalizes residual alignment against all pairwise source text difference vectors: $\(\mathcal{L}_{\mathrm{inter}} = \frac{1}{C(C-1)} \sum_{c=1}^C \sum_{c' \neq c} \cos(\mathbf{r}_c^t, \, \mathbf{z}_{c'}^t - \mathbf{z}_c^t)\)$ The combined anchor objective \(\mathcal{L}_{\mathrm{anchor}} = \mathcal{L}_{\mathrm{BCE}} + \lambda_{\mathrm{anchor}}(\mathcal{L}_{\mathrm{align}} + \mathcal{L}_{\mathrm{intra}} + \mathcal{L}_{\mathrm{inter}})\) maintains structural cohesion across modalities under minimal supervisory feedback.
3. Memory-Driven Self-Refinement: Converting Temporal Discrepancies into Delayed Rectifications
In continuous test-time adaptation, misclassified instances pose severe risks if ignored or blindly suppressed without target direction. Directly discarding false predictions (\(v_{\mathbf{x}}=0\)) squanders valuable human verification effort, while naive repulsive updates can distort neighboring boundaries. The Amend step introduces a memory bank \(\mathcal{M}\) acting as a slow controller. Each verified false prediction is stored as \((\mathbf{x}_i, \hat{y}_i, \mathcal{H}_i, \ell_i)\), caching the test sample, erroneous pseudo-label \(\hat{y}_i\), initial prediction entropy \(\mathcal{H}_i\), and a remaining lifetime counter \(\ell_0=3\). In subsequent cycles, the updated model re-evaluates these cached hard cases. When a sample's predicted class flips (\(\hat{y}' \neq \hat{y}\)) and its entropy drops (\(\mathcal{H}' < \mathcal{H}\)), the model has autonomously resolved its previous ambiguity. This temporal discrepancy is harvested into a rectification subset \(\Delta\), producing a self-refinement distillation loss: $\(\mathcal{L}_{\mathrm{amend}} = -\frac{1}{|\Delta|} \sum_{(\mathbf{x}, \hat{y}') \in \Delta} \log p'_{\hat{y}'}\)$ Unresolved entries decrement their lifetime counter (\(\ell \leftarrow \ell - 1\)) and are purged upon expiry. This asynchronous filtering loop rectifies past errors and curbs error accumulation at zero extra annotation expense.
Loss & Training¶
The overall adaptation objective per mini-batch combines the immediate anchor loss with the delayed memory refinement loss: $\(\mathcal{L} = \mathcal{L}_{\mathrm{anchor}} + \lambda_{\mathrm{amend}} \mathcal{L}_{\mathrm{amend}}\)$ Hyper-parameters are set to \(\lambda_{\mathrm{anchor}} = 0.1\) and \(\lambda_{\mathrm{amend}} = 0.1\). The visual and textual transformer backbones remain completely frozen; optimization is restricted exclusively to visual LayerNorm parameters \(\theta^v_{\mathrm{LN}}\) and per-class text residual vectors \(\{\mathbf{r}_c^t\}_{c=1}^C\). Following TPT, 63 augmented views per test instance are sampled during inference and aggregated for prediction stability. Batch size is set to 128, the per-cycle query budget is \(\kappa = 16\) for central and boundary sets respectively (32 binary queries per batch), prototype cache capacity is \(H=3\), and initial memory lifetime is \(\ell_0 = 3\). Adaptation executes on a single consumer-grade NVIDIA RTX 3090 GPU.
Key Experimental Results¶
Main Results¶
Experiments are conducted on the Natural Distribution Shifts benchmark (ImageNet and four robustness variants: ImageNet-A, ImageNet-V2, ImageNet-R, ImageNet-Sketch) and the Cross-Domain benchmark comprising 10 fine-grained visual classification datasets.
| Method | Type | Visual Backbone | ImageNet | ImageNet-A | ImageNet-V2 | ImageNet-R | ImageNet-S | Average OOD | Overall Average |
|---|---|---|---|---|---|---|---|---|---|
| CLIP (Zero-shot) | Non-adapted | ResNet-50 | 58.16 | 21.83 | 51.41 | 56.15 | 33.37 | 44.18 | 40.69 |
| TPT (NeurIPS'22) | Non-Active TTA | ResNet-50 | 60.74 | 26.67 | 54.70 | 59.11 | 35.09 | 47.26 | 43.89 |
| TDA (CVPR'24) | Non-Active TTA | ResNet-50 | 61.35 | 30.29 | 55.54 | 62.58 | 38.12 | 49.58 | 46.63 |
| DPE (NeurIPS'24) | Non-Active TTA | ResNet-50 | 63.41 | 30.15 | 56.72 | 63.72 | 40.03 | 50.81 | 47.66 |
| BoostAdapter (NeurIPS'24) | Non-Active TTA | ResNet-50 | 62.14 | 35.47 | 56.31 | 62.61 | 38.80 | 51.07 | 48.30 |
| SimATTA (ICLR'24) | Active TTA | ResNet-50 | 60.31 | 35.15 | 56.34 | 62.63 | 39.47 | 50.78 | 48.40 |
| EATTA (CVPR'25) | Active TTA | ResNet-50 | 60.99 | 35.68 | 56.29 | 62.71 | 39.02 | 50.94 | 48.43 |
| PAA (Ours) | Active TTA | ResNet-50 | 63.42 | 37.27 | 56.77 | 67.37 | 41.03 | 53.15 | 50.61 |
| CLIP (Zero-shot) | Non-adapted | ViT-B/16 | 66.73 | 47.87 | 60.86 | 73.98 | 46.09 | 59.11 | 57.20 |
| TPT (NeurIPS'22) | Non-Active TTA | ViT-B/16 | 68.98 | 54.77 | 63.45 | 77.06 | 47.94 | 62.44 | 60.81 |
| TDA (CVPR'24) | Non-Active TTA | ViT-B/16 | 69.51 | 60.11 | 64.67 | 80.24 | 50.54 | 65.01 | 63.89 |
| ZERO (NeurIPS'24) | Non-Active TTA | ViT-B/16 | 71.17 | 62.75 | 65.23 | 80.75 | 50.59 | 66.10 | 64.83 |
| BoostAdapter (NeurIPS'24) | Non-Active TTA | ViT-B/16 | 69.37 | 64.53 | 65.51 | 80.95 | 51.28 | 66.33 | 65.57 |
| BATCLIP (ICCV'25) | Non-Active TTA | ViT-B/16 | 70.60 | 62.03 | 65.18 | 80.66 | 51.49 | 65.99 | 64.84 |
| SimATTA (ICLR'24) | Active TTA | ViT-B/16 | 67.07 | 63.63 | 64.42 | 79.30 | 50.12 | 64.91 | 64.37 |
| EATTA (CVPR'25) | Active TTA | ViT-B/16 | 67.10 | 64.85 | 64.52 | 79.59 | 49.79 | 65.17 | 64.69 |
| PAA (Ours) | Active TTA | ViT-B/16 | 72.03 | 66.59 | 65.51 | 82.93 | 53.89 | 68.19 | 67.23 |
On the 10 cross-domain datasets (ViT-B/16), PAA achieves an overall top-1 accuracy of 73.60%, outperforming zero-shot CLIP (63.58%), TDA (67.53%), DPE (69.40%), SimATTA (71.34%), and EATTA (71.49%). The performance gains are particularly remarkable on EuroSAT (90.93% vs. 42.01% zero-shot), indicating superior cross-domain adaptation under extreme distribution gaps.
Ablation Study¶
Systematic ablations evaluate the contribution of individual objective functions and parameter modules on the Natural Distribution Shifts benchmark using ViT-B/16.
| Config | ImageNet | ImageNet-A | ImageNet-V2 | ImageNet-R | ImageNet-S | Average (%) | Drop vs. Full |
|---|---|---|---|---|---|---|---|
| Full Model | 72.03 | 66.59 | 65.51 | 82.93 | 53.89 | 68.19 | - |
| w/o \(\mathcal{L}_{\mathrm{align}}\) | 71.61 | 66.11 | 65.08 | 82.65 | 53.53 | 67.80 | -0.39% |
| w/o \(\mathcal{L}_{\mathrm{intra}}\) | 71.82 | 66.16 | 65.31 | 82.63 | 53.69 | 67.92 | -0.27% |
| w/o \(\mathcal{L}_{\mathrm{inter}}\) | 71.75 | 66.48 | 65.37 | 82.84 | 53.70 | 68.03 | -0.16% |
| w/o \(\mathcal{L}_{\mathrm{amend}}\) | 71.79 | 66.27 | 65.50 | 82.68 | 53.67 | 67.98 | -0.21% |
| Freeze \(\theta^v_{\mathrm{LN}}\) | 71.55 | 65.24 | 65.14 | 81.85 | 52.61 | 67.28 | -0.91% |
| Freeze \(\{\mathbf{r}_c^t\}\) | 71.27 | 65.08 | 65.12 | 82.18 | 53.26 | 67.38 | -0.81% |
| \(\{\mathbf{r}_c^t\} \to \theta^t_{\mathrm{LN}}\) | 71.76 | 65.33 | 65.22 | 82.47 | 53.28 | 67.61 | -0.58% |
Key Findings¶
- Cross-Modal Alignment is the Anchor's Backbone: Disabling \(\mathcal{L}_{\mathrm{align}}\) incurs the sharpest decline among all loss ablations (-0.39% average accuracy), demonstrating that maintaining mutual information alignment between visual prototypes and text embeddings is essential to prevent representation collapse under sparse supervision.
- Decoupled Visual and Text Adaptation is Essential: Freezing visual LayerNorm or text residuals degrades accuracy by 0.91% and 0.81%, respectively. Furthermore, replacing per-class text residuals with global text LayerNorm parameters lowers accuracy to 67.61%, proving that class-specific residual vectors provide necessary semantic granularity while safeguarding general linguistic priors.
- Binary Verification Rivals Full Categorical Supervision: Under identical budgets, binary verification achieves 68.19% average accuracy, lagging only 0.65% behind full 1000-class categorical ground truth (68.84%). This highlights the efficiency of the prototype-anchored loss in extracting maximal supervisory utility from simple yes/no signals.
- Superior Compute-Accuracy Pareto Frontier: On ImageNet-A, while TPT requires >2 hours and DiffTPT requires >4 hours, PAA finishes in only 17 minutes—matching the speed of BoostAdapter (17 min) and SimATTA (18 min)—while delivering a substantial accuracy boost from 63.63%~64.53% to 66.59%.
Highlights & Insights¶
- From Local Corrections to Structured State Evolution: Prior ATTA methods treat queried instances as localized gradient penalties. PAA re-envisions class prototypes as an evolving state, translating sparse binary feedback into global adjustments across visual and textual manifolds.
- Dual-Anchor Active Exploration: The Probe mechanism breaks the bias of entropy-only querying by jointly sampling center instances for anti-forgetting stability and boundary instances for shift tracking, balanced by an online global-local budget scheduler.
- Self-Healing Asynchronous Memory Loop: The Amend module converts past negative feedback into an asset, waiting for temporal confidence flips to generate high-precision self-distillation targets without requiring human re-annotation.
Limitations & Future Work¶
- Back-Propagation Overhead vs. Training-Free Methods: Although an order of magnitude faster than prompt-tuning approaches, updating LayerNorm and text residuals still incurs back-propagation graph latency compared to training-free cache lookup methods (e.g., ZERO).
- Idealized Noise-Free Oracle: Experiments assume 100% accurate human binary verifications. Practical deployments will encounter human fatigue and label noise, motivating the integration of noise-tolerant verification losses.
- Ultra-Lightweight Adaptation Modules: Future research could explore hypernetwork-driven parameter generators to achieve the benefits of PAA without explicit backward passes on edge devices.
Related Work & Insights¶
- vs. TDA / ZERO: Training-free cache and zero-temperature forward methods are computationally lightweight but vulnerable to confirmation bias and error accumulation under adversarial distribution shifts (e.g., TDA reaches only 60.11% on ImageNet-A). PAA leverages sparse binary checks to steer prototype caches safely, attaining 66.59%.
- vs. SimATTA / EATTA: Established ATTA schemes target single-modality vision models and rely on standard cross-entropy updates that risk distorting multimodal representation geometry. PAA customizes active querying and triple-force regularization specifically to preserve joint visual-textual manifold structures.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates active test-time adaptation as a cyclic evolution of cross-modal prototypes.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarks spanning 5 distribution shift sets and 10 cross-domain datasets with deep efficiency and ablation analysis.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous conceptual formulation, clear mathematical modeling, and seamless narrative progression.
- Value: ⭐⭐⭐⭐⭐ Provides an efficient, label-frugal paradigm for deploying foundation vision-language models reliably in dynamic production environments.