SyncVL: Synchronizing Vision โท Language Using Unsupervised Adaptation¶
Conference: ECCV 2026
Paper: ECCV 2026 Official
PDF: ECCV Paper PDF
Area: Multimodal VLM / Segmentation
Keywords: Vision-Language Models, Unsupervised Adaptation, Modality Synchronization, Group Relative Policy Optimization (GRPO), Bidirectional Distillation
TL;DR¶
Addressing the modality gap and domain shift in contrastive vision-language models without target labels, SyncVL freezes pretrained backbones and introduces bidirectional lightweight Transformer synchronizers optimized through group relative policy optimization (GRPO) consistency and alignment rewards alongside latent mutual distillation.
Background & Motivation¶
Large-scale contrastive vision-language models (such as CLIP, EVA-CLIP, ImageBind, and SigLIP) learn joint multimodal representations across paired image-text datasets via dual-encoder architectures, exhibiting remarkable zero-shot generalization capabilities. However, due to intrinsic heterogeneity between visual and textual modalities, an inherent modality gap persists within their latent representation space. Crucially, when adapting to specialized visual domains, severe distribution shifts, or degraded scenarios, this pre-learned synchronization deteriorates significantly, limiting the discriminative capacity and transferability of learned representations on unlabeled target data.
Existing methods aiming to improve cross-modal alignment generally fall into supervised tuning and unsupervised adaptation. The former fine-tunes prompt vectors or adds parameter adapters using target domain labels; however, annotated samples are often cost-prohibitive or inaccessible, and supervised fine-tuning on scarce target sets easily overfits to spurious biases and induces catastrophic forgetting of general representations. Conversely, unsupervised adaptation approaches typically rely on self-supervised pseudo-labeling to update backbone parameters. Updating large-scale encoders inevitably erodes pre-trained multimodal priors, while errors in pseudo-labels accumulate through iterative cycles, exacerbating confirmation bias. Furthermore, standard absolute loss formulations induce severe gradient conflicts between intra-instance augmentation consistency and cross-modal contrastive alignment, destabilizing the optimization dynamics.
To resolve these tensions, this work proposes a label-free bidirectional representation synchronization paradigm that completely circumvents backbone updating. By introducing lightweight bidirectional mappings on top of frozen representations and leveraging group-relative advantage normalization, it eliminates competing gradient interference. Core idea: introduce bidirectional Vision-to-Text (V2T) and Text-to-Vision (T2V) Transformer synchronizers over frozen VLM encoders, formalizing unsupervised modality synchronization as a Group Relative Policy Optimization (GRPO) process with view consistency, pseudo-class alignment, and latent mutual distillation rewards to construct a unified, symmetric shared subspace.
Method¶
Overall Architecture¶
The end-to-end execution pipeline of SyncVL operates as follows. Given target unlabeled images, multiple augmented views are generated to form visual input groups, and initial visual embeddings are extracted using the frozen image encoder. Concurrently, multiple prompt templates per target class are processed by the frozen text encoder to generate textual embeddings and corresponding class centers. Two lightweight encoder-decoder synchronizers then operate cooperatively: the Vision-to-Text synchronizer (\(f_{v2t}\)) projects visual augmentation groups into the textual semantic space guided by intra-group consistency and contrastive alignment rewards; subsequently, the Text-to-Vision synchronizer (\(f_{t2v}\)) maps prompt template groups into the visual space anchored by pseudo-class centers generated by \(f_{v2t}\), under the supervision of consistency, alignment, and latent knowledge distillation rewards. Both synchronizers alternate across iterative rounds with dynamic pseudo-label regrouping. At inference time, only the lightweight encoders of both synchronizers are retained to map test images and textual classes into the shared latent subspace for cosine matching.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Unlabeled Images & Class Prompt Templates"] --> B["Frozen Backbone Feature Extraction<br/>Multi-view Augmentations & Multi-template Encoding"]
B --> C["Bidirectional Lightweight Synchronizers<br/>V2T & T2V Transformer Mappings"]
C --> D["GRPO-Driven Unsupervised Reward Optimization<br/>Intra-Group Distribution Consistency + Pseudo-Class Alignment"]
D --> E["Latent Subspace Mutual Distillation & Iterative Refinement<br/>V2T Teacher Guides Symmetric T2V Latent Space"]
E --> F["Output: Aligned Common Subspace<br/>Preserving Dual-Encoder Zero-Shot Structure"]
Key Designs¶
1. Bidirectional Lightweight Synchronizers: Decoupling Frozen Backbones from Symmetrical Cross-Modal Projections Standard adaptation techniques directly fine-tune encoder weights (\(f_\theta\) and \(g_\phi\)), risking the degradation of pre-trained multimodal representations under unsupervised settings. SyncVL strictly freezes all parameters of both visual and textual encoders, appending two lightweight Transformer-based encoder-decoder synchronizers (each featuring 2 Transformer blocks), denoted as Vision-to-Text \(f_{v2t}: \mathbb{R}^d \to \mathbb{R}^d\) and Text-to-Vision \(f_{t2v}: \mathbb{R}^d \to \mathbb{R}^d\). While \(f_{v2t}\) progressively adapts visual embeddings toward textual semantic centers, \(f_{t2v}\) projects textual templates toward visual clusters. This bidirectional architecture breaks the conventional asymmetry of unidirectional visual-to-text projection, establishing a balanced reciprocal anchor for mutual calibration.
2. GRPO-Driven Unsupervised Reward Optimization: Eliminating Destructive Gradient Conflicts Directly combining consistency loss across augmented views with contrastive alignment loss under standard gradient descent leads to severe gradient conflict, where their gradient cosine similarity is persistently negative. SyncVL reformulates feature synchronization using Group Relative Policy Optimization (GRPO), treating the synchronizer as an implicit policy \(\pi_\theta\) over continuous representations. For an image \(x_j\), \(k\) augmented view embeddings \(\{v_{j,i}\}_{i=1}^k\) mapped via \(v_{j,i}^{t} = f_{v2t}(v_{j,i})\) constitute a group. Two cooperative rewards supervise the group: - Distributional Consistency Reward (\(L_{v2t}^{con}\)): Computes the softmax class distribution \(p_{j,i}\) between each projected instance and text class centers \(\bar{t}_c\), minimizing the KL-divergence against the group mean distribution \(\bar{p}_j = \frac{1}{k}\sum_{i=1}^k p_{j,i}\): $\(L_{v2t}^{con}(j, i) = D_{\text{KL}}(\bar{\mathbf{p}}_j \,\|\, \mathbf{p}_{j,i})\)$ - Contrastive Alignment Reward (\(L_{v2t}^{aln}\)): Establishes a positive pseudo-class \(c^+\) via majority voting across the group and designates the second-highest class as hard negative \(c^-\), optimizing contrastive likelihood: $\(L_{v2t}^{aln}(j, i) = -\log \frac{\exp(s_{(j,i,c^+)} / \tau)}{\exp(s_{(j,i,c^+)} / \tau) + \exp(s_{(j,i,c^-)} / \tau)}\)$ GRPO normalizes rewards relative to the group mean and standard deviation: \(A_{j,i} = (r_{j,i} - \bar{r}_j) / (\sigma_j + \epsilon)\), optimizing a clipped objective. This normalizer decouples competing objectives and turns the gradient cosine similarity from negative conflict into strong positive synergy (e.g., from -0.24 to +0.65 on CIFAR-100).
3. Latent Subspace Mutual Distillation & Iterative Refinement: Progressive Teacher-Student Symmetry for Inference Upon training \(f_{v2t}\), it serves as an informative teacher guiding the reverse synchronizer \(f_{t2v}\). First, pseudo-labels predicted by \(f_{v2t}\) on the unlabeled target data are aggregated to form visual pseudo-class centroids \(\bar{v}_c = \frac{1}{n_c}\sum_{j=1}^{n_c} v_{j}^t\). Second, the latent representation of the \(f_{v2t}\) encoder provides direct supervision to the \(f_{t2v}\) encoder through a knowledge distillation reward: $\(L_{t2v}^{KD}(c, i) = f_{v2t}^e(\bar{\mathbf{v}}_c) \cdot f_{t2v}^e(\mathbf{t}_{c,i})\)$ where \(f_{v2t}^e\) and \(f_{t2v}^e\) denote the encoder segments of each synchronizer. In later iterations, group formation transitions from instance augmentations to pseudo-label clustering (grouping all images assigned to the same pseudo-class). During inference, decoders are discarded, and only the two lightweight encoders \(f_{v2t}^e\) and \(f_{t2v}^e\) are executed, projecting inputs into the shared aligned subspace with zero inter-model communication overhead.
Loss & Training¶
The overall learning pipeline follows a multi-stage progressive optimization: 1. Initial V2T Phase: Keeping VLM backbones frozen, each unlabeled image generates \(k=7\) augmentations. \(f_{v2t}\) is trained with the AdamW optimizer (learning rate \(1\times 10^{-4}\), weight decay 0.01) under GRPO on \(L_{v2t} = L_{v2t}^{con} + L_{v2t}^{aln}\) for 100 epochs. 2. Reverse T2V Phase: Visual pseudo-centroids \(\bar{v}_c\) are derived using trained \(f_{v2t}\). For each class, \(m\) text templates form a group, optimizing \(f_{t2v}\) using \(L_{t2v} = L_{t2v}^{con} + L_{v2t}^{aln} + L_{t2v}^{KD}\) for 100 epochs. 3. Iterative Refinement: Pseudo-labels are updated to regroup the dataset, refining both synchronizers alternately. At inference time, only \(f_{v2t}^e\) and \(f_{t2v}^e\) are deployed for zero-shot testing.
Key Experimental Results¶
Main Results¶
On standard unsupervised adaptation (UA) benchmarks with the CLIP ViT-B/32 backbone across 14 diverse datasets (encompassing generic object recognition, fine-grained categories, textures, human actions, and remote sensing), SyncVL significantly outperforms existing state-of-the-art methods:
| Method | Avg. | CIFAR-10 | CIFAR-100 | DTD | Cars | UCF101 | EuroSAT | Food101 | Flowers | Aircraft | ImageNet |
|---|---|---|---|---|---|---|---|---|---|---|---|
| CLIP (B/32) Baseline | 65.04 | 91.30 | 65.10 | 44.50 | 64.50 | 49.40 | 82.40 | 87.00 | 66.46 | 21.20 | 63.20 |
| UPL (ICLR 22) | 64.05 | 91.26 | 67.41 | 45.37 | 62.04 | 51.88 | 84.25 | 83.84 | 67.40 | 17.07 | 62.12 |
| POUF (ICML 23) | 64.82 | 90.50 | 62.00 | 46.10 | 61.20 | 62.90 | 82.10 | 87.80 | 67.80 | 18.20 | 60.00 |
| CPL (ICML 24) | 68.54 | 94.68 | 77.30 | 56.30 | 71.00 | 82.20 | 83.18 | 87.62 | 76.70 | 19.76 | - |
| TransCLIP (NeurIPS 24) | 68.42 | 91.18 | 67.38 | 50.40 | 59.00 | 81.50 | 89.00 | 74.40 | 68.70 | 20.30 | 67.44 |
| LafTer (NeurIPS 23) | 69.90 | 95.80 | 74.60 | 46.10 | 73.90 | 82.45 | 84.93 | 71.00 | 68.20 | 19.86 | 64.50 |
| ReCLIP (WACV 24) | 70.08 | 94.84 | 72.26 | 53.88 | 67.01 | 70.80 | 84.19 | 87.49 | 72.63 | 18.87 | 73.89 |
| DPA (WACV 25) | 72.29 | 95.97 | 76.47 | 55.69 | 68.49 | 80.04 | 84.76 | 90.71 | 75.56 | 20.67 | 68.13 |
| SyncVL (Ours) | 76.48 | 98.65 | 81.14 | 63.28 | 74.50 | 84.14 | 86.98 | 93.96 | 79.84 | 25.06 | 74.84 |
Furthermore, evaluation on dense prediction tasks demonstrates that SyncVL effectively models fine-grained localization and structural segmentation:
| Backbone | Task / Dataset | PASCAL VOC (AP50) | PASCAL VOC (mIoU) | MS COCO (AP50) | MS COCO (mIoU) |
|---|---|---|---|---|---|
| CLIP (B/16) | Baseline | 30.48 | 16.20 | 27.80 | 4.40 |
| CLIP (B/16) | + SyncVL (Ours) | 35.09 (+4.61) | 27.65 (+11.45) | 31.40 (+3.60) | 6.01 (+1.61) |
| TULIP (B/16) | Baseline | 28.70 | 15.84 | 21.32 | 3.72 |
| TULIP (B/16) | + SyncVL (Ours) | 32.14 (+3.44) | 17.95 (+2.11) | 25.04 (+3.72) | 5.69 (+1.97) |
Ablation Study¶
Systematic ablations analyze module combinations, objective formulations, and individual loss components:
| Study Category | Variant Configuration | CIFAR-10 | CIFAR-100 | DTD | Impact & Findings |
|---|---|---|---|---|---|
| Module Hierarchy | Baseline (No synchronizer) | 91.30 | 65.10 | 44.50 | Pretrained dual-encoder zero-shot performance |
| Unidirectional V2T (w/o GRPO) | 86.37 | 68.54 | 50.18 | Gradient conflicts degrade accuracy on CIFAR-10 | |
| Unidirectional V2T (+ GRPO) | 93.06 | 71.86 | 53.24 | Relative advantage normalization recovers stability | |
| Bidirectional V2T + T2V (w/o GRPO) | 95.86 | 74.18 | 56.72 | Dual-path mapping improves cross-modal representation | |
| Full SyncVL (Bidirectional + GRPO) | 98.65 | 81.14 | 63.28 | Optimal synergy across all components | |
| Reward Composition | Consistency Reward \(L_{con}\) only | 94.18 | 73.69 | 55.24 | Stabilizes instance groupings but lacks cross-modal margin |
| Alignment Reward \(L_{aln}\) only | 94.64 | 73.24 | 55.40 | Enforces discrimination but susceptible to perturbation noise | |
| Joint \(L_{con} + L_{aln}\) | 96.32 | 77.16 | 57.06 | Complementary views improve clustering boundary | |
| Full Rewards (\(L_{con} + L_{aln} + L_{KD}\)) | 98.65 | 81.14 | 63.28 | Latent KD bridges bidirectional latent symmetry |
Key Findings¶
- GRPO Resolves Gradient Conflict: Gradient cosine similarity between consistency and alignment rewards shifts from negative without GRPO (-0.10, -0.24, -0.09 on CIFAR-10, CIFAR-100, DTD) to strongly positive with GRPO (+0.63, +0.65, +0.56), verifying that relative policy normalization decouples conflicting objectives into cooperative gradients.
- Bidirectional Symmetry Outperforms Unidirectional Mapping: Moving from single-branch V2T to bidirectional V2T+T2V yields a 5-10% accuracy surge, proving that backward T2V projection anchors textual semantics in target visual density.
- Universal Backbone Transferability: SyncVL consistently improves performance across 6 diverse backbones (CLIP, SigLIP, ImageBind, EVA-CLIP, SigLIP-2, TULIP), yielding gains from +2.5% to +13.5% without modifying encoder parameters.
Highlights & Insights¶
- Adapting Policy Optimization to Continuous Multimodal Embeddings: Instead of restricting RLHF/GRPO to discrete token decoding, SyncVL innovatively treats continuous feature generation as a latent policy, leveraging group advantage normalization to supervise embedding spaces.
- Strict Parameter Freezing with Minimal Overhead: Keeping heavy dual encoders frozen eliminates catastrophic forgetting and drastically minimizes trainable parameters to two compact 2-block Transformers.
- Self-Healing Capability under Poor Initializations: On low-performing initializations where baseline accuracy is below 25% (Aircraft, GTSRB), iterative majority voting within groups enables SyncVL to recover by +15% to +17% without ground truth labels.
Limitations & Future Work¶
- Admitted Limitations: When initial zero-shot predictions on fine-grained datasets are overwhelmingly dominated by incorrect labels, group majority voting can initially propagate noisy pseudo-supervision, slowing convergence.
- Identified Boundaries: The training pipeline requires multi-stage sequential scheduling (initial V2T training, centroid aggregation, T2V distillation, and iterative regrouping), which is computationally more complex than one-step prompt adapters. Additionally, the textual side relies on curated prompt templates.
- Promising Extensions: Integrating LLM-generated descriptive attribute expansions could enrich prompt template distributions dynamically, enhancing T2V synchronization.
Related Work & Insights¶
- vs DPA (WACV 2025): DPA introduces dual prototype alignment via direct backbone adaptation; SyncVL keeps encoders frozen, adopts bidirectional Transformers with GRPO, and outperforms DPA by 4.19% on average across 14 benchmarks.
- vs LafTer (NeurIPS 2023): LafTer optimizes zero-shot classifiers via unidirectional text generation; SyncVL introduces symmetrical T2V projection and latent space distillation, avoiding confirmation bias.
- vs CoOp / MaPLe (CVPR / IJCV): While CoOp and MaPLe require 8-shot or 16-shot supervised target annotations, unsupervised SyncVL exceeds 16-shot MaPLe accuracy on 5 out of 6 standard benchmarks without any labels.
Rating¶
- Novelty: โญโญโญโญโญ First formulation of continuous VLM embedding synchronization via GRPO relative rewards with bidirectional Transformer adapters.
- Experimental Thoroughness: โญโญโญโญโญ Evaluated across 24 benchmarks covering classification, OOD shifts, clustering, cross-modal retrieval, detection, and segmentation.
- Writing Quality: โญโญโญโญโญ Clear theoretical motivation on gradient interference, rigorous mathematical formulation, and consistent experimental structure.
- Value: โญโญโญโญโญ Provides an effective and scalable paradigm for adapting frozen foundational models to label-scarce real-world target distributions.