Skip to content

DICE: Disentangled Instance-Class knowlEdge prompt tuning via SAE for Vision-Language Models

Conference: ECCV 2026
Paper: ECCV Official
Area: Multimodal VLM
Keywords: Prompt Learning, Sparse Autoencoder, Vision-Language Model, Few-Shot Generalization, Feature Disentanglement

TL;DR

Addressing the misalignment and visual ungroundedness of predefined LLM class descriptions, DICE reformulates the Sparse Autoencoder (SAE) into a learnable concept dictionary and proposes a co-activation score to dynamically disentangle and select instance-specific concept vectors for prompt fusion, setting a new benchmark for few-shot prompt tuning.

Background & Motivation

Vision-language models (VLMs) like CLIP exhibit remarkable zero-shot transfer capabilities across a wide spectrum of visual tasks, but adapting them in low-resource few-shot regimes often leads to catastrophic overfitting. Prompt tuning has emerged as a parameter-efficient fine-tuning alternative to mitigate this vulnerability. Recently, numerous approaches have leveraged Large Language Models (LLMs) to synthesize descriptive textual promptsโ€”such as querying "What does a [class] look like?"โ€”to introduce rich semantic priors that help structure the embedding space and anchor novel classes.

However, existing LLM-driven prompt tuning frameworks suffer from three fundamental limitations. First, predefined LLM descriptions are generated in isolation from actual visual evidence and applied uniformly across all instances within a category; in practice, intra-class variations in lighting, pose, background, and sub-species mean that generic descriptions often clash with the specific image content (e.g., describing visual features absent in the scene), degrading cross-modal alignment. Second, representations in VLMs and LLMs are dense and entangled (polysemantic), lacking an explicit mechanism to isolate and extract the precise semantic factors directly relevant to an individual image. Finally, current pipelines lack a principled mechanism to unify global class-level priors with local instance-level visual cues into a coherent prompt representation, frequently collapsing into coarse class-average semantics that struggle on fine-grained visual distinctions.

The core premise of this work is that class-level semantics provide global structural consistency while instance-level cues deliver granular discrimination, requiring both to be dynamically disentangled and synthesized. Core idea: repurpose a Sparse Autoencoder (SAE) as an interpretable cross-modal concept dictionary, using a class-guided co-activation score between image and class representations to dynamically select instance-specific monosemantic concept units, fusing them with class embeddings to construct rich, adaptive prompts.

Method

Overall Architecture

DICE is organized into two sequential stages: offline pretraining and online prompt tuning. During offline pretraining, DICE jointly trains a projection autoencoder (ProjectAutoencoder, which aligns and maps 3072-dimensional LLM raw class name embeddings into the CLIP text embedding space) and a Top-k Sparse Autoencoder (SAE, which structures dense CLIP representations into sparse, disentangled monosemantic concept bases). During prompt tuning, DICE employs a dual-branch architecture consisting of a Class Context Module and an Instance Context Module, culminating in a lightweight meta-network (MetaNet) that projects the fused context into an intermediate layer of the text encoder.

The end-to-end processing pipeline and component interactions are illustrated below:

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Data<br/>Class Name + Image"] --> B["Stage 1: Cross-Modal Autoencoder Pretraining<br/>Projection Mapping and Sparse Concept Disentanglement"]
    B --> C["Class Context Module<br/>Projected LLM Class Embedding"]
    B --> D["Instance Context Module<br/>SAE Co-Activation and Concept Selection"]
    C --> E["Unified Context Injection<br/>Residual Fusion and MetaNet Projection"]
    D --> E
    E --> F["Downstream Classification and Task Adaptation<br/>Knowledge-Guided Consistency Optimization"]

Key Designs

1. Cross-Modal Autoencoder Pretraining: Aligning and Constructing a Disentangled Concept Dictionary To circumvent the noise and verbosity of LLM-generated descriptions, DICE directly extracts embeddings from raw class names using OpenAI's text-embedding-3-large model (\(l \in \mathbb{R}^{d_{\text{llm}}}\)). The projection autoencoder consists of a linear encoder \(E_{\text{proj}}\) and decoder \(D_{\text{proj}}\), trained via a cosine projection loss \(\mathcal{L}_{\text{proj}} = 1 - \cos(t_{\text{llm}}, t_{\text{clip}})\) to map LLM features into the CLIP space, alongside an LLM reconstruction loss \(\mathcal{L}_{\text{llm\_rec}} = 1 - \cos(l, \hat{l})\) to prevent semantic distortion. Concurrently, a Top-k Sparse Autoencoder (SAE) applies centering bias \(b_{\text{pre}}\) to CLIP embeddings to produce sparse activations; each row vector of its decoder weight matrix \(D_{\text{sae}} \in \mathbb{R}^{d_{\text{hid}} \times d}\) naturally acts as an independent, monosemantic concept direction, forming a learnable concept dictionary expanded to dimension \(d_{\text{hid}} = 32 \times d\).

2. Instance Context Module: Class-Guided Co-Activation Scoring Mechanism To extract the precise visual concepts relevant to an input image from the vast concept dictionary, DICE introduces a dual-activation alignment mechanism. Given projected class embedding \(t_{\text{llm}}\) and visual image embedding \(f\), both are centered using learned bias \(b_{\text{pre}}\) and processed by the shared SAE encoder, yielding sparse activation vectors \(s_{\text{class}}\) and \(s_{\text{inst}}\). Their element-wise addition forms the co-activation score: $\(s_{\text{coact}} = s_{\text{class}} \oplus s_{\text{inst}}\)$ This scoring highlights latent dimensions simultaneously reinforced by both visual evidence and textual class identity, effectively filtering out spurious visual noise. The top-\(k\) concept indices from \(s_{\text{coact}}\) (optimally \(k=4\)) are selected to retrieve the corresponding concept basis vectors from \(D_{\text{sae}}\), assembling the instance representation \(W_{\text{inst}} \in \mathbb{R}^{C \times k \times d}\).

3. Unified Context Injection: Multi-Level Prompt Injection via Meta-Network To ensure instance-specific details enrich rather than disrupt the global geometric topology of class representations, DICE performs element-wise addition between broadcasted class embeddings \(W_{\text{class}}\) and instance concepts \(W_{\text{inst}}\). The combined tensor is transformed by a lightweight meta-network \(\mathcal{M}\) (two linear layers with an intermediate QuickGELU activation) to yield the unified context \(uc = \mathcal{M}(W_{\text{class}} \oplus W_{\text{inst}})\). This vector is injected into Layer 8 of the CLIP text encoder to replace standard learnable prefix tokens, enabling deep multi-head attention blocks to dynamically balance global class boundaries with localized visual traits.

4. Knowledge-Guided Consistency Optimization: Anti-Forgetting Regularization During few-shot adaptation, to prevent the prompt parameters from drifting away from CLIP's pre-trained semantic geometry, DICE incorporates a knowledge-guided consistency loss (\(\mathcal{L}_{\text{kg}}\)). This regularizer penalizes the Euclidean distance between task-adapted prompt embeddings \(t_{\text{dice}}\) and frozen zero-shot CLIP text embeddings \(t_{\text{clip}}\), combined in a composite multi-task objective: $\(\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{ce}} + \lambda_{\text{kg}} \cdot \mathcal{L}_{\text{kg}} + \lambda_{\text{llm}} \cdot \mathcal{L}_{\text{llm\_rec}}\)$ This formulation retains the broad generalization capabilities of pre-trained vision-language foundation models while facilitating robust specialization to few-shot target distributions.

Key Experimental Results

Main Results

On 11 standard few-shot benchmarks spanning general object recognition, fine-grained classification, textures, satellite imagery, and action recognition under the 16-shot base-to-novel generalization benchmark, DICE establishes a new state-of-the-art among all single-textual-prompt (tp) methods and delivers substantial gains when plugged into existing baselines.

Method Type Prompt Architecture Base Avg Novel Avg Harmonic Mean H
CoOp tp Static Textual Prompt 82.69% 63.22% 71.66%
CoCoOp tp Conditional Textual Prompt 80.47% 71.69% 75.83%
KART tp Knowledge-Aware Textual Prompt 78.42% 70.52% 74.26%
TCP tp Class-Aware Textual Prompt 84.13% 75.36% 79.51%
MaPLe mp Multi-modal Deep Prompt 81.54% 72.30% 76.65%
PSRC mp Self-Regulating Multi-modal Prompt 84.26% 76.10% 79.97%
HPT mp Hierarchical Textual Prompt 84.32% 76.86% 80.23%
CoPrompt mp Consistency Multi-modal Prompt 84.00% 77.23% 80.48%
DICE (Ours) tp Disentangled Concept Textual Prompt 84.45% 76.77% 80.43%
CoOp w/ DICE tp Plug-and-play Enhancement 83.09% 72.54% 77.45% (+5.79%)
CoCoOp w/ DICE tp Plug-and-play Enhancement 82.70% 73.19% 77.65% (+1.82%)
PSRC w/ DICE mp Plug-and-play Enhancement 85.76% 77.29% 81.31% (+1.34%)

Ablation Study

A comprehensive component-wise ablation across 11 benchmark datasets validates the design choices behind the dual-branch setup, the SAE concept dictionary, and the co-activation scoring mechanism.

Config / Variant Evaluation Note Base (%) Novel (%) Harmonic Mean H (%)
DICE (Full Model) Dual-Branch + SAE Disentanglement 84.45 76.77 80.43
Class Context Module Only Without instance context branch 84.28 75.78 79.80 (-0.63)
Instance Context Module Only Without class prior branch 82.52 74.34 78.22 (-2.21)
Replace SAE with Equivalent MLP Loss of sparse monosemantic property 81.31 73.82 77.38 (-3.05)
Use CLIP Text Embeddings for LLM Ablating rich LLM semantic priors 83.75 75.77 79.56 (-0.87)
Concept Selection via Image \(s_{\text{inst}}\) Only Lacks class-guided co-activation 82.10 74.60 78.16 (-2.27)
Selected Concepts \(k=2\) Insufficient concept coverage 84.48 75.84 79.93 (-0.50)
Selected Concepts \(k=8\) Redundant concepts introduce noise 84.57 76.12 80.12 (-0.31)

Key Findings

  • Crucial Role of SAE Disentanglement: Replacing the SAE module with an MLP of equivalent parameter capacity causes the harmonic mean H to drop by 3.05%, accompanied by severe degradation on both Base and Novel classes. This confirms that performance gains stem from sparse concept disentanglement and explicit selection rather than increased parameter capacity.
  • Superior Gains on Fine-Grained Benchmarks: DICE delivers the most pronounced improvements on fine-grained datasets with subtle inter-class differences, such as FGVC-Aircraft (CoOp H of 28.75% rises to 41.56% with DICE, and 43.57% with PSRC w/ DICE) and StanfordCars.
  • Co-Activation Restructures Feature Space: Analyzing activation similarity shows that \(s_{\text{coact}}\) increases intra-class sample similarity (e.g., Aircraft rises from 0.07 to 0.14) while dramatically reducing inter-class similarity (from 0.62 to 0.18), establishing well-separated decision boundaries.

Highlights & Insights

  • First Learnable SAE Concept Dictionary: Rather than using SAE purely as an ex-post interpretability tool, DICE pioneers its use as an active concept dictionary in multimodal prompt learning, enabling monosemantic visual-textual concepts to emerge without manual annotation.
  • Raw Class Embeddings Over Hallucinated Prompts: By encoding raw class names directly via OpenAI's text embedding models and using a projection autoencoder, DICE avoids the hallucinations, semantic mismatch, and prompt engineering overhead of free-form LLM descriptions.
  • Seamless Plug-and-Play Integration with White-Box Interpretability: DICE effortlessly enhances existing prompting frameworks (CoOp, CoCoOp, PSRC) while offering transparent interpretability: practitioners can trace active concepts back to explicit lexical terms (e.g., "freckle", "truck bed", "dashboard").

Limitations & Future Work

  • Inference Computational Overhead: Because prompts are conditioned on instance-level features, DICE requires an independent forward pass through the CLIP text encoder for each input image. Like CoCoOp, this increases latency and memory footprint during high-throughput inference.
  • Cross-Modal Concept Transfer Gap: The framework assumes that an SAE trained on text embeddings directly transfers to image embeddings due to CLIP's shared space. While broadly effective, extreme domain shifts (such as satellite imagery in EuroSAT) can reveal domain-specific alignment gaps.
  • Future Directions: Developing hypernetworks or caching mechanisms to reduce per-image text forward passes, and training joint cross-modal sparse autoencoders tailored to specialized visual domains.
  • vs CoOp / CoCoOp: CoOp suffers from severe overfitting on base classes; CoCoOp introduces image conditioning via an unconstrained meta-network without semantic structure. DICE introduces structured SAE concept units and co-activation scoring, achieving superior generalizability.
  • vs KART / HPT / CoPrompt: These methods rely on verbose, static LLM descriptions prone to hallucination and visual mismatch. DICE leverages raw class embeddings and dynamic SAE concept decomposition for robust grounding.
  • vs TCP: TCP injects class embeddings into intermediate text encoder layers to enhance class separability but omits instance-level cues. DICE integrates both global class priors and disentangled instance concepts within the intermediate layer.

Rating

  • Novelty: โญโญโญโญโญ Pioneering repurposing of Sparse Autoencoders as an unsupervised concept dictionary for vision-language prompt tuning.
  • Experimental Thoroughness: โญโญโญโญโญ Thorough evaluation across 11 benchmarks, rigorous ablations, parameter sensitivity tests, and mechanistic concept visualizations.
  • Writing Quality: โญโญโญโญโญ Exceptionally clear narrative, disciplined mathematical formulation, and complete implementation pseudocode.
  • Value: โญโญโญโญโญ Unlocks a compelling and interpretable direction at the intersection of mechanistic interpretability and parameter-efficient multimodal learning.