Unlocking Few-Shot Capabilities in LVLMs via Prompt Conditioning and Head Selection¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/AdhemarDeSenneville/HEC
Area: Multimodal VLM
Keywords: Large Vision-Language Models, Few-Shot Image Classification, Zero-Shot Classification, Prompt Conditioning, Attention Head Selection
TL;DR¶
Proposes Head Ensemble Classifiers (HEC), a training-free framework that conditions LVLM multimodal feature distributions using tiered text prompts and identifies a sparse set of discriminative attention heads via Gaussian Discriminant Analysis and soft accuracy ranking, bridging the longstanding classification performance gap between LVLMs and CLIP.
Background & Motivation¶
Current Large Vision-Language Models (LVLMs) demonstrate exceptional versatility and general visual reasoning capabilities across a broad spectrum of zero-shot multimodal tasks, including image captioning, complex visual question answering, document transcription, and open-world grounding. Operating with unified frozen weights, these models process vision and text seamlessly. However, when evaluated on classical, highly structured image classification benchmarks in zero-shot or few-shot regimes, these foundation models perform surprisingly poorly, consistently lagging behind contrastive dual-encoder architectures like CLIP. This performance deficit is counter-intuitive and paradoxical, given that contemporary LVLMs directly inherit their vision encoders from high-capacity pretrained CLIP or SigLIP backbones; in standard auto-regressive evaluations, the LVLMโs generative predictions underperform the raw classification capacity of its underlying visual trunk.
This fundamental tension stems from divergent architectural paradigms and decision-making mechanisms. Standard CLIP architectures feature completely isolated visual and textual encoders, restricting cross-modal interaction to a final late-stage dot-product similarity computation; this hardwires a bias toward literal class-name matching while preventing visual representations from dynamically adapting to rich linguistic context. Conversely, LVLMs fuse visual patch tokens and text prompt tokens in a shared transformer decoder, enabling deep multi-layer cross-attention; nevertheless, because standard auto-regressive next-token prediction objectives are not explicitly aligned with geometric classification boundaries, the discriminative signal embedded in the final summary token is diluted. Prior attempts to reconcile this discrepancy have predominantly relied on generating verbose text captions with LVLMs to feed external CLIP classifiers or executing parameter-heavy, task-specific fine-tuning.
This work re-examines the internal representation hierarchy of LVLMs to reveal that their low generative accuracy masks highly discriminative latent geometries. By steering internal feature distributions via multi-level prompt conditioning and using Gaussian Discriminant Analysis to isolate the small subset of attention heads that exhibit optimal class separability, one can construct state-of-the-art classifiers entirely training-free. Core idea: condition LVLM internal feature distributions via task- and domain-specific text prompts at inference, and select the top discriminative vision and text attention heads via Gaussian Discriminant Analysis with soft-accuracy scoring to construct training-free Head Ensemble Classifiers (HEC).
Method¶
Overall Architecture¶
HEC converts any off-the-shelf LVLM into an accurate few-shot and zero-shot classifier without adjusting any parameters. The framework takes support set images, optional textual class names, and query images as input. Visual tokens from an image encoder are concatenated with hierarchical text prompt tokens and ingested by the shared LLM decoder. The framework extracts attention vectors from the final summary token across all layers and attention heads. Next, using Gaussian Discriminant Analysis (GDA) coupled with a temperature-scaled soft accuracy metric on the support set, HEC scores and identifies the optimal vision-head subset \(\mathcal{H}^V\) and text-head subset \(\mathcal{H}^T\). Finally, the posterior class probabilities derived from these selected heads are averaged and convexly combined into robust ensembled predictions for the query images.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Data<br/>Support Images / Class Names / Query Images"] --> B["Multi-Level Prompt Conditioning<br/>Hierarchical Task/Domain/Class Guidance"]
B --> C["Attention Head Feature Extraction<br/>Summary Token Layer-wise Attention Vectors"]
C --> D["Head Selection via GDA & Soft-Accuracy<br/>Unbiased Covariance Shrinkage & Metric Ranking"]
D --> E["Bimodal Head Ensemble Classification<br/>Averaging Vision & Text Head Probabilities"]
E --> F["Output Predicted Class Distribution"]
Key Designs¶
1. Multi-Level Prompt Conditioning: Guiding Decoder Latent Geometries Traditional dual-encoder models evaluate visual features independently of upstream context, making fine-grained domain adaptation challenging without parameter tuning. To overcome this, HEC reformulates the input structure by directly concatenating the visual token sequence with a structured textual prompt \(\pi\) before passing it through the LLM decoder: \([\mathbf{f}^v(x_i); \pi]\). Conditioning is structured across three progressive tiers: first, Task Conditioning uses structural constraints such as "What is the object in the image? Answer in one word." to force the summary token to compress global discriminative information; second, Domain Conditioning explicitly specifies fine-grained visual categories (e.g., "What breed is that dog?"), steering the attention heads to focus on domain-specific diagnostic cues rather than spurious background textures; third, Class Conditioning appends known candidate class names directly to the prompt (e.g., "Between: boxer, yorkshire, beagle or havanese."). This joint cross-modal processing drives internal feature distributions toward linearly separable, domain-aligned manifolds during standard forward inference.
2. Attention Head Feature Extraction: Mining Specialized Token Subspaces Because standard output representations suffer from objective misalignment, taking the final summary token directly yields suboptimal classification accuracy. HEC instead probes the internal multi-head attention structure across all \(L\) layers and \(H\) heads per layer (\(M = L \times H\) heads total). For the final summary token in the sequence, let \(\mathbf{q}_m\) represent its query vector in head \(m\), and let \(\mathbf{K}_m, \mathbf{V}_m\) be the corresponding key and value matrices of dimension \(D\). The attention vector \(\mathbf{h}_m\) is extracted as: $$ \mathbf{h}_m = \operatorname{softmax}\left(\frac{\mathbf{q}_m \mathbf{K}_m^\top}{\sqrt{D}}\right) \mathbf{V}_m $$ Each attention vector is \(L_2\)-normalized such that inner products correspond directly to cosine similarities. Empirical analysis demonstrates that while the vast majority of heads track background noise or syntactic dependencies, a sparse set of specialized heads possesses exceptional discriminative power, outperforming the model's raw generative output by more than 10% on fine-grained visual benchmarks.
3. Head Selection via GDA & Soft-Accuracy: Mitigating Low-Shot Saturation & Overfitting In low-shot regimes (e.g., 4-shot), fitting linear probes or evaluating raw hard accuracy causes severe overfitting and performance saturation: on small support sets, many heads trivially score 100% training accuracy, destroying the ability to rank and isolate truly generalizable representations. HEC addresses this through Gaussian Discriminant Analysis (GDA). Assuming that the vision attention vectors for class \(c\) follow a class-conditional Gaussian distribution \(\mathcal{N}(\boldsymbol{\mu}_{m,c}, \boldsymbol{\Sigma}_m)\) with a shared covariance matrix \(\boldsymbol{\Sigma}_m\) across all classes, the shared covariance captures common anisotropic variations and avoids ill-conditioned sample covariance estimates. Using an empirical Bayes shrinkage estimator for precision matrix inversion \(\widehat{\boldsymbol{\Sigma}}_m^{-1}\), the class logit \(\ell_{i,m,c}\) is computed. To resolve the saturation problem, a temperature-scaled Soft-Accuracy Ranking (SAR) metric is defined: $$ s_m^{(\text{v})} = \frac{1}{KN} \sum_{i=1}^{KN} p_{i,m,y_i}^{(\text{v})}, \quad \text{where} \quad p_{i,m,c}^{(\text{v})} = \frac{\exp(\ell_{i,m,c}/\tau)}{\sum_{j=1}^N \exp(\ell_{i,m,j}/\tau)} $$ By setting a finite temperature \(\tau\), the metric measures the average assigned probability to the ground-truth class, preserving continuous confidence differentials to select the top \(k\) vision heads \(\mathcal{H}^V\). For the zero-shot text heads \(\mathcal{H}^T\), the same soft accuracy ranking is calculated via cross-modal dot-product logits over a shared task bank (such as ImageNet support tasks) to establish a universal, domain-transferable head selection offline.
4. Bimodal Head Ensemble Classification: Training-Free Robust Probability Fusion Once the top head sets \(\mathcal{H}^V\) and \(\mathcal{H}^T\) are determined, HEC avoids complex parameterized heads or metric-learning projections, relying instead on a non-parametric probabilistic ensemble. The few-shot visual classifier HEC-V averages predicted class probabilities across all heads in \(\mathcal{H}^V\), while the text zero-shot classifier HEC-T averages probabilities across \(\mathcal{H}^T\). When both support samples and textual class candidates are provided, the hybrid vision-text classifier HEC-VT combines their outputs through a weighted convex combination: $$ \bar{p}^{(\mathrm{HEC\text{-}VT})}{q,c} = \frac{\alpha \bar{p}^{(\mathrm{HEC\text{-}V})} $$ Averaging posterior class probabilities across sparse, diverse heads provides inherent variance reduction and resilience against outliers without requiring additional optimization during test time.} + \bar{p}^{(\mathrm{HEC\text{-}T})}_{q,c}}{\alpha + 1
Key Experimental Results¶
Main Results¶
The framework was evaluated across 12 diverse image classification benchmarks spanning fine-grained object categories, textures, human actions, and natural scenes (PETS, ESAT, UCF, SUN, CAL, DTD, AIR, FOOD, FLWR, CARS, BIRD, SIGN).
The table below reports 4-shot Vision-Few-Shot classification accuracy (%) without knowing class names:
| Model Backbone | Classification Method | PETS | ESAT | UCF | SUN | CAL | DTD | AIR | FOOD | FLWR | CARS | BIRD | SIGN | AVG |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DINOv1 | Probing | 81.9 | 81.6 | 71.8 | 50.8 | 86.9 | 49.6 | 25.5 | 37.9 | 88.1 | 28.3 | 54.2 | 46.3 | 58.6 |
| DINOv2 | Probing | 78.5 | 73.6 | 67.6 | 63.4 | 87.0 | 51.6 | 29.9 | 43.0 | 97.0 | 35.9 | 65.5 | 35.6 | 60.7 |
| DINOv3 | Probing | 88.0 | 77.2 | 82.8 | 72.7 | 95.4 | 63.6 | 56.9 | 74.3 | 99.5 | 79.4 | 77.3 | 53.1 | 76.7 |
| OpenAI CLIP | Probing | 72.9 | 73.6 | 81.5 | 73.2 | 90.2 | 55.8 | 28.2 | 73.4 | 89.5 | 56.9 | 53.6 | 56.5 | 67.1 |
| DFN | Probing | 84.5 | 80.8 | 79.8 | 73.9 | 94.4 | 62.8 | 40.0 | 77.5 | 96.8 | 85.6 | 66.9 | 68.2 | 75.9 |
| Qwen2-VL (7B) | Probing (TC) | 84.9 | 75.6 | 81.4 | 77.2 | 93.4 | 58.2 | 40.9 | 81.1 | 98.1 | 79.8 | 65.1 | 71.3 | 75.6 |
| Qwen2-VL (7B) | Probing (DC) | 92.0 | 74.0 | 82.3 | 79.8 | 94.2 | 66.9 | 60.7 | 82.7 | 98.2 | 89.2 | 69.5 | 69.2 | 79.9 |
| Qwen2-VL (7B) | SAVsโ (DC) | 91.0 | 72.0 | 80.5 | 81.7 | 94.4 | 70.5 | 59.3 | 84.9 | 97.8 | 89.5 | 69.8 | 68.1 | 80.0 |
| Qwen2-VL (7B) | HEC-Vโ (DC) | 92.2 | 78.8 | 85.0 | 82.4 | 95.5 | 71.8 | 62.2 | 85.3 | 98.5 | 89.8 | 72.0 | 75.5 | 82.4 |
In the full Vision-Text-Few-Shot setting (4-shot), where both class names and labeled support images are accessible, performance comparisons against state-of-the-art methods are summarized below:
| Backbone | Method | PETS | ESAT | UCF | SUN | CAL | DTD | AIR | FOOD | FLWR | CARS | BIRD | SIGN | AVG |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DFN | Zero-Shotโ | 92.0 | 51.6 | 63.4 | 79.5 | 95.6 | 51.1 | 29.6 | 87.2 | 82.0 | 92.1 | 78.0 | 28.1 | 69.2 |
| DFN | TipAdapter | 92.3 | 69.2 | 77.4 | 80.5 | 95.8 | 64.4 | 40.2 | 87.2 | 97.0 | 92.7 | 78.7 | 50.0 | 77.1 |
| DFN | GDA | 92.8 | 78.0 | 84.0 | 83.2 | 96.5 | 70.0 | 46.3 | 87.2 | 98.5 | 93.6 | 80.3 | 67.8 | 81.5 |
| DFN | ProKeR | 91.2 | 82.4 | 83.9 | 82.2 | 97.0 | 66.2 | 43.1 | 87.8 | 98.3 | 93.6 | 81.5 | 64.6 | 81.0 |
| Qwen2-VL | Summary Token | 55.0 | 48.8 | 30.8 | 65.5 | 73.3 | 32.0 | 30.1 | 72.4 | 8.2 | 45.5 | 6.1 | 30.9 | 41.5 |
| Qwen2-VL | Probing | 92.0 | 74.0 | 82.3 | 79.8 | 94.2 | 66.9 | 60.7 | 82.7 | 98.2 | 89.2 | 69.5 | 69.2 | 79.9 |
| Qwen2-VL | SAVsโ | 91.0 | 72.0 | 80.5 | 81.7 | 94.4 | 70.5 | 59.3 | 84.9 | 97.8 | 89.5 | 69.8 | 68.1 | 80.0 |
| Qwen2-VL | HEC-Vโ | 92.2 | 78.8 | 85.0 | 82.4 | 95.5 | 71.8 | 62.2 | 85.3 | 98.5 | 89.8 | 72.0 | 75.5 | 82.4 |
| Qwen2-VL | HEC-VT | 92.8 | 82.0 | 85.0 | 83.3 | 95.6 | 72.7 | 62.3 | 85.7 | 98.6 | 90.1 | 72.1 | 76.2 | 83.0 |
Ablation Study¶
To isolate the individual contributions transitioning from the previous state of the art (SAVs) to HEC-VT, the authors systematically ablated each component across 10-way 4-shot and 16-shot tasks:
| Configuration | Domain Cond. (DC) | Soft-Accuracy (SAR) | GDA Model (GDA) | Probability Ens. (PE) | Text-Head (HEC-T) | 4-shot Acc. (%) | 16-shot Acc. (%) | Note |
|---|---|---|---|---|---|---|---|---|
| In-Context Baseline | - | - | - | - | - | 84.98 | N/A | Support examples placed in prompt; exceeds context at 16-shot |
| SAVs Baseline | 87.03 | 94.45 | Nearest-centroid hard selection | |||||
| + DC | โ | 88.13 | 95.23 | Domain prompt guidance (+1.10% / +0.78%) | ||||
| + SAR | โ | โ | 93.28 | 94.99 | Soft accuracy resolves saturation (+5.15% at 4-shot) | |||
| + GDA | โ | โ | 83.91 | 95.06 | GDA without SAR severely overfits in low-data regimes | |||
| Full Ranking | โ | โ | โ | 93.99 | 96.03 | GDA precision matrix coupled with SAR ranking | ||
| Pure Ensemble | โ | โ | โ | 93.28 | 95.12 | Uniform probability ensembling | ||
| HEC-V w/o DC | โ | โ | โ | 93.23 | 95.70 | Head selection without domain prompts | ||
| HEC-V | โ | โ | โ | โ | 94.09 | 96.17 | Complete vision-head classifier | |
| HEC-VT (Full) | โ | โ | โ | โ | โ | 94.42 | 96.23 | Combining vision and text head distributions |
Key Findings¶
- Soft-Accuracy Ranking (SAR) is crucial for low-shot regimes: In 4-shot tasks, transitioning from SAVs' hard matching to SAR delivers a massive +5.15% accuracy gain. Hard accuracy frequently hits 100% on tiny sets, blinding head selectors to generalization capability, whereas continuous soft probabilities accurately preserve confidence gradients.
- GDA Covariance Estimation scales with data density: In the 16-shot setup, GDA's shared covariance estimation models the anisotropic feature manifold accurately, raising classification accuracy to 96.03% and complementing SAR effectively.
- Hierarchical prompt conditioning monotonically reduces error: Stepping from unconditioned inputs to Task, Domain, and Class conditioning reduces error rates (Error Reduction) in HEC-V by 10.31%, 15.85%, and 18.84% respectively, confirming that textual guidance shapes latent visual separability.
- Text-heads exhibit universal cross-task transferability: Evaluating the 20 text-heads selected once on ImageNet across downstream datasets with less than 13% class overlap produced consistent accuracy gains across all 12 benchmarks, confirming the existence of persistent multimodal semantic alignment circuits inside LVLM decoders.
Highlights & Insights¶
- Rebuts the necessity of costly fine-tuning for LVLM classification: Proves that poor out-of-the-box accuracy in LVLMs is an artifact of unaligned generative decoding heads rather than deficient visual representations, unlocking competitive few-shot capabilities training-free.
- Unifies prompt engineering with statistical discriminant analysis: Demonstrates that natural language prompts can serve as geometric manifold adjusters, shifting latent features into domain-discriminative subspaces prior to downstream classification.
- Excels on fine-grained and non-object-centric distributions: Delivers notable performance margins on specialized domains like aircraft identification (AIR) and road signs (SIGN), outperforming classic OpenAI CLIP and DFN baselines by leveraging deep contextual knowledge.
Limitations & Future Work¶
- Requires white-box model access: Computing attention vectors \(\mathbf{h}_m\) requires intermediate layer activations and Key/Value states, limiting applicability to open-weight models and precluding closed-source commercial APIs.
- Class conditioning scalability is constrained: When the candidate class count \(N\) is very large (e.g., ImageNet-1k), packing all category names into the textual prompt exceeds the LLM context budget and causes attention diffusion.
- Inference latency exceeds dual-tower encoders: Running the full 7B LLM decoder forward pass incurs significantly higher compute latency and GPU memory overhead than forwarding a lightweight ViT, suggesting a need for offline head pruning and distillation in future work.
Related Work & Insights¶
- vs SAVs (Mitra et al., ICCV 2025): While SAVs pioneered internal head extraction for LVLMs using nearest centroid classifiers, its hard metric selection saturates in low-shot regimes and lacks prompt-conditioned guidance; HEC introduces GDA covariance modeling and temperature-scaled soft accuracy, beating SAVs by over 2.4% in 4-shot accuracy.
- vs GDA / TipAdapter (Wang et al., ICLR 2024 / Zhang et al., ECCV 2022): Conventional training-free CLIP adapters operate on static dual-encoder embeddings without generative joint modeling; HEC extends Gaussian discriminant principles to generative multimodal decoders, outperforming dual-tower adaptation limits across diverse domains.
Rating¶
- Novelty: โญโญโญโญโญ Formulates an original, elegant synthesis of prompt conditioning, Gaussian discriminant analysis, and attention head selection to activate dormant LVLM classification capabilities.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluations across 12 benchmarks, 2 leading LVLMs (Qwen2-VL and LLaVA-OV), and meticulous multi-shot component ablations.
- Writing Quality: โญโญโญโญโญ Structured and rigorous exposition with intuitive visual diagrams, crisp mathematical formulations, and thorough analyses.
- Value: โญโญโญโญโญ Establishes a compelling parameter-free baseline for few-shot multimodal learning, prompting deeper investigation into the latent interpretability of transformer decoders.