Decompose, Compare, and Decide: Multimodal LLMs are Implicit Few-Shot Learners¶
Conference: ECCV 2026
Paper: ECCV Official Page
Code: https://github.com/yunhanwang1105/DeCoDe
Area: Multimodal VLM
Keywords: few-shot learning, multimodal large language models, in-context learning, pairwise binary decomposition, out-of-domain generalization
TL;DR¶
Proposes DeCoDe, a training-free inference framework that decomposes few-shot classification into independent pairwise support-query binary comparisons and ranks candidate classes via the unnormalized "Yes" token logit, resolving the vulnerability of off-the-shelf MLLMs to context distraction and label semantic priors.
Background & Motivation¶
Multimodal Large Language Models (MLLMs, e.g., Qwen2.5-VL, Qwen3-VL, InternVL3) have achieved notable breakthroughs in open-vocabulary recognition and complex visual reasoning. However, transferring these powerful capabilities to few-shot image classification without fine-tuning remains surprisingly difficult. Under the standard multimodal in-context learning protocol, all support demonstrations and the query image are concatenated into a single long-context prompt. Counter-intuitively, off-the-shelf MLLMs frequently perform worse in 1-shot in-context setups than in their zero-shot baseline (e.g., InternVL3 drops sharply from 90.7% zero-shot accuracy to 72.2% 1-shot accuracy across standard benchmarks). Multi-image visual token attention and position biases introduce severe cross-instance distraction.
The core tension behind this breakdown is that standard multimodal in-context prompting relies overwhelmingly on semantic language priors induced by textual class names rather than true visual-to-visual correspondence. When semantic labels are removed and replaced with abstract numerical identifiers (e.g., Class 1, Class 2), few-shot accuracy across mainstream MLLMs collapses toward random guessing (around 20% to 28% for 5-way classification). Existing remedies either require extensive supervised fine-tuning (SFT) or extract internal representations to train separate classifiers (like KNNs), sacrificing the flexibility and native generative reasoning of large multimodal models.
This paper tackles the challenge by discarding the brittle practice of stacking multiple demonstrations into one prompt and restructuring few-shot classification into an intuitive pairwise verification task. Core idea: decompose multi-class classification into independent pairwise binary comparisons (Decompose), score visual similarity via the model's affirmative "Yes" token output logit (Compare), decide the predicted class by aggregating candidate support scores (Decide), and incorporate high-level domain concepts to achieve state-of-the-art few-shot performance purely at inference time.
Method¶
Overall Architecture¶
DeCoDe (Decompose, Compare, and Decide) replaces single-pass multi-image prompting with an entirely training-free, pairwise inference framework. Given an \(N\)-way \(K\)-shot task, DeCoDe decomposes the classification problem into \(N \times K\) isolated binary prompts. Each prompt presents only one candidate support image alongside the query image, querying whether the two images depict the same class or concept. The unnormalized logit of the "Yes" token is retrieved as the pairwise match score, which is then averaged across all \(K\) shots per class to determine the final predicted label via argmax.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Few-shot input<br/>Query image x_q and support set S"] --> B["Task decomposition & domain prompting<br/>Split into NรK independent dual-image prompts"]
B --> C["Pairwise binary comparison<br/>Feed (x^s, x^q) pairs and extract Yes logit"]
C --> D["Class decision & score aggregation<br/>Average over K shots per class and argmax"]
D --> E["Output predicted class y_hat"]
Key Designs¶
1. Task decomposition and pairwise binary prompting: eliminating multi-image interference and label bias Standard in-context prompting concatenates \(N \times K\) support images with the query image, causing visual tokens to compete and cross-contaminate in self-attention layers while suffering from position bias. DeCoDe breaks the support set \(\mathcal{S} = \{(x_{n,k}^s, c_n)\}\) down into an ensemble of pairwise prompts \(\mathcal{P} = \{p_{n,k} \mid n=1,\dots,N, k=1,\dots,K\}\), where each prompt evaluates exactly two visual instances: $\(p_{n,k} = x_{n,k}^s \oplus x^q \oplus \text{Prompt}(c_n)\)$ In the anonymized regime, the query becomes: "Are the two images depicting the same class? Answer Yes or No." This structural formulation restricts attention exclusively to the query and a single reference image, eliminating spurious label correlations and forcing the model to perform direct visual correspondence even without semantic class cues.
2. Continuous similarity scoring via affirmative token logits: robust unnormalized confidence extraction Relying on greedily decoded text tokens ("Yes" or "No") yields discrete, uncalibrated binary outputs that cannot resolve ties among multiple candidate classes; asking the model to verbalize an explicit numeric confidence score similarly leads to heavy verbalization distortion. Drawing inspiration from generative alignment scoring, DeCoDe extracts the unnormalized logit assigned to the "Yes" token at the first generation step as a continuous, fine-grained visual similarity metric: $\(s_{n,k}^{\mathrm{pair}} = \text{logit}(\text{Yes} \mid p_{n,k})\)$ For an \(N\)-way \(K\)-shot problem, the scores of all \(K\) support examples for class \(n\) are averaged to make the final prediction: $\(\hat{y} = \arg\max_{n \in \{1,\dots,N\}} \frac{1}{K} \sum_{k=1}^K s_{n,k}^{\mathrm{pair}}\)$ This continuous score exhibits monotonic consistency, and empirical ablations show that it matches or exceeds complex subtraction-based calibrations while greatly outperforming generated confidence ratings.
3. High-level domain information guidance: contextualizing target conceptual spaces In open-ended visual comparison, the generic word "class" is overly ambiguous, leading the model to inadvertently align backgrounds, viewpoints, or textures instead of the primary semantic category. DeCoDe introduces dataset-level domain information (\(D_{\mathrm{info}}\)), replacing the abstract term "class" with a dataset-appropriate high-level concept descriptor (e.g., "yoga pose" for Yoga, "Egyptian hieroglyph" for Hiero, "industrial product" for Industrial, and "insect species" for Insects). This high-level descriptor steers the vision-language attention toward task-relevant visual dimensions without revealing granular class-level ground truth, yielding an additional 3% to 8% accuracy boost on out-of-domain benchmarks.
A Worked Example¶
Consider a 5-way 1-shot episode on ancient Egyptian hieroglyph recognition. In standard in-context prompting, concatenating 5 hieroglyph exemplars with an unfamiliar query produces only 30.0% accuracy on Qwen3-VL (barely better than random guess) because the model lacks lexical priors for ancient symbols. Under DeCoDe: 1. Five independent binary prompts are constructed, each containing one reference hieroglyph \(x_n^s\) and the query \(x^q\); 2. The prompt asks: "Are the two images depicting the same Egyptian hieroglyph? Answer Yes or No."; 3. In a single batched forward pass, the model evaluates each pair and yields the following Yes logits: Class 1 = 12.4, Class 2 = 8.1, Class 3 = 24.8, Class 4 = 5.2, Class 5 = 9.7; 4. The decision module executes argmax, decisively picking Class 3 based on the highest affirmative logit.
Key Experimental Results¶
Main Results¶
The framework is evaluated across 12 datasets comprising 6 standard few-shot benchmarks (mini-ImageNet, UCF101, CUB, Aircraft, Dogs, DomainNet) and 6 newly curated, highly specialized out-of-domain (Novel OOD) benchmarks (Lego, Industrial, Yoga, Egyptian Hieroglyphs, Insects, Arabic Sign Language) in 5-way 1-shot settings, under both semantic and anonymized label protocols.
| Method / Setting | Standard Avg | Novel OOD Avg | Total Avg | Description / Configuration |
|---|---|---|---|---|
| CLIP (ViT-B/32, 0-shot) | 86.7 | 35.2 | 60.9 | Standard contrastive VLM representation |
| ProKeR (1-shot) | 89.3 | 53.5 | 71.4 | Training-free kernel regression over CLIP |
| SAVs (Qwen2.5-VL feats, 1-shot) | 92.9 | 60.3 | 76.6 | Linear/KNN classifier on MLLM representations |
| InternVL3 (In-context 1-shot) | 72.2 | 49.5 | 60.8 | Severe degradation due to multi-image prompt clutter |
| Qwen3-VL (In-context 1-shot) | 84.9 | 76.9 | 80.9 | Standard multimodal in-context baseline |
| Qwen3-VL (SFT 1-shot baseline) | 95.7 | 80.8 | 88.2 | Fine-tuned with LoRA on mini-ImageNet |
| Qwen3-VL (DeCoDe 1-shot, Ours) | 95.5 | 79.1 | 87.3 | Completely training-free pairwise inference |
| Qwen3-VL (DeCoDe+Dinfo 1-shot, Ours) | 96.1 | 85.2 | 90.6 | Decomposed inference with domain descriptor |
| Qwen3-VL (Anonymous In-context) | 28.9 | 26.6 | 27.8 | Standard in-context collapses when class names are masked |
| Qwen3-VL (Anonymous SFT baseline) | 56.3 | 54.6 | 55.5 | Supervised tuning remains impaired by label removal |
| Qwen3-VL (Anonymous DeCoDe+Dinfo, Ours) | 92.5 | 85.8 | 89.2 | Strong performance preserved without labels (+33.7 vs SFT) |
Ablation Study¶
The table below isolates the influence of binary decomposition, prompt calibration, and scoring formulations for 5-way 1-shot tasks across standard and novel representative subsets.
| Config / Variant | Standard Subset Avg (mini/CUB/Dogs) | Novel Subset Avg (Lego/Yoga/Hiero) | Total Accuracy (%) | Mechanism & Insight |
|---|---|---|---|---|
| In-context 1-shot baseline (Qwen3-VL) | 85.5 | 74.6 | 80.0 | Full concatenation with position & label bias |
| In-context + PMI calibration (Qwen3-VL) | 93.4 | 78.9 | 86.2 | Subtracts blank-query prior \(P(y\|S)\) |
| 0-shot pairwise decomposition (Qwen3-VL) | 96.0 | 57.6 | 76.8 | Text-only binary prompt without visual support |
| DeCoDe 1-shot (Qwen3-VL) | 97.3 | 81.2 | 89.2 | Support-query pairwise visual alignment |
| DeCoDe 1-shot + Dinfo (Qwen3-VL) | 97.8 | 83.8 | 90.8 | Direct high-level concept specification |
| DeCoDe (Generated Confidence Text) | 81.5 | 62.1 | 71.8 | Verbalized confidence tokens suffer verbalization noise |
| DeCoDe (Score(Yes) - Score(No)) | 95.3 | 74.8 | 85.1 | Subtracting negative logit matches direct Yes score |
| DeCoDe (Direct Yes Token Logit) | 95.3 | 75.0 | 85.1 | Elegant, uncalibrated raw logit scoring |
Key Findings¶
- Context concatenation introduces negative transfer: On standard datasets, off-the-shelf MLLMs perform worse under 1-shot in-context concatenation than under 0-shot evaluation (e.g., InternVL3 drops by 18.5%). This confirms that concatenating multiple support images disrupts internal zero-shot representations.
- Label anonymization exposes pseudo in-context learning: Anonymizing class names reduces standard in-context accuracy from 84.9% to 27.8% (near random). This proves that vanilla prompting leans on textual category knowledge rather than learning from image exemplars. In contrast, DeCoDe maintains 89.2% accuracy under anonymized labels, establishing authentic visual-to-visual reasoning.
- Superior scaling with Shot and Way count: When expanding to 5-shot episodes, standard in-context accuracy degrades from 73.9% to 50.2% due to extended context clutter. Conversely, DeCoDe improves from 80.6% to 85.8%. Similarly, in 20-way evaluations, DeCoDe exhibits graceful accuracy retention where vanilla prompting falls off sharply.
Highlights & Insights¶
- Reformulating multi-way classification as pairwise verification: Breaking an \(N\)-way decision into decoupled pairs leverages the model's native cross-attention for direct fine-grained image comparison while avoiding contextual crosstalk.
- The power of affirmative token logits over text generation: Bypassing surface autoregressive token generation in favor of the unnormalized "Yes" logit provides a continuous, monotonic, and low-variance similarity score.
- Guiding representation geometry with domain descriptors (\(D_{\mathrm{info}}\)): Injecting a high-level concept (e.g., "yoga pose" or "hieroglyph") without revealing individual class names focuses feature alignment on task-critical dimensions without leaking class identity.
Limitations & Future Work¶
- Linear computational scaling with candidate count: An \(N\)-way \(K\)-shot problem requires \(N \times K\) forward evaluations. Although each pass processes only two images and is embarrassingly parallel across batch or GPU workers, latency increases by roughly 1.7x in 20-way setups on a single accelerator.
- Dependency on token logit access: The framework requires access to the model's vocabulary logits at the first decoding step, making it incompatible with black-box commercial endpoints that restrict logprob extraction.
- Future directions: Integrating a coarse-to-fine cascade (e.g., lightweight CLIP retrieval to narrow candidate classes down to Top-\(M\) followed by DeCoDe pairwise verification) could dramatically enhance scalability for large-\(N\) tasks.
Related Work & Insights¶
- vs GFSL [25]: While GFSL identifies label memorization and introduces anonymization during supervised tuning, it remains tied to multi-image concatenation; DeCoDe surpasses GFSL and SFT baselines in an entirely training-free manner through structured pairwise inference.
- vs SAVs [27]: SAVs extracts intermediate MLLM layer representations to train external classifiers (e.g., KNNs); DeCoDe retains the end-to-end multimodal language generation pathway, avoiding feature caching or external classification heads.
- vs VQAScore [24]: While VQAScore uses affirmative token logits to measure text-to-image synthesis quality, DeCoDe demonstrates that logit scoring is equally powerful for cross-image visual few-shot comparison.
Rating¶
- Novelty: โญโญโญโญโ Elegantly reframes few-shot MLLM inference as pairwise binary verification, bypassing long-context attention degradation.
- Experimental Thoroughness: โญโญโญโญโญ Rigorously tested across 12 datasets (including 6 newly introduced OOD benchmarks), 3 leading MLLMs, SFT baselines, and masked-label controls.
- Writing Quality: โญโญโญโญโญ Clear exposition, lucid problem formulation, and sharp empirical insights into the pitfalls of multimodal in-context learning.
- Value: โญโญโญโญโญ Provides an effective, training-free paradigm for real-world few-shot adaptation while exposing critical evaluation blind spots in current VLM benchmarks.