Show Me Examples: Inferring Visual Concepts from Image Sets¶
Conference: ECCV 2026
Paper: ECCV 2026
Area: Multimodal VLM
Keywords: visual in-context learning, concept inference, image set reasoning, set learner, flow matching
TL;DR¶
Addressing the failure of modern VLMs to infer implicit concepts purely from visual examples without text, this paper introduces the Visual Concept Inference from Sets (VICIS) benchmark along with a framework combining a Set Learner, subspace projection, and rectified flow matching for nonverbal concept induction and diverse target synthesis.
Background & Motivation¶
Vision-language models (VLMs) have advanced rapidly in complex text instruction following, visual question answering, and multimodal generation, with much of their power stemming from in-context learning principles popularized by LLMs. However, a foundational pillar of human cognition operates nonverbally: when humans inspect a small collection of images sharing an implicit commonality, they naturally isolate the relevant visual dimension without verbal definitions or explicit annotations, subsequently projecting that concept onto novel instances. Current frontier generative models struggle fundamentally with this type of reasoning. When presented with implicit visual context sets, they frequently ignore the visual evidence, collapse into simply replicating the query image, or exhibit strong biases toward dominant object categories rather than following the visual prompt.
The underlying tension resides in the representation bottleneck of standard visual in-context formulations. Existing methods generally treat visual context either as global pixel-level concatenation for inpainting or as generic conditioning tokens, lacking any principled mechanism to decouple the inferred conceptual subspace from query-specific irrelevant attributes. When an ambiguous query presents multiple features, current models cannot selectively isolate the single attribute specified by the context set while allowing other degrees of freedom to vary freely. Furthermore, evaluating visual context understanding without a query-target grounding mechanism risks turning the benchmark into a trivial sample memorization or reconstruction task.
The authors address this challenge by introducing a paired query-target setup within the Visual Concept Inference from Sets (VICIS) benchmark, coupled with an end-to-end concept induction and flow matching diffusion architecture. Core idea: infer concept subspace direction vectors from an image context set via a dedicated Set Learner, project the query image embedding onto this subspace to strip away unaligned visual attributes, and condition a rectified flow diffusion model on the clean concept token to achieve accurate, nonverbal concept-guided generation.
Method¶
Overall Architecture¶
The VICIS task takes as input a small context set of example images \(X_{\text{set}}\) sharing an unlabeled concept and a query image \(x_{\text{query}}\), aiming to generate new target images \(x_{\text{target}}\) that preserve the query's specific instantiation of the inferred concept while varying all irrelevant attributes. The framework operates through three sequential stages: first, a frozen vision transformer backbone encodes the context images, which are fed into a Set Learner to infer orthogonal concept direction vectors; second, an Instantiation Module projects the query token onto this concept subspace, discarding background clutter and unrelated semantics; third, a rectified flow matching diffusion model ingests the projected concept token directly via its timestep embeddings to generate target images from noise.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input context set and query image<br/>X_set and x_query"] --> B["Set Learner<br/>infer concept direction basis across images"]
B --> C["Instantiation Module<br/>project query embedding to filter irrelevant visual details"]
C --> D["Flow Matching Diffusion Model<br/>condition timestep embeddings to generate target images"]
Key Designs¶
1. Set Learner: infer concept direction basis across images To overcome the inability of standard vision backbones to isolate shared concept dimensions from unannotated collections, the Set Learner is tasked with identifying concept-specific directions in feature space. Each image in \(X_{\text{set}}\) is first processed independently by a pretrained Vision Transformer encoder \(E\) (such as DINOv2) to yield token embeddings. The Set Learner, parameterized as a dedicated Vision Transformer, takes the concatenated tokens of the entire set as input, performing cross-image self-attention to discover latent semantic commonalities. It then outputs a compact set of unit direction vectors \(D_c = \{d_1, d_2, \dots, d_k\}\) that define the subspace of the shared concept, effectively spanning the space of all possible concept instantiations.
2. Instantiation Module: project query embedding to filter irrelevant visual details To ensure the generative model only inherits the specific concept instantiation designated by the context while remaining invariant to extraneous query properties (such as pose, color, or background), the system employs an orthogonal projection operation. Given \(x_{\text{query}}\), its [CLS] token embedding \(e_{\text{query}}\) is extracted via encoder \(E\). The module computes inner products between \(e_{\text{query}}\) and each concept basis vector \(d_i \in D_c\) to obtain coordinates \(s_i = \langle e_{\text{query}}, d_i \rangle\). The concept-conditioned token is constructed via linear combination: $\(e_{\text{query}}^{\text{proj}} = \sum_{i=1}^k s_i d_i\)$ This projection serves as a strict semantic bottleneck: any visual variation present in the query that falls outside the subspace spanned by \(D_c\) is mathematically eliminated, retaining strictly the concept-specific attributes.
3. Flow Matching Diffusion Model: condition timestep embeddings to generate target images The generative backbone \(DM_\theta\) adopts the rectified flow framework to synthesize target images. Rectified flow forms smooth, straight probability paths between standard Gaussian noise and data distributions, enabling stable vector field regression without large-batch contrastive objectives. The projected concept token \(e_{\text{query}}^{\text{proj}}\) is added directly to the timestep embedding of \(DM_\theta\), providing global conditioning throughout denoising. The entire system is trained end-to-end minimizing the flow matching objective: $\(\mathcal{L}_{\text{FM}}(\theta) = \mathbb{E}_{t \sim U[0,1], x_t, e_{\text{query}}^{\text{proj}}} \left[ \| DM_\theta(x_t, t, e_{\text{query}}^{\text{proj}}) - u(x_t, t) \|_2^2 \right]\)$ where \(u(x_t, t) = x_{\text{target}} - x_0\) represents the linear velocity field. This loss provides reliable gradient flow even under modest training batch sizes, yielding samples that accurately mirror the inferred concept while preserving high visual diversity.
Key Experimental Results¶
Main Results¶
The method is benchmarked on an ImageNet hierarchy constructed from WordNet subtrees, cross-modality ImageNet-Sketch data, and unseen ImageNet21k classes. Metrics include mean accuracy per concept (per con. %), mean accuracy per query instantiation (per inst. %), and a hierarchical relative entropy diversity score (Diversity).
| Method | Dataset / Setup | Per Con. Acc (%) | Per Inst. Acc (%) | Diversity Score (Div.) |
|---|---|---|---|---|
| Copy Query Baseline | Hierarchy Evaluation | โ | โ | 0.47 |
| ILLUME+ 3B | Hierarchy Evaluation | 37.15 | 55.00 | 0.57 |
| BAGEL 7B-MoT | Hierarchy Evaluation | 26.05 | 33.76 | 0.46 |
| Visual Prompting | Hierarchy Evaluation | 26.02 | 45.89 | 0.70 |
| Ours | Hierarchy Evaluation | 46.34 | 54.46 | 0.81 |
| Visual Prompting | Sketch Queries | 23.76 | 44.57 | 0.68 |
| Ours | Sketch Queries | 32.60 | 47.11 | 0.72 |
| Visual Prompting | Sketch Context | 27.67 | 47.04 | 0.71 |
| Ours | Sketch Context | 41.50 | 52.24 | 0.79 |
| Visual Prompting | Sketch C+Q | 25.69 | 44.45 | 0.68 |
| Ours | Sketch C+Q | 31.90 | 46.47 | 0.71 |
| Visual Prompting | 21k Context (Unseen Classes) | 29.07 | 49.95 | 0.76 |
| Ours | 21k Context (Unseen Classes) | 39.63 | 51.40 | 0.79 |
Ablation Study¶
The authors investigate the impact of context set size and noisy context corruption on the ImageNet hierarchy validation split:
| Config | Setup / Corruption | Per Con. Acc (%) | Per Inst. Acc (%) | Diversity Score (Div.) |
|---|---|---|---|---|
| Context Set Size Ablation | Set Size = 2 | 44.74 | 48.92 | 0.719 |
| Context Set Size Ablation | Set Size = 3 | 45.26 | 53.51 | 0.801 |
| Context Set Size Ablation (Default) | Set Size = 5 | 46.34 | 54.46 | 0.811 |
| Context Set Size Ablation | Set Size = 7 | 46.56 | 54.74 | 0.824 |
| Noise Sensitivity Ablation | Clean Context (5/5 clean) | 46.34 | 54.46 | 0.811 |
| Noise Sensitivity Ablation | Noisy Context (1/5 random) | 37.62 | 46.85 | 0.812 |
| Noise Sensitivity Ablation | Noisy Context (2/5 random) | 26.01 | 36.42 | 0.906 |
Key Findings¶
- Context Set Scaling: Increasing context examples from 2 to 5 significantly improves both per-concept accuracy and generation diversity, after which performance plateaus at 7 examples, verifying that small visual sets offer strong induction signals.
- Graceful Degradation under Noise: Corrupting 1 out of 5 context images reduces accuracy moderately to 37.62% without causing diversity collapse (0.812), demonstrating the resilience of cross-image attention against outlier instances.
- Comparison with Proprietary VLMs: In simplified dual-concept disambiguation benchmarks, proprietary VLMs struggle severelyโNano Banana achieves only 46% accuracy and Gemini 2.5 Flash Image reaches 40% (often defaulting to copying query subjects), whereas the proposed model achieves 93% accuracy.
Highlights & Insights¶
- Subspace Projection Bottleneck: Inferring an explicit concept direction basis and projecting query embeddings mathematically forces the removal of unshared visual cues, resolving the pervasive query-copying failure mode in visual context learning.
- Scalable Weakly-Supervised Training via WordNet: Exploiting lexical tree hierarchies enables constructing thousands of concept-inference episodes automatically, bypassing tedious manual attribute annotations.
- Zero-Shot Cross-Modality Generalization: Without sketch training data, the model effectively extrapolates to sketch contexts and queries, demonstrating that the learned subspace captures high-level semantic abstractions rather than low-level texture cues.
Limitations & Future Work¶
- Ontological Hierarchy Bias: The real-world evaluation relies primarily on taxonomic semantic categories (WordNet nouns); continuous non-taxonomic concepts like physical material friction or complex lighting dynamics remain less explored.
- Inherent Ambiguity in Multi-Concept Sets: When context images simultaneously share multiple unannotated attributes (e.g., all being black cats at night), the model lacks an interactive disambiguation interface to query user intent.
- Future Directions: Integrating this nonverbal concept inference formulation into conversational interactive agents for photo album editing and reference-driven visual design.
Related Work & Insights¶
- vs Visual Prompting (Bar et al.): Visual Prompting arranges images in an inpainting canvas grid, which entangles pixel geometry with task instructions; the proposed method separates concept subspace inference from generative diffusion, achieving higher diversity.
- vs General-Purpose VLMs (ILLUME+, BAGEL, Gemini 2.5 Flash): General VLMs heavily prioritize language instructions and struggle to perform inductive visual reasoning over image sets, whereas this work provides a dedicated architectural inductive bias for nonverbal visual in-context learning.
Rating¶
- Novelty: โญโญโญโญโญ Formulates the nonverbal VICIS task and introduces an elegant subspace projection mechanism.
- Experimental Thoroughness: โญโญโญโญโญ Evaluates synthetic setups, WordNet hierarchies, cross-modality sketches, out-of-distribution classes, and closed-source SOTA comparisons.
- Writing Quality: โญโญโญโญโญ Rigorous task formalization and comprehensive empirical validation.
- Value: โญโญโญโญโญ Crucial step toward expanding visual in-context learning beyond text-prompted regimes.