Skip to content

Do Vision Language Models Recognize Visual Ambiguity?

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/tadeephuy/visual-semantic-entropy
Area: Multimodal VLM
Keywords: vision-language models, uncertainty estimation, visual ambiguity, semantic entropy, prototype aggregation

TL;DR

Addressing the issue where vision-language models produce overconfident predictions on visually ambiguous inputs, this paper proposes Visual Semantic Entropy (VSE), which perturbs only the image and measures mass-weighted dispersion among semantic answer prototypes, substantially outperforming prior uncertainty estimators across five VLMs and five VQA benchmarks.

Background & Motivation

In Visual Question Answering (VQA), the reliability of vision-language models (VLMs) fundamentally hinges on their ability to perceive ambiguity in visual evidence. However, existing uncertainty estimation (UE) techniques are largely inherited from text-only LLMs, heavily relying on output diversity under stochastic decoding (such as Semantic Entropy, SE) or verbalized self-confidence. When confronted with visually biased or ambiguous images, the visual encoder of a VLM often produces overconfident visual embeddings that force autoregressive decoding into a sharply peaked, homogeneous answer distribution. Even when the prediction is wrong, standard SE mistakenly yields a low uncertainty score.

To elicit output diversity, recent approaches incorporate textual paraphrasing or joint text-image perturbations (e.g., C&U, VL-Uncertainty). Yet this introduces subtle modality confounding: natural language is discrete and sensitive, causing textual paraphrasing to induce large, non-local semantic shifts in multimodal embedding space. Purity analysis demonstrates that variability under joint perturbation is overwhelmingly dominated by textual changes (Text Purity consistently exceeds Image Purity), causing the estimated uncertainty to reflect prompt sensitivity rather than true visual ambiguity. Furthermore, directly aggregating pairwise semantic distances across raw text outputs (e.g., SNNE) misinterprets syntactic and wording variations among semantically identical answers as factual disagreement, artificially inflating uncertainty.

To overcome these three failure modes, the authors argue that perturbations must be strictly confined to the continuous visual input domain to probe local visual instability, while answers must first be grouped semantically to eliminate wording noise. Core Idea: Propose Visual Semantic Entropy (VSE), which fixes the text query, perturbs only the input image with semantics-preserving local Gaussian noise, clusters generated answer samples into semantic prototypes, and quantifies visual uncertainty via mass-weighted pairwise dispersion across prototype embeddings.

Method

Overall Architecture

The evaluation pipeline of Visual Semantic Entropy (VSE) comprises three sequential stages: local visual perturbation, cross-view answer sampling, and prototype semantic aggregation (ProtoSem). First, semantics-preserving local variants are generated for the input image while keeping the question query strictly fixed. Next, each perturbed view along with the original question is fed into the VLM to sample alternative candidate answers. Finally, a semantic similarity model groups the candidate responses into distinct semantic clusters, extracts a representative prototype for each cluster, and computes the mass-weighted pairwise semantic dispersion between prototypes as the uncertainty score.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image & Fixed Question Query"] --> B["Local Visual Perturbation<br/>Generate M semantics-preserving local views"]
    B --> C["Cross-View Answer Sampling<br/>VLM stochastic decoding generates candidate answers"]
    C --> D["Prototype Semantic Aggregation<br/>Hierarchical clustering & prototype extraction"]
    D --> E["Mass-Weighted Prototype Dispersion<br/>Output visual uncertainty score VSE"]

Key Designs

1. Local Visual Perturbation: Isolating Prompt Sensitivity and Probing Local Visual Ambiguity

Prior techniques utilizing question paraphrasing or joint multimodal noise trigger non-local semantic shifts due to prompt sensitivity, masking true visual ambiguity. To ensure uncertainty is strictly conditioned on the original pair \((q, v)\), VSE introduces an image perturbation operator \(\mathcal{T}(v; \xi_m)\) that injects zero-mean Gaussian noise with standard deviation \(\sigma = 20\) in pixel space, producing \(M\) controlled local views \(v_m = \mathcal{T}(v; \xi_m)\) while keeping question \(q\) unchanged. Unlike previous methods that progressively scale up noise levels, VSE adheres to a single small perturbation scale. This avoids pushing inputs outside the local neighborhood manifold of \(v\), ensuring that the observed answer variance faithfully reflects the model's prediction instability around local visual decision boundaries.

2. Cross-View Answer Sampling: Overcoming Decoding Suppression from Overconfident Embeddings

On ambiguous images, visual tokens projected into the vocabulary space often exhibit sharp, low-entropy distributions, causing regular stochastic sampling to collapse into a single mode. VSE first generates the primary prediction \(a_0\) via greedy decoding (\(T=0.0\)) on the original input. Subsequently, the \(M\) perturbed image views are passed to the VLM under temperature \(T=1.0\) to collect a candidate response set \(\{a_m\}_{m=1}^M\). Gentle pixel perturbations destabilize the overfitted visual representations, prompting the model to expose alternative visual interpretations (such as eliciting "bag" or "pocket" alongside an overconfident "pouch").

3. Prototype Semantic Aggregation: Eliminating Wording Inflation and Measuring True Semantic Disagreement

Directly aggregating pairwise distances over discrete textual outputs (as in SNNE) penalizes stylistic variations between semantically equivalent answers (e.g., "The dog is eating a cucumber" vs. "The puppy is biting into a chunk of cucumber"), inflating uncertainty without genuine conflict. VSE introduces ProtoSem: using a pretrained NLI model (DeBERTa-v2-xlarge-mnli) as the distance metric \(d(\cdot, \cdot)\), candidate answers are grouped into \(K\) semantic clusters \(\{c_k\}_{k=1}^K\) via hierarchical clustering. Within each cluster, the answer minimizing the average distance to other cluster members is selected as the representative prototype \(p_k = \arg\min_{a \in c_k} \sum_{a' \in c_k} d(a, a')\). Defining the empirical mass of each cluster as \(w_k = \frac{|c_k|}{M}\), the uncertainty score is computed as the expected semantic distance between independently drawn prototypes:

\[\tilde{u} = \sum_{k \neq k'} w_k w_{k'} d(p_k, p_{k'})\]

This formulation absorbs wording variations within intra-cluster boundaries, yielding high uncertainty only when multiple high-mass prototypes are distinctly separated in semantic space.

Key Experimental Results

Main Results

The Area Under the ROC Curve (AUC) for distinguishing incorrect predictions from correct ones serves as the primary evaluation metric. On standard VQA benchmarks (AOKVQA, OKVQA, MMVet), VSE is compared against verbalized confidence (Verb-U), logit-based metrics, consistency-based approaches (SCG, C&U), and entropy-based estimators (SE, SNNE, KLE, VL-U).

Dataset Model VSE (Ours) Best Baseline (Method / AUC) AUC Gain
AOKVQA Qwen2.5-VL-7B 0.783 KLE (0.774) +0.009
AOKVQA Gemma3-4B 0.778 SE (0.732) +0.046
AOKVQA LLaVA-NeXt-8B 0.724 SNNE (0.651) +0.073
AOKVQA Qwen3-VL-8B 0.798 KLE (0.772) +0.026
AOKVQA Intern3.5-VL-8B 0.792 VL-U (0.781) +0.011
OKVQA Qwen2.5-VL-7B 0.767 SCG-Pr (0.767) / SE (0.758) +0.009 (vs SE)
OKVQA LLaVA-NeXt-8B 0.749 AvgEnt (0.729) / C&U (0.718) +0.020
MMVet Qwen2.5-VL-7B 0.781 SNNE (0.758) +0.023
MMVet Gemma3-4B 0.778 SNNE (0.740) +0.038

Ablation Study & Adversarial Analysis

To verify the isolated contributions of visual perturbation (\(+T\)) and prototype semantic aggregation (ProtoSem), component ablations were conducted on AOKVQA. Furthermore, evaluations on visually adversarial and biased benchmarks (VILP and VLM-are-biased) assessed the limits of VLM visual ambiguity detection.

Configuration / Dataset Qwen2.5-VL Gemma3 Note
SE (Base) 0.702 0.732 Standard semantic entropy without image perturbation
SE + \(T\) (Visual Perturbation) 0.761 (+0.059) 0.775 (+0.043) Image perturbation alone markedly boosts SE
SNNE (Base) 0.700 0.651 Pairwise text distances, distorted by wording variance
SNNE + \(T\) (Visual Perturbation) 0.720 (+0.020) 0.746 (+0.095) Visual perturbation benefits distance-based metrics
ProtoSem (Base, No Perturbation) 0.711 0.745 Prototype aggregation prevents wording inflation
VSE (ProtoSem + \(T\), Full Model) 0.783 (+0.072) 0.778 (+0.033) Combined perturbation & prototype aggregation achieves best AUC
Adversarial: VILP (SE Baseline) 0.535 0.660 Standard SE degrades near random guessing on adversarial samples
Adversarial: VILP (VSE Ours) 0.650 (+0.115) 0.665 (+0.005) Recovers robust uncertainty estimates on visual adversaries
Bias Benchmark: VLM-are-biased (VL-U) 0.783 0.758 Joint perturbation baseline on biased images
Bias Benchmark: VLM-are-biased (VSE Ours) 0.826 (+0.043) 0.776 (+0.018) Avoids language artifacts and reliably captures visual ambiguity

Key Findings

  • Visual perturbation breaks visual overconfidence: Appending visual perturbation \(+T\) to SE and SNNE yields universal AUC improvements ranging from +2.0% to +9.5%, confirming that overconfident visual representations suppress stochastic decoding unless activated by image variations.
  • ProtoSem eliminates wording variance inflation: Under identical perturbation conditions, ProtoSem consistently outperforms raw pairwise text distances (SNNE) (e.g., 0.783 vs 0.720 on Qwen2.5-VL), confirming that intra-cluster equivalence grouping removes superficial linguistic divergence.
  • Substantial gains on adversarial vision benchmarks: On the VILP benchmark containing challenging visual ambiguities, traditional SE achieves only 0.535 AUC on Qwen2.5-VL, whereas VSE boosts it to 0.650 (+11.5%), demonstrating superior robustness when visual evidence is intentionally perturbed or biased.

Highlights & Insights

  • Uncovering textual contamination in multimodal UE: Through kernel density estimation and cluster purity analysis, this paper rigorously demonstrates that joint text-image perturbations are predominantly driven by text sensitivity, formalizing the imperative to keep prompts fixed when evaluating visual uncertainty.
  • Two-tier aggregation paradigm via prototype dispersion: By decomposing output variance into intra-cluster paraphrasing and inter-cluster disagreement, VSE bridges discrete categorical entropy with continuous semantic distance metrics.
  • Model-agnostic black-box applicability: VSE operates entirely on input image manipulation and output answer clustering without accessing model internals, attention weights, or logits, facilitating straightforward real-world deployment.

Limitations & Future Work

  • Computational sampling overhead: Evaluating uncertainty requires \(M+1\) forward visual passes and autoregressive generations (\(M=10\)), along with an auxiliary DeBERTa clustering step, posing latency challenges for real-time applications.
  • Single perturbation modality: The current implementation relies solely on additive Gaussian noise, leaving structured visual variations (e.g., occlusion, viewpoint changes, motion blur) unexplored.
  • Hyperparameter dependency: The choice of temperature \(T\) and the clustering threshold directly govern prototype formation, warranting broader exploration across very large multimodal foundation models and extended reasoning tasks.
  • vs. Semantic Entropy (SE): SE assumes uncertainty originates solely from language decoding randomness and fails under confident visual tokens; VSE actively elicits visual variance and measures dispersion across semantic prototypes.
  • vs. VL-Uncertainty / C&U: These methods alter prompts or apply multi-level joint noise, falling prey to prompt sensitivity; VSE adopts a vision-only perturbation rule to eliminate textual confounding.
  • vs. SNNE / KLE: SNNE computes distances between all text pairs, mistakenly penalizing synonymous phrasings; VSE summarizes each cluster into an invariant prototype before measuring semantic divergence.

Rating

  • Novelty: โญโญโญโญโ˜† Pinpoints the critical blind spots of visual overconfidence and prompt contamination in VLM uncertainty estimation with sharp analytical insights.
  • Experimental Thoroughness: โญโญโญโญโญ Evaluates 5 contemporary VLMs across 5 standard and adversarial VQA benchmarks with comprehensive purity metrics and ablations.
  • Writing Quality: โญโญโญโญโญ Exceptionally clear narrative structure, linking failure hypotheses, formal propositions, and algorithmic designs seamlessly.
  • Value: โญโญโญโญโ˜† Offers a robust, plug-and-play uncertainty evaluation benchmark for reliable VLM deployment and hallucination mitigation.