Disentangling Hallucinations: Orthogonal Semantic Projection for Robust Interpretability¶
Conference: ECCV 2026
Paper: CVF / ECCV Poster
Code: https://emirhanbilgic.github.io/Orthogonal-Semantic-Projection
Area: Interpretability / Multimodal VLM
Keywords: vision-language models, attribution methods, explanation hallucination, Orthogonal Matching Pursuit, semantic disentanglement
TL;DR¶
Discloses that the root cause of widespread "explanation hallucination" in multimodal XAI saliency methods is linear semantic leakage in high-dimensional embedding spaces, and proposes Orthogonal Semantic Projection (OSP), a plug-and-play geometric intervention based on OMP residuals that eliminates ghost signals by orthogonalizing query vectors against distractor concepts.
Background & Motivation¶
Vision-Language Models (such as CLIP and SigLIP) and multimodal diffusion models (such as Stable Diffusion) have become foundational across safety-critical domains including autonomous driving, medical image synthesis, and zero-shot robotics. To unpack these black-box models and enable trustworthy diagnostics, researchers widely deploy post-hoc Explainable AI (XAI) methods like GradCAM, LeGrad, CheferCAM, or DAAM to generate spatial attribution heatmaps (saliency maps). However, when existing saliency methods are applied to multimodal architectures, a pervasive and severe failure mode emerges: "explanation hallucination." Attribution maps frequently highlight prominent foreground regions even when prompted with textual concepts completely absent from the image (e.g., highlighting a dog when prompted with "cat" or "sheep"). Such explanations appear convincing yet fail to ground properly in the queryโmaking the model seem "right for the wrong reasons" and fundamentally undermining trust in interpretability.
Traditional literature often attributes model hallucinations to visual encoder bottlenecks or generative decoding flaws, attempting concept decomposition within image representation space (such as TextSpan or Concept Bottleneck Models). This work identifies that explanation hallucination is not an issue of poor spatial localization, but rather an inevitable consequence of the geometry of multimodal embedding spaces, termed "Linear Semantic Leakage." In high-dimensional representations trained via contrastive objectives, semantically related concepts naturally share directional components; their pairwise cosine similarities are virtually never zero (e.g., across all 499,500 unique pairs of ImageNet-1k classes, the pairwise similarity never drops to zero, exhibiting a mean of 0.5925 and a minimum of 0.0967). Because mainstream discriminative and generative attribution methods exhibit linear or locally linear dependencies on text embeddings, the projection of a target query onto non-target distractor directions inevitably produces a "ghost signal" in the saliency map that is proportional to the genuine foreground object.
The paper's key insight is that since explanation hallucination stems from non-orthogonal semantic overlap in the continuous representation space, remedies should intervene directly on the text embeddings prior to attribution rather than patching pixel maps post-hoc. Core idea: formulate semantic purification as a sparse dictionary decomposition problem, exploit the mathematical residual orthogonality of Orthogonal Matching Pursuit (OMP) to strip away shared distractor semantics, and guide attribution using the purified orthogonal residual, rendering the attribution mechanism provably immune to shared ghost artifacts.
Method¶
Overall Architecture¶
Orthogonal Semantic Projection (OSP) is a plug-and-play geometric intervention requiring no model retraining or parameter fine-tuning, compatible with both discriminative dual-tower encoders (CLIP, SigLIP) and cross-attention generative diffusion models (Stable Diffusion). The overall pipeline proceeds through three main phases: first, constructing a structured dictionary of semantics containing distractor concepts for the queried target; second, executing Orthogonal Matching Pursuit in the text embedding space to obtain a purified semantic residual strictly orthogonal to all selected distractor atoms; and finally, substituting this purified representation into downstream discriminative attribution weights or diffusion cross-attention key spaces to produce faithful, hallucination-free saliency maps.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Target Text Query and Image Input"] --> B["Semantic Dictionary Construction<br/>visual confusers / co-occurring context / taxonomy"]
B --> C["Sparse Dictionary Decomposition<br/>OMP greedy selection & orthogonal projection"]
C --> D["Orthogonal Residual Projection<br/>purified query vector strictly orthogonal to distractors"]
D --> E["Ghost-Free Spatial Attribution<br/>plugged into discriminative weights or diffusion key space"]
Key Designs¶
1. Linear Semantic Leakage: formal geometric derivation of attribution hallucination
Standard discriminative attribution methods (GradCAM, LeGrad, CheferCAM, AttentionCAM) and generative cross-attention attribution (DAAM) can be abstracted into an interaction operator between visual representations and textual queries:
Suppose an image contains only class \(A\). When querying an absent class \(B\), because contrastive text embeddings are non-orthogonal (\(\langle a_A^{\text{txt}}, a_B^{\text{txt}} \rangle = \cos(\theta_{AB}) \neq 0\)), the query vector \(a_B^{\text{txt}}\) decomposes via orthogonal projection into an aligned component and an orthogonal component \(a_{B\perp A}^{\text{txt}}\):
By applying the linearity of the attribution operator \(f\) with respect to text embeddings, the resulting saliency map becomes \(V_B(x) = \cos(\theta_{AB}) V_A(x) + f(\bar{a}^{\text{img}}, a_{B\perp A}^{\text{txt}})\). The leading term \(\cos(\theta_{AB}) V_A(x)\) is the "ghost signal" that directly triggers explanation hallucination. As long as the cosine similarity is non-zero, the saliency map for absent class \(B\) will inevitably light up the region occupied by class \(A\), establishing non-orthogonality as the root cause of explanation hallucination.
2. OMP-Based Orthogonal Semantic Projection: mathematical elimination of ghost leakage
To eliminate ghost signals, OSP models text feature purification as a sparse decomposition problem. Given a unit target embedding \(a_{\text{target}}^{\text{txt}} \in \mathbb{R}^d\) and a dictionary of semantic distractor embeddings \(\mathbf{D} = [a_{D_1}^{\text{txt}}, \dots, a_{D_k}^{\text{txt}}] \in \mathbb{R}^{d \times k}\), the embedding is decomposed into shared semantics and unique semantics: \(a_{\text{target}}^{\text{txt}} = \mathbf{D}\boldsymbol{\alpha} + \mathbf{r}\). OSP employs Orthogonal Matching Pursuit (OMP), iteratively selecting the atom with the largest absolute inner product with the current residual to update the active set \(\Lambda\):
At step \(t\), the residual is updated via orthogonal least-squares projection: \(\mathbf{r}^{(t)} = a_{\text{target}}^{\text{txt}} - \mathbf{D}_\Lambda (\mathbf{D}_\Lambda^\top \mathbf{D}_\Lambda)^{-1} \mathbf{D}_\Lambda^\top a_{\text{target}}^{\text{txt}}\). By fundamental properties of OMP, at termination step \(T\), the final residual is strictly orthogonal to every selected distractor atom in \(\Lambda\), satisfying \(\langle \mathbf{r}^{(T)}, a_{D_i}^{\text{txt}} \rangle = 0, \forall i \in \Lambda\). Normalizing the residual to unit norm \(\mathbf{r} = \mathbf{r}^{(T)} / \|\mathbf{r}^{(T)}\|_2\) and substituting it for the original query vector ensures that ghost saliency components corresponding to distractor directions vanish analytically.
3. Universal Multimodal Adaptation: bridging discriminative and diffusion architectures
OSP functions as a universal, plug-and-play module across diverse model families. For discriminative vision-language encoders (CLIP, SigLIP), OSP directly swaps the raw text feature \(a_c^{\text{txt}}\) for the purified residual \(\mathbf{r}\), yielding \(V_c^{\text{OSP}}(x) = f(\bar{a}^{\text{img}}, \mathbf{r})\) without modifying the underlying network. For generative diffusion models, where spatial attributions are derived from cross-attention maps (DAAM), OSP intervenes in the key space of layer \(l\) and head \(h\). By defining the projection of target key vector \(K_{(l,h)}^{\text{target}}\) onto normalized distractor keys \(\hat{K}_{(l,h)}^{D_i}\), the purified key vector is calculated as:
where \(\beta\) regulates suppression intensity. Because attention weights remain linear with respect to keys prior to softmax normalization, dynamic orthogonalization suppresses attention alignment with distractor directions, cleanly confining diffusion attributions to genuine prompt tokens.
4. Structured Semantic Dictionary Construction: multi-tier distractor modeling
The effectiveness of orthogonal projection relies on the coverage and relevance of the dictionary \(\mathbf{D}\). The paper investigates four accessible dictionary strategies: ImageNet class names, ImageNet combined with WordNet semantic hierarchies, and custom dictionaries synthesized via Large Language Models (Gemini 3 Flash, GPT-OSS 120B). The top-performing LLM strategy organizes distractor atoms into three structured categories: 15 visual confusers (sharing salient visual appearance with the target), 15 co-occurring contextual concepts (frequently co-present in background scenes), and 10 hierarchical semantic concepts (5 hypernyms and 5 hyponyms). In addition, atoms with excessive cosine similarity to the target are filtered out via thresholding to prevent orthogonal projection from eroding the core identity of the target concept.
Key Experimental Results¶
Main Results¶
Evaluated on the ImageNet-Segmentation benchmark (4,276 images with pixel-level ground truth masks) under Positive prompts (queried concept present) and Negative prompts (queried concept absent). AUROC measures the ranking fidelity between positive and negative attributions (directly quantifying hallucination mitigation), alongside mIoU and mAP assessing spatial localization accuracy.
| Method | Base Model | Prompt | mIoU (Base / +OSP / \(\Delta\)) | mAP (Base / +OSP / \(\Delta\)) | AUROC (Base / +OSP / \(\Delta\)) |
|---|---|---|---|---|---|
| LeGrad | CLIP | Pos \(\uparrow\) / Neg \(\downarrow\) | 58.66 / 61.33 (+2.67) | 82.49 / 85.74 (+3.25) | 79.62 / 80.64 (+1.02) |
| LeGrad | SigLIP | Pos \(\uparrow\) / Neg \(\downarrow\) | 49.51 / 51.18 (+1.67) | 78.32 / 78.78 (+0.46) | 74.42 / 74.73 (+0.31) |
| CheferCAM | CLIP | Pos \(\uparrow\) / Neg \(\downarrow\) | 48.71 / 51.32 (+2.61) | 80.36 / 82.60 (+2.24) | 77.63 / 80.14 (+2.51) |
| CheferCAM | SigLIP | Pos \(\uparrow\) / Neg \(\downarrow\) | 37.66 / 39.60 (+1.94) | 73.49 / 75.95 (+2.46) | 55.47 / 62.64 (+7.17) |
| AttentionCAM | CLIP | Pos \(\uparrow\) / Neg \(\downarrow\) | 40.14 / 47.96 (+7.82) | 70.34 / 76.62 (+6.28) | 52.68 / 67.73 (+15.05) |
| AttentionCAM | SigLIP | Pos \(\uparrow\) / Neg \(\downarrow\) | 50.01 / 52.12 (+2.11) | 80.20 / 80.81 (+0.61) | 80.49 / 83.45 (+2.96) |
| GradCAM | CLIP | Pos \(\uparrow\) / Neg \(\downarrow\) | 44.68 / 51.43 (+6.75) | 74.94 / 79.80 (+4.86) | 65.39 / 71.82 (+6.43) |
| GradCAM | SigLIP | Pos \(\uparrow\) / Neg \(\downarrow\) | 38.69 / 43.26 (+4.57) | 71.43 / 73.61 (+2.18) | 67.07 / 73.01 (+5.94) |
| DAAM | Stable Diffusion 2 | Pos \(\uparrow\) / Neg \(\downarrow\) | 65.71 / 66.34 (+0.63) | 88.55 / 89.76 (+1.21) | 83.07 / 85.48 (+2.41) |
Ablation Study: Dictionary Strategy Comparison¶
Averaged across all 9 method-model pairs on ImageNet-Segmentation:
| Dictionary Strategy | Prompt | mIoU (Base / +OSP / \(\Delta\)) | mAP (Base / +OSP / \(\Delta\)) | AUROC (Base / +OSP / \(\Delta\)) |
|---|---|---|---|---|
| ImageNet Classes | Pos \(\uparrow\) / Neg \(\downarrow\) | 48.20 / 51.96 (+3.76) | 77.79 / 80.32 (+2.53) | 70.64 / 74.08 (+3.44) |
| ImageNet + WordNet | Pos \(\uparrow\) / Neg \(\downarrow\) | 47.98 / 51.40 (+3.42) | 77.79 / 80.32 (+2.53) | 70.43 / 73.90 (+3.47) |
| Gemini 3 Flash | Pos \(\uparrow\) / Neg \(\downarrow\) | 48.20 / 51.61 (+3.41) | 77.79 / 80.41 (+2.62) | 70.65 / 75.52 (+4.87) |
| GPT-OSS 120B | Pos \(\uparrow\) / Neg \(\downarrow\) | 48.20 / 51.34 (+3.14) | 77.79 / 80.29 (+2.50) | 70.65 / 76.26 (+5.61) |
Key Findings¶
- Robust hallucination suppression without sacrificing localization: OSP consistently improves AUROC across all 9 model-method pairs (e.g., +15.05 for AttentionCAM on CLIP, +7.17 for CheferCAM on SigLIP). Crucially, positive prompt localization does not degrade; mIoU and mAP increase by +0.6% to +7.8% because removing shared semantic blur sharpens object contours.
- Independence from LLMs: Dictionaries constructed purely from ImageNet classes or WordNet ontologies achieve strong AUROC gains (+3.44 to +3.47), demonstrating that the orthogonal projection mechanism is fundamentally robust. Incorporating structured LLM dictionaries further amplifies gains to +4.87 and +5.61.
- High computational efficiency: OSP requires only greedy residual projections over small dictionaries, running in under 1 second per sample across all evaluated setups.
Highlights & Insights¶
- Root-cause mitigation in text representation space: Rather than attempting spatial thresholding or visual attention pruning, OSP proves mathematically that non-orthogonal text embeddings generate ghost signals through linear attribution operators, solving hallucination cleanly via geometric projection.
- Inverted utilization of OMP residuals: Conventional sparse approximation focuses on the active dictionary coefficients; OSP cleverly repurposes the OMP residual vector as the purified query signal, leveraging guaranteed mathematical orthogonality to eliminate distractor components.
- Fine-grained disambiguation for CBMs: The geometric purification translates seamlessly to concept bottleneck models, accurately teasing apart sub-category visual attributes (e.g., fine-grained bird species features) where vanilla XAI methods bleed broadly across whole objects.
Limitations & Future Work¶
- Handling open-vocabulary complex queries: Predefined or LLM-prompted dictionaries target discrete visual concepts; constructing dynamic, comprehensive distractor sets for arbitrary long-form descriptive prompts remains challenging.
- Risk of over-projection: When distractor atoms lie excessively close to the target in angle, orthogonal projection risks stripping away legitimate core semantics, necessitating careful cosine similarity filtering.
- Extension to autoregressive MLLMs: The study centers on dual-tower encoders and diffusion cross-attention, leaving open the adaptation of orthogonal projection to multi-token autoregressive generation in vision-language models like LLaVA.
Related Work & Insights¶
- vs TextSpan (ICLR 2024): TextSpan decomposes image representations into discrete text concepts to interpret visual features, whereas OSP performs continuous geometric orthogonalization in the text embedding space as a pre-attribution purification module.
- vs CHILI (TMLR 2025): CHILI disentangles image embeddings to localize concepts in Concept Bottleneck Models, while OSP establishes a formal mathematical foundation for explanation hallucination in general attribution maps and operates plug-and-play across discriminative and generative architectures.
Rating¶
- Novelty: โญโญโญโญโญ Formulates linear semantic leakage as the geometric cause of explanation hallucination and introduces an elegant OMP-based purification paradigm.
- Experimental Thoroughness: โญโญโญโญโญ Evaluates 3 foundational backbones, 5 attribution methods, 4 dictionary strategies, and complements quantitative benchmarks with a 200-participant user study.
- Writing Quality: โญโญโญโญโญ Clear mathematical derivations, coherent visual diagrams, and rigorous ablation analyses.
- Value: โญโญโญโญโญ Plug-and-play with sub-second execution overhead (<1s), providing immediate diagnostic utility for trustworthy multimodal vision systems.