GridVQA-X: A Diagnostic Framework for Evaluating Multimodal Explainability Methods¶
Conference: ECCV 2026
Paper: CVF Open Access
Code: Hugging Face Dataset & Models
Area: Interpretability
Keywords: multimodal explainability, shortcut learning, diagnostic benchmarking, cross-modal synergy, visual question answering
TL;DR¶
Addressing the lack of ground-truth causal validation in multimodal explainable AI (MxAI), this paper introduces GridVQA-X, the first closed-world diagnostic framework featuring mathematically guaranteed unique ground-truth explanations and paired models (\(M_{\text{pure}}\) vs. \(M_{\text{spur}}\)), revealing that prominent explainers fail to distinguish genuine spatial reasoning from shallow Bag-of-Words shortcuts.
Background & Motivation¶
The rapid progression of Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) has established remarkable capabilities across vision-text understanding. However, this success is achieved via deeply entangled cross-modal representations that render their internal decision-making opaque, severely impeding their deployment in high-stakes domains such as healthcare and automated decision auditing. In response, a range of post-hoc, model-agnostic multimodal explainable AI (MxAI) methods have emerged, claiming to attribute predictions to individual multimodal features and uncover cross-modal synergies. Nonetheless, existing explainability benchmarks remain restricted to unimodal tasks or rely on noisy, real-world vision-language datasets where human annotations lack causally defined ground truth and introduce confounding statistical priors. Consequently, the fundamental faithfulness of post-hoc MxAI methods remains unverified.
This challenge is acutely exacerbated by shortcut learning in multimodal architectures. Vision-language models frequently exploit shallow cross-modal shortcuts, such as Bag-of-Words attribute matching—counting visual objects matching specified color or shape tokens while completely bypassing spatial relational directives like "left of" or "above". Because existing evaluation protocols lack a gold-standard ground truth for the model's actual internal reasoning pathway, it is fundamentally impossible to discern whether an explainer faithfully reflects true multi-hop compositional reasoning or merely hallucinates a plausible rationale for a model operating as a trivial feature detector. If an explanation tool cannot differentiate between genuine cross-modal synergy and shallow heuristics, it creates a dangerous illusion of interpretability.
This paper tackles this foundational blind spot by constructing a fully observable, closed-world synthetic universe with provable causal properties. The core idea is to synthesize closed-world visual scenes with mathematically guaranteed unique ground-truth causal masks and train an identical-architecture model pair—\(M_{\text{pure}}\) (which learns robust spatial-relational reasoning) and \(M_{\text{spur}}\) (which is structurally forced to rely on cross-modal shortcuts)—establishing the GridVQA-X diagnostic benchmark to evaluate whether explainers reliably differentiate true reasoning from shortcut learning.
Method¶
Overall Architecture¶
GridVQA-X establishes a rigorous testbed by abstracting multimodal reasoning into a discrete 2D grid universe where every object is fully observable. By applying Pearl's do-calculus, the causal effect of background distractors is mathematically driven to zero, establishing a provably unique ground-truth causal explanation mask defined strictly by target and anchor objects. The data generation engine algorithmically bifurcates into two parallel datasets: \(D_{\text{pure}}\), which destroys spatial shortcuts and partial logic heuristics by actively overloading spatial confuser regions with adversarial distractors, and \(D_{\text{spur}}\), which restricts attribute sampling so that shallow keyword matching yields a 100% predictive success rate. Across these environments, a unified transformer architecture (MDETR) is optimized via explanation-guided dynamics into two reference models: \(M_{\text{pure}}\) and \(M_{\text{spur}}\). Finally, local attribution heatmaps and global synergy metrics from modern MxAI algorithms are benchmarked across three evaluation regimes using adapted fidelity metrics and an additive fallacy check.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multimodal Query Parameters & Atomic Entity Set"] --> B["Closed-World Causal Space & 4D Taxonomy Parameterization"]
B --> C["Adversarial Confuser Construction & Robust Intersection Theorem"]
C --> D["Dual-Track Counterfactual Synthesis & Paired Diagnostic Models"]
D --> E["Multimodal Attribution Protocol & Additive Fallacy Diagnosis"]
E --> F["Diagnostic Evaluation: Local Heatmap Collapse & Global Synergy Hallucination"]
Key Designs¶
1. Closed-World Causal Space & 4D Taxonomy Parameterization: Eliminating Confounders with Provably Unique Causal Ground Truth
Real-world multimodal tasks cannot serve as zero-ambiguity explainability benchmarks due to entangled feature correlations and lack of causal ground truth. GridVQA-X abstracts scenes into a discrete \(S \times S\) spatial grid \(V\), where each object is represented as an independent tuple \(o_i = (c_i, s_i, p_i)\) denoting color, shape, and discrete coordinate position. A multimodal query \(Q\) specifies attribute constraints for target entities \(T\), reference anchor entities \(A\), and the directional relationships connecting them. Let the distractor set be \(\text{Dist} = V \setminus (A \cup T)\). Applying Pearl's do-calculus, intervening on any distractor \(o_d \in \text{Dist}\) satisfies:
Because the causal effect of \(\text{Dist}\) on the prediction \(y\) is mathematically zero, the set \(A \cup T\) is proven to constitute the unique, unambiguous ground-truth causal explanation for \(Q\). Any explainer allocating attribution mass to objects in \(\text{Dist}\) is provably unfaithful. Samples are parameterized across four axes \(\mathcal{T} = (D, Q, F, \rho)\): relational depth \(D \in \{1, 2, 3\}\) scales multi-hop anchor intersections; question type \(Q\) spans Attribute-Only, Shape-Only, Color-Only, Mixed, and Comparison queries; task form \(F \in \{0, 1\}\) alternates between counting and existence without altering the underlying causal graph; and grid density \(\rho \in \{d0.3, d0.7\}\) models varying visual clutter.
2. Adversarial Confuser Construction & Robust Intersection Theorem: Neutralizing Partial Logic and Shortcut Reliance
In multi-hop compositional reasoning, models frequently collapse into a partial logic heuristic, evaluating only a subset of directional constraints (e.g., dropping anchor \(k\) to compute the relaxed intersection \(\bigcap_{i \neq k} V_i\)) while still obtaining high empirical accuracy. To definitively prevent this shortcut in the pure environment \(D_{\text{pure}}\), the authors formulate the Generalized Robust Intersection Guarantee: let \(V_i\) denote the valid spatial region for anchor \(i\), \(V_{-k} = \bigcap_{i \neq k} V_i\) the relaxed region, and \(V_{\text{target}} = \bigcap_{i=1}^M V_i\) the true target intersection. The confuser region \(C_k\) is the relative complement:
Theorem 1 proves that by algorithmically guaranteeing that \(C_k\) is non-empty and actively populated with at least one adversarial object sharing the target's visual attributes, any model dropping anchor \(k\) is guaranteed to detect false positives in \(C_k\). The model strictly overcounts or falsely confirms existence, incurring an immediate optimization loss penalty. During procedural generation of \(D_{\text{pure}}\), the engine deterministically computes \(C_k\) for every anchor and overloads it with target-matching distractors, compelling the network to compute the full multi-hop spatial intersection.
3. Dual-Track Counterfactual Synthesis & Paired Diagnostic Models: Isolating Identifiable Model Behaviors
Evaluating whether explainers diagnose shortcuts requires reference models whose internal reasoning mechanisms are verifiably known. GridVQA-X leverages a unified modulated detection architecture (MDETR) trained with explanation-guided learning across the divergent datasets. In \(D_{\text{spur}}\), target-matching objects are strictly confined to the valid target region, ensuring \(P(Y_{\text{ans}} \mid \text{Target\_Attrs}) = 1.0\) and embedding the Bag-of-Words shortcut. Conversely, in \(D_{\text{pure}}\), adversarial confusers collapse the shortcut predictive probability to 36.07%.
The models are trained under a two-phase regime: Phase 1 optimizes visual grounding via bounding box alignment (\(L_{\text{ground}} = \lambda_{L_1} L_1 + \lambda_{\text{giou}} L_{\text{giou}}\)), and Phase 2 optimizes answer prediction using a dynamically weighted cross-entropy loss \(L_{\text{QA}}\) to neutralize answer prior bias (Case-0). When tested on \(D_{\text{pure}}\), the shortcut-trained model \(M_{\text{spur}}\) catastrophically fails on multi-hop relational queries (dropping to 8% accuracy on Depth-2 Mixed and 14% on Depth-3 Mixed queries), yet maintains 100% on non-relational Attribute-Only queries. This verified behavioral divergence creates an absolute benchmark: a faithful explainer must produce divergent attributions for \(M_{\text{pure}}\) and \(M_{\text{spur}}\), exposing the latter's complete disregard of spatial relations.
4. Multimodal Attribution Protocol & Additive Fallacy Diagnosis: Measuring Precision and Interaction Monotonicity
To evaluate diverse explainers, GridVQA-X adapts local and global metrics into a unified protocol. For local attribution heatmaps (Dime, MultiSHAP, MultiViz), in addition to Otsu-binarized Intersection over Union (IoU), the framework evaluates Relevance Mass Accuracy (RMA):
A faithful explainer must assign high RMA on ground-truth masks for \(M_{\text{pure}}\), while allocating minimal mass to valid anchors when evaluating \(M_{\text{spur}}\). For global explainers producing a cross-modal synergy scalar \(S_{\text{score}}\) (EMAP, InterSHAP, PID), the authors introduce the Additive Fallacy Check. True cross-modal synergy must monotonically increase with relational depth (\(D=1 \to 3\)) for \(M_{\text{pure}}\), as each additional anchor demands higher compositional integration, while remaining low and depth-invariant for \(M_{\text{spur}}\). This diagnostic tests whether global explainers measure actual cross-modal synergy or merely capture unimodal feature sensitivity.
Key Experimental Results¶
Main Results¶
The framework benchmarks state-of-the-art local attribution algorithms across visual and textual modalities. Evaluations are conducted under Pure Evaluation (\(M_{\text{pure}}\) on \(D_{\text{pure}}\)), Spurious Evaluation (\(M_{\text{spur}}\) on \(D_{\text{spur}}\)), and Cross-Evaluation (\(M_{\text{spur}}\) on \(D_{\text{pure}}\)). The following table summarizes the performance on visual and textual modalities:
| Modality | Method | Pure Eval RMA (%) (↑) | Pure Eval IoU (%) (↑) | Spur. Eval RMA (%) (↑) | Spur. Eval IoU (%) (↑) |
|---|---|---|---|---|---|
| Visual | Dime | \(28.0 \pm 13.9\) | \(20.5 \pm 10.2\) | \(27.2 \pm 14.5\) | \(18.2 \pm 10.0\) |
| Visual | MultiSHAP | \(56.0 \pm 22.7\) | \(50.0 \pm 22.6\) | \(64.5 \pm 21.8\) | \(46.2 \pm 23.1\) |
| Visual | MultiViz | \(44.0 \pm 16.3\) | \(4.7 \pm 3.3\) | \(42.2 \pm 20.0\) | \(6.7 \pm 4.5\) |
| Textual | Dime | \(43.0 \pm 26.8\) | \(12.3 \pm 12.5\) | \(37.5 \pm 25.9\) | \(10.3 \pm 11.5\) |
| Textual | MultiSHAP | \(66.8 \pm 17.3\) | \(46.7 \pm 18.9\) | \(67.5 \pm 16.1\) | \(40.2 \pm 19.9\) |
Ablation Study & Complexity Scaling Analysis¶
To evaluate whether global explainability metrics scale faithfully with relational complexity, EMAP's synergy scalar \(S_{\text{score}}\) is tracked across relational depths (\(D \in \{1, 2, 3\}\)). In addition, the predictive accuracy of paired diagnostic models and open-source foundation VLMs on multi-hop mixed queries (Mixed, Form 0, \(\rho=d0.7\)) is benchmarked:
| Evaluation Regime / Model Config | Depth 1 Metric | Depth 2 Metric | Depth 3 Metric | Trend / Analysis |
|---|---|---|---|---|
| EMAP Pure Eval Synergy (%) | \(78.6 \pm 16.7\) | \(74.7 \pm 23.8\) | \(64.1 \pm 28.2\) | Violates additive check: decreases as reasoning depth increases |
| EMAP Spur. Eval Synergy (%) | \(69.2 \pm 23.9\) | \(65.7 \pm 24.4\) | \(68.8 \pm 18.9\) | Hallucinates deep synergy on a pure Bag-of-Words shortcut model |
| EMAP Cross Eval Synergy (%) | \(59.4 \pm 27.3\) | \(54.0 \pm 26.2\) | \(65.0 \pm 17.6\) | Assigns ~60% synergy even when the model's accuracy collapses |
| Reference Model \(M_{\text{pure}}\) Acc (%) | 100.0 | 100.0 | 100.0 | Robustly learns true multi-hop spatial geometry |
| Reference Model \(M_{\text{spur}}\) Acc (%) | 6.0 | 2.0 | 0.0 | Fails completely on out-of-distribution adversarial spatial distractors |
| Open-Source Qwen3-VL-30B Acc (%) | 14.0 | 10.0 | 10.0 | Large foundation VLM degrades under dense spatial distractors |
| Open-Source Llava-1.5-7B Acc (%) | 6.0 | 28.0 | 30.0 | Weak spatial grounding in multi-hop compositional reasoning |
Key Findings¶
- Model Blindness and Accidental Faithfulness: Current local explainers fail to diagnose shortcut learning. MultiViz yields nearly identical RMA (\(\approx 44\%\) vs. \(42\%\)) across \(M_{\text{pure}}\) and \(M_{\text{spur}}\), displaying total model blindness. Dime produces highly diffuse, unconstrained heatmaps that accidentally overlap with target areas, generating an illusion of faithfulness.
- Game-Theoretic Preference for Unary Feature Detectors: MultiSHAP paradoxically achieves higher RMA on the shortcut model \(M_{\text{spur}}\) (64.5%) than on the faithful model \(M_{\text{pure}}\) (56.0%). Because Shapley value marginals align naturally with independent 1-to-1 feature matching, game-theoretic methods structurally favor shallow shortcuts over non-linear geometric intersections.
- Failure of the Additive Fallacy Check in Global Estimators: EMAP's estimated synergy on \(M_{\text{pure}}\) paradoxically deteriorates from 78.6% at Depth 1 down to 64.1% at Depth 3, falsely implying that tri-anchor queries require less cross-modal interaction. On \(M_{\text{spur}}\), it hallucinates \(\sim 69\%\) synergy despite the model bypassing directional tokens. InterSHAP displays an identical collapse (dropping from 0.921 to 0.629).
- Distractor Leakage on Existence Tasks: When evaluating existence queries (Form 1), MultiSHAP's RMA suffers severe degradation because it highlights all objects matching target attributes regardless of whether they violate spatial constraints, acting as a blind attribute detector rather than a relational explainer.
Highlights & Insights¶
- Mathematically Guaranteed Ground-Truth Explanations: By formulating scene generation through Pearl's do-calculus, the framework eliminates subjective human annotation ambiguity and provides the first provably exact causal mask for multimodal interactions.
- Controlled Behavioral Divergence via Paired Models: Rather than guessing what a black-box model is thinking, the authors train identical architectures into provably divergent reasoning regimes (\(M_{\text{pure}}\) vs. \(M_{\text{spur}}\)), creating a definitive diagnostic ground truth for XAI methods.
- Deconstruction of the Multimodal Interpretability Illusion: The findings demonstrate that popular explainability algorithms frequently generate plausible-looking rationales on models operating entirely via shallow heuristics, highlighting an urgent need for causally faithful evaluation benchmarks.
Limitations & Future Work¶
- Limitations Admitted by the Authors: GridVQA-X operates in a synthetic 2D geometric universe that lacks the complex textures and continuous semantic distributions of natural images. Furthermore, the taxonomy is confined to spatial-relational composition and basic attribute binding, excluding temporal dynamics, physical commonsense, or mathematical logic.
- Additional Limitations: The evaluation is centered on MDETR, an object-centric detection transformer; it does not directly evaluate the internal attention rationales or generative chain-of-thought (CoT) text tokens produced by autoregressive MLLMs.
- Future Directions: Extending the adversarial confuser placement logic into continuous 3D embodied simulation environments and developing new explainability algorithms capable of passing this zero-ambiguity diagnostic testbed.
Related Work & Insights¶
- vs CLEVR-XAI: CLEVR-XAI provides ground-truth masks for synthetic 3D scenes but evaluates only unimodal visual neural networks, lacking multimodal interaction metrics and paired shortcut models; GridVQA-X evaluates cross-modal synergy with mathematical guarantees.
- vs OpenXAI: OpenXAI is a pioneering transparent XAI benchmark for tabular and image domains, but is strictly unimodal; GridVQA-X fills the void for cross-modal interaction benchmarking.
- vs Dime / MultiSHAP / MultiViz: While these methods claim to isolate multimodal interactions from unimodal contributions, GridVQA-X empirically proves that they fail to distinguish between true spatial reasoning and Bag-of-Words shortcuts.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First diagnostic framework providing mathematically guaranteed ground truth and paired models for multimodal explainability.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 6 state-of-the-art local and global explainers across vision and text modalities.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous causal formulations, clean mathematical proofs, and well-structured empirical analyses.
- Value: ⭐⭐⭐⭐⭐ Exposes fundamental blind spots in multimodal XAI and sets a new rigorous standard for evaluating cross-modal faithfulness.