Skip to content

GridVQA-X: A Diagnostic Framework for Evaluating Multimodal Explainability Methods

Conference: ECCV 2026
Paper: CVF Open Access
Code: Hugging Face Dataset & Models
Area: Interpretability
Keywords: multimodal explainability, shortcut learning, diagnostic benchmarking, cross-modal synergy, visual question answering

TL;DR

Addressing the lack of ground-truth causal validation in multimodal explainable AI (MxAI), this paper introduces GridVQA-X, the first closed-world diagnostic framework featuring mathematically guaranteed unique ground-truth explanations and paired models (\(M_{\text{pure}}\) vs. \(M_{\text{spur}}\)), revealing that prominent explainers fail to distinguish genuine spatial reasoning from shallow Bag-of-Words shortcuts.

Background & Motivation

The rapid progression of Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) has established remarkable capabilities across vision-text understanding. However, this success is achieved via deeply entangled cross-modal representations that render their internal decision-making opaque, severely impeding their deployment in high-stakes domains such as healthcare and automated decision auditing. In response, a range of post-hoc, model-agnostic multimodal explainable AI (MxAI) methods have emerged, claiming to attribute predictions to individual multimodal features and uncover cross-modal synergies. Nonetheless, existing explainability benchmarks remain restricted to unimodal tasks or rely on noisy, real-world vision-language datasets where human annotations lack causally defined ground truth and introduce confounding statistical priors. Consequently, the fundamental faithfulness of post-hoc MxAI methods remains unverified.

This challenge is acutely exacerbated by shortcut learning in multimodal architectures. Vision-language models frequently exploit shallow cross-modal shortcuts, such as Bag-of-Words attribute matching—counting visual objects matching specified color or shape tokens while completely bypassing spatial relational directives like "left of" or "above". Because existing evaluation protocols lack a gold-standard ground truth for the model's actual internal reasoning pathway, it is fundamentally impossible to discern whether an explainer faithfully reflects true multi-hop compositional reasoning or merely hallucinates a plausible rationale for a model operating as a trivial feature detector. If an explanation tool cannot differentiate between genuine cross-modal synergy and shallow heuristics, it creates a dangerous illusion of interpretability.

This paper tackles this foundational blind spot by constructing a fully observable, closed-world synthetic universe with provable causal properties. The core idea is to synthesize closed-world visual scenes with mathematically guaranteed unique ground-truth causal masks and train an identical-architecture model pair—\(M_{\text{pure}}\) (which learns robust spatial-relational reasoning) and \(M_{\text{spur}}\) (which is structurally forced to rely on cross-modal shortcuts)—establishing the GridVQA-X diagnostic benchmark to evaluate whether explainers reliably differentiate true reasoning from shortcut learning.

Method

Overall Architecture

GridVQA-X establishes a rigorous testbed by abstracting multimodal reasoning into a discrete 2D grid universe where every object is fully observable. By applying Pearl's do-calculus, the causal effect of background distractors is mathematically driven to zero, establishing a provably unique ground-truth causal explanation mask defined strictly by target and anchor objects. The data generation engine algorithmically bifurcates into two parallel datasets: \(D_{\text{pure}}\), which destroys spatial shortcuts and partial logic heuristics by actively overloading spatial confuser regions with adversarial distractors, and \(D_{\text{spur}}\), which restricts attribute sampling so that shallow keyword matching yields a 100% predictive success rate. Across these environments, a unified transformer architecture (MDETR) is optimized via explanation-guided dynamics into two reference models: \(M_{\text{pure}}\) and \(M_{\text{spur}}\). Finally, local attribution heatmaps and global synergy metrics from modern MxAI algorithms are benchmarked across three evaluation regimes using adapted fidelity metrics and an additive fallacy check.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multimodal Query Parameters & Atomic Entity Set"] --> B["Closed-World Causal Space & 4D Taxonomy Parameterization"]
    B --> C["Adversarial Confuser Construction & Robust Intersection Theorem"]
    C --> D["Dual-Track Counterfactual Synthesis & Paired Diagnostic Models"]
    D --> E["Multimodal Attribution Protocol & Additive Fallacy Diagnosis"]
    E --> F["Diagnostic Evaluation: Local Heatmap Collapse & Global Synergy Hallucination"]

Key Designs

1. Closed-World Causal Space & 4D Taxonomy Parameterization: Eliminating Confounders with Provably Unique Causal Ground Truth

Real-world multimodal tasks cannot serve as zero-ambiguity explainability benchmarks due to entangled feature correlations and lack of causal ground truth. GridVQA-X abstracts scenes into a discrete \(S \times S\) spatial grid \(V\), where each object is represented as an independent tuple \(o_i = (c_i, s_i, p_i)\) denoting color, shape, and discrete coordinate position. A multimodal query \(Q\) specifies attribute constraints for target entities \(T\), reference anchor entities \(A\), and the directional relationships connecting them. Let the distractor set be \(\text{Dist} = V \setminus (A \cup T)\). Applying Pearl's do-calculus, intervening on any distractor \(o_d \in \text{Dist}\) satisfies:

\[P(y \mid do(o_d \to o'_d)) = P(y), \quad \forall o_d \in \text{Dist}\]

Because the causal effect of \(\text{Dist}\) on the prediction \(y\) is mathematically zero, the set \(A \cup T\) is proven to constitute the unique, unambiguous ground-truth causal explanation for \(Q\). Any explainer allocating attribution mass to objects in \(\text{Dist}\) is provably unfaithful. Samples are parameterized across four axes \(\mathcal{T} = (D, Q, F, \rho)\): relational depth \(D \in \{1, 2, 3\}\) scales multi-hop anchor intersections; question type \(Q\) spans Attribute-Only, Shape-Only, Color-Only, Mixed, and Comparison queries; task form \(F \in \{0, 1\}\) alternates between counting and existence without altering the underlying causal graph; and grid density \(\rho \in \{d0.3, d0.7\}\) models varying visual clutter.

2. Adversarial Confuser Construction & Robust Intersection Theorem: Neutralizing Partial Logic and Shortcut Reliance

In multi-hop compositional reasoning, models frequently collapse into a partial logic heuristic, evaluating only a subset of directional constraints (e.g., dropping anchor \(k\) to compute the relaxed intersection \(\bigcap_{i \neq k} V_i\)) while still obtaining high empirical accuracy. To definitively prevent this shortcut in the pure environment \(D_{\text{pure}}\), the authors formulate the Generalized Robust Intersection Guarantee: let \(V_i\) denote the valid spatial region for anchor \(i\), \(V_{-k} = \bigcap_{i \neq k} V_i\) the relaxed region, and \(V_{\text{target}} = \bigcap_{i=1}^M V_i\) the true target intersection. The confuser region \(C_k\) is the relative complement:

\[C_k = V_{-k} \setminus V_k = \left( \bigcap_{i \neq k} V_i \right) \setminus V_k\]

Theorem 1 proves that by algorithmically guaranteeing that \(C_k\) is non-empty and actively populated with at least one adversarial object sharing the target's visual attributes, any model dropping anchor \(k\) is guaranteed to detect false positives in \(C_k\). The model strictly overcounts or falsely confirms existence, incurring an immediate optimization loss penalty. During procedural generation of \(D_{\text{pure}}\), the engine deterministically computes \(C_k\) for every anchor and overloads it with target-matching distractors, compelling the network to compute the full multi-hop spatial intersection.

3. Dual-Track Counterfactual Synthesis & Paired Diagnostic Models: Isolating Identifiable Model Behaviors

Evaluating whether explainers diagnose shortcuts requires reference models whose internal reasoning mechanisms are verifiably known. GridVQA-X leverages a unified modulated detection architecture (MDETR) trained with explanation-guided learning across the divergent datasets. In \(D_{\text{spur}}\), target-matching objects are strictly confined to the valid target region, ensuring \(P(Y_{\text{ans}} \mid \text{Target\_Attrs}) = 1.0\) and embedding the Bag-of-Words shortcut. Conversely, in \(D_{\text{pure}}\), adversarial confusers collapse the shortcut predictive probability to 36.07%.

The models are trained under a two-phase regime: Phase 1 optimizes visual grounding via bounding box alignment (\(L_{\text{ground}} = \lambda_{L_1} L_1 + \lambda_{\text{giou}} L_{\text{giou}}\)), and Phase 2 optimizes answer prediction using a dynamically weighted cross-entropy loss \(L_{\text{QA}}\) to neutralize answer prior bias (Case-0). When tested on \(D_{\text{pure}}\), the shortcut-trained model \(M_{\text{spur}}\) catastrophically fails on multi-hop relational queries (dropping to 8% accuracy on Depth-2 Mixed and 14% on Depth-3 Mixed queries), yet maintains 100% on non-relational Attribute-Only queries. This verified behavioral divergence creates an absolute benchmark: a faithful explainer must produce divergent attributions for \(M_{\text{pure}}\) and \(M_{\text{spur}}\), exposing the latter's complete disregard of spatial relations.

4. Multimodal Attribution Protocol & Additive Fallacy Diagnosis: Measuring Precision and Interaction Monotonicity

To evaluate diverse explainers, GridVQA-X adapts local and global metrics into a unified protocol. For local attribution heatmaps (Dime, MultiSHAP, MultiViz), in addition to Otsu-binarized Intersection over Union (IoU), the framework evaluates Relevance Mass Accuracy (RMA):

\[\text{RMA}(I_{\text{map}}, \text{Mask}) = \frac{\sum_{p \in \text{Mask}} |I_{\text{map}}(p)|}{\sum_{p \in V} |I_{\text{map}}(p)|}\]

A faithful explainer must assign high RMA on ground-truth masks for \(M_{\text{pure}}\), while allocating minimal mass to valid anchors when evaluating \(M_{\text{spur}}\). For global explainers producing a cross-modal synergy scalar \(S_{\text{score}}\) (EMAP, InterSHAP, PID), the authors introduce the Additive Fallacy Check. True cross-modal synergy must monotonically increase with relational depth (\(D=1 \to 3\)) for \(M_{\text{pure}}\), as each additional anchor demands higher compositional integration, while remaining low and depth-invariant for \(M_{\text{spur}}\). This diagnostic tests whether global explainers measure actual cross-modal synergy or merely capture unimodal feature sensitivity.

Key Experimental Results

Main Results

The framework benchmarks state-of-the-art local attribution algorithms across visual and textual modalities. Evaluations are conducted under Pure Evaluation (\(M_{\text{pure}}\) on \(D_{\text{pure}}\)), Spurious Evaluation (\(M_{\text{spur}}\) on \(D_{\text{spur}}\)), and Cross-Evaluation (\(M_{\text{spur}}\) on \(D_{\text{pure}}\)). The following table summarizes the performance on visual and textual modalities:

Modality Method Pure Eval RMA (%) (↑) Pure Eval IoU (%) (↑) Spur. Eval RMA (%) (↑) Spur. Eval IoU (%) (↑)
Visual Dime \(28.0 \pm 13.9\) \(20.5 \pm 10.2\) \(27.2 \pm 14.5\) \(18.2 \pm 10.0\)
Visual MultiSHAP \(56.0 \pm 22.7\) \(50.0 \pm 22.6\) \(64.5 \pm 21.8\) \(46.2 \pm 23.1\)
Visual MultiViz \(44.0 \pm 16.3\) \(4.7 \pm 3.3\) \(42.2 \pm 20.0\) \(6.7 \pm 4.5\)
Textual Dime \(43.0 \pm 26.8\) \(12.3 \pm 12.5\) \(37.5 \pm 25.9\) \(10.3 \pm 11.5\)
Textual MultiSHAP \(66.8 \pm 17.3\) \(46.7 \pm 18.9\) \(67.5 \pm 16.1\) \(40.2 \pm 19.9\)

Ablation Study & Complexity Scaling Analysis

To evaluate whether global explainability metrics scale faithfully with relational complexity, EMAP's synergy scalar \(S_{\text{score}}\) is tracked across relational depths (\(D \in \{1, 2, 3\}\)). In addition, the predictive accuracy of paired diagnostic models and open-source foundation VLMs on multi-hop mixed queries (Mixed, Form 0, \(\rho=d0.7\)) is benchmarked:

Evaluation Regime / Model Config Depth 1 Metric Depth 2 Metric Depth 3 Metric Trend / Analysis
EMAP Pure Eval Synergy (%) \(78.6 \pm 16.7\) \(74.7 \pm 23.8\) \(64.1 \pm 28.2\) Violates additive check: decreases as reasoning depth increases
EMAP Spur. Eval Synergy (%) \(69.2 \pm 23.9\) \(65.7 \pm 24.4\) \(68.8 \pm 18.9\) Hallucinates deep synergy on a pure Bag-of-Words shortcut model
EMAP Cross Eval Synergy (%) \(59.4 \pm 27.3\) \(54.0 \pm 26.2\) \(65.0 \pm 17.6\) Assigns ~60% synergy even when the model's accuracy collapses
Reference Model \(M_{\text{pure}}\) Acc (%) 100.0 100.0 100.0 Robustly learns true multi-hop spatial geometry
Reference Model \(M_{\text{spur}}\) Acc (%) 6.0 2.0 0.0 Fails completely on out-of-distribution adversarial spatial distractors
Open-Source Qwen3-VL-30B Acc (%) 14.0 10.0 10.0 Large foundation VLM degrades under dense spatial distractors
Open-Source Llava-1.5-7B Acc (%) 6.0 28.0 30.0 Weak spatial grounding in multi-hop compositional reasoning

Key Findings

  • Model Blindness and Accidental Faithfulness: Current local explainers fail to diagnose shortcut learning. MultiViz yields nearly identical RMA (\(\approx 44\%\) vs. \(42\%\)) across \(M_{\text{pure}}\) and \(M_{\text{spur}}\), displaying total model blindness. Dime produces highly diffuse, unconstrained heatmaps that accidentally overlap with target areas, generating an illusion of faithfulness.
  • Game-Theoretic Preference for Unary Feature Detectors: MultiSHAP paradoxically achieves higher RMA on the shortcut model \(M_{\text{spur}}\) (64.5%) than on the faithful model \(M_{\text{pure}}\) (56.0%). Because Shapley value marginals align naturally with independent 1-to-1 feature matching, game-theoretic methods structurally favor shallow shortcuts over non-linear geometric intersections.
  • Failure of the Additive Fallacy Check in Global Estimators: EMAP's estimated synergy on \(M_{\text{pure}}\) paradoxically deteriorates from 78.6% at Depth 1 down to 64.1% at Depth 3, falsely implying that tri-anchor queries require less cross-modal interaction. On \(M_{\text{spur}}\), it hallucinates \(\sim 69\%\) synergy despite the model bypassing directional tokens. InterSHAP displays an identical collapse (dropping from 0.921 to 0.629).
  • Distractor Leakage on Existence Tasks: When evaluating existence queries (Form 1), MultiSHAP's RMA suffers severe degradation because it highlights all objects matching target attributes regardless of whether they violate spatial constraints, acting as a blind attribute detector rather than a relational explainer.

Highlights & Insights

  • Mathematically Guaranteed Ground-Truth Explanations: By formulating scene generation through Pearl's do-calculus, the framework eliminates subjective human annotation ambiguity and provides the first provably exact causal mask for multimodal interactions.
  • Controlled Behavioral Divergence via Paired Models: Rather than guessing what a black-box model is thinking, the authors train identical architectures into provably divergent reasoning regimes (\(M_{\text{pure}}\) vs. \(M_{\text{spur}}\)), creating a definitive diagnostic ground truth for XAI methods.
  • Deconstruction of the Multimodal Interpretability Illusion: The findings demonstrate that popular explainability algorithms frequently generate plausible-looking rationales on models operating entirely via shallow heuristics, highlighting an urgent need for causally faithful evaluation benchmarks.

Limitations & Future Work

  • Limitations Admitted by the Authors: GridVQA-X operates in a synthetic 2D geometric universe that lacks the complex textures and continuous semantic distributions of natural images. Furthermore, the taxonomy is confined to spatial-relational composition and basic attribute binding, excluding temporal dynamics, physical commonsense, or mathematical logic.
  • Additional Limitations: The evaluation is centered on MDETR, an object-centric detection transformer; it does not directly evaluate the internal attention rationales or generative chain-of-thought (CoT) text tokens produced by autoregressive MLLMs.
  • Future Directions: Extending the adversarial confuser placement logic into continuous 3D embodied simulation environments and developing new explainability algorithms capable of passing this zero-ambiguity diagnostic testbed.
  • vs CLEVR-XAI: CLEVR-XAI provides ground-truth masks for synthetic 3D scenes but evaluates only unimodal visual neural networks, lacking multimodal interaction metrics and paired shortcut models; GridVQA-X evaluates cross-modal synergy with mathematical guarantees.
  • vs OpenXAI: OpenXAI is a pioneering transparent XAI benchmark for tabular and image domains, but is strictly unimodal; GridVQA-X fills the void for cross-modal interaction benchmarking.
  • vs Dime / MultiSHAP / MultiViz: While these methods claim to isolate multimodal interactions from unimodal contributions, GridVQA-X empirically proves that they fail to distinguish between true spatial reasoning and Bag-of-Words shortcuts.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ First diagnostic framework providing mathematically guaranteed ground truth and paired models for multimodal explainability.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 6 state-of-the-art local and global explainers across vision and text modalities.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous causal formulations, clean mathematical proofs, and well-structured empirical analyses.
  • Value: ⭐⭐⭐⭐⭐ Exposes fundamental blind spots in multimodal XAI and sets a new rigorous standard for evaluating cross-modal faithfulness.