Skip to content

Wavelet-based Intra-video Counterfactual Reasoning for Video Question Grounding

Conference: ECCV 2026
Paper: ECCV 2026
Area: VLM Reasoning
Keywords: Video Question Grounding / Causal Intervention / Intra-video Counterfactual Mining / Wavelet Dynamic Modeling / Weakly-supervised Learning

TL;DR

To tackle intra-video temporal bias caused by visually identical yet non-causal frames in weakly-supervised Video Question Grounding, WICI introduces wavelet-based dynamic modeling to capture fine-grained motion cues and curriculum-guided intra-video counterfactual mining with margin-based supervision for faithful evidence localization and robust reasoning.

Background & Motivation

Video Question Grounding (VideoQG) requires models to answer natural language queries while temporally localizing the supporting visual evidence within videos. Under the practical weakly-supervised setting where explicit temporal boundary annotations are unavailable during training, existing methods predominantly rely on pretrained vision-language backbones and cross-modal attention mechanisms. These approaches typically employ multiple instance learning or cross-modal contrastive alignment to discover evidence intervals indirectly. However, because pretrained visual encoders focus primarily on static appearance semantics, models are prone to learning superficial correlations and dataset-level statistical shortcuts rather than performing genuine multi-modal reasoning.

While prevailing debiasing techniques target cross-sample statistical correlations (such as language priors or cross-modal co-occurrence), they consistently overlook a fundamental vulnerability: intra-video temporal bias. A video inherently consists of consecutive or recurring frames sharing near-identical visual appearances, backgrounds, and entities, yet their causal relevance to answering a specific question varies substantially. In standard benchmarks like NeXT-GQA, irrelevant background or misleading distractor segments dominate the majority of video durations. When passed through appearance-oriented vision-language encoders, causally essential frames and non-causal distractors become indistinguishable in the embedding space. Formulated under Structural Causal Models (SCM), this entanglement creates a spurious shortcut path from non-causal video segments to the final prediction (\(N \to V \to a\)), which explains why models often remain indifferent when video frame order is randomly shuffled.

Eliminating this spurious dependency through conventional representation-level orthogonal disentanglement or rigid invariance constraints is challenging, as the high visual similarity among temporally adjacent frames inside a single video introduces severe optimization instability. The core idea is to shift from rigid representation disentanglement to decision-centric preference modeling and frequency-domain dynamic enhancement: proposing the WICI framework, which applies multi-scale wavelet decomposition to isolate subtle inter-frame motion details from static contexts, while employing curriculum-guided intra-video counterfactual mining coupled with a margin-based discrimination objective to enforce a strict prediction confidence hierarchy among causal, reference, and counterfactual video inputs.

Method

Overall Architecture

The WICI pipeline takes uniformly sampled video frames and natural language questions with candidate choices as input, and outputs the predicted answer along with the optimal temporal grounding interval. Visual frames are embedded via a frozen CLIP-ViT, while RoBERTa independently encodes the question query and concatenated question-choice candidates. To overcome the static appearance bias of the visual backbone, frame embeddings are first processed by the Wavelet-based Dynamic Modeling (WDM) module, which decomposes temporal features into multi-level frequency bands to isolate fine-grained motion dynamics and re-injects them into the representations. A temporal Transformer encoder coupled with a Gaussian smoothing grounding module then computes query-conditioned temporal attention over the video frames. Using this attention distribution, the Intra-video Counterfactual Mining (ICM) module constructs curriculum-guided boundary counterfactual proposals and a temporal-shuffled counterfactual branch. Finally, a margin-based discrimination loss quantifies prediction margins across causal, full-video reference, and counterfactual inputs, optimized jointly with task classification and cross-modal contrastive objectives.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Input: Video frames & question-choice text"] --> WDM["Wavelet-based Dynamic Modeling<br/>Multi-scale DWT separating dynamics and context"]
    WDM --> TempEnc["Temporal Encoding & Gaussian Grounding<br/>Global temporal modeling & attention estimation"]
    TempEnc --> ICM["Intra-video Counterfactual Mining<br/>Curriculum-guided boundary & shuffled counterfactuals"]
    ICM --> MarginLoss["Margin-based Causal Discrimination<br/>Quantifying prediction margins across segments"]
    MarginLoss --> Out["Output: Predicted answer & compact grounding interval"]

Key Designs

1. Wavelet-based Dynamic Modeling: Multi-scale temporal decomposition via discrete wavelet filtering Standard visual encoders emphasize static visual semantics and suppress subtle temporal transitions, leaving the model blind to the dynamic shifts that distinguish causal action moments from visually identical non-causal states. To capture critical motion cues with minimal computational overhead, WDM introduces a lightweight temporal decomposition layer based on the Discrete Wavelet Transform (DWT). Implemented via 1D temporal group convolution initialized with Haar wavelet filters and a stride of 2, WDM separates video features into a low-frequency approximation signal and a high-frequency detail signal. The approximation component is recursively decomposed across hierarchical levels to capture short-range and long-range dynamics, while detail signals across all levels are up-sampled via interpolation and fused into the original frame embeddings with a residual weight of 0.3. This explicit frequency separation enriches the video representation with temporal transition cues while maintaining rich static semantic context.

2. Intra-video Counterfactual Mining: Curriculum-guided progressive spatiotemporal counterfactuals Cross-instance negatives fail to eliminate intra-video shortcuts because non-causal frames inside the same video share identical scenes and actors. ICM constructs two complementary intra-video counterfactual interventions: frame-based boundary counterfactuals and temporal-shuffled counterfactuals. For frame-based counterfactuals, segments immediately adjacent to the causal interval share visual context and motion residues, presenting the hardest distractors; introducing them early during training disrupts initial localization convergence. ICM computes the temporal center and variance of the estimated causal interval from frame attention weights, anchors counterfactual proposal windows at the video boundary edges, and schedules their inward shift toward the causal center via a curriculum pacing function \(\rho(e) = (e / E_{\max})^\gamma\). Combined with a parallel temporal-shuffled counterfactual branch that breaks chronological ordering while preserving frame appearance, ICM isolates non-causal factors without manual boundary annotations.

3. Margin-based Causal Discrimination: Quantitative ranking hierarchy against non-causal shortcuts Existing causal debiasing methods frequently rely on loose inequality objectives or strict representation invariance, which suffer from numerical instability when intra-video segments are visually homogenous. To provide robust and quantifiable supervisory signals, WICI establishes a strict posterior probability hierarchy: the prediction confidence conditioned on the causal segment must strictly exceed that of the full video reference, which in turn must strictly surpass that of counterfactual segments. Formulating this preference via cross-entropy losses under causal (\(v_t\)), reference (\(v\)), and counterfactual (\(v_{cf}\)) representations, the margin-based discrimination objective is defined as: $\(\mathcal{L}_{\mathrm{mdl}} = \max(0, \mathcal{L}_{ce}(v_t) - \mathcal{L}_{ce}(v) + \lambda_1) + \max(0, \mathcal{L}_{ce}(v) - \mathcal{L}_{ce}(v_{cf}) + \lambda_2)\)$ Enforcing \(\lambda_1 < \lambda_2\) (set to 0.15 and 0.20 empirically) guarantees a stronger separation margin between causal evidence and counterfactual distractors than between causal and global reference, penalizing shortcut paths and driving the model toward compact, reliable grounding.

Loss & Training

WICI is optimized end-to-end under weak supervision using a composite objective: $\(\mathcal{L} = \mathcal{L}_{qa} + \alpha \mathcal{L}_{align} + \beta \mathcal{L}_{mdl}\)$ where \(\mathcal{L}_{qa}\) represents the cross-entropy classification loss on the causal representation \(v_t\), \(\mathcal{L}_{align}\) is an InfoNCE contrastive alignment loss aligning grounded video features with linguistic representations across in-batch negatives, and \(\alpha, \beta\) are balancing hyperparameters (\(\alpha = 1.0, \beta = 0.3\)). The model is trained using AdamW with an initial learning rate of \(1 \times 10^{-5}\) on an NVIDIA RTX 4090 GPU. The RoBERTa text encoder, WDM layer, and temporal Transformer are optimized while the CLIP-L visual backbone remains frozen. Wavelet decomposition depth is set to 2, and the curriculum pacing parameter is \(\gamma = 0.5\).

Key Experimental Results

Main Results

Quantitative evaluations on the weakly-supervised VideoQG benchmarks NeXT-GQA (test set) and STAR (validation set) demonstrate the clear superiority of WICI:

Dataset Method Vision Backbone Text Backbone Acc@GQA (%) Acc@VQA (%) mIoP (%) [email protected] (%) [email protected] (%)
NeXT-GQA SeViLA* ViT-G (BLIP-2) Flan-T5-XL 16.6 68.1 29.5 22.9 13.8
NeXT-GQA IGV ResNet BERT 10.2 50.1 21.4 18.9 9.6
NeXT-GQA TempCLIP ViT-L RoBERTa 16.0 60.2 25.7 25.5 8.9
NeXT-GQA TimeCraft ViT-L RoBERTa 18.2 - 28.1 27.8 9.6
NeXT-GQA TempCLIP ViT-L RoBERTa 18.2 61.1 28.6 28.5 10.6
NeXT-GQA WICI (Ours) ViT-L RoBERTa 19.0 61.5 28.8 29.4 10.2
STAR SeViLA* ViT-G (BLIP-2) Flan-T5-XL 17.0 45.5 - 37.6 19.5
STAR IGV ResNet BERT 13.1 31.7 - 42.8 1.7
STAR TempCLIP ViT-L RoBERTa 24.4 57.3 - 41.4 4.7
STAR TempCLIP ViT-L RoBERTa 26.8 58.6 - 44.5 5.5
STAR WICI (Ours) ViT-L RoBERTa 29.2 58.8 - 48.5 5.9

Ablation Study

Ablation experiments on NeXT-GQA isolate the impact of individual architectural components and design choices:

Config Acc@GQA (%) Acc@VQA (%) [email protected] (%) [email protected] (%) Note
Full model (WICI) 19.0 61.5 29.4 10.2 Full framework with all components
w/o WDM 17.8 (โ†“1.2) 60.5 (โ†“1.0) 28.6 (โ†“0.8) 9.5 (โ†“0.7) Omits fine-grained wavelet dynamics
w/o ICM 17.7 (โ†“1.3) 60.8 (โ†“0.7) 28.4 (โ†“1.0) 7.1 (โ†“3.1) Disables counterfactual sample mining
Baseline Align 17.3 (โ†“1.7) 60.1 (โ†“1.4) 28.1 (โ†“1.3) 7.5 (โ†“2.7) Standard cross-modal alignment baseline
w/o CM (fixed \(\rho=0.5\)) 18.2 (โ†“0.8) 60.7 (โ†“0.8) 28.8 (โ†“0.6) 8.7 (โ†“1.5) Replaces curriculum with static distance
w/o TO-CF 18.2 (โ†“0.8) 60.9 (โ†“0.6) 28.8 (โ†“0.6) 8.9 (โ†“1.3) Removes temporal-shuffled branch
w/o ML 18.3 (โ†“0.7) 61.2 (โ†“0.3) 29.0 (โ†“0.4) 8.6 (โ†“1.6) Replaces margin loss with contrastive loss

Key Findings

  • Dual-Level Synergy: Removing ICM drops Acc@GQA by 1.3%, while disabling WDM drops Acc@GQA by 1.2%. The substantial compound improvement confirms that representation-level dynamic enrichment and decision-level causal supervision are mutually reinforcing.
  • Necessity of Curriculum Pacing: Switching from curriculum-guided counterfactual mining to a fixed sampling distance drops Acc@GQA by 0.8%, proving that aggressive near-evidence counterfactual perturbations early in training destabilize attention convergence.
  • True Temporal Sensitivity: Under test-time frame shuffling, WICI incurs a 4.0% drop in Acc@GQA (19.0% โ†’ 15.0%) compared to a 3.4% drop for the baseline. This heightened sensitivity confirms that WICI genuinely relies on temporal event progression rather than static appearance shortcuts.
  • Compact Causal Evidence: WICI substantially reduces Bias Error (from 28.5% to 27.1%) and Unfaithful predictions (from 41.4% to 39.6%). By penalizing non-causal distractors, the model actively prioritizes high-confidence, minimal-evidence intervals, leading to notable IoP gains over broad-window IoU metrics.

Highlights & Insights

  • From Cross-dataset to Intra-video Debiasing: Pioneered the formalization of intra-video temporal bias in VideoQG, targeting the critical dilemma where visually identical frames in the same video exert vastly distinct causal influences on question answering.
  • Frequency-Domain Fusion with Causal Ranking: Seamlessly integrated classical 1D discrete wavelet filtering into modern vision-language representations and formulated causal prioritization as a numerically stable margin-ranking hierarchy, bypassing the optimization instability of rigid feature disentanglement.
  • Transferable Causal Paradigm: The combination of curriculum-guided boundary mining and multi-tier margin supervision is model-agnostic and readily adaptable to video moment retrieval, temporal action localization, and video LLM hallucination mitigation.

Limitations & Future Work

  • Single Continuous Evidence Assumption: The current boundary counterfactual sampling formulation assumes that questions depend on a single continuous temporal interval, which may limit performance on complex multi-hop or multi-event queries.
  • Sensitivity to High-Frequency Noise: Increasing wavelet decomposition depth beyond 2 levels degrades global semantic coherence due to excessive high-frequency amplification, requiring careful hyperparameter tuning.
  • Future Directions: Exploring mixture-based multi-span counterfactual generation and dynamic frequency-band weighting for open-ended video LLM reasoning.
  • vs CRA (CVPR 2025): CRA targets cross-modal and textual dataset-level bias via front-door and back-door adjustments; WICI focuses on the orthogonal challenge of intra-video temporal bias using wavelet dynamics and counterfactual mining, outperforming CRA with identical backbones.
  • vs IGV / VCSR: IGV enforces strict representation invariance or explicit causal/non-causal feature separation, which is prone to collapse on homogenous adjacent frames; WICI avoids rigid feature orthogonalization and instead enforces stable preference ranking in prediction space.

Rating

  • Novelty: โญโญโญโญโ˜† Innovatively identifies intra-video temporal bias and bridges wavelet dynamics with curriculum causal reasoning.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluations across NeXT-GQA and STAR benchmarks, supported by temporal perturbation and error decomposition studies.
  • Writing Quality: โญโญโญโญโญ Clear mathematical formulation, intuitive causal DAG illustrations, and coherent methodology.
  • Value: โญโญโญโญโ˜† Offers an efficient, plug-and-play debiasing framework for weakly-supervised temporal video-language reasoning.