Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: PDF
Area: Multimodal VLM
Keywords: Zero-Shot Video Moment Retrieval, Temporal Self-Similarity, Multimodal Large Language Models, Modality Gap, Language-Style Gap
TL;DR¶
Addressing the vulnerability of zero-shot video moment retrieval to both modality and language-style gaps in query-video matching, this paper proposes Self-Similarity-based Moment proposal and Scoring (Self-SiMS), which leverages intrinsic intra-video self-similarity to generate proposals and score contextual consistency, followed by MLLM-based query-aware binary re-ranking to achieve new state-of-the-art performance across benchmarks.
Background & Motivation¶
Natural language-driven video moment retrieval (VMR) aims to localize specific temporal segments in untrimmed videos corresponding to a natural language query. Traditional fully supervised approaches rely heavily on precise human-annotated start and end boundaries, which are labor-intensive, expensive, and difficult to scale to diverse real-world domains. Consequently, training-free zero-shot video moment retrieval (ZMR) utilizing pre-trained vision-language models (VLMs) and multimodal large language models (MLLMs) has attracted substantial research interest. However, existing ZMR methods predominantly rely on similarity scores between the query and video contents for candidate span proposal and scoring, rendering them highly sensitive to cross-modal and linguistic representation discrepancies.
This bottleneck stems from two fundamentally unresolved representation gaps: first, methods calculating direct similarity between queries and visual frame features suffer from the modality gap, where visual and textual embeddings occupy distinct representations; even within ground-truth intervals, similarity signals remain low, flat, and indistinguishable from background noise, frequently producing excessively long and coarse proposals. Second, approaches that convert visual frames into dense captions to evaluate text-to-text similarity introduce an equally detrimental language-style gap; human-written queries operate at an abstract intent level, whereas MLLM-generated captions focus on disparate granular details across adjacent frames (such as alternating between "holding a tool" and "reaching outward"), causing similarity trajectories to fluctuate sharply and collapse into fragmented short intervals. Quantitative analysis through the Inner-to-Outer Ratio (IOR)—which measures similarity within ground-truth moments relative to surrounding backgrounds—demonstrates that existing methods fail to maintain stable discriminative margins.
This paper tackles the problem from a distinct angle: since query-dependent similarity signals are inherently corrupted during proposal generation, temporal event partitioning should be decoupled from the query text and anchored entirely in the intrinsic temporal coherence of the video itself. Continuous human activities exhibit natural internal consistency across neighboring frames, providing a robust, query-agnostic prior for event transitions. Core idea: propose Self-Similarity-based Moment proposal and Scoring (Self-SiMS), which derives candidate temporal boundaries solely from a joint visual-caption self-similarity matrix via contrastive kernel filtering, evaluates proposals via peak query matching combined with keyframe-anchored internal consistency, and refines final rankings through query-aware MLLM binary reasoning.
Method¶
Overall Architecture¶
Given an untrimmed video \(V = \{v_i\}_{i=1}^L\) with \(L\) sampled frames and a natural language query \(Q\), the framework executes in three training-free stages: first, frame visual features and MLLM-generated frame-level caption features are extracted to construct a unified Temporal Self-similarity Matrix (TSM), which is filtered with a contrastive kernel to detect transition boundaries and produce candidate spans; second, each span receives a composite score fusing its top-\(k_S\) peak query-caption cosine similarities with an internal self-similarity consistency score centered on the keyframe; finally, the top-\(k_C\) candidate spans undergo query-aware verification where sampled representative frames are evaluated by an MLLM via Yes/No binary prompt logits to determine the final re-ranked interval.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Untrimmed Video & Text Query"] --> B["Self-Sim based Span Generation<br/>Joint Visual+Caption TSM & Contrastive Kernel"]
B --> C["Scoring with Self-Sim Awareness<br/>Top-kS Peak Matching & Keyframe Self-Consistency"]
C --> D["Query-Aware MLLM Re-Ranking<br/>Sample Representative Frames for Yes/No Verification"]
D --> E["Output High-Confidence Moment Interval"]
Key Designs¶
1. Self-similarity-based candidate span generation: decoupling query signals to eliminate modality and style bias Prior candidate proposal generators rely on query-frame or query-caption similarity trajectories, rendering proposal boundaries vulnerable to cross-modal misalignment or caption phrasing noise. To fundamentally avoid query corruption, candidate spans are derived exclusively from intrinsic video relationships. Frame visual features \(F^f\) and MLLM caption embeddings \(F^c\) are \(\ell_2\)-normalized to construct frame and caption temporal self-similarity matrices: \(M^f = \hat{F}^f (\hat{F}^f)^\top\) and \(M^c = \hat{F}^c (\hat{F}^c)^\top\). These are averaged into a unified matrix \(M = \frac{1}{2}(M^f + M^c) \in \mathbb{R}^{L \times L}\). To detect natural event boundaries, a contrastive kernel \(K \in \mathbb{R}^{N_K \times N_K}\) is applied along the diagonal of \(M\): positive weights in the top-left and bottom-right quadrants reward intra-segment coherence, negative weights in the off-diagonal quadrants penalize cross-boundary similarity, and the central cross is zeroed. Convolving \(K\) with local patches \(P_i\) centered at \((i, i)\) yields a boundary score \(b_i = \sum (P_i \odot K)\). Dynamic thresholds \(\mathcal{T}\) conditioned on video visual variance extract local maxima as robust temporal boundaries, forming non-overlapping candidate spans \(\mathcal{E} = \{E_1, E_2, \dots, E_{N_s}\}\) that match natural event transitions without query contamination.
2. Self-similarity aware span scoring: balancing peak semantic alignment and contextual consistency Conventional methods average frame-level query-caption similarities across an entire candidate span. Due to the language-style gap, valid moment frames frequently receive low scores when caption details diverge from the query phrasing, unfairly depressing the entire span's average. To resolve this, the Query-Matching Span Score (QMS) computes the mean of only the top-\(k_S\) highest frame-level similarities within span \(E_n\): $\(S^Q_n = \frac{1}{k_S} \sum_{i \in I_{k_S}^{E_n}} S^f_i\)$ This preserves the most salient semantic alignment signals while filtering out caption style noise. However, relying solely on isolated peak frames risks selecting overly broad intervals containing semantic drift. To enforce temporal cohesiveness, the Self-Matching Span Score (SMS) identifies the peak frame as a keyframe anchor \(\kappa_n = \arg\max_{i \in E_n} S^f_i\), and computes the average self-similarity between \(\kappa_n\) and all other frames in \(E_n\) via the unified TSM \(M\): $\(S^S_n = \frac{1}{|E_n|} \sum_{j \in E_n} M_{\kappa_n, j}\)$ Spans that extend into irrelevant scenes suffer lower self-similarity, effectively penalizing temporal over-extension. The composite score combines both terms via weighting parameter \(\alpha\) (set to 0.1): \(S_n = (1 - \alpha) S^Q_n + \alpha S^S_n\), ensuring both strong semantic relevance and tight internal coherence.
3. Query-aware MLLM re-ranking: deep cross-modal binary verification Shallow cosine similarities over pre-trained embeddings struggle with fine-grained visual reasoning and complex compositional queries. To perform rigorous cross-modal verification, an MLLM-based re-ranking stage is applied to the top-\(k_C\) (default \(k_C = 5\)) candidate spans. From each selected span, \(N_R\) representative frames are adaptively sampled according to their query-caption scores (with \(N_R\) bounded between 10 and 30 frames). Each frame is paired with the query and evaluated by an instruction-tuned vision-language model (e.g., LLaMA-3.2-11B-Vision-Instruct) prompted to answer whether the image depicts the query description. Applying Softmax to the output logits extracts the "Yes" generation probability \(S^r_i\). The mean of the top-\(k_R\) frame probabilities forms the span re-ranking score \(S_{(m)}^R\), which linearly blends with initial score \(S_{(m)}\) using \(\beta = 0.5\): $\(S_{(m)}^* = (1 - \beta) S_{(m)} + \beta S_{(m)}^R\)$ Direct end-to-end multimodal reasoning bridges shallow representation discrepancies and provides definitive discrimination among competing candidate spans.
Loss & Training¶
The framework is completely training-free and operates without parameter optimization. Key operational parameters include: - Sampling Rates: 0.5 fps for QVHighlights; 1.0 fps for Charades-STA, ActivityNet-Captions, and TVR. - Hyperparameter Settings: Peak frame counts \(k_S = 3, k_R = 3\); candidate span pool \(k_C = 5\); fusion weights \(\alpha = 0.1, \beta = 0.5\). - Runtime Efficiency: Non-MLLM operations (TSM construction, contrastive filtering, proposal generation, and composite scoring) require only 0.194 seconds per video, demonstrating negligible computational overhead before MLLM evaluation.
Key Experimental Results¶
Main Results¶
Performance comparison against fully supervised (FS), weakly supervised (WS), and zero-shot (ZS) baselines across QVHighlights, Charades-STA, and ActivityNet-Captions:
| Dataset | Method | Setting | [email protected] | [email protected] | [email protected] | [email protected] | mAP@avg / mIoU |
|---|---|---|---|---|---|---|---|
| QVHighlights (val) | Moment-DETR [15] | FS | - | 54.2 | 33.4 | 55.4 | 31.1 (mAP@avg) |
| QVHighlights (val) | TFVTG‡ [41] | ZS | - | 21.0 | 7.4 | 19.2 | 7.3 (mAP@avg) |
| QVHighlights (val) | Moment-GPT [38] | ZS | - | 58.9 | 38.6 | 55.7 | 35.9 (mAP@avg) |
| QVHighlights (val) | Self-SiMS (Ours) | ZS | - | 61.0 | 43.2 | 60.5 | 39.3 (mAP@avg) |
| QVHighlights (test) | Moment-DETR [15] | FS | - | 52.9 | 33.0 | 54.8 | 30.7 (mAP@avg) |
| QVHighlights (test) | UMT [21] | FS | - | 56.4 | 40.8 | 53.1 | 35.4 (mAP@avg) |
| QVHighlights (test) | TFVTG‡ [41] | ZS | - | 21.6 | 8.0 | 20.1 | 7.9 (mAP@avg) |
| QVHighlights (test) | Moment-GPT [38] | ZS | - | 58.3 | 37.7 | 55.1 | 35.0 (mAP@avg) |
| QVHighlights (test) | Self-SiMS (Ours) | ZS | - | 59.7 | 42.2 | 59.2 | 38.3 (mAP@avg) |
| Charades-STA | Moment-DETR [15] | FS | 62.1 | 48.2 | 25.3 | - | 42.3 (mIoU) |
| Charades-STA | TFVTG‡ [41] | ZS | 64.8 | 30.0 | 10.1 | - | 38.4 (mIoU) |
| Charades-STA | Moment-GPT [38] | ZS | 58.2 | 38.4 | 21.6 | - | 36.5 (mIoU) |
| Charades-STA | Self-SiMS (Ours) | ZS | 62.7 | 39.7 | 21.0 | - | 41.9 (mIoU) |
| ActivityNet-Captions | Moment-DETR [15] | FS | 52.6 | 32.5 | 15.3 | - | 37.8 (mIoU) |
| ActivityNet-Captions | TFVTG‡ [41] | ZS | 49.5 | 26.5 | 12.3 | - | 34.0 (mIoU) |
| ActivityNet-Captions | Moment-GPT [38] | ZS | 48.1 | 31.1 | 14.9 | - | 30.8 (mIoU) |
| ActivityNet-Captions | Self-SiMS (Ours) | ZS | 49.9 | 28.2 | 13.8 | - | 34.7 (mIoU) |
Ablation Study¶
Candidate proposal quality evaluation (Oracle mIoU on QVHighlights val): | Method | Proposal Generation Basis | Avg. # Candidate Spans | Oracle-mIoU (%) | Note | |--------|---------------------------|------------------------|-----------------|------| | TFVTG [41] | Cross-modal Query-Frame matching | 8.38 | 38.16 | Modality gap causes flat signals and overly long spans | | Moment-GPT [38] | Cross-modal Query-Caption matching | 7.62 | 65.01 | Language-style gap causes unstable peaks and fragmented spans | | Self-SiMS (Ours) | Intra-video Self-Similarity (Query-Agnostic) | 6.46 | 71.13 | Achieves highest upper bound (+6.12%) with fewest proposals |
Scoring and re-ranking component breakdown (QVHighlights val): | Config | [email protected] | [email protected] | [email protected] | mAP@avg | Note | |--------|--------|--------|---------|---------|------| | Mean scoring | 55.7 | 40.9 | 56.7 | 37.1 | Baseline: average across all frames in span | | + Query-Matching Span Score (QMS) | 59.5 | 42.1 | 59.4 | 38.6 | Top-\(k_S\) peak frames eliminate caption style noise | | + Self-Matching Span Score (SMS) | 60.3 | 42.9 | 59.6 | 38.7 | Keyframe self-similarity penalizes context drift | | + MLLM Re-ranking (Full) | 61.0 | 43.2 | 60.5 | 39.3 | Query-aware Yes/No probability verification |
Key Findings¶
- Proposal quality dictates the upper bound: Self-SiMS achieves 71.13% Oracle mIoU with only 6.46 candidate spans on average, outperforming Moment-GPT's 65.01% with 7.62 spans. Generating proposals via self-similarity without query corruption yields tighter, more faithful temporal boundaries.
- Complementary scoring dynamics: Moving from uniform mean scoring to QMS provides the largest performance jump (+3.8% [email protected], +2.7% [email protected]), validating that peak frames are far more resilient to captioning noise than global averages. Adding SMS further boosts strict localization metrics (+0.8% [email protected]), penalizing inconsistent background frames.
- Architectural efficiency: Moment-GPT requires three separate models for query rephrasing, captioning, and re-ranking (totaling 22B parameters). In contrast, Self-SiMS unifies captioning and binary re-ranking under a single 11B LLaMA-3.2-Vision backbone, cutting model footprint in half while re-ranking only a small set of representative frames in 13.0 seconds.
Highlights & Insights¶
- Decoupled proposal formulation: Prior ZMR methods routinely incorporated query similarity during span generation, introducing modality and style corruption into the boundaries. Framing proposal generation as query-agnostic intra-video self-similarity change detection represents an elegant, principled decoupling.
- Internal self-consistency regularization: Instead of relying solely on external query matching, the keyframe-to-span self-matching score (SMS) acts as an unsupervised temporal consistency regularizer, penalizing semantically drifted spans without manual boundary heuristics.
- Broad cross-task transferability: The contrastive kernel boundary detector operates on pre-computed feature similarity matrices without training and completes in less than 0.2 seconds per video, making it directly portable to unsupervised video summarization, scene segmentation, and temporal action proposal generation.
Limitations & Future Work¶
- Author-admitted limitations: While re-ranking is restricted to sampled frames across top-5 spans, processing ultra-long videos (e.g., full television episodes in TVR) still incurs noticeable end-to-end latency from dense frame captioning and MLLM prompting.
- Potential blind spots: The unified self-similarity matrix weights visual features and caption features equally (0.5 / 0.5), lacking dynamic context-aware weighting for scenarios dominated by static backgrounds or rapid visual scene transitions.
- Future directions: Exploring frequency-domain filtering on the self-similarity matrix could smooth high-frequency visual noise, while packaging sampled frames into structured multi-image grid layouts could reduce MLLM re-ranking queries into a single inference pass.
Related Work & Insights¶
- vs TFVTG [41]: TFVTG computes direct cosine similarities between text queries and visual frames using pre-trained VLMs and searches spans via sliding windows. It suffers heavily from the modality gap, producing flat similarity curves and overly broad intervals; Self-SiMS decouples proposal generation via self-similarity and utilizes MLLM binary reasoning, significantly improving localization accuracy and mIoU.
- vs Moment-GPT [38]: Moment-GPT leverages MLLM-generated frame captions to circumvent the modality gap, but exposes retrieval to the language-style gap, resulting in erratic similarity curves and fragmented intervals; Self-SiMS avoids query bias during proposal generation and stabilizes scoring via peak-matching and keyframe self-similarity.
Rating¶
- Novelty: ⭐⭐⭐⭐ [First work to formally analyze both modality and language-style gaps in training-free ZMR and introduce intra-video self-similarity for proposal generation and scoring]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across QVHighlights, Charades-STA, ActivityNet-Captions, and TVR, complete with IOR distributional analysis, Oracle mIoU comparisons, and runtime benchmarks]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear problem formulation, rigorous mathematical definitions, cohesive narrative structure, and informative visual diagrams]
- Value: ⭐⭐⭐⭐⭐ [Provides an efficient, plug-and-play paradigm for training-free video grounding that bridges the gap between vision-language representations and temporal localization]