QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/human-analysis/QSVideo
Area: Video Understanding
Keywords: video understanding / multimodal retrieval / temporal localization / vision-language model / long and streaming videos
TL;DR¶
To resolve biased relevance estimation, limited visual diversity, and temporal collapse in long and streaming video understanding, QSVideo decouples query intent into Object/Action/Location semantic importance via QSRanker and leverages QSRetrieval with L2 visual distance and adaptive temporal windowing to achieve state-of-the-art video QA under strict frame budgets.
Background & Motivation¶
Vision-language models (VLMs) have made tremendous strides in dynamic video understanding, yet their reasoning accuracy consistently degrades as video length scales up. Simply feeding more candidate frames into the LLM context window does not reliably improve downstream performance; instead, it incurs prohibitive GPU memory and compute footprints. Irrelevant background frames dilute the language model's attention allocation, triggering context distraction and visual hallucinations in long-form reasoning.
In contrast, human visual cognition adopts an active, hypothesis-driven retrieval paradigm: humans first dissect the core intent of a question, search for key supporting evidence across video segments, deliberately avoid re-examining visually redundant frames, and collect complementary clues across distant time periods. Prior retrieval and frame-selection techniques attempt to replicate this behavior, but suffer from three fundamental limitations: first, biased relevance estimation, where using original questions (especially multiple-choice questions packed with distractors) directly as queries misleads the ranker toward irrelevant frames; second, limited diversity, where selecting frames solely based on high relevance scores causes selections to cluster into visually near-identical shots; and third, temporal collapse, where selected evidence concentrates in narrow temporal pockets while critical clues remain dispersed across distant timestamps.
This work addresses these challenges by decoupling question understanding from visual relevance scoring, and by jointly optimizing relevance, diversity, and temporal coverage. The core idea is to transcribe arbitrary questions into clean queries with Object/Action/Location importance weights, score candidate frames using a fine-tuned 3D semantic ranker (QSRanker), and retrieve non-redundant frames via L2 visual distance across non-overlapping temporal windows tailored for global anchor expansion in long videos and recency-first traversal in streaming videos (QSRetrieval).
Method¶
Overall Architecture¶
The QSVideo framework comprises three key stages: a Question Analyzer that translates raw queries into retrieval-friendly prompts and determines semantic dimension weights; a Semantic Ranker that evaluates frame relevance across entity attributes, dynamic actions, and spatial locations; and a QSRetrieval engine that enforces non-overlapping temporal window partitions, balances relevance with L2 visual diversity, and dynamically navigates time for either long offline videos or real-time streaming videos. The selected informative frames are subsequently fed into standard, plug-and-play video VLMs (e.g., MiniCPM-o 2.6, InternVL2, LLaVA-OneVision) for accurate question answering.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Raw Question & Video Frame Sequence Input"] --> D1["Question Analysis and Semantic Importance Modeling<br/>LLM transcribes query and predicts Object/Action/Location weights"]
In --> D2["Fine-Grained 3D Semantic Ranking<br/>Lightweight VLM evaluates Object/Action/Location relevance scores"]
D1 & D2 --> D3["Relevance-Diversity Temporal Window Retrieval<br/>Combines L2 visual distance with non-overlapping temporal partition"]
D3 --> D4["Adaptive Temporal Alignment for Long & Streaming Videos<br/>Global anchor expansion for long videos / Recency-first traversal for streaming"]
D4 --> Out["Selected Informative Frames Fed to Video VLM for Inference"]
Key Designs¶
1. Question Analysis and Semantic Importance Modeling: Intent Disentanglement and Adaptive Weighting
Feeding raw questions containing distractor options directly into multimodal retrievers introduces spurious correlations. To eliminate this issue, a frozen text LLM (Qwen3-4B) acts as the Question Analyzer \(W\). It reformulates the raw prompt \(\hat{q}\) into an intent-purified query \(q\) and predicts discrete importance weights \(w_o, w_a, w_l \in \{1, 2, 3, 4, 5\}\) along Object, Action, and Location dimensions:
These importance values adaptively adjust the contribution of different visual aspects depending on the specific question (e.g., focusing on clothing favors Object, motion speed favors Action, environment favors Location). The discrete weights are normalized via Softmax to yield \((\tilde{w}_o, \tilde{w}_a, \tilde{w}_l)\), preventing distractor options from polluting the downstream scoring process.
2. Fine-Grained 3D Semantic Ranking: Disentangled Multimodal Scoring and SFT Alignment
Off-the-shelf VLMs typically generate binary yes/no answers or flat scalar scores that cannot disentangle complex cross-modal semantics. QSVideo introduces a lightweight vision-language ranker \(R_\theta\) (fine-tuned on Qwen3-VL-4B-Instruct) that takes the query-frame pair \((q, v)\) and predicts structured relevance scores \((o, a, l) \in \{1, 2, 3, 4, 5\}^3\). The unified ranking score is computed as:
To train this ranker, the authors curate Video-Ranker-35K (35K video-query pairs from LLaVA-Video-178K) and employ a strong teacher model (Qwen-VL-30B) to produce automated frame-level triplet annotations. The semantic ranker is fine-tuned using LoRA on the vision tower and merger modules while freezing the LLM backbone to preserve language reasoning, optimized via Cross-Entropy loss \(\mathcal{L}_{\mathrm{SFT}}\) across the 5 discrete score tokens.
3. Relevance-Diversity Temporal Window Retrieval: L2 Distance and Non-Overlapping Partitions
Greedy top-K selection purely based on semantic relevance leads to duplicate frames from identical scene angles. QSVideo extracts global visual embeddings \(g = f(v)\) via mean pooling across all visual tokens from the ranker's vision encoder, utilizing Euclidean (L2) distance \(d_{ij} = \|g_i - g_j\|_2\) to quantify visual diversity. L2 distance captures both vector orientation and feature norm variations, offering superior sensitivity to subtle layout and appearance shifts over standard cosine similarity. Candidates are chosen by balancing relevance and visual diversity with parameter \(\lambda\):
To definitively eliminate temporal collapse, the entire video frame set \(\mathcal{V}\) is divided into \(K\) disjoint temporal windows \(\{\mathcal{V}_i\}_{i=1}^K\), restricting the selection to at most one optimal frame per window.
4. Adaptive Temporal Alignment for Long and Streaming Videos: Global Anchor Expansion and Recency-First Traversal
Long offline videos and real-time streaming videos require distinct evidence traversal patterns: - Long Videos (Global Anchor and Local Expansion): Maintains unselected window indices \(\mathcal{U}\) (initialized to all \(K\) windows). At step \(t\), the global anchor frame \(a^{(t)}\) with the highest overall relevance score is identified across all remaining windows. Then, within a temporal radius \(r\) around anchor window \(i^{(t)}\), local neighboring windows \(\mathcal{N}^{(t)} = \{i \in \mathcal{U} \mid |i - i^{(t)}| \le r\}\) are traversed, selecting frames using the anchor as the L2 visual reference: $\(v_i^* = \arg\max_{v \in \mathcal{V}_i} \left\{ \lambda \cdot \mathrm{Score}(q, v) + (1 - \lambda) \|f(v) - f(a^{(t)})\|_2 \right\}\)$ Visited windows are removed from \(\mathcal{U}\), and this anchor-and-expand cycle repeats until \(K\) frames are selected, effectively retrieving key clues dispersed across distant video segments. - Streaming Videos (Recency-First Traversal): Traverses windows backwards from the newest window to the past. The most recent window \(\mathcal{V}_K\) selects its highest-scoring frame \(v_K^*\). The search then iteratively moves to earlier windows \(i = K-1, \dots, 1\), using the immediately preceding selected frame \(v_{i+1}^*\) as the visual distance reference: $\(v_i^* = \arg\max_{v \in \mathcal{V}_i} \left\{ \lambda \cdot \mathrm{Score}(q, v) + (1 - \lambda) \|f(v) - f(v_{i+1}^*)\|_2 \right\}\)$ This design enforces chronological priority on recent events while filtering out repetitive visual frames across adjacent streaming time bins.
Key Experimental Results¶
Main Results¶
QSVideo was evaluated on LVBench (103 long videos averaging 67 minutes across 6 tasks) and StreamingBench (real-time streaming evaluation covering 10 tasks) under strict frame budgets (8, 16, and 32 frames).
| Benchmark / Model | LLM Size | Frames | Overall Accuracy (%) | Entity Recog. ER / Object Perc. OP | Key Info. KIR / Event Under. EU | Reasoning Rea / Causal Reason. CR |
|---|---|---|---|---|---|---|
| LVBench | ||||||
| MiniCPM-o 2.6 (Baseline) | 8B | 16 | 38.9 | 37.1 | 37.5 | 46.3 |
| MiniCPM-o 2.6 (Baseline) | 8B | 32 | 42.0 | 41.4 | 39.2 | 48.3 |
| InternVL2-40B | 34B | 16 | 39.6 | 37.4 | 43.4 | 42.5 |
| Qwen2-VL-72B | 72B | 48 | 41.3 | 38.0 | 38.3 | 46.5 |
| SEAL | 34B | 16 | 45.9 | 47.9 | 51.5 | 43.3 |
| GLM-4V-Plus | Proprietary | \(\le 300\) | 48.7 | 46.2 | 54.1 | 46.5 |
| GPT-4o | Proprietary | 60 | 48.9 | 48.9 | 48.1 | 50.3 |
| QSVideo (Ours) | 8B | 16 | 45.8 (+6.9โ) | 47.3 | 49.8 | 50.8 |
| QSVideo (Ours) | 8B | 32 | 49.1 (+7.1โ) | 51.1 | 54.6 | 53.2 |
| StreamingBench | ||||||
| MiniCPM-o 2.6 (Baseline) | 8B | 8 | 64.44 | 67.30 | 71.43 | 71.09 |
| InternVL2 | 8B | 8 | 68.52 | 76.02 | 68.94 | 60.94 |
| LLaVA-OV | 7B | 8 | 72.12 | 79.29 | 72.67 | 72.66 |
| QSVideo (Ours) | 8B | 8 | 77.68 (+13.24โ) | 83.11 | 80.12 | 79.69 |
| MiniCPM-o 2.6 (Baseline) | 8B | 16 | 65.12 | 69.21 | 72.67 | 77.34 |
| InternVL2 | 8B | 16 | 70.52 | 75.20 | 75.16 | 64.84 |
| LLaVA-OV | 7B | 16 | 73.76 | 79.56 | 73.29 | 79.69 |
| Claude 3.5 Sonnet | Proprietary | 20 | 74.04 | 82.45 | 76.39 | 73.77 |
| StreamForest | 7B | 1fps | 77.26 | 83.11 | 77.50 | 82.81 |
| QSVideo (Ours) | 8B | 16 | 79.36 (+14.24โ) | 84.47 | 81.99 | 82.81 |
Ablation Study¶
1. Sensitivity to Diversity Weight \(\lambda\) (LVBench)
The trade-off parameter \(\lambda\) balances semantic relevance and L2 visual distance. Under 16 and 32 frame budgets, the model exhibits robust performance across wide intervals:
| Frame Budget | \(\lambda\) | Overall Accuracy (%) | ER (%) | EU (%) | KIR (%) | TG (%) | Rea (%) | Sum (%) |
|---|---|---|---|---|---|---|---|---|
| 16 | 0.1 | 46.1 | 46.7 | 44.2 | 49.8 | 37.3 | 50.2 | 32.8 |
| 16 | 0.3 | 45.6 | 47.3 | 43.3 | 49.1 | 33.6 | 48.3 | 31.0 |
| 16 | 0.5 | 46.0 | 47.6 | 42.3 | 51.2 | 35.9 | 50.7 | 31.0 |
| 16 | 0.7 | 45.6 | 47.0 | 41.7 | 50.9 | 35.0 | 50.2 | 31.0 |
| 16 | 0.9 | 45.8 | 47.3 | 41.7 | 49.8 | 35.5 | 50.8 | 32.8 |
| 32 | 0.1 | 48.2 | 49.0 | 45.1 | 53.6 | 38.2 | 51.7 | 34.5 |
| 32 | 0.3 | 48.3 | 51.1 | 44.4 | 54.3 | 36.4 | 51.2 | 31.0 |
| 32 | 0.5 | 48.2 | 51.3 | 44.5 | 54.0 | 35.5 | 52.2 | 29.3 |
| 32 | 0.7 | 48.7 | 51.3 | 45.0 | 55.0 | 39.1 | 51.7 | 32.8 |
| 32 | 0.9 | 49.1 | 51.1 | 45.6 | 54.6 | 39.6 | 53.2 | 32.8 |
2. Plug-and-Play Generalization across Diverse Video VLMs (StreamingBench)
Integrating QSVideo frame selections into diverse open-source video backbones provides consistent performance gains across 8, 16, and 32 frames:
| Backbone Model | Frames | Baseline Accuracy (%) | + QSVideo Accuracy (%) | Absolute Gain \(\Delta\) |
|---|---|---|---|---|
| MiniCPM-V 2.6 | 8 | 66.36 | 70.96 | +4.60% |
| MiniCPM-V 2.6 | 16 | 70.24 | 71.84 | +1.60% |
| MiniCPM-V 2.6 | 32 | 70.80 | 72.96 | +2.16% |
| InternVL2 | 8 | 68.52 | 71.28 | +2.76% |
| InternVL2 | 16 | 70.52 | 72.24 | +1.72% |
| InternVL2 | 32 | 71.52 | 72.92 | +1.40% |
| LLaVA-OneVision | 8 | 72.12 | 74.48 | +2.36% |
| LLaVA-OneVision | 16 | 73.76 | 74.68 | +0.92% |
| LLaVA-OneVision | 32 | 73.56 | 74.40 | +0.84% |
3. Latency and Throughput on a 34-Minute Long Video (2048 Candidate Frames)
| Method | Parameters | Candidate Frames | Selected Frames | End-to-End Latency (s) | Throughput |
|---|---|---|---|---|---|
| SEAL | 34B | 2045 | 16 | 1702.24s | 1.2 fps |
| AdaReTaKe | 72B | 1024 | - | 178.15s | 5.8 fps |
| QSVideo (16 frames) | 8B | 2048 | 16 | 88.73s | 23.1 fps |
| QSVideo (32 frames) | 8B | 2048 | 32 | 91.17s | 22.5 fps |
Key Findings¶
- Superiority of Disentangled Evidence Ranking: Evaluating temporal evidence retrieval on annotated LVBench time references shows QSRanker outperforms flat-scoring rankers across all Recall@K thresholds: Recall@10 reaches 33.20% vs. 32.43%, Recall@50 reaches 50.65% vs. 46.77%, and Recall@100 reaches 58.20% vs. 52.78%.
- Critical Value of Retrieval in Short Streaming Windows: Even within narrow 60-second streaming contexts where uniform sampling is conventionally presumed sufficient, QSVideo achieves substantial improvements (+11.32% for 8 frames, +14.24% for 16 frames), indicating that semantic frame selection is crucial even in short horizons.
- High Computational Throughput: QSVideo is 18.7ร to 19.2ร faster than SEAL. By leveraging data parallelism and batch inference, QSRanker scoring latency over 2048 candidate frames decreases from 608.40s to 61.32s.
Highlights & Insights¶
- Decoupled Semantic Dimension Weighting: Instead of relying on monolithic text-visual matching, QSVideo decomposes visual evidence into Object, Action, and Location components, dynamically weighting them via an LLM. This mitigates distractor interference and improves semantic interpretability.
- Cognitive-Aligned Temporal Alignment: Mapping long-form video retrieval to global anchor exploration and streaming retrieval to recency-first backtracking provides an intuitive and effective temporal search procedure.
- High-Performance Efficiency Frontier: Achieving 49.1% on 67-minute videos using an 8B model with only 32 frames proves that targeted visual evidence localization is more compute- and memory-efficient than scaling context windows or parameters.
Limitations & Future Work¶
- Author-Acknowledged Limitations: Ground-truth semantic ranking supervision relies on teacher model pseudo-annotations (Qwen-VL-30B), which may inherit biases from the teacher. Initial candidate sampling from dense uniform frames still requires substantial initial feature extraction time.
- Observed Constraints: Performance gains on whole-video summarization (Sum) are comparatively limited, suggesting that aggressive temporal pruning may discard diffuse contextual background necessary for broad thematic summaries.
- Future Directions: Integrating hierarchical retrieval with fast temporal action localization (TAL) to prune candidate frames before ranker invocation, and developing hybrid memory representations that combine dense sparse anchors with lightweight global summary tokens.
Related Work & Insights¶
- vs. SeViLA / FRAG / Frame Selector: Previous frame selectors directly feed raw questions into pretrained VLMs to obtain yes/no probabilities, which degrades on complex multiple-choice prompts. QSVideo purifies queries and calculates fine-grained multi-dimensional relevance.
- vs. SEAL: SEAL deploys separate heavy models for object detection, action recognition, and scene classification, yielding slow inference (1.2 fps). QSVideo unifies scoring within a single lightweight VLM, executing up to 19ร faster with superior QA accuracy.
- vs. StreamForest / QueryStream: Instead of building complex episodic memory forests or specialized streaming architectures, QSVideo functions as a fully plug-and-play retriever compatible with any existing video VLM.
Rating¶
- Novelty: โญโญโญโญ [Well-motivated semantic dimension decomposition coupled with human-inspired temporal traversal]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive evaluations across LVBench and StreamingBench with complete ablations and latency benchmarks]
- Writing Quality: โญโญโญโญโญ [Clear problem formulation, structured mathematical modeling, and crisp narrative]
- Value: โญโญโญโญโญ [Provides a practical, highly efficient retrieval blueprint for long and streaming video understanding]