Driving Video Retrieval for Complex Queries with Structured Grounding¶
Conference: NeurIPS2026
arXiv: 2606.09109
Area: Autonomous Driving
Keywords: driving video retrieval, structured grounding, rule calibration, weakly supervised learning to rank, multi-source fusion
TL;DR¶
STRIVE-D calibrates motion rules against a weak-label ranking objective on separate driving videos, reuses or adapts those rules per query, and fuses them with visual and lexical retrieval, raising DrivingDojo Acc@1 from the main table's strongest dense baseline of 14.6% to 26.8%; it does not simply ask an LLM to inspect videos and determine true geometry.
Background & Motivation¶
Driving-log retrieval must find not only scenes containing a bus, but also brief events in which a bus cuts into the ego lane. Global video embeddings from a vision-language model (VLM) can mix lateral displacement and relative-distance changes with background semantics, while lexical retrieval depends on captions explicitly naming the event. Retrieving the right object therefore does not establish that the right behavior was retrieved. The paper calls this dilution of dynamic evidence by scene content โdilution.โ
Structured rules can directly test tracked-object distance, lane, and lateral offset, but face another failure: a large language model (LLM) understands the approximate meaning of approach, hard brake, and cut-in without knowing the noise, quantization, or change magnitudes produced by the deployed perception pipeline. A linguistically plausible threshold may fire almost everywhere or never fire with a particular depth estimator and tracker. This is miscalibration; even a calibrated finite event library leaves a coverage gap for open-vocabulary queries.
Rather than relying on a larger video model, the paper separates semantic proposals from empirical calibration: rules capture explicit motion evidence, independent auxiliary videos establish numerical scales, and visual and lexical branches supplement scene content and long-tail entities that rules struggle to express. Core Idea: derive weak relevance labels from auxiliary-video captions, optimize the ranking induced by rules on actual perception records, and combine reusable calibrated rules with query adaptation and multi-source retrieval.
Method¶
Overall Architecture¶
The input is a natural-language query and a candidate corpus of driving videos; the output is a relevance ranking, not a vehicle-control action or a proven accident cause. STRIVE-D preprocesses videos into per-frame object records and computes motion relevance by executing rules on those records. A dense branch uses visual embeddings, a sparse branch uses video text, and fusion is followed by reranking the top 20 candidates.
The four designs are Structured Motion Evidence, Calibrated Rule Library, Confidence-Aware Adaptation, and Multi-Source Rank Fusion. Offline construction calibrates the library on an independent Nexar auxiliary set. Online retrieval decides whether to reuse or adapt a rule for a new query. Auxiliary captions construct offline weak labels; they neither replace per-frame perception records nor constitute evaluation ground truth.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
V["Auxiliary and candidate videos"] --> S["Structured Motion Evidence"]
C["Weak labels from auxiliary captions"] -.->|Offline supervision| L["Calibrated Rule Library"]
S -.->|Offline execution on auxiliary records| L
L --> A["Confidence-Aware Adaptation"]
Q["New query"] --> A
A -->|Execute rule on candidate records| F["Multi-Source Rank Fusion"]
S -->|Candidate motion evidence| F
D["Candidate visual embeddings<br/>and video text"] --> F
F --> O["Rerank top 20<br/>Return video ranking"]
Key Designs¶
1. Structured Motion Evidence: separate object identity, admissibility, and event strength
Rules consume object time series generated by iFinder's perception pipeline, not true positions measured by the LLM. Upstream components include OWL-V2 detection, ByteTrack tracking, OMR lane detection, Metric3D depth, SAM foreground masks, and camera calibration, pose estimation, attribute recognition, and 3D detection. Distance, lateral offset, and lane relations can all contain errors; these errors help explain why thresholds must be calibrated to the perception system's output distribution. Qwen2.5-VL-7B captions auxiliary videos during library construction rather than acting as an online event classifier for every candidate.
A rule first uses entity selectors to restrict object classes, optionally select the closest object, and determine how multiple tracks are aggregated. Hard gates test necessary conditions, such as the existence of an object closer than a threshold; every gate must hold, or the video's motion score is zero. Soft scores measure the strength of field changes within specified windows and combine evidence terms through a weighted average. Events attributable to one vehicle use any aggregation. Ego motion can instead use coherent relative motion across surrounding objects through consensus aggregation, rather than treating one object's lateral movement as an ego turn.
Appendix A.3 defines the unsigned soft-change operator and final score as:
The window restricts the interval over which changes are examined, while the magnitude scale determines when the continuous score saturates. A specified direction replaces absolute change with a one-sided change. In the appendix, any takes the maximum object score, while consensus averages the maximum and the mean object score. This retains strong evidence while making group agreement affect ranking; it should not be reduced to a trigger on an arbitrary track.
Two implementation boundaries deserve explicit attention. First, hard gates and soft scores are aggregated independently at video level: object A can satisfy a distance condition while object B contributes motion strength. The operator does not require one object to satisfy all evidence terms. Second, the D.3 prompt describes soft scores as 0.5 at the threshold and 1 at twice the threshold, and describes consensus as an average gated by a passing fraction. These descriptions do not fully match A.3. This note follows A.3 for the formal operators; the prompt differences cannot be resolved by guessing the implementation.
2. Calibrated Rule Library: let AP, rather than linguistic intuition, select numerical constraints
The offline auxiliary set contains 1,147 Nexar videos disjoint from all three evaluation corpora. A captioner describes motion changes, and an LLM extracts events. A.2 specifies a predefined inventory of 21 categories, rather than unconstrained discovery of arbitrary events. For each category, videos whose captions mention it become weak positives, and the rest become weak negatives. A video can belong to several categories, so event counts cannot be summed as distinct videos. Hard braking has 1,068 associated videos, whereas reverse and sharp turn have only 1 each.
Section 3.4 explicitly fixes a human-specified rule form for each event and jointly searches hard-gate thresholds, soft-score magnitudes, weights, and window lengths. The LLM receives auxiliary-record field means, standard deviations, and quantiles, then proposes parameters. An executor scores auxiliary perception records and computes average precision (AP) against the weak labels. Field statistics help proposals use appropriate scales but do not enter the optimization objective. The LLM is a black-box proposer; AP selects candidate assignments.
AP measures precision at each weak positive's position in the rule-induced ranking and averages over all weak positives:
The loop feeds the current AP, best-so-far rule, and score distribution back to the LLM. It stops at a target AP or after at most 20 iterations, retaining the best parameters. The objective is to rank weak positives above weak negatives, not to regress exact speed or human-annotated geometric thresholds; it also requires no back-propagation through the video encoder. Aggregation across examples can reduce individual caption errors but does not guarantee correction of systematic omissions, especially for events with very few positives.
The D.3 initial prompt emits a complete JSON rule, and D.4 permits revision of the whole object, which is broader wording than the fixed-rule-form description in the main text. The demonstrated contribution should be understood as numerical calibration and distinguished from online rule adaptation. These materials do not establish fully automatic discovery and optimization of rule structure.
3. Confidence-Aware Adaptation: reuse high-confidence matches and adapt from calibrated exemplars otherwise
At retrieval time, a GPT-4.1 matcher maps the query to the closest event category and returns a matching confidence. A confidence of at least 0.6 reuses the library rule; a lower confidence uses the closest calibrated rule as an in-context exemplar to generate a query-specific rule. This confidence estimates category matching: it is neither the motion score nor a probability that a particular video contains the event. The high-confidence path avoids rule generation, not necessarily an LLM call for matching; A.2's โwithout invoking the LLMโ wording needs to be read in conjunction with D.5.
Adaptation is not merely copying the original rule. D.6 supplies unlabeled target-dataset field statistics and permits changes to object classes and numerical constraints, so online adaptation may use target perception-distribution information. Avoiding target relevance labels is different from never accessing target data. For queries modified by weather or cause, such as loss of control in snow, the matcher prompt caps confidence at 0.55 to force adaptation rather than collapse the request into a generic loss-of-control event.
The advantage is a data-calibrated starting point instead of guessing scales again from language alone. However, the new rule is not recalibrated with AP on target human labels. Borrowing a neighboring rule cannot guarantee reliable threshold transfer or create causal, environmental, or event-order evidence absent from the perception fields.
4. Multi-Source Rank Fusion: use motion rules to complement visual retrieval, not replace it
STRIVE-D produces symbolic-rule rankings, Qwen3-VL-Embedding-8B dense rankings, and BM25 lexical rankings. A fusion router assigns nonnegative weights according to the query and each branch's input capabilities: rules capture relative object motion, vision captures scene content and salient pixel changes, and text captures colors or locations explicitly named in summaries. Its routing uses driving-geometry priors encoded in the prompt; LLM reasoning does not prove that the actual geometry holds.
Because raw scores have incompatible scales, the system uses query-conditioned weighted reciprocal rank fusion (RRF), rather than adding raw scores:
Ranks start at 1, and 60 is the smoothing constant. The system sorts by fused score and reranks the top 20 candidates. This pool size comes from the sweep in B.1, which reports peaks or plateaus near that value for the evaluated methods and fixes the same reranking budget in the main experiments.
The sparse branch's text source remains ambiguous: Section 3.5 and D.7 describe per-candidate video summaries, while Section 4.1 says BM25 is built over auxiliary captions. A.1 also says its video-captioning module is used only for library construction. Since auxiliary and evaluation corpora are separate, these descriptions do not establish how the actual index maps to evaluation candidates. This note preserves the lexical mechanism without conflating auxiliary weak labels, candidate text indexing, and evaluation ground truth.
A Worked Example¶
For a query asking for a car cutting into the ego lane, the matcher first selects cut-in: a confidence of at least 0.6 reuses the calibrated rule; otherwise it adapts from a related rule. The main-text example selects car/truck, applies a distance gate below 40.0, and examines lateral-offset changes within 4.0 seconds with magnitude scale 14.0. These are calibrated values for that rule, not universal physical laws for all vehicles and cameras.
To illustrate A.3's operator, suppose a candidate track has a maximum lateral change of 7.0: its single soft term is 0.5; a change of 14.0 saturates at 1. These two track-change values are illustrative calculations in this note, not measurements reported by the paper. A video failing the gate still receives zero. Passing videos merely have rule evidence and must still undergo visual/lexical fusion and top-20 reranking. If different objects supply the gate and lateral change, the published operator can still assign a high score; this example is not strict same-object cut-in verification.
Loss & Training¶
The paper introduces no end-to-end video-model training loss. Its central training-like step is per-event black-box numerical search using weak-label AP. The construction-time captioner is Qwen2.5-VL-7B; language components default to GPT-4.1; the confidence threshold is 0.6; each event uses at most 20 iterations; and experiments use one 48 GB NVIDIA RTX A6000. Auxiliary calibration excludes evaluation queries, labels, and videos, while online adaptation prompts may access unlabeled target-dataset field statistics.
Key Experimental Results¶
Main Results¶
Main-table values are percentages; the table below preserves Table 1's reporting. DrivingDojo has 583 candidate videos and 41 queries, with 236 videos relevant to at least one query and 1โ138 relevant videos per query. CarCrashDataset uses 110 videos with accident-reason annotations and 45 query categories. MM-AU uses 1,953 test videos and 49 accident-description categories. These tasks differ in difficulty and relevance density.
Acc@k is the fraction of queries with at least one ground-truth video in the top k, not per-video classification accuracy or Recall@k. MRR averages the reciprocal rank of the first relevant item, while mAP averages AP across queries.
| Dataset | Method | MRR | mAP | Acc@1 | Acc@3 | Acc@5 | Acc@10 |
|---|---|---|---|---|---|---|---|
| DrivingDojo | Qwen3-VL-8B | 20.1 | 2.2 | 14.6 | 31.7 | 51.2 | 68.3 |
| DrivingDojo | NSVS-TL | 31.8 | 4.9 | 22.0 | 34.2 | 43.9 | 53.7 |
| DrivingDojo | STRIVE-D | 38.7 | 8.6 | 26.8 | 48.8 | 53.7 | 73.2 |
| MM-AU | Qwen3-VL-8B | 32.7 | 6.3 | 20.4 | 38.8 | 55.1 | 61.2 |
| MM-AU | NSVS-TL | 18.0 | 2.8 | 6.1 | 20.4 | 36.7 | 53.1 |
| MM-AU | STRIVE-D | 40.5 | 9.5 | 26.5 | 53.1 | 59.2 | 65.3 |
| CarCrashDataset | Qwen3-VL-8B | 40.1 | 33.5 | 28.9 | 46.7 | 62.2 | 66.7 |
| CarCrashDataset | NSVS-TL | 23.1 | 19.3 | 11.1 | 24.4 | 37.8 | 53.3 |
| CarCrashDataset | STRIVE-D | 45.0 | 37.7 | 28.9 | 60.0 | 73.3 | 75.6 |
DrivingDojo's 26.8% versus 14.6% is a gain of 12.2 percentage points and 83.6% relative improvement, not an 84% gain over every baseline. Against NSVS-TL, the Acc@1 gain is only 4.8 percentage points. CarCrashDataset Acc@1 is 28.9%, tying Qwen3-VL-8B; the method does not strictly win every metric.
The source contains numerical inconsistencies: DrivingDojo Qwen3-VL-8B Acc@1 is 14.6 in Table 1 but 7.3 in Table 8's +Reranker row. CarCrash Qwen3-VL-2B +Reranker MRR/mAP is 32.1/38.3 in Table 8 versus 38.3/32.1 in Table 1. No values are silently swapped or reconciled here. Main results use Table 1, with reduced certainty about the precise gain accounting.
Ablation Study¶
The following values come from Section 4.3's text and Appendix Table 6. Absolute scores for removing hard gates and entity selectors are calculated from the reported percentage-point drops, not estimated from plots. All results are DrivingDojo percentages.
| Config | Acc@1 | mAP | Note |
|---|---|---|---|
| Full STRIVE-D | 26.8 | 8.6 | full model |
| Remove dense branch | 9.76 | Not reported | Section 4.3 explicitly states Acc@1 |
| Binarize soft scores | 12.2 | Not reported | Drop of 14.6 percentage points |
| Remove hard gates | 19.5 | Not reported | Calculated from a 7.3-point drop |
| Remove entity selector | 22.0 | Not reported | Calculated from a 4.8-point drop |
| Replace rules with NSVS-TL; retain fusion and reranking | 12.2 | 3.9 | Table 6; surrounding retrieval pipeline fixed |
Unprinted plot values in cached Figures 3/4 are not guessed. Textual analyses describe trends for removing the calibrated library, symbolic branch, and sparse branch, but do not justify inventing complete quantitative rows.
Key Findings¶
- The largest overall degradation comes from removing dense retrieval: 26.8% falls to 9.76%. This contradicts an interpretation in which rules replace vision; the improvement depends on complementarity.
- Within the rule, continuous strength matters most: binarization yields only 12.2%. Gates and object selection further reduce mismatches. These ablations fix the remaining parameters without recalibration, measuring dependencies of this fixed system rather than proving the same ordering for every possible rule form.
- In Table 6, replacing the symbolic component with NSVS-TL under identical fusion and reranking reduces Acc@1 from 26.8% to 12.2% and mAP from 8.6% to 3.9%. This supports calibrated graded rules over the substituted structured branch, but does not completely isolate calibration from continuous scoring.
- Five full-pipeline reruns report DrivingDojo Acc@1 of 26.8ยฑ1.3 and CarCrash of 28.9ยฑ1.2. Reported stability does not resolve the baseline-number conflicts between the main table and appendix.
Online timings come from Tables 2/7, not the combined cost of offline perception, captioning, and calibration:
| Method or stage | Mean online latency | Evidence |
|---|---|---|
| Qwen3-VL-Embedding-8B | 128.54 ms | Table 2; also the dense branch in Table 7 |
| Full STRIVE-D | 3.10 s | Table 2; Table 7 totals 3,096.66 ms |
| NSVS-TL | 4,868.05 s | Table 2 |
| STRIVE-D rule synthesis | 1,987.43 ms | Table 7; invoked only on the adaptation branch |
| STRIVE-D multi-source fusion | 725.26 ms | Table 7 |
The two LLM stages account for about 87% of the reported total. The complete method must not be described as taking 128.54 ms. The approximately 1,500-fold speedup over NSVS-TL does not mean it is faster than dense retrieval. The table does not separately itemize matcher and reranker costs, leaving the full deployment-cost boundary in need of clarification.
Highlights & Insights¶
- Separate language knowledge from numerical measurement. An LLM may propose which fields to inspect, but ranking performance on perception records must select their scales, rather than treating linguistic priors as physical measurements.
- Use continuous rule scores rather than only satisfaction verdicts. Retrieval asks which candidate should rank first; strength differences near an event boundary are precisely what binary outputs discard.
- Treat calibrated rules as reusable retrieval components. Similar designs may apply to industrial or sports videos with object time series, but transfer requires reassessing field scales, weak-label quality, and rule coverage.
Limitations & Future Work¶
- The authors explicitly acknowledge the absence of temporal sequencing such as event A followed by event B and scoped negation. Such queries fall back to other branches or approximate adaptation; the current rule system is not a complete temporal-logic framework.
- Calibration depends on noisy caption-derived event labels. Extreme category imbalance and single-positive events limit AP reliability. Independent small-scale human checks, weak-label confidence modeling, and rare-event uncertainty reporting are more informative than merely increasing proposal iterations.
- Video-level gates and soft evidence can come from different objects or times. Same-track, same-window evidence constraints are a concrete extension for reducing high-scoring candidates in which individual conditions hold but the requested event does not.
- Perception errors, occlusion, and identity switches propagate into rule scores. The paper does not comprehensively evaluate robustness and cost under changed perception systems, long videos, or fleet-scale online services.
- Sparse-index provenance, formal-operator/prompt differences, baseline conflicts, and the absence of independent calibration for adapted target rules affect reproducibility and precision of conclusions. Deployment also requires audits of perception bias, video authorization, and privacy boundaries.
Related Work & Insights¶
- vs Qwen3-VL / SigLIP families: Dense methods retrieve through shared visual-text representations; STRIVE-D adds explicit motion records and calibrated rules. Dense retrieval remains essential, and additional structured precision costs seconds of online processing plus offline preprocessing.
- vs NSVS-TL: NSVS-TL organizes video verification through temporal logic, whereas this method uses a flatter entity/gate/continuous-score schema with data calibration. Table 6's same-pipeline substitution controls surrounding factors, but STRIVE-D also has more restricted expressive power.
- vs iFinder: The method reuses its perception structures rather than reinventing geometric extraction. Its contribution is weakly supervised numerical calibration, rule reuse/adaptation, and multi-source ranking, not an LLM acquiring direct access to true geometry.
- vs LLM-as-Optimizer: The method connects language proposals to executable rules and AP feedback. A useful research question is whether random search or conventional black-box optimization reaches equivalent calibration under identical weak labels and budgets; the paper does not provide that comparison.
Rating¶
- Novelty: 4/5 โ Grounds rule-parameter calibration in a weakly supervised ranking objective with a reasonably clear contribution boundary.
- Experimental Thoroughness: 3/5 โ Three datasets, structured substitution, and five reruns provide useful evidence, weakened by reporting conflicts and reproduction details.
- Writing Quality: 3/5 โ Failure modes map clearly to components, but formal definitions, prompts, and appendix statements require clarification.
- Value: 4/5 โ Offers practical ideas for driving-scenario mining while retaining perception, online LLM, and calibration-maintenance costs.