Skip to content

Fine-Grained Text-to-Video Retrieval for Camera-Trap Data

Conference: ECCV 2026
Paper: ECCV Official
Project Page: Project Page
Area: Video Understanding
Keywords: Text-to-Video Retrieval, Camera-Traps, Spatiotemporal Action Localization, LLM Agent, Computational Ecology

TL;DR

This paper introduces Prompting-MammAlps, the first fine-grained text-to-video retrieval benchmark for camera-trap wildlife data, and proposes an interpretable framework that decouples offline spatiotemporal action localization (SALMA) from online LLM-guided predicate code generation, avoiding hallucinations while delivering robust set-based retrieval.

Background & Motivation

Camera-traps are an indispensable remote sensing tool for wildlife ecology and biodiversity conservation, passively recording vast amounts of high-resolution video with minimal human disturbance. However, transforming this raw footage into ecological insights remains severely bottlenecked by manual annotation time. While deep learning has streamlined false-positive image filtering (e.g., via MegaDetector) and coarse species classification, fine-grained visual events such as behavioral repertoires, courtship displays, parent-offspring interactions, and environmental responses still depend on tedious manual labeling by field biologists or citizen science platforms. Text-to-video retrieval (TVR) offers a direct path to query complex ecological events using free-form natural language queries.

Existing TVR methods based on generic video-language foundation models (VLMs) fail drastically in wildlife camera-trap domains. First, foundation VLMs are predominantly pretrained on web videos (e.g., YouTube, Kinetics) and lack ecological domain knowledge, compounded by the scarcity of dedicated benchmarks assessing fine-grained animal behaviors. Second, mainstream TVR relies on projecting sampled video frames and text prompts into a joint latent embedding space to compute cosine similarities. This latent similarity paradigm is a black box, lacks long-term spatiotemporal reasoning, and struggles to disambiguate multi-individual interactions, temporal sequences, or conditional logic. Directly prompting multimodal LLMs to transcribe videos into descriptive prose before retrieval introduces severe hallucinations and context-length bottlenecks.

To address these tensions, this paper decouples perceptual spatiotemporal feature extraction from symbolic linguistic reasoning: candidate videos are processed offline by an end-to-end spatiotemporal action localization model to extract animal trajectories and attribute trees stored as structured JSON representations, while incoming user queries are translated online by an LLM coding agent into executable boolean filter functions calling a custom primitive library. Core idea: decouple camera-trap video retrieval into offline spatiotemporal action localization with structured JSON extraction and online LLM code generation, where a fine-tuned video transformer (SALMA) builds structured text representations and an LLM coding agent translates natural language queries into executable boolean predicate functions over a predefined parsing library, eliminating hallucinations while providing full interpretability.

Method

Overall Architecture

The proposed end-to-end fine-grained retrieval architecture completely separates video feature extraction from query parsing. In the offline stage, raw camera-trap candidate videos are fed into SALMA (Spatiotemporal Action Localization for MammAlps), an encoder-decoder transformer that detects and tracks animal bounding boxes across frames while predicting species, high-level activities, low-level actions, deer age, and adult deer sex. The extracted spatiotemporal trajectories and metadata are recorded into structured .json files. In the online retrieval stage, a natural language query describing target ecological events or cross-video comparisons is given to an LLM code generation agent. Governed by a dedicated library of 23 primitive parsing operations, the agent synthesizes a deterministic Python predicate function that executes across all candidate video JSON files to return matching videos as a discrete retrieved set.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Raw Camera-Trap Video Pool"] --> B["SALMA Spatiotemporal Action Localization & Curriculum Learning<br/>VideoMAE Encoder + Evolving Object Queries with Contrastive Loss"]
    B --> C["Structured JSON Trajectory Representations & Spatiotemporal Aggregation<br/>Track-level Attribute Pooling + Frame-level Bbox/Action Alignments"]
    D["User Ecological Query Prompt<br/>Complex Multi-Attribute / Temporal / Cross-Video Query"] --> E["Grammar-Constrained LLM Code Agent Retrieval<br/>Synthesizes Deterministic Boolean Predicate Function via 23 Primitives"]
    C --> F["Sandbox Execution & Set-Based Filtering<br/>Execute Predicate Function against Each Candidate JSON File"]
    E --> F
    F --> G["Retrieved Video Set Returned to User<br/>Transparent Python Source Code + Visualized SALMA Predictions"]

Key Designs

1. SALMA Spatiotemporal Action Localization & Curriculum Learning: Full Attribute Decoupling Standard object detectors or clip classifiers fail to maintain spatiotemporal consistency across erratic animal movements under harsh infrared illumination. SALMA adapts the DETR and MOTR architectures into a unified transformer framework. The VideoMAE vision backbone encodes video frames into spatiotemporal tokens \(f_{i,t}\), while a transformer decoder progressively refines learnable object queries \(q_{i,t}\) across successive time steps via cross-attention. Each active query connects to six parallel MLP prediction heads responsible for bounding box regression, track activation score \(s_{i,t,j}\), species, low-level actions, high-level activity, deer age, and adult deer sex. To ensure training stability, SALMA employs a two-stage curriculum: the first stage freezes attribute heads and optimizes box coordinates with L1 and gIoU losses, activation scores with binary cross-entropy, and query identity temporal consistency via an InfoNCE contrastive objective: $$ \mathcal{L}{\text{InfoNCE}} = -\log \frac{\exp(q $$ This contrastive loss penalizes feature drift for identical animal tracks across consecutive frames, drastically suppressing identity swaps in dense herds or heavy occlusion. The second stage unfreezes the classification MLP heads and trains them with categorical cross-entropy while joint fine-tuning maintains tracking quality. During inference, inputs are processed on two overlapping square-cropped views matched via the Hungarian algorithm to overcome the 1:1 training versus 16:9 recording aspect ratio mismatch.} \cdot q_{i,t+1,j} / \tau)}{\sum_{k} \exp(q_{i,t,j} \cdot q_{i,t+1,k} / \tau)

2. Structured JSON Trajectory Representations & Spatiotemporal Aggregation: Preserving Frame-Level Semantics Directly converting long video sequences into free-form textual captions via VLMs leads to severe loss of temporal boundaries and spatial coordinate relations. The authors introduce a hierarchical JSON schema as a lossless semantic bridge: the root section captures global video metadata including site ID, camera ID, resolution, and environmental weather; the frames list contains frame-indexed animal detections, each tagged with its active track_id, bounding box coordinates, and current predicted attributes. For time-invariant attributes across an individual track (species, age, adult sex), a majority voting aggregation is performed across all track detections to eliminate instantaneous frame flicker. Dynamic actions and activities preserve their exact per-frame temporal boundaries. This schema turns unconstrained visual pixels into a standardized symbolic graph, shielding downstream retrieval from perceptual noise.

3. Grammar-Constrained LLM Code Agent Retrieval: Synthesizing Deterministic Boolean Predicates Directly passing video descriptions into LLM prompts causes severe attention dilution, hallucinations, and inability to audit decisions. The framework leverages the smolagents library to build a code generation agent that is strictly prohibited from viewing candidate video contents. Instead, the agent is provided with an API specification of 23 type-safe primitive functions (e.g., check_tracks_contains_species(tracks, species_name), get_unique_activities_from_tracks(tracks), get_tracks_from_file_id(file_id)), an exact vocabulary of valid attribute labels, and three in-context reasoning demonstrations. The agent's sole objective is translating the semantic intent of the query into a deterministic Python function check_file(file_id). The generated code is executed within an isolated sandbox against a mock JSON object. Any reference to non-existent functions or out-of-vocabulary attribute strings raises an execution exception returned to the agent, prompting self-correction across up to 10 trials. The verified function executes across all candidate JSON files, producing a verifiable binary decision set while allowing users to inspect the exact Python code for full algorithmic transparency.

Loss & Training

SALMA is trained on 2,090 videos from MammAlps-S2 for 500 epochs per curriculum stage using AdamW with linear warmup and cosine decay. Inference runs at a temporal stride of 4 frames with forward-filling. The LLM agent framework was evaluated across four open-source LLMs: Qwen3-8B, Llama-3.1-8B-Instruct, Mistral-7B-Instruct, and Apertus-8B-Instruct. Qwen3-8B delivered the highest function generation fidelity and serves as the primary code agent.

Key Experimental Results

Main Results

Retrieval performance is evaluated on the Prompting-MammAlps test split (775 candidate videos, 135 ecological queries). Because the method yields discrete binary retrieval decisions rather than ranking scores, the benchmark evaluates set-based F1-Scores (including queries where no videos match). Comparisons are drawn against zero-shot and fine-tuned video foundation models (CLIP4Clip, SigLIP4Clip, GRAM, InternVideo2.0).

Method Training Setup All Queries (135) w/o Vid. Comp. (113) Rare (48) Courtship (27) Other Social (36) Cam. Reaction (28) Common (2)
CLIP4Clip Zero-shot n/a 0.17 0.16 0.13 0.26 0.14 0.40
SigLIP4Clip Zero-shot n/a 0.17 0.17 0.12 0.25 0.12 0.37
GRAM Zero-shot n/a 0.16 0.16 0.08 0.25 0.12 0.35
InternVideo2.0 Zero-shot n/a 0.18 0.16 0.15 0.28 0.12 0.40
CLIP4Clipโ€  Fine-tuned (MammAlps-S2) n/a 0.26 0.29 0.22 0.30 0.18 0.70
Agent + SALMAโ€  (Ours) Supervised SALMA + Code Agent 0.34 0.37 0.43 0.34 0.35 0.21 0.87
Agent + Oracle (Upper Bound) Ground-Truth JSON + Code Agent 0.87 0.89 0.94 0.79 0.90 0.82 0.98

Note: โ€  indicates supervised training/fine-tuning on MammAlps-S2. n/a denotes that latent similarity models cannot natively process two-step cross-video comparison queries.

Ablation Study

The evaluation decouples the tracking/attribute extraction performance of SALMA from the reasoning capacity of the LLM code agent.

1. SALMA Multi-Object Tracking & Attribute Classification Ablation

Model & Configuration HOTA โ†‘ IDF1 โ†‘ DetA โ†‘ IDswp โ†“ Species F1 Activity F1 Actions F1 Deer Age F1 Adult Sex F1 Weather F1
SALMA-MOT w/o InfoNCE 67.5 75.4 64.9 3485 - - - - - -
SALMA-MOT (Stage 1 Full) 72.9 84.5 67.7 192 - - - - - -
SALMA (Two-Stage Full Model) 69.0 81.1 62.6 202 0.48 0.40 0.29 0.70 0.75 0.76
MegaDetector v5a + ByteTrack 76.2 81.4 67.2 93 - - - - - -

2. Qwen3-8B Code Agent Components Ablation (on Ground-Truth Oracle JSON)

Configuration Retrieval F1-Score Gain Note
Vanilla Agent 0.44 Baseline Free-form code generation with JSON schema & labels
+ Custom Parsing Library 0.78 +0.34 Constrained composition of 23 predefined primitives
+ 3 In-context Examples 0.84 +0.06 Demonstrating query decomposition & nested logic
+ Label Space Constraint 0.87 +0.03 Automatic runtime exceptions for out-of-vocabulary labels

Key Findings

  • LLM code generation exhibits extraordinary reasoning bounds: Applied to ground-truth structured annotations (Agent + Oracle), Qwen3-8B reaches an F1-Score of 0.87 (1.00 on single attributes, 0.94 on rare events), confirming that code generation over domain primitives reliably translates complex natural language ethograms into executable logic.
  • Perception remains the critical operational bottleneck: The gap between Agent + SALMA (0.34) and Agent + Oracle (0.87) underscores that extracting fine-grained animal actions under night vision, severe motion blur, and vegetation occlusions (e.g., Action F1 of 0.29) is the true performance barrier rather than linguistic reasoning.
  • Foundation VLMs struggle with fine-grained ecological domains: Zero-shot foundation VLMs achieved F1-Scores under 0.18, and even supervised fine-tuning (CLIP4Clipโ€  at 0.26) lagged significantly behind our structured trajectory extraction and symbolic retrieval pipeline.

Highlights & Insights

  • Decoupled code synthesis eliminates hallucination risk: Banning the LLM from processing raw video features or full JSON files directly prevents hallucinated matches across large candidate sets. The LLM acts purely as an intent compiler producing deterministic Python code.
  • White-box explainability for scientific practitioners: Instead of opaque similarity scores, researchers receive inspectable Python predicate functions and frame-level SALMA bounding boxes. Domain experts can verify whether a false retrieval stemmed from a perceptual misclassification or logical translation error, and can interactively tweak the code.
  • Support for non-matching queries and cross-video comparisons: Prompting-MammAlps explicitly benchmarks queries associated with zero matching videos and comparative queries referencing an external exemplar video, scenarios where conventional dot-product retrieval models fundamentally fail.

Limitations & Future Work

  • Binary filtering lacks continuous relevance ranking: The current system provides a discrete boolean set decision rather than continuous ranking scores. When queries involve soft spatial-temporal proximities (e.g., "a juvenile following an adult female"), rigid boolean logic can cause abrupt recall drops.
  • Perceptual fragility in fine-grained actions: SALMA achieves only 0.29 macro-F1 on subtle low-level actions. Extreme darkness, infrared lighting artifacts, and animal occlusions in dense alpine underbrush challenge vision transformer tracking heads.
  • Future directions: The authors outline incorporating soft score outputs into primitive operators to support ranking-based TVR, as well as integrating synchronized environmental audio streams (e.g., rustling leaves, vocalizations) to disambiguate occluded visual events.
  • vs CLIP4Clip / InternVideo2.0: Conventional retrieval embeds text and video into a shared metric space. While computationally direct, it fails on fine-grained ethological queries requiring temporal sequence verification. This paper's offline structured extraction paired with online code execution substantially outperforms latent embeddings on multi-entity and compositional reasoning.
  • vs X-CoT Re-ranking: X-CoT performs pairwise video comparisons using LLM Chain-of-Thought, creating an \(O(k^2)\) computational bottleneck that is highly prone to hallucinated explanations. This work compiles the query once into a lightweight Python script that executes over thousands of videos in milliseconds with zero hallucination risk.

Rating

  • Novelty: โญโญโญโญโญ First dedicated text-to-video retrieval benchmark for wildlife camera-traps; elegant code-synthesis retrieval paradigm.
  • Experimental Thoroughness: โญโญโญโญโญ Densely annotated 18-hour field dataset, 135 complex queries, thorough tracking/attribute evaluations, and multi-LLM ablations.
  • Writing Quality: โญโญโญโญโญ Rigorous methodology, crisp motivation, and clear empirical diagnostic breakdowns.
  • Value: โญโญโญโญโญ Exceptional value for computational ecology, remote wildlife monitoring, and neuroethology data management.