SE-DETR: Explicit Semantic Exploration for Generalizability and Distinguishability in Video Temporal Grounding¶
Conference: ECCV 2026
Paper: ECCV 2026 Official Page
Code: https://github.com/hu-cheng-yang/ECCV26-SE-DETR
Area: Object Detection
Keywords: Video Temporal Grounding, Cross-Modal Alignment, Semantic Representation, DETR Architecture, Self-Information Modulation
TL;DR¶
SE-DETR resolves the limitations of word fragmentation and critical sparse semantic dilution in video temporal grounding by introducing Semantics-Proxied Alignment (SPA) and self-information-based Temporal Sparsity Modulation (TSM), establishing new state-of-the-art results across five public benchmarks.
Background & Motivation¶
Video Temporal Grounding (VTG) requires models to localize specific temporal intervals within untrimmed videos according to natural language prompts, encompassing three primary sub-tasks: Video Moment Retrieval (VMR), Highlight Detection (HD), and Video Summarization (VS). Ever since Moment-DETR unified the field with detection transformers, follow-up methods such as QD-DETR, CG-DETR, Keyword-DETR, and R2-Tuning have focused on optimizing cross-attention operations between text tokens and visual patches. However, these methods implicitly assume an ideal one-to-one mapping between individual query words and visual concepts, ignoring two fundamental semantic characteristics inherent in video-language grounding.
The first bottleneck is inter-video semantic generalizability. In natural language queries, different lexical terms often refer to essentially the same visual concept—for example, "meals", "hamburgers", and "food" share overlapping visual semantics. Conventional word-level alignment treats each word token independently, which fragments shared conceptual knowledge and undermines the model's ability to generalize across synonyms and taxonomical variations. The second bottleneck is intra-video semantic distinguishability. A typical query often contains static background references (such as "woman" or "in the car") that persist across nearly the entire video, whereas the truly informative action specifying the temporal boundaries (such as "wearing a mask around her chin" or "waterfall diving") is temporally sparse and transient. When treating all query tokens uniformly, these sparse yet pivotal semantic signals are easily overshadowed and diluted by dominant background features.
Confronted with the dual dilemmas of lexical isolation and sparse temporal cues being swamped by background noise, this work shifts from merely deepening attention networks to explicit semantic modeling. Core idea: introduce dynamically updated learnable semantic proxies to anchor shared visual concepts across vocabularies (SPA), and deploy information-theoretic self-information to quantify temporal sparsity and reweight word-guided visual features (TSM), yielding a unified detection transformer with enhanced generalizability and distinguishability.
Method¶
Overall Architecture¶
The overall architecture of SE-DETR comprises feature projection, Semantics-Proxied Alignment (SPA), Temporal Sparsity Modulation (TSM), multi-layer Adaptive Feature Fusion (AFF), and a multi-task Transformer decoder with dedicated prediction heads. The input consists of patch features \(V \in \mathbb{R}^{B \times T \times (P+1) \times D_V}\) extracted from intermediate layers of a CLIP vision encoder and text query features \(Q \in \mathbb{R}^{B \times L \times D_Q}\). Both modalities are first mapped to a common dimension \(D=256\) via MLPs. Next, the SPA module aggregates word-guided visual features and performs concept-level alignment against global semantic proxies. The TSM module estimates the occurrence probability of word-related visual components and dynamically modulates their weights according to temporal self-information. For the multi-layer extension SE-DETR+, the AFF module integrates multi-layer representations via top-down learnable gating. Finally, the Transformer decoder aggregates temporal representations to predict moment intervals and highlight saliency scores.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multimodal Inputs<br/>CLIP Video Patches V and Text Query Q"] --> B["Feature Projection & Word-Guided Aggregation<br/>MLP mapping to dimension D and computing V_Q"]
B --> C["Semantics-Proxied Alignment (SPA)<br/>Proxy assignment and cross-modal contrastive alignment"]
C --> D["Temporal Sparsity Modulation (TSM)<br/>Cross-frame similarity self-information & dynamic reweighting"]
D --> E["Adaptive Feature Fusion (AFF)<br/>Top-down adaptive gated multi-layer fusion"]
E --> F["Transformer Decoder & Multitask Heads<br/>VMR boundary regression, HD saliency and VS focal loss"]
Key Designs¶
1. Semantics-Proxied Alignment (SPA): Bridging Lexical Gaps via Shared Concept Anchors
To address the semantic fragmentation caused by directly aligning visual features with discrete word tokens, SPA introduces a set of learnable continuous semantic proxies \(C \in \mathbb{R}^{C_N \times D}\). To avoid trivial solutions from arbitrary initialization, the proxies are initialized using K-means cluster centroids computed over the global [CLS] embeddings of all training queries, providing balanced coverage of the textual feature space. During training, visual patch features \(V_P\) interact with query tokens \(Q\) to form word-guided visual representations \(V_Q \in \mathbb{R}^{B \times T \times L \times D}\). SPA then optimizes representations via a two-step scheme: in the assignment step, positive video clips \(V_Q^{POS}\) are sampled from the ground-truth interval, and each query word is assigned to its nearest proxy based on cosine similarity, obtaining target assignments \(A = \arg\max_{C_N}(\text{Softmax}(\cos(Q, C)))\); in the optimization step, cross-entropy contrastive objectives pull both the textual word tokens and positive visual representations toward their assigned proxy while repelling other proxies:
Because semantic proxies serve strictly as optimization regularizers during training, they are completely discarded during inference, introducing zero inference overhead while ensuring that related vocabularies (e.g., "meals" and "food") naturally converge toward unified concept representations.
2. Temporal Sparsity Modulation (TSM): Suppressing Temporal Redundancy with Self-Information
Static background concepts (e.g., "in the car") often dominate feature variance due to their persistent presence across video frames, obscuring transient target moments (e.g., "wearing a mask"). TSM introduces an unsupervised formulation grounded in Shannon's information theory to measure concept unexpectedness. For each word-aggregated visual component, TSM computes the cross-frame cosine similarity matrix \(\mathcal{A}_{V_Q} \in [0, 1]^{B \times L \times T \times T}\) and uses the average similarity of the upper triangular elements as an empirical occurrence probability. The self-information score is formulated as the negative logarithm of this probability:
To prevent spurious over-amplification in naturally volatile videos, TSM measures the overall temporal consistency \(\mathcal{S}_{CLS} \in [0, 1]^B\) using visual [CLS] tokens. When background consistency is high (\(\mathcal{S}_{CLS} \to 1\)), the risk of salient events being drowned out is highest, and the modulation weights \(\mathcal{W}_{V_Q} = \text{Softmax}(\mathcal{S}_{V_Q} \cdot \mathcal{S}_{CLS})\) strongly boost sparse semantic components; conversely, when the video displays high intrinsic variability (\(\mathcal{S}_{CLS} \to 0\)), the modulation smoothly relaxes to a uniform distribution. The modulated visual stream is further constrained by a margin-based triplet loss \(\mathcal{L}_T\) (\(m=0.5\)) applied to positive and negative clips sampled relative to ground-truth boundaries \([t_s, t_e]\), reinforcing sharp temporal distinguishability.
3. Adaptive Feature Fusion (AFF): Top-Down Layer-Wise Feature Integration
To exploit the complementary benefits of hierarchical representations, SE-DETR is extended to SE-DETR+ using multi-layer CLIP features (\(K=4\)), where intermediate layers preserve fine spatial details while deeper layers encode abstract semantics. AFF fuses features progressively from deep to shallow layers (\(K \to 1\)). For the \(k\)-th layer feature \(V_{O_k}\) and the accumulated representation \(\hat{V}_{O_{k+1}}\), AFF concatenates them along the channel dimension and computes an adaptive frame-level gating weight \(g = \text{Sigmoid}(\text{MLP}([V_{O_k}, \hat{V}_{O_{k+1}}]))\). The blended features \(g \cdot V_{O_k} + (1-g) \cdot \hat{V}_{O_{k+1}}\) are subsequently refined through a Transformer decoder layer. Layer-wise contrastive loss \(\mathcal{L}_{K-Align}\) ensures multimodal alignment across all feature depths, preventing high-frequency noise from corrupting semantic localization.
Loss & Training¶
The framework is optimized end-to-end under a multi-task objective. The VMR task employs smooth L1 regression \(\mathcal{L}_{VMR} = |t_s - \hat{t}_s| + |t_e - \hat{t}_e|\); the HD task applies InfoNCE contrastive loss \(\mathcal{L}_{HD}\) on positive and negative clip-query similarities (\(\tau=0.07\)); and the VS task utilizes Focal Loss \(\mathcal{L}_{VS}\) (\(\beta=0.9, \gamma=0.2\)). The total loss is formulated as:
The task loss weights are set to \(\lambda_{VMR}=1.0\), \(\lambda_{VS}=0.2\), \(\lambda_{HD}=0.1\), and \(\lambda_{Align}=0.1\). The number of proxies \(C_N\) is configured as 1000 for QVHighlight, 100 for Charades-STA/TACoS/Youtube-HL, and 10 for TVSum.
Key Experimental Results¶
Main Results¶
SE-DETR was benchmarked against leading methods across five public datasets spanning all three VTG tasks. The table below details performance on the primary QVHighlight benchmark test and validation splits.
| Method | Feature | Test VMR Avg mAP | Test VMR [email protected] | Test VMR [email protected] | Test HD mAP | Test HD HIT@1 | Val VMR Avg mAP | Val HD mAP |
|---|---|---|---|---|---|---|---|---|
| Moment-DETR (NeurIPS 21) | CLIP+SlowFast | 30.73 | 52.89 | 33.02 | 35.69 | 55.60 | 32.20 | 35.65 |
| UMT (CVPR 22) | CLIP+SlowFast | 36.12 | 56.23 | 41.18 | 38.18 | 59.99 | 38.59 | 39.85 |
| QD-DETR (CVPR 23) | CLIP+SlowFast | 39.86 | 62.40 | 44.98 | 38.94 | 62.40 | 41.22 | 39.13 |
| UniVTG (ICCV 23) | CLIP+SlowFast | 35.47 | 58.86 | 40.86 | 38.20 | 60.96 | 36.13 | 38.83 |
| CG-DETR (NeurIPS 23) | CLIP+SlowFast | 42.90 | 65.40 | 48.40 | 40.30 | 66.20 | 44.90 | 40.80 |
| Keyword-DETR (AAAI 25) | CLIP+SlowFast | 45.69 | 66.86 | 51.23 | 40.94 | 64.79 | 47.69 | 41.67 |
| R2-Tuning (ECCV 24) | CLIP (K=4) | 46.17 | 68.03 | 49.35 | 40.75 | 64.20 | 47.86 | 39.45 |
| SE-DETR (Ours K=1) | CLIP (K=1) | 45.73 | 68.03 | 48.88 | 40.61 | 65.84 | 47.57 | 40.04 |
| SE-DETR+ (Ours K=4) | CLIP (K=4) | 46.42 | 68.22 | 49.96 | 40.82 | 66.40 | 48.29 | 40.27 |
On dedicated VMR benchmarks, SE-DETR+ achieves 51.2 mIoU on Charades-STA and 36.2 mIoU on TACoS, outperforming both 2D-TAN and UniVTG. On Youtube-HL highlight detection, it achieves 77.0 average mAP without audio inputs, surpassing multimodal baselines such as UMT (74.9). On TVSum video summarization, it reaches 89.6 Top-5 mAP, establishing a new state of the art.
Ablation Study¶
The contribution of each component was systematically evaluated on the QVHighlight validation split (removing modules leaves standard DETR with average feature pooling):
| Config | SPA | TSM | AFF | VMR [email protected] | VMR [email protected] | VMR Avg mAP | HD mAP | HD HIT@1 | Note |
|---|---|---|---|---|---|---|---|---|---|
| Baseline | - | - | - | 62.19 | 46.26 | 42.15 | 37.40 | 61.35 | Standard DETR without semantic modules |
| +SPA | ✓ | - | - | 67.74 | 50.77 | 45.29 | 40.02 | 65.41 | Solely adding SPA brings substantial +5.55% gain |
| +TSM | - | ✓ | - | 64.97 | 49.15 | 44.86 | 38.68 | 63.47 | Solely adding TSM validates sparsity importance |
| +AFF | - | - | ✓ | 63.47 | 47.61 | 43.18 | 37.47 | 61.09 | Multi-layer adaptive feature gating |
| SPA + TSM | ✓ | ✓ | - | 68.80 | 52.68 | 47.63 | 40.08 | 66.02 | Dual explicit semantic modules |
| SPA + AFF | ✓ | - | ✓ | 68.73 | 51.65 | 46.72 | 39.88 | 65.51 | Concept proxies with multi-layer features |
| TSM + AFF | - | ✓ | ✓ | 64.26 | 49.03 | 45.00 | 38.05 | 62.39 | Lacks concept generalization |
| Full model (SE-DETR+) | ✓ | ✓ | ✓ | 69.51 | 53.61 | 48.29 | 40.27 | 66.97 | Synergistic combination achieving peak scores |
Key Findings¶
- SPA is the primary driver of performance gains: Adding SPA alone to the baseline boosts VMR [email protected] from 62.19% to 67.74% (+5.55%) and HD mAP from 37.40% to 40.02%. This confirms that token-level cross-attention suffers from severe lexical fragmentation and that concept-level proxies provide essential semantic anchoring.
- TSM and SPA are highly synergistic: Once SPA establishes robust semantic prototypes, TSM sharpens temporal focus on transient actions, elevating VMR [email protected] from 50.77% to 52.68%.
- Proxy capacity exhibits an inverted U-shape curve: Varying proxy count \(C_N\) from 100 to 1000 steadily improves QVHighlight VMR Avg mAP from 46.90 to 48.29. However, setting \(C_N=2000\) degrades performance to 46.92 due to long-tail under-optimization of rarely assigned proxies.
Highlights & Insights¶
- Self-Information as an Unsupervised Saliency Metric: Treating cross-frame feature similarity as an empirical probability distribution and deriving self-information dynamically penalizes pervasive background features and elevates transient action cues without requiring explicit supervision.
- Zero-Inference-Cost Semantic Proxies: Semantic proxies guide multimodal contrastive alignment purely during training and are discarded at inference, delivering concept-level generalization without adding model parameters or computational latency.
- Global Consistency Gating Prevents Over-Amplification: Modulating local self-information with visual [CLS] temporal consistency prevents erratic weight amplification in naturally fast-paced, multi-scene videos.
Limitations & Future Work¶
- Manual Heuristic Proxy Sizing: The number of proxies \(C_N\) is currently treated as a fixed discrete hyperparameter tuned to dataset vocabulary size, lacking dynamic adaptability for open-world streaming video.
- Lack of Multi-Action Sequential Reasoning: TSM computes sparsity independently per word, leaving sequential compositional relationships between multiple actions unmodeled.
- Extending Explicit Semantics to Video-LLMs: Integrating explicit semantic proxies and temporal self-information into the visual token reduction layers of multimodal large language models presents a promising avenue for reducing prompt length and hallucination.
Related Work & Insights¶
- vs R2-Tuning (ECCV 2024): R2-Tuning incorporates frozen multi-layer CLIP features via standard cross-attention. SE-DETR introduces concept-level alignment and temporal sparsity reweighting, demonstrating superior precision under the same 4-layer backbone.
- vs Keyword-DETR (AAAI 2025): Keyword-DETR relies on an auxiliary keyword classification head with hard-coded linguistic heuristics. SE-DETR utilizes self-information derived directly from video dynamics, offering a more robust and unsupervised alternative.
- vs CG-DETR (NeurIPS 2023): CG-DETR injects dummy tokens to prevent passive cross-modal contamination. SE-DETR directly structures the latent space via concept proxies and temporal modulation to resolve semantic interference at its source.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Principled introduction of semantic proxies and temporal self-information modulation.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 5 benchmarks, 3 distinct tasks, and thorough ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Well-structured, mathematically rigorous, and clearly motivated.
- Value: ⭐⭐⭐⭐☆ Provides an elegant, lightweight paradigm for fine-grained video-language grounding.