TRINITY: A Multi-Perspective Benchmark for Personal-Style Video Highlight Detection¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://huggingface.co/datasets/vanilladucky/TRINITY
Area: Interpretability
Keywords: Video Highlight Detection, Personal Videos, Multi-Perspective Learning, Shared-Private Architecture, Heterogeneous Saliency Modeling
TL;DR¶
Addressing the limitation of traditional event-centric video highlight detection, this paper introduces TRINITY, a multi-perspective benchmark that decomposes saliency into Event, Emotion, and Nature, along with a shared-backbone multi-branch architecture featuring view-specific experts that achieves new state-of-the-art results on Mr. HiSum and YouTube Highlights.
Background & Motivation¶
Video highlight detection aims to automatically localize moments in untrimmed videos that significantly capture human attention and memory. Despite rapid advancements, existing benchmarks predominantly adopt a scenario-bound and narrowly defined conception of saliency. Most datasets are constructed around sports competitions, broadcast television programs, or query-guided moment retrieval, where highlights are equated almost exclusively with salient semantic events or dramatic narrative peaks. While this single-perspective formulation has driven strong progress in controlled domains, it inherently constrains the conceptual scope of what constitutes a highlight and fails to generalize to unconstrained personal videos.
The core tension is that personal-style recordings—such as daily lifelogs, family records, and social media vlogs—naturally lack scripted plot structures or professionally edited climaxes. In these unconstrained videos, engaging moments arise from diverse, multi-causal, and perspective-dependent factors, such as affective facial expressions during interpersonal interactions or visually stunning scenic compositions along a journey. Collapsing these heterogeneous cues into a single scalar highlight score blurs the underlying supervisory signals, causing conflicting gradients and feature ambiguity in temporal models. Furthermore, existing pipelines lack the capacity to account for dramatic temporal variances across distinct saliency types, such as transient affective reactions versus continuous scenic vistas.
To move beyond the constraints of scenario-specific single-score supervision, video highlight detection requires explicitly disentangling heterogeneous saliency mechanisms within a unified framework. Core idea: decompose personal video highlight saliency into three complementary, mutually orthogonal dimensions—Event, Emotion, and Nature—supported by scalable automated annotation pipelines, and deploy a shared-backbone multi-branch Transformer architecture with view-specific experts for parallel multi-perspective highlight localization.
Method¶
Overall Architecture¶
The framework models personal video highlight detection as a multi-task temporal segment scoring task across three distinct perspectives: Event (\(t=e\)), Emotion (\(t=m\)), and Nature (\(t=n\)). Given an untrimmed input video, it is segmented into non-overlapping 5-second intervals to form a sequence of \(N\) temporal segments, with each segment encoded into a 512-dimensional feature embedding via a frozen CLIP ViT-B/32 backbone. The temporal architecture comprises one globally shared temporal backbone \(T_s\) and three perspective-specific private backbones \(\{T_p^t\}_{t \in \{e, m, n\}}\), all instantiated as 5-layer Transformer encoders equipped with 1D Rotary Position Embeddings (RoPE). The shared backbone captures task-agnostic temporal continuity, while the private experts isolate view-specific temporal patterns. The segment-level outputs of the shared and private streams are concatenated and routed to view-specific linear prediction heads to produce calibrated highlight scores.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Video Segment Sequence<br/>CLIP 512-d Features"] --> B["Multi-Perspective Saliency Benchmark Construction<br/>Event/Emotion/Nature"]
B --> C["Shared-Private Dual-Stream Temporal Encoding<br/>Shared Backbone + View-Specific Experts"]
C --> D["Task-Conditioned Deterministic Routing & MTL<br/>Round-Robin Sampling & Combined Backprop"]
D --> E["Parallel Multi-Perspective Highlight Scores<br/>Event/Emotion/Nature Predictions"]
Key Designs¶
1. Multi-Perspective Saliency Benchmark Construction: Disentangling Heterogeneous Highlight Annotations Conventional benchmarks conflate disparate saliency triggers into an ambiguous single ground truth. TRINITY addresses this by establishing structured, reproducible annotation protocols across three orthogonal dimensions. For the Event dimension, the benchmark adopts crowd-sourced replay intensity from Mr. HiSum, aggregating frame-level statistics into 5-second intervals via average pooling while retaining local maximum values to prevent peak dilution. For the Emotion dimension, a two-stage cascaded pipeline is introduced: Stage-1 samples frames at 1 fps and applies EmotiEffLib to categorize 7 basic facial expressions, selecting stable non-neutral segments of at least three consecutive frames to suppress transient micro-expression noise; Stage-2 pads candidate segments with a \(\pm 2\)-second temporal buffer and prompts Qwen2-VL-7B to perform context-aware multimodal verification, retaining only clips where both models reach consensus. For the Nature dimension, Stage-1 employs YOLOv8n to filter out videos where human presence exceeds 30%, followed by CLIP similarity scoring against tailored landscape prompts; Stage-2 utilizes the Everypixel UGC quality assessment model to score sharpness, exposure, and composition, outputting normalized continuous aesthetic scores. This automated protocol yields high-quality, mutually complementary supervision across tens of thousands of videos.
2. Shared-Private Dual-Stream Temporal Encoding: Balancing Global Context with Perspective Specialization Because highlight patterns differ fundamentally across perspectives—emotion-driven highlights are sharp and concentrated (averaging 4.87 seconds), whereas nature-driven highlights are sustained and smooth (averaging 10.34 seconds)—a monolithic temporal encoder suffers representation collapse. The proposed architecture disentangles shared temporal dynamics from task-specific activations by feeding the segment sequence \(X \in \mathbb{R}^{N \times 512}\) into independent linear projections for the shared stream and task-private streams: $$ \begin{aligned} Q_s &= X W_Q^{(s)}, \quad K_s = X W_K^{(s)}, \quad V_s = X W_V^{(s)} \ Q_p^t &= X W_Q^{(p, t)}, \quad K_p^t = X W_K^{(p, t)}, \quad V_p^t = X W_V^{(p, t)} \end{aligned} $$ Both the shared backbone \(T_s\) and private backbones \(T_p^t\) incorporate 1D RoPE within self-attention layers to inject relative temporal order. The shared stream captures overarching contextual progression across the entire timeline, while each private expert sharpens temporal activations specific to its domain. The resulting representations are concatenated at the segment level and mapped to a highlight probability through an independent multi-layer perceptron head \(C^t\): $$ \hat{y}_i^t = \sigma\left(C^t\left([T_s(x_i), T_p^t(x_i)]\right)\right) $$
3. Task-Conditioned Deterministic Routing & Multi-Task Training: Mitigating Gradient Interference In multi-task video learning, combining diverse supervision types (continuous replay counts, binary emotion labels, and aesthetic regression values) under standard soft MoE gating often causes routing instability and destructive gradient conflicts. This work adopts a deterministic, task-conditioned routing strategy: three independent DataLoaders are maintained for the three perspectives. In each training iteration, the model samples mini-batches from each loader in a round-robin manner. A sample from task \(t\) activates only the shared backbone \(T_s\) and the corresponding private branch \(T_p^t\). The binary cross-entropy (BCE) loss is computed independently for each task: $$ \mathcal{L}t = \frac{1}{BN} \sum^t\right) $$ The overall objective aggregates the task losses with equal weights, }^B \sum_{i=1}^N \mathrm{BCE}\left(y_{b, i}^t, \hat{y}_{b, i\(\mathcal{L}_{total} = \sum_{t \in \{e,m,n\}} \lambda_t \mathcal{L}_t\) with \(\lambda_e = \lambda_m = \lambda_n = 1\). Gradients are accumulated across all three task passes before performing a single parameter update. Empirical cosine similarity measurements confirm that cross-task gradients quickly stabilize near zero, preventing negative transfer and catastrophic forgetting.
Loss & Training¶
The entire network is optimized using AdamW during joint pretraining on TRINITY. The training configuration sets the initial learning rate to \(5 \times 10^{-5}\), weight decay to 0.01, mini-batch size to 16, and dropout ratio to 0.2. Early stopping is triggered if validation performance fails to improve across all three tasks for three consecutive epochs. When adapting to downstream benchmarks, the shared temporal backbone and non-target private experts are frozen, fine-tuning only the target private backbone and its prediction head.
Key Experimental Results¶
Main Results¶
The model demonstrates superior highlight localization capabilities across diverse public benchmarks spanning event-centric (Mr. HiSum, YouTube Highlights) and emotion-centric (VEATIC) tasks.
| Dataset / Benchmark | Evaluation Metric | Ours | Prev. SOTA | Gain |
|---|---|---|---|---|
| Mr. HiSum | \(\text{mAP}_{\rho=15\%}\) | 40.98 | 33.83 (SummDiff) | +7.15 |
| Mr. HiSum | \(\text{mAP}_{\rho=50\%}\) | 69.06 | 65.44 (SummDiff) | +3.62 |
| YouTube Highlights (Average) | mAP | 83.82 | 73.00 (PLD-VHD) | +10.82 |
| YouTube Highlights (Gymnastics) | mAP | 91.89 | 75.80 (RASL) | +16.09 |
| YouTube Highlights (Skating) | mAP | 83.95 | 72.50 (SL-Module) | +11.45 |
| YouTube Highlights (Surfing) | mAP | 87.33 | 79.00 (PLD-VHD) | +8.33 |
| VEATIC | SAGR ↑ | 0.8185 | 0.8040 (VDFS) | +0.0145 |
| VEATIC | RMSE ↓ | 0.1749 | 0.1820 (VDFS) | -0.0071 |
Under the TRINITY multi-perspective evaluation benchmark, the proposed architecture consistently outperforms competitive baselines across all three dimensions:
| Perspective & Metric | Ours | PGL-SUM | SL-Module | CSTA | Qwen3-VL-235B (Zero-shot) |
|---|---|---|---|---|---|
| Nature \(\text{mAP}_{\rho=15\%}\) | 58.74 | 46.40 | 38.68 | 36.47 | 27.91 |
| Nature \(\text{mAP}_{\rho=50\%}\) | 68.05 | 61.84 | 62.23 | 39.67 | 58.29 |
| Emotion \(\text{mAP}_{\rho=15\%}\) | 67.42 | 62.47 | 50.55 | 64.96 | 26.42 |
| Emotion \(\text{mAP}_{\rho=50\%}\) | 71.32 | 73.81 | 67.57 | 67.29 | 56.90 |
| Event \(\text{mAP}_{\rho=15\%}\) | 40.98 | 33.61 | 31.46 | 35.97 | 18.10 |
| Event \(\text{mAP}_{\rho=50\%}\) | 69.06 | 61.84 | 62.71 | 40.77 | 52.82 |
Ablation Study¶
The architectural ablation (settings S1–S4) and dataset contribution ablation (settings D1–D3) on TRINITY quantify the benefits of structural factorization and multi-perspective supervision:
| Config | Backbone & Mechanism | Event 15% | Nature 15% | Emotion 15% | Mean 15% | Mean 50% |
|---|---|---|---|---|---|---|
| S1: Single-Task Baseline | 1 shared backbone, no MTL | 32.85 | 43.16 | 61.05 | 45.69 | 68.06 |
| S2: Shared-Backbone MTL | 1 shared backbone + 3 heads | 37.97 | 44.41 | 65.20 | 49.19 | 68.97 |
| S3: Independent Experts | 3 separate backbones, isolated | 37.54 | 44.78 | 63.27 | 48.53 | 69.70 |
| S4: Shared-Private Experts (Ours) | 1 shared + 3 private experts | 40.98 | 58.74 | 67.42 | 55.71 | 69.48 |
| D1: Event-Only Data | S4 architecture trained on Event | 40.25 | 31.25 | 40.83 | 37.44 | 66.93 |
| D2: Event + Nature Data | S4 architecture trained on Event+Nature | 39.98 | 57.89 | 45.98 | 48.25 | 64.32 |
| D3: Full TRINITY Data (Ours) | S4 architecture trained on all three | 40.98 | 58.74 | 67.42 | 55.71 | 69.48 |
Key Findings¶
- Shared Context and Private Specialization are Complementary: Moving from independent experts (S3) to the shared-private architecture (S4) produces substantial gains across all perspectives, boosting average \(\text{mAP}_{\rho=15\%}\) from 48.53 to 55.71 (+7.18). The Nature perspective benefits most dramatically (+13.96 points), indicating that scenic aesthetic saliency requires specialized temporal feature filters layered onto global context.
- Multi-Perspective Supervision Exhibits Positive Transfer: In the dataset ablation, incorporating Nature supervision (D2) increases Nature mAP from 31.25 to 57.89 (+26.64) without degrading Event performance (39.98 vs. 40.25). Adding Emotion data (D3) further boosts performance across all dimensions, confirming that the formulated perspectives are mutually reinforcing.
- Gradient Orthogonality Prevents Interference: Gradient cosine similarity tracking demonstrates that task gradients rapidly decouple after initial warmup, remaining stably near zero throughout training and confirming the efficacy of deterministic round-robin updates.
Highlights & Insights¶
- Paradigm Shift from Monolithic Saliency to Heterogeneous Decomposition: Rather than forcing models to predict a single scalar score across conflicting visual stimuli, TRINITY formalizes video saliency as a multi-dimensional construct comprising event dynamics, emotional expressions, and natural aesthetics.
- Scalable Cascaded Annotation: The multi-stage automated labeling pipeline leverages specialized perceptual models (EmotiEffLib, YOLOv8n) alongside contextual vision-language reasoning (Qwen2-VL-7B) and aesthetic evaluators, ensuring rigorous quality without prohibitive manual annotation costs.
- Inherent Temporal Filtering: Temporal attention maps reveal that perspective-specific backbones selectively focus on semantically relevant moments (conversations, smiles, landscapes) while uniformly suppressing uninformative segments like motion blur and rapid camera pans.
Limitations & Future Work¶
- Author-Acknowledged Limitations: The Emotion perspective relies primarily on visible facial expressions, leaving subtle affect transitions and complex vocal cues underexplored. Additionally, the benchmark does not address highly structured cinematic videos governed by professional montage rules.
- Prospective Extensions: The current framework operates exclusively on visual embeddings (frozen CLIP ViT-B/32). Incorporating audio features (speech cadence, background music, ambient sound) represents a vital next step, as acoustic signals often trigger human emotional and narrative engagement in personal video editing.
Related Work & Insights¶
- vs Mr. HiSum / YouTube Highlights: Existing benchmarks treat highlight detection as a single-label problem, leading to feature ambiguity in unstructured videos. TRINITY offers explicit multi-perspective ground truth for fine-grained video summarization.
- vs QVHighlights / UMT / TR-DETR: Query-conditioned grounding methods require explicit natural language descriptions during inference. TRINITY operates in a text-agnostic setting, generating multi-dimensional saliency curves automatically.
- vs Standard Multi-Task MoE: Unlike dynamic soft-gating architectures prone to expert collapse and gradient competition, the proposed deterministic routing and round-robin optimization preserve expert specialization while sharing temporal context.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Pioneering multi-perspective benchmark tailored for unstructured personal videos with orthogonal saliency definitions.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluations across multiple established benchmarks, rigorous multi-task ablations, gradient conflict tracking, and attention visualizations.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear conceptual formulation, robust methodological descriptions, and detailed empirical analyses.
- Value: ⭐⭐⭐⭐⭐ Highly impactful for intelligent mobile vlogging, automated album compilation, and personalized video summarization.