HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization¶
Conference: ECCV 2026
Paper: ECCV Official Page
PDF: Open Access PDF
Area: Multimodal VLM
Keywords: Video Summarization, Multimodal LLM, Attention Steering, Training-free Inference Control, Temporal Saliency
TL;DR¶
To overcome the context fragmentation and irreversible information loss caused by hard keyframe selection, HAS introduces an inference-time continuous attention steering mechanism that smoothly biases frozen video MLLM cross-attention heads toward salient moments using prompt-conditioned highlight distributions while preserving complete background context.
Background & Motivation¶
With the explosive growth of video generative models, long video archives, and interactive embodied agents, visual data volume has drastically outpaced human consumption capacity, making video summarization essential for rapid navigation, indexation, and content retrieval. Recently, multimodal large language models (Video-MLLMs) have emerged as powerful video reasoners by treating sampled frames as visual token sequences and producing coherent summaries via cross-attention mechanisms. However, existing video summarization frameworks (e.g., LLMVS, AKeyS) predominantly follow a two-stage "score-then-select" paradigm that discretely isolates keyframes or candidate segments before feeding them into downstream language generators.
This hard truncation paradigm suffers from severe structural tensions. First, natural human perception during video summarization involves reviewing the continuous stream and subsequently recalling salient highlights while anchoring them against smooth background transitions, rather than inspecting disjointed image slices in isolation. Second, hard selection permanently eliminates lower-scored background frames that carry vital causal, spatial, and narrative context, creating an irreversible information bottleneck where omitted evidence can never be retrieved downstream regardless of the output token budget. Finally, manually filtering frames at the input layer wastes the intrinsic capacity of large multimodal models to evaluate temporal importance through internal self- and cross-attention mechanisms.
To resolve these dilemmas, this work explores lightweight inference-time control: can we keep the entire continuous video input intact while introducing an external highlight distribution to steer computation and attention softly? Core idea: calibrate prompt-conditioned temporal highlight scores into a continuous distribution, vectorize it into a token-level steering vector in log-space, and inject it as a gated additive bias into a small subset of visual cross-attention heads during decoding to guide MLLMs toward highlight moments without dropping background context.
Method¶
Overall Architecture¶
HAS comprises two primary stages: (i) constructing and calibrating a prompt-conditioned temporal continuous highlight distribution from the input video, and (ii) converting this distribution into visual token-level steering vectors and injecting them into frozen MLLM cross-attention heads via a learned sparse steering policy during autoregressive decoding. Given an input video frame sequence and a user text prompt, an off-the-shelf highlight detector predicts temporal saliency curves, which are normalized and interpolated into a continuous bounded distribution. This distribution is then lifted to token space and applied as an additive logit bias on selected attention heads, enabling the frozen foundation model to softly accentuate salient moments while retaining background context.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Video Frames F + Text Prompt q"] --> B["Stage 1: Continuous Highlight Distribution Calibration<br/>Off-the-shelf Saliency โ Linear Interpolation โ MinMax Normalization"]
B --> C["Stage 2: Token-level Steering Vector Construction<br/>Frame-level Broadcast across P Tokens with Logarithmic Smoothing"]
C --> D["Stage 3: Gated Cross-Attention Steering<br/>Sparse Attention Head Logit Bias while Keeping Backbone Frozen"]
D --> E["Output: Faithful & Coherent Summary O*"]
Key Designs¶
1. Continuous Highlight Distribution Calibration: Resolving Irreversible Information Loss Off-the-shelf moment retrieval and highlight detection modules (e.g., QVHighlights, QD-DETR) typically generate raw temporal relevance curves that exhibit discrete temporal window artifacts and noisy oscillations. Instead of thresholding these scores to drop non-highlight frames, HAS calibrates the raw sequence \(\hat{\mathbf{h}} = \mathcal{H}(\mathbf{F}, q)\) into a smooth, bounded distribution. It first performs 1D linear temporal interpolation to match the exact target video length \(T\), followed by Min-Max normalization to constrain scores strictly within \([0, 1]\). The resulting continuous distribution \(\mathbf{h} = [h_1, \dots, h_T]\) retains non-zero weights across all temporal steps, ensuring every temporal slice remains accessible to the generator for uninterrupted causal reasoning.
2. Token-level Steering Vector Construction: Bridging Frame Priors to Multimodal Token Spaces Because multimodal LLM vision encoders project each video frame into \(P\) distinct visual tokens (spatial patches or spatio-temporal tubelets), scalar frame-level scores cannot be directly integrated into token-level cross-attention matrices. HAS lifts the scalar highlight curve \(\mathbf{h}\) to the token level by broadcasting each frame score \(h_t\) across its corresponding \(P\) visual tokens, computing the steering vector in logarithmic space: $\(\mathbf{V} = \log\left(\mathrm{Repeat}(\mathbf{h}, P) + \epsilon\right) \in \mathbb{R}^{TP}\)$ where \(\epsilon\) is a small positive constant preventing numerical instability and \(-\infty\) collapse. Crucially, applying an additive bias in log-space to pre-softmax attention logits is mathematically equivalent to multiplicative reweighting of unnormalized attention weights. The non-zero floor \(\epsilon\) guarantees that low-salience peace-time frames retain non-zero attention mass, enacting the core design philosophy of "attenuating without forgetting."
3. Gated Cross-Attention Steering: Lightweight Inference-time Intervention Rather than performing costly parameter updates on multi-billion-parameter backbones, HAS intervenes exclusively on a sparse subset of visual cross-attention heads during decoding. For each selected candidate head \((\ell, m) \in \mathcal{S}\) at layer \(\ell\) and head index \(m\), let \(\mathbf{A}^{(\ell, m)}\) represent the pre-softmax attention logits between the current text query and all visual tokens. HAS introduces a continuous sigmoid gate \(g_{\ell, m} = \sigma(a_{\ell, m}) \in [0, 1]\) and a head-specific steering strength scalar \(\beta_{\ell, m}\), adding a row-wise bias to the visual attention logits: $\(\tilde{\mathbf{A}}^{(\ell, m)}_{i, :} = \mathbf{A}^{(\ell, m)}_{i, :} + g_{\ell, m} \beta_{\ell, m} \mathbf{V}\)$ All non-selected heads and self-attention operations proceed entirely unaltered. This gated modulation dynamically directs the generation process toward evidence-rich frames during key descriptive phases while preventing structural degeneration or hallucinations caused by hard token masking.
Loss & Training¶
HAS adopts a completely non-invasive, training-free scheme for the base model: both the video MLLM backbone and the upstream highlight generator remain frozen throughout. The only trainable parameters are the ultra-lightweight policy parameters \(\Theta = \{a_{\ell, m}, \beta_{\ell, m}\}\). These scalar parameters are calibrated once on a small held-out validation set by minimizing the standard teacher-forcing negative log-likelihood (NLL) with an \(L_1\) sparsity regularizer on the active gates: $\(\min_{a, \beta} \mathcal{L}_{\mathrm{NLL}}(a, \beta) + \lambda_s \sum_{(\ell, m) \in \mathcal{S}} g_{\ell, m}\)$ where \(\lambda_s\) governs head activation sparsity. Optimized using Adam in a single lightweight pass, the calibrated policy parameters \(\{a, \beta\}\) are fixed during deployment, requiring zero gradient computation and negligible memory overhead during forward-only inference.
Key Experimental Results¶
Main Results¶
The authors evaluated HAS across three major benchmark families: V2V temporal ranking on SumMe and TVSum, unified cross-modal summarization on VideoXum, and grounded scientific long-form talk summarization on VISTA. All experiments were conducted on Nvidia L40S GPUs.
On human-aligned temporal importance estimation (SumMe and TVSum), evaluated via Kendallโs \(\tau\) and Spearmanโs \(\rho\) rank correlation against human consensus:
| Category | Method | SumMe \(\tau\) | SumMe \(\rho\) | TVSum \(\tau\) | TVSum \(\rho\) |
|---|---|---|---|---|---|
| Baseline | Random | 0.000 | 0.000 | 0.000 | 0.000 |
| Visual-only | VASNet | 0.160 | 0.170 | 0.160 | 0.170 |
| Visual-only | CSTA | 0.246 | 0.274 | 0.194 | 0.255 |
| Visual + Text | CLIP-It | โ | โ | 0.108 | 0.147 |
| Visual + Text | A2Summ | 0.108 | 0.129 | 0.137 | 0.165 |
| LLM-centric | LLMVS (CVPR 2025) | 0.253 | 0.282 | 0.211 | 0.275 |
| LLM-centric | V2Xum-LLaMA | 0.296 | 0.378 | 0.222 | 0.293 |
| Ours | HAS | 0.298 (\(\pm0.02\)) | 0.335 (\(\pm0.03\)) | 0.224 (\(\pm0.02\)) | 0.299 (\(\pm0.03\)) |
On the comprehensive VideoXum cross-modal benchmark, evaluating video-to-text (V2T), video-to-video (V2V), and joint cross-modal semantic alignment (V2VT):
| Method | V2T: BLEU-4 | V2T: ROUGE-L | V2T: CIDEr | V2V: F1 | V2V: Spearman | V2VT: FCLIP | V2VT: Cross-FCLIP |
|---|---|---|---|---|---|---|---|
| Frozen-BLIP | 0.0 | 1.4 | 0.0 | 16.1 | 0.011 | โ | โ |
| Vid2Seq-HCV | 2.7 | 19.8 | 8.3 | 25.1 | โ | 0.899 | 0.200 |
| VTSUM-BLIP | 5.8 | 25.1 | 23.1 | 23.5 | 0.258 | 0.894 | 0.247 |
| V2Xum-LLaMA-7B | 5.8 | 26.3 | 26.9 | 29.0 | 0.298 | 0.931 | 0.253 |
| V2Xum-LLaMA-13B | 5.7 | 26.2 | 25.3 | 31.6 | 0.276 | 0.957 | 0.251 |
| HAS (Ours) | 6.0 (\(\pm0.03\)) | 26.7 (\(\pm0.3\)) | 28.0 (\(\pm0.8\)) | 32.0 (\(\pm0.5\)) | 0.304 (\(\pm0.02\)) | 0.963 (\(\pm0.01\)) | 0.258 (\(\pm0.007\)) |
Ablation Study¶
On the VISTA long-form scientific presentation benchmark, measuring text quality alongside video-grounded alignment (VideoScore) and factual consistency (FactVC):
| Configuration / Backbone | RLsum | BERTScore | VideoScore | FactVC (Factual Consistency) |
|---|---|---|---|---|
| LLaVA-NeXT-Interleave (Zero-shot) | 22.68 | 81.40 | 1.73 | 40.12 |
| mPLUG-Owl3 (Zero-shot) | 22.84 | 81.39 | 1.77 | 42.07 |
| Plan-mPLUG-Owl3 (Zero-shot) | 22.97 | 81.45 | 1.86 | 47.37 |
| GPT-o1 (Zero-shot) | 24.37 | 82.63 | 2.17 | 51.36 |
| Gemini 2.0 (Zero-shot) | 24.29 | 82.64 | 2.02 | 52.02 |
| HAS (Zero-shot, averaged across backbones) | 23.94 | 82.05 | 2.03 | 50.11 |
| mPLUG-Owl3 (Full Fine-tuning) | 32.91 | 84.22 | 3.28 | 71.94 |
| Plan-mPLUG-Owl3 (Full Fine-tuning) | 33.25 | 84.37 | 3.33 | 75.41 |
In zero-shot transfer from SumMe to MR.HiSum (evaluating 50 out-of-domain videos without retraining), VASNet obtained \(\tau=0.364 / \rho=0.364\) and LLMVS achieved \(\tau=0.440 / \rho=0.440\), whereas HAS reached \(\tau=0.450 / \rho=0.450\), demonstrating superior out-of-domain robustness.
Key Findings¶
- Delayed Saturation Demonstrates Continuous Context Retention: As analyzed in Fig. 2, under varying frame budgets \(T\), hard selection baselines plateau prematurely because discarded evidence cannot be recovered downstream. In contrast, HAS sustains steady performance gains in ROUGE-Lsum and transcript fact recall across mid-to-high budgets, proving that soft attention steering preserves distributed evidence essential for long-form reasoning.
- Backbone-Agnostic Plug-and-Play Generalization: Applying HAS across six distinct open-source architectures (mPLUG-Owl3, LLaVA-NeXT-Interleave, Video-LLaVA, LLaMA-VID, Video-ChatGPT, Video-LLaMA) yielded consistent positive \(\Delta \text{FactVC}\) gains (Fig. 3). The consistent factual improvement confirms that the steering mechanism operates as a general control layer independent of specific tokenizer or pretraining recipes.
Highlights & Insights¶
- From Discrete Truncation to Continuous Attention Modulation: Instead of treating temporal saliency as a crude binary gate at the input level, HAS elevates it to an internal attention steering vector in log-space, bridging the gap between external highlight priors and internal MLLM attention dynamics.
- Non-Invasive Inference-time Efficiency: Bypassing backbone finetuning entirely, HAS steers generation through sparse scalar gating parameters optimized once on a small calibration set, introducing virtually zero inference latency overhead.
- Generalizability to Broad Multimodal Reasoning: The concept of translating 1D/2D continuous importance priors into logit biases on cross-attention heads offers an adaptable blueprint for other dense perception tasks, such as grounded document QA, audio moment reasoning, and robotic trajectory summarization.
Limitations & Future Work¶
- Dependency on Upstream Highlight Quality: HAS operates as an attention mediator rather than an end-to-end highlight extractor; noisy or biased initial curves from off-the-shelf detectors can skew attention allocation and omit salient details.
- Absence of Coarse-to-Fine Adaptive Boundaries: While continuous temporal smoothing preserves global context, it may dilute sharp focus needed for transient, highly localized sub-second events. Future work could combine continuous soft steering with local adaptive multi-scale token zooming.
Related Work & Insights¶
- vs LLMVS (CVPR 2025): LLMVS employs a discrete score-then-select workflow that fragments narrative context across isolated frame captions; HAS preserves continuous video context and applies smooth attention bias, avoiding irreversible information loss.
- vs PASTA & FarSight: While PASTA focuses on text-span reweighting and FarSight intervenes on decoding causal masks to suppress hallucinations, HAS is the first to introduce query-conditioned continuous temporal highlight distributions into visual cross-attention heads for generative video summarization.
Rating¶
- Novelty: โญโญโญโญโ (Pioneering application of continuous temporal highlight distributions as inference-time cross-attention steering vectors in video MLLMs)
- Experimental Thoroughness: โญโญโญโญโญ (Rigorous validation spanning V2V, V2T, and V2VT benchmarks across 6 open-source foundation model backbones)
- Writing Quality: โญโญโญโญโ (Cohesive mathematical formulation connecting continuous priors to log-space additive attention modulation)
- Value: โญโญโญโญโญ (Completely training-free, highly practical, and effectively solves the trade-off between highlight focus and context preservation)