Skip to content

H2SVC: Head-aware Heterogeneous Streaming Video Cache for Online Video Understanding

Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Area: Video Understanding
Keywords: Streaming Video, KV Cache Compression, Attention Heterogeneity, Semantic Trajectory Curvature Eviction, Multimodal Large Language Models

TL;DR

Reveals the intrinsic temporal receptive field heterogeneity among MLLM attention heads, proposing a training-free Head-aware Heterogeneous Streaming Video Cache (H2SVC) framework that combines offline static routing with geometric trajectory curvature eviction to achieve streaming SOTA while slashing KV memory by over 50%.

Background & Motivation

Deploying Multimodal Large Language Models (MLLMs) in real-world streaming video scenarios—such as continuous autonomous driving, active robotics, and live video surveillance—encounters an intractable memory bottleneck. The autoregressive Key-Value (KV) cache is indispensable for eliminating redundant computation across sequential tokens, yet as incoming video streams extend indefinitely over time, the visual KV cache scales linearly. In practice, a standard ten-minute video at just 1 frame per second (FPS) quickly exhausts the onboard VRAM of most high-end GPUs, precipitating out-of-memory failures and severely constraining continuous real-time video understanding.

To alleviate this memory pressure, existing techniques split into architectural modifications and training-free cache compression. Architectural approaches (such as TimeChat-Online and StreamForest) implement learned visual token-dropping or persistent event tree structures; however, they require computationally expensive supervised fine-tuning over large-scale streaming datasets and fundamentally alter model architectures, rendering them incompatible with frozen off-the-shelf MLLMs. Meanwhile, training-free KV cache compression strategies (such as FIFO sliding windows, similarity-based merging, or prompt-guided retrieval) uniformly apply an identical compression policy across all attention heads within a layer. This uniform design treats every head as an identical temporal observer, overlooking the fact that individual attention heads inherently serve distinct functional roles. Furthermore, runtime attention-guided retrieval relies on materializing full \(O(N^2)\) attention matrices, which breaks compatibility with IO-aware hardware kernels like FlashAttention.

Through systematic Spatio-temporal Attention Profiling (SAP), this work demonstrates that attention heads naturally and robustly partition into specialized temporal roles that remain input-agnostic and invariant across diverse videos and prompts: instantaneous heads capturing current spatial details, short-term heads tracking recent motions, and episodic heads maintaining long-range video history. The core idea is to exploit the intrinsic temporal heterogeneity of attention heads by classifying and routing each head's KV cache offline into dedicated memory banks with tailored temporal budgets, while applying online Semantic Trajectory Curvature Eviction (STCE) on Value states in the episodic bank to continuously preserve sharp semantic transitions while evicting predictable redundant frames in full compatibility with FlashAttention.

Method

Overall Architecture

H2SVC is a training-free KV cache compression framework tailored for streaming video processing with infinite sequence length. The operational pipeline proceeds in two primary phases: an offline profiling and classification phase, followed by an online inference phase with static routing and geometry-aware eviction. Prior to deployment, Spatio-temporal Attention Profiling (SAP) evaluates the look-back distance of each KV head across visual prefilling and text querying stages on a small set of sample videos, producing a frozen head-classification mapping. During streaming inference, each head's KV vectors are routed into one of three dedicated memory banks based on this mapping. For the global memory bank of episodic heads, Value states are tracked as a geometric trajectory to continuously evict predictable transitional frames at runtime.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Streaming Video Frames Input<br/>Frame-by-frame visual encoding"] --> B["Spatio-temporal Attention Profiling & Offline KV Routing<br/>TAD score ranking & 3-way head partition"]
    B --> C["Heterogeneous Dedicated Memory Banks<br/>Specialized cache with distinct temporal windows"]
    C -->|Bottom 10% Instantaneous| D1["Ultra-Short Sliding Cache<br/>FIFO window Ws=4 frames"]
    C -->|Middle 40% Short-term| D2["Streaming Sliding Cache<br/>FIFO window Ww=16 frames"]
    C -->|Top 50% Episodic| D3["Semantic Trajectory Curvature Eviction<br/>Global bounded capacity Ne=64 frames"]
    D1 & D2 & D3 --> E["Multi-Head Attention Concatenation<br/>FlashAttention accelerated decode"]
    E --> F["Streaming Temporal Reasoning Output"]

Key Designs

1. Spatio-temporal Attention Profiling & Offline KV Routing: breaking the uniform intra-layer assumption Existing compression approaches compress KV caches either uniformly or purely across layers (Layer-wise), completely disregarding functional specialization among heads within the same layer. To rigorously characterize the temporal receptive field of each KV head \(h\), H2SVC introduces two profiling metrics: Vision-to-Vision Temporal Attention Distance (\(\text{TAD}_{\text{V2V}}^{(h)}\)) and Text-to-Vision Temporal Attention Distance (\(\text{TAD}_{\text{T2V}}^{(h)}\)). During visual prefilling, \(\text{TAD}_{\text{V2V}}^{(h)}\) measures the expected look-back distance from newly arriving query frames to past visual tokens: $\(\text{TAD}_{\text{V2V}}^{(h)} = \mathbb{E}_{v \in \mathcal{V}} \left[ \frac{1}{N_v} \sum_{i=1}^{N_v} \mathbb{E}_{q \in F_i} \sum_{j=1}^{i} (i - j) \cdot A_{q, F_j}^{(h)} \right]\)$ where \(F_i\) and \(F_j\) are the current query frame and historical key frame, respectively, and \(A_{q, F_j}^{(h)}\) is the normalized attention weight. Similarly, \(\text{TAD}_{\text{T2V}}^{(h)}\) computes the expected look-back distance of text query tokens across the entire visual sequence. Empirical profiling reveals that \(\text{TAD}_{\text{V2V}}\) and \(\text{TAD}_{\text{T2V}}\) correlate strongly, while their standard deviation across different input videos is remarkably low, demonstrating that temporal receptive preferences are intrinsic to the pre-trained weights. H2SVC averages their min-max normalized values into a unified Temporal Receptive Score \(S(h) = \frac{1}{2}(\mathcal{N}(\text{TAD}_{\text{V2V}}^{(h)}) + \mathcal{N}(\text{TAD}_{\text{T2V}}^{(h)}))\), sorting all heads into three categories offline: Instantaneous Heads (bottom 10%), Short-term Heads (middle 40%), and Episodic Heads (top 50%). This classification is executed only once offline, incurring zero dynamic profiling overhead during inference.

2. Heterogeneous Dedicated Memory Banks: tailored allocation for varying receptive fields Because different heads exhibit vastly different attention spans, maintaining identical cache budgets causes either wasteful storage of redundant background tokens or premature eviction of long-range contexts. H2SVC designs three dedicated memory banks to accommodate distinct head behaviors: - Ultra-Short Cache (Instantaneous Bank): Tailored for instantaneous heads. Profiling demonstrates that over 90% of their attention mass concentrates on intra-frame visual tokens of the current frame, with negligible historical look-back. A tiny FIFO sliding window (\(W_s = 4\) frames by default) is allocated, drastically reducing memory usage. - Streaming Cache (Short-term Bank): Tailored for short-term heads. These heads track recent temporal dynamics and local motion continuity. A standard FIFO sliding window (\(W_w = 16\) frames by default) captures short-term temporal interactions. - Global Cache (Episodic Bank): Tailored for episodic heads. These heads require a comprehensive global temporal view to resolve complex multi-event queries. They receive a larger bounded capacity (\(N_e = 64\) frames by default). Rather than relying on FIFO eviction, this bank applies geometry-aware eviction to retain critical semantic transitions indefinitely. During Multi-Head Attention, queries interact with their respective head-specific KV banks, and the resulting representations are concatenated along the head dimension, introducing zero architectural alterations.

3. Semantic Trajectory Curvature Eviction: geometric inflection identification in Value space Within the episodic global bank, uniform time downsampling (Uniform Eviction) ignores semantic content density, while pairwise cosine similarity (Similarity Eviction) drops key frames in continuous motions or retains predictable intermediate frames along smooth paths. H2SVC models the sequence of cached frame representations as a geometric trajectory in latent feature space. Frames positioned along a straight, smooth path are semantically predictable by linear interpolation, whereas frames located at sharp directional turns indicate semantic inflection points or scene event boundaries. Crucially, this trajectory must be constructed strictly using Value states (\(V\)) rather than Key states (\(K\)). Rotary Position Embeddings (RoPE) inject absolute temporal position into Key states, thereby introducing a rotational phase bias that warps latent geometric distances and corrupts angle measurements. Given cached Value vectors \(\{V_j\}_{j=1}^{N_e}\), we define the incoming velocity \(\vec{v}_{\text{in}} = V_j - V_{j-1}\) and outgoing velocity \(\vec{v}_{\text{out}} = V_{j+1} - V_j\). The importance score \(I(j)\) for frame \(F_j\) is formulated as: $\(I(j) = \underbrace{\left[1 - \text{CosSim}(\vec{v}_{\text{in}}, \vec{v}_{\text{out}})\right]}_{\text{Angular Curvature}} \cdot \underbrace{\min\left(\|\vec{v}_{\text{in}}\|_2, \|\vec{v}_{\text{out}}\|_2\right)}_{\text{Magnitude Gating}} \cdot \underbrace{\log(t_{j+1} - t_{j-1})}_{\text{Temporal Weight}}\)$ Here, Angular Curvature captures the directional change of the trajectory (sharp turns yield higher scores); Magnitude Gating suppresses artificial high-curvature noise arising from minute feature fluctuations in static scenes; and Temporal Weight (where \(t_j\) is the original frame index) penalizes closely clustered adjacent frames to preserve temporal diversity. When the bank is full, the frame with the lowest score \(I(j)\) is evicted upon frame arrival. Evicting frame \(F_j\) only invalidates the scores of its immediate neighbors \(F_{j-1}\) and \(F_{j+1}\), requiring only two local score updates with constant \(O(1)\) complexity per step, scaling effortlessly to infinite streams.

A Worked Example

Consider a 1000-frame streaming surveillance feed under the default budget configuration (\(N_e=64, W_w=16, W_s=4\)): 1. Streaming Ingestion & Routing: When frame 65 arrives, the bottom 10% instantaneous heads retain only frames 62–65 in their Ultra-Short Cache. The middle 40% short-term heads retain frames 50–65 in their Streaming Cache. 2. Global Cache Overflow: The top 50% episodic heads have filled their capacity of 64 frames, queueing frame 65 for insertion into the Global Cache. 3. Curvature Calculation & Eviction: Scores are computed across cached Value states. Between frames 20 and 35, a vehicle drives straight at constant speed; velocity vectors \(\vec{v}_{\text{in}}\) and \(\vec{v}_{\text{out}}\) align almost collinearly, producing near-zero angular curvature. At frame 40, the vehicle abruptly brakes and turns, causing a sharp angular turn in Value space. Frame 28 along the smooth path receives the lowest score and is evicted. 4. Local Update: Only frames 27 and 29 are rescored locally against their new adjacent neighbors, taking under 0.1 ms. The pruned KV states across all banks are immediately ingested into FlashAttention for text generation.

Key Experimental Results

Main Results

H2SVC was extensively evaluated on standard online streaming benchmarks (StreamingBench, OVO-Bench, ODV-bench) across three foundational backbones (LLaVA-OneVision-7B, LLaVA-Video-7B, and Qwen2.5-VL-7B). It consistently sets new state-of-the-art results among training-free offline-to-online methods, surpassing full-context offline baselines on long-range backward retrieval tasks.

Backbone & Method Method Category Sampling Rate / Budget StreamingBench (Real-Time) OVO-Bench (Real-Time) OVO-Bench (Backward) OVO-Bench (Avg.) ODV-Bench (Avg.)
LLaVA-OV-7B (Full Context) Offline Baseline 32 frames 71.1 61.5 43.4 52.5 51.3
+ ReKV Training-free Streaming 0.5 fps 69.1 64.4 46.7 55.6 50.9
+ LiveVLM Training-free Streaming 0.5 fps 72.9 - - 53.5 50.8
+ StreamKV Training-free Streaming 0.5 fps 68.8 - - - -
+ H2SVC (Ours) Training-free Streaming 0.5 fps 73.9 65.7 56.3 61.0 52.4
LLaVA-Video-7B (Full Context) Offline Baseline 64 frames 75.4 63.0 52.7 57.8 52.6
+ ReKV Training-free Streaming 1 fps 72.9 - - 57.1 52.1
+ LiveVLM Training-free Streaming 1 fps 76.9 - - 59.0 51.9
+ H2SVC (Ours) Training-free Streaming 1 fps 77.5 67.2 58.0 62.6 53.1
Qwen2.5-VL-7B (Full Context) Offline Baseline 1 fps 72.6 64.7 48.3 56.5 57.4
+ LiveVLM Training-free Streaming 1 fps 73.2 - - 56.3 57.3
+ InfiniPot-V Training-free Streaming 1 fps 76.4 65.9 47.6 56.8 -
+ H2SVC (Ours) Training-free Streaming 1 fps 77.0 67.7 52.8 60.3 58.7

Ablation Study

Ablations on LLaVA-Video-7B across classification granularity, head allocation ratios, bank capacities, and eviction components evaluated on Video-MME and OVO-Bench:

Dimension / Setting Visual KV Size/Layer Video-MME (Medium) Video-MME (Long) Video-MME (All) OVO-Bench (Real-Time) OVO-Bench (Backward) OVO-Bench (Avg.)
Granularity: Layer-wise Routing 7K 58.8 52.0 60.5 66.1 57.7 61.9
Granularity: Head-wise Routing (Ours) 7K 60.7 53.4 63.0 67.2 58.0 62.6
Head Ratio: All Instantaneous (1.00 : 0.00 : 0.00) 0.7K 49.0 45.7 49.1 70.1 54.5 62.3
Head Ratio: All Short-term (0.00 : 1.00 : 0.00) 3K 51.6 48.7 54.0 68.0 57.1 62.6
Head Ratio: All Episodic (0.00 : 0.00 : 1.00) 12K 61.8 53.9 64.0 63.6 55.5 59.5
Head Ratio: Default Ratio (0.10 : 0.40 : 0.50) 7K 60.7 53.4 63.0 67.2 58.0 62.6
Eviction: Uniform Eviction 7K 60.9 51.7 61.7 65.7 57.5 61.6
Eviction: Similarity Eviction 7K 60.5 52.6 62.0 66.1 56.3 61.2
Eviction: STCE with Key Trajectory 7K 60.4 53.1 62.7 66.5 56.7 62.1
Eviction: STCE w/o Magnitude Gating 7K 60.2 52.7 62.6 65.5 57.3 61.4
Eviction: STCE w/o Temporal Weight 7K 60.3 52.8 62.7 68.0 57.0 62.5
Eviction: STCE Full Design (Ours) 7K 60.7 53.4 63.0 67.2 58.0 62.6
Full Context (Offline Baseline) 12K 61.6 52.3 63.3 63.0 52.7 57.8

Key Findings

  • Head-wise classification markedly beats Layer-wise routing: Averaging TAD scores across an entire layer drops Video-MME performance by 2.5 points (60.5 vs. 63.0). This confirms that intra-layer head variance is substantial, and conflating heterogeneous heads under a single layer-level policy degrades temporal representation.
  • Complementarity across memory banks: Allocating all heads to the episodic bank drops OVO-Bench to 59.5 due to lost short-term fine-grained cues. Conversely, allocating all heads to instantaneous or short-term banks decimates long-form reasoning (Video-MME plunging to 49.1 and 54.0). The 10%:40%:50% ratio optimally harmonizes local motion tracking with long-range reasoning.
  • Value-space curvature modeling is critical: Switching from Value to Key trajectory modeling drops OVO-Bench by 0.5 points, corroborating that RoPE position bias distorts geometric distances. Removing magnitude gating drops performance by 1.2 points, demonstrating the necessity of filtering noise in static scenes.
  • High throughput and bounded VRAM scaling: On an NVIDIA H800 GPU processing a 512-second stream, full KV cache consumption surges to 76.2 GB while throughput degrades to 26.8 tok/sec. In contrast, H2SVC limits VRAM to 19.0 GB (a 75.1% memory reduction) and maintains a steady decoding throughput of 42.7 tok/sec (+59.3% speedup).

Highlights & Insights

  • Empirical discovery of head temporal heterogeneity: Uncovers that attention heads naturally specialize into distinct, input-agnostic temporal receptive roles across pre-trained weights, disproving the long-standing assumption of intra-layer head homogeneity.
  • Geometry-aware curvature metric in Value space: Astutely avoids Key-space RoPE positional distortion, leveraging geometric velocity vectors, angular curvature, and magnitude gating on Value states to filter out predictable frames in \(O(1)\) update time.
  • Native FlashAttention compatibility: Eliminates runtime dependency on dynamic attention matrices and query-guided retrieval, operating purely on KV tensors to fully exploit hardware-accelerated IO kernels.

Limitations & Future Work

  • Static heuristic thresholds: The head distribution ratio (10%:40%:50%) and bank budget parameters (\(N_e=64, W_w=16, W_s=4\)) are determined via offline grid search rather than learned dynamically according to varying video complexity.
  • Discontinuous camera dynamics: In highly erratic, shaky egocentric video captures, rapid camera panning may disrupt latent trajectory smoothness, causing minor non-event motions to be misinterpreted as semantic inflections.
  • Future directions: Exploring lightweight adaptive gating mechanisms or reinforcement learning agents to dynamically scale head capacities according to incoming visual entropy.
  • vs LiveVLM / ReKV: Prior training-free methods rely on attention-guided retrieval or query-aware clustering, requiring explicit materialization of attention matrices during inference. H2SVC operates strictly on raw KV states via static routing and local curvature updates, maintaining full hardware compatibility with FlashAttention.
  • vs StreamingLLM / H2O: LLM-centric token eviction methods target discrete text sequences, struggling to distinguish genuine video events from smooth redundant frames. H2SVC models continuous geometric trajectories tailored specifically to continuous visual media.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering discovery of head temporal heterogeneity in MLLMs; elegant formulation of geometric curvature eviction on Value states.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous validation across 7 online and offline benchmarks, 3 model families, with comprehensive scaling and ablation studies.
  • Writing Quality: ⭐⭐⭐⭐⭐ Precise mathematical formulations, cohesive prose structure, and clear visual illustrations.
  • Value: ⭐⭐⭐⭐⭐ Plug-and-play, training-free, and hardware-friendly, establishing an impactful blueprint for real-time infinite video streaming inference.