S3-Prune: Stability-Aware Token Budgeting for Long-Form Video-Language Models¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Area: Video Understanding
Keywords: Video-LLM, Token Pruning, Long-form Video Understanding, Adaptive Budgeting, Kalman Filter
TL;DR¶
To tackle transient observational noise and cumulative budget misallocation in long-form video pruning, S3-Prune formulates token budgeting as state estimation with Kalman filtering, sustaining full baseline accuracy even under extreme 90% token pruning.
Background & Motivation¶
Long-form video-language models (Video-VLMs) have emerged as the standard paradigm for complex temporal reasoning, causal analysis, and grounded event comprehension. However, processing long video sequences requires ingesting dozens or hundreds of high-resolution frames, causing an explosive expansion in the number of visual tokens. Because the self-attention mechanism in large language models exhibits quadratic complexity \(O((F \cdot N)^2)\) with respect to sequence length, this token deluge quickly overflows the context window and incurs prohibitive computational latency and KV-cache memory overhead during inference.
Existing efficiency approaches attempt to thin visual tokens either at the vision encoder output or within intermediate LLM layers by exploiting temporal redundancy across adjacent frames. Nonetheless, prevailing pruning techniques rely on static, frame-level heuristics (such as local entropy or inter-frame cosine dissimilarity) that exhibit three fundamental control failures over long horizons: first, observational stochasticity, where transient measurement noise (such as camera jitter, occlusion, or sudden lighting fluctuations) misleads deterministic metrics into pruning critical dynamic actions; second, accumulated decision bias, where myopic, greedy budgeting over-allocates tokens to early noisy intervals and irreversibly starves subsequent critical frames; and third, non-stationarity of information density, where static pruning rates fail to dynamically expand budget capacity during rapid scene transitions or abrupt action bursts.
The fundamental tension lies in mapping erratic, noisy local observations onto an irreversible, strictly finite token budget across long video horizons. Rather than adding more empirical heuristics, this paper treats token budgeting as a continuous state estimation problem subject to observation noise. Core idea: model temporal information demand as a latent dynamic state combining local spatial uncertainty and segment transitions, apply Kalman filtering with stability accumulation to filter transient observational noise, and dynamically allocate token budgets with uncertainty margins through hierarchical two-stage pruning.
Method¶
Overall Architecture¶
The S3-Prune framework operates through three integrated stages in a completely training-free manner: joint spatio-temporal demand modeling, stability accumulation via Kalman filtering, and two-stage hierarchical token selection. The entire pipeline seamlessly plugs into frozen pre-trained Video-VLMs during inference without auxiliary fine-tuning.
Given an input long video stream, the vision encoder extracts patch-level representations and frame-level CLS tokens. S3-Prune measures patch-level representation shifts across adjacent frames to capture spatial uncertainty, while simultaneously tracking semantic centroid distances across temporal segments to quantify segment transitions. These dual multi-scale measurements are normalized and fused into an integrated observation signal. Next, a linear state-space model equipped with Kalman filtering estimates the true latent information demand and posterior variance, effectively filtering high-frequency noise spikes. Finally, an uncertainty-aware dynamic budget controller assigns frame-level token quotas, steering a coarse pre-LLM spatial pruning stage followed by a query-guided attention refinement inside intermediate LLM layers.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Long Video Stream<br/>F high-resolution frames"] --> B["Spatio-Temporal Demand Modeling<br/>Spatial uncertainty + Segment transition"]
B --> C["Stability Accumulation<br/>Kalman filter state estimation"]
C --> D["Uncertainty-Aware Budget Allocation<br/>Dynamic safety margin per-frame budget Bt"]
D --> E["Two-Stage Hierarchical Selection<br/>Pre-LLM vision coarse filter + Intra-LLM query refinement"]
E --> F["Compressed Salient Token Sequence<br/>Input to deep LLM reasoning layers"]
Key Designs¶
1. Spatio-Temporal Demand Modeling: Multi-scale fusion of patch perturbation and segment transition
Information density in long videos fluctuates non-stationarily across both micro and macro scales: focusing solely on global frame differences overlooks fine-grained local motions, while inspecting only isolated patch discrepancies is easily hijacked by sensor noise. S3-Prune establishes a dual-scale metric combining Spatial Uncertainty (SU) and Segment Transition (ST). For patch \(i\) in frame \(t\) with embedding \(\mathbf{x}_{t,i}\), its spatial uncertainty relative to the prior frame is formulated as: $\(\mathrm{SU}_{t,i} = \frac{\|\mathbf{x}_{t,i} - \mathbf{x}_{t-1,i}\|_2}{\|\mathbf{x}_{t,i}\|_2 + \epsilon}\)$ The \(L_2\) norm in the denominator normalizes across feature scales to ensure consistent sensitivity. Frame-level observational demand \(m_t\) is then captured via the ensemble mean \(\mu_t\) and standard deviation \(\sigma_t\) as \(m_t = \mu_t + \sigma_t\), reflecting both the overall magnitude of change and intra-frame spatial dispersion. Simultaneously, for each temporal segment \(s\), the segment centroid \(\mathbf{c}_s = \frac{1}{T_s}\sum_{t\in\mathcal{F}_s}\mathbf{v}_t\) is computed over frame CLS tokens \(\mathbf{v}_t\). The segment transition score \(\mathrm{ST}_s = 1 - \frac{\mathbf{c}_s^\top\mathbf{c}_{s-1}}{\|\mathbf{c}_s\|_2 \|\mathbf{c}_{s-1}\|_2 + \epsilon}\) quantifies macroscopic semantic shifts between consecutive segments, flagging when scene transitions require expanded token retention.
2. Stability Accumulation: Suppressing transient spikes and preventing error compounding via Kalman control
Directly converting raw, fluctuating observation metrics into token counts causes acute budget instability: momentary occlusions or motion blurs trigger erratic budget spikes, starving subsequent narrative climaxes over long horizons. S3-Prune introduces Stability Accumulation (SA), treating the normalized fused measurement \(z_t = \tilde{m}_t + \widetilde{\mathrm{ST}}_t\) as a noisy observation of the true latent demand \(d_t\) governed by a linear state-space model: $\(d_t = d_{t-1} + w_t, \quad w_t \sim \mathcal{N}(0, Q)\)$ $\(z_t = d_t + n_t, \quad n_t \sim \mathcal{N}(0, R)\)$ At each time step, the prior state \(\hat{d}_t^- = \hat{d}_{t-1}\) and prior covariance \(P_t^- = P_{t-1} + Q\) are updated via the Kalman gain \(K_t = \frac{P_t^-}{P_t^- + R}\). Under severe measurement variance, a small \(K_t\) suppresses noisy observation spikes; under persistent, systematic information influx, \(K_t\) dynamically increases to track genuine state shifts. The posterior demand estimate \(\hat{d}_t = \hat{d}_t^- + K_t(z_t - \hat{d}_t^-)\) and updated variance \(P_t = (1 - K_t)P_t^-\) provide a mathematically stabilized control trajectory, eliminating decision bias accumulation across extended temporal contexts.
3. Uncertainty-Aware Budget Allocation: Elastic budgeting with adaptive safety margins
Relying strictly on the posterior mean \(\hat{d}_t\) risks dropping critical content during sudden, high-uncertainty events. S3-Prune constructs an uncertainty-augmented demand metric \(\phi_t = \hat{d}_t \cdot (1 + \sqrt{P_t})\), where the standard deviation term \(\sqrt{P_t}\) provides an explicit safety margin that conservatively inflates token quotas during volatile intervals. To respect online temporal causality during streaming inference, the token budget \(B_t\) for frame \(t\) is normalized against the running maximum demand \(\Phi_t = \max(\Phi_{t-1}, \phi_t)\): $\(B_t = \left\lceil \rho \cdot \frac{\phi_t}{\Phi_t} \cdot N \right\rceil\)$ where \(\rho\) denotes the target global retention ratio and \(N\) is the uncompressed token count per frame. This dynamic quota allocation enables causal, streaming adaptability while scaling resources elastically with genuine visual complexity.
4. Two-Stage Hierarchical Selection: Pre-LLM coarse filtering and intra-LLM task-aware refinement
To achieve aggressive compression without sacrificing query relevance, S3-Prune deploys a two-stage hierarchical selection mechanism. First, prior to LLM injection, the Top-\(B_t\) visual tokens with the highest spatial uncertainty \(\mathrm{SU}_{t,i}\) are gathered into candidate set \(\mathcal{C}_t\). This vision-only coarse filtering sheds massive background redundancy at negligible computational cost, drastically shrinking the prompt length during LLM prefilling. Second, at an intermediate LLM layer \(M\), self-attention maps between text query tokens and visual candidates are leveraged to compute query relevance \(a_{t,j} = \max_{1 \le i \le N_q} \mathbf{A}^{(M)}_{q,t}(i, j)\) for each candidate \(j \in \mathcal{C}_t\). The final retained set \(\mathcal{R}_t\) selects the top \(\lfloor \alpha \cdot B_t \rfloor\) tokens ranked by \(a_{t,j}\), ensuring that the tokens entering deeper LLM layers possess both prominent visual dynamics and high task semantic relevance.
Key Experimental Results¶
Main Results¶
The method was evaluated across four diverse Video-VLM backbones (LLaVA-OV-7B, LLaVA-VID-7B, Qwen-2.5-VL-7B, Qwen-3-VL-8B) on seven recognized video understanding benchmarks, including MVBench, EgoSchema, LongVideoBench, and VideoMME. Even under aggressive 10% and 25% token retention ratios, S3-Prune consistently outperforms state-of-the-art compression frameworks.
| Model & Method | Retention Ratio | MVBench | EgoSchema | LongVideoB | VideoMME (w/ sub) | Average Score (Avg %) |
|---|---|---|---|---|---|---|
| LLaVA-OV-7B (Vanilla) | 100% | 58.28 | 60.34 | 42.18 | 61.81 | 60.7 (100.0%) |
| FastVid (NeurIPS'25) | 25% | 58.23 | 59.23 | 42.92 | 61.11 | 60.1 (98.9%) |
| HoliTom (NeurIPS'25) | 25% | 58.25 | 60.94 | 42.55 | 61.96 | 60.8 (100.1%) |
| S3-Prune (Ours) | 25% | 58.50 | 61.11 | 43.12 | 62.34 | 61.1 (100.6%) |
| FastVid (NeurIPS'25) | 10% | 57.50 | 58.61 | 42.35 | 60.15 | 59.2 (97.5%) |
| HoliTom (NeurIPS'25) | 10% | 57.40 | 59.96 | 42.63 | 60.33 | 59.6 (98.0%) |
| S3-Prune (Ours) | 10% | 57.63 | 60.25 | 42.44 | 62.03 | 59.7 (98.3%) |
| LLaVA-VID-7B (Vanilla) | 100% | 60.43 | 57.18 | 57.58 | 70.14 | 65.9 (100.0%) |
| HoliTom (NeurIPS'25) | 10% | 56.63 | 53.78 | 57.51 | 69.96 | 63.1 (95.8%) |
| S3-Prune (Ours) | 10% | 59.95 | 54.09 | 57.65 | 70.04 | 63.9 (96.9%) |
Regarding runtime efficiency on Qwen-2.5-VL-7B at 10% token retention, S3-Prune reduces LLM inference time to 26.0% and total runtime to 71.7% of the unpruned baseline while trimming peak GPU memory to 95.9%, maintaining a 61.47 score on VideoMME (95.5% of vanilla performance).
Ablation Study¶
Ablation experiments analyze the contribution of each demand signal, the necessity of Kalman filtering, and compatibility across backbones.
| Configuration | SU | ST | Kalman Filter (KF) | MVBench | VideoMME | Avg Score (%) | Note |
|---|---|---|---|---|---|---|---|
| Segment transition only | - | โ | โ | 60.83 | 58.53 | 59.7 | Captures scene shifts but misses fine-grained dynamics |
| Spatial uncertainty only | โ | - | โ | 61.00 | 58.75 | 59.9 | Detects motion patches but lacks macro scene transition cues |
| Full dual-signal demand | โ | โ | โ | 61.33 | 59.10 | 60.2 | Optimal synergy of multi-scale dynamics |
| w/o Kalman Filter (LLaVA-OV) | โ | โ | โ | 56.40 | 55.90 | 53.6 | Severe budget oscillation from transient observation noise |
| w/ Kalman Filter (LLaVA-OV) | โ | โ | โ | 57.60 | 57.30 | 54.4 | +0.8% gain from smooth demand and uncertainty margins |
| w/o Kalman Filter (Qwen-2.5) | โ | โ | โ | 67.10 | 60.80 | 59.6 | Susceptible to observational bias accumulation |
| w/ Kalman Filter (Qwen-2.5) | โ | โ | โ | 68.40 | 61.50 | 60.2 | Consistent +0.6% gain across benchmarks |
Key Findings¶
- Kalman filtering is essential for long-horizon stability: Across all evaluated backbones, adding Kalman filtering produces an immediate +0.6% to +0.8% overall improvement. Qualitative trajectory analysis confirms that Kalman filtering removes false-positive budget spikes caused by brief camera shakes while safeguarding critical narrative scenes via variance expansion.
- Superior scalability over ultra-long frame sequences: When scaling input frames from 32 to 768 on VideoMME and VideoMMMU, conventional greedy heuristics suffer from compounding bias as sequence length increases. In contrast, S3-Prune exhibits monotonic accuracy gains, widening its competitive margin notably in the 512โ768 frame regime.
- Strong standalone performance and modular compatibility: In an ablation isolating the pre-LLM stage without intra-LLM attention refinement, S3-Prune still surpasses FastVid by +0.5% on average. Furthermore, plugging S3-Prune's budgeting into existing segmentation pipelines (FastVid and HoliTom) delivers uniform performance gains (+0.4% to +0.5%), proving that advantages stem from principled budget stabilization rather than ad-hoc partitioning heuristics.
Highlights & Insights¶
- Control-theoretic formulation for token budgeting: Formulating the allocation of finite computational tokens as a noisy state-space estimation problem provides a mathematically grounded, recursive solution to long-standing cumulative error drift in long-video understanding.
- Uncertainty margins as adaptive safety buffers: Utilizing the posterior variance \(\sqrt{P_t}\) to dynamically expand budget capacity during volatile intervals elegantly bridges the gap between aggressive sparsification and robust retention of unpredictable critical events.
- Architecture-agnostic plug-and-play efficiency: S3-Prune operates entirely without training or parameter tuning, enabling direct zero-shot deployment across diverse Video-VLM families and streaming video workloads.
Limitations & Future Work¶
- Handcrafted noise covariance priors: The process covariance \(Q\) and measurement covariance \(R\) are currently chosen based on empirical heuristics, which may not dynamically adapt to videos with extreme genre shifts (e.g., fast-paced sports montage versus slow-paced security feeds).
- Lag risk under continuous cut scenes: Under extreme, continuous jump cuts violating first-order Markov temporal smoothness, the Kalman state update may exhibit minor tracking latency during the initial frame of a scene shift.
- Future directions: Investigating adaptive covariance estimation (e.g., via Extended or Unscented Kalman Filters) conditioned on video domain metadata, and extending stability-aware budgeting to temporal action localization and multimodal generative tasks.
Related Work & Insights¶
- vs FastVid (NeurIPS 2025): FastVid performs dynamic density pruning based on inter-frame distance but lacks temporal noise filtering, causing early budget exhaustion in noisy long videos; S3-Prune introduces recursive Kalman smoothing to eliminate error accumulation without requiring explicit segmentation indexing.
- vs HoliTom (NeurIPS 2025): HoliTom relies on offline dynamic programming to determine segment merging, which incurs non-trivial preprocessing overhead and cannot operate causally; S3-Prune maintains running causal normalization, facilitating genuine online streaming inference.
- vs VidCom2 (EMNLP 2025): VidCom2 allocates budgets via single-stage frame uniqueness without text query interaction; S3-Prune pairs pre-LLM visual filtering with intra-LLM cross-modal attention refinement, ensuring higher semantic density and task relevance.
Rating¶
- Novelty: โญโญโญโญโ (Creative and rigorous introduction of Kalman filtering and state-space control to long-form video token scheduling)
- Experimental Thoroughness: โญโญโญโญโญ (Evaluated across 4 backbones, 7 benchmarks, scaling from 32 to 768 frames with comprehensive ablations)
- Writing Quality: โญโญโญโญโญ (Clear mathematical formulation, cohesive problem framing, and intuitive experimental validation)
- Value: โญโญโญโญโญ (Provides a highly practical, training-free, and architecture-agnostic speedup tool for long-form Video-VLM deployment)