Skip to content

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/cvsp-lab/SFPruner
Area: Multimodal VLM / Model Compression
Keywords: visual token pruning, high-resolution MLLMs, ridge leverage score, directional masking, single-forward pass

TL;DR

To overcome the severe latency bottleneck caused by iterative greedy searches in subset-optimization visual token pruning for high-resolution MLLMs, this paper proposes Single-Forward Pruner (SFPruner), which attenuates global covariance-level redundancy via semantics-guided ridge leverage scoring and eliminates local spatial overlap through parallel ranking-based directional masking, slashing token selection time from 112.4 ms to 2.5 ms while preserving up to 99.0% relative performance.

Background & Motivation

As Multimodal Large Language Models (MLLMs, e.g., LLaVA-NeXT and Qwen2.5-VL) scale to ultra-high resolutions and long video sequences, dynamic multi-patch slicing and dense spatiotemporal sampling produce thousands or even tens of thousands of visual tokens per query. Because the self-attention mechanism in Transformer architectures scales quadratically with sequence length, this visual token explosion incurs prohibitive end-to-end inference latency and GPU memory burdens. Training-free visual token pruning has emerged as a critical technique to sparsify visual representations before entering the costly LLM backbone.

However, existing token pruning paradigms suffer from a pronounced divergence between execution efficiency and task accuracy. On one hand, heuristic pruning approaches rely on local attention scores or isolated similarity rankings. While computationally lightweight, they fail to explicitly model structural redundancy across tokens, frequently clustering selections within dominant salient patches and triggering severe performance degradation under aggressive pruning budgets. On the other hand, subset-optimization approachesโ€”such as determinantal point processes (DPP), diversity maximization, and maximum weight independent set (MWIS) formulationsโ€”jointly optimize instruction relevance and feature diversity. Because these formulations are NP-hard, existing methods invariably rely on greedy sequential search loops to update residual states after each selection. When token counts scale to thousands in high-resolution or multi-frame regimes, this sequential loop introduces severe CPU/GPU synchronization and kernel-launch overheads, so that theoretical FLOPs reductions fail to materialize as practical wall-clock speedups.

This paper tackles the issue from a structural perspective: redundancy control does not necessitate combinatorial greedy search, but can be algebraically embedded directly into the scoring space in a single forward pass. Core idea: reformulate token redundancy modeling into global covariance-level energy attenuation via semantics-guided ridge leverage scores and local pairwise exclusion via parallel ranking-based directional masking, achieving non-iterative, single-forward visual token pruning with minimal selection overhead.

Method

Overall Architecture

The SFPruner pipeline takes as input normalized visual token embeddings from the vision encoder and projected instruction text representations, and outputs an informative, non-redundant subset of visual tokens for LLM generation. The entire workflow bypasses iterative state updates and is composed strictly of parallel matrix operations structured in three consecutive stages: first, instruction relevance and visual saliency are fused into a normalized semantic guidance base score; second, a ridge-regularized leverage score (SG-RLS) is computed over the feature covariance matrix to attenuate dominant feature subspaces; finally, an asymmetric ranking-based directional masking operation penalizes lower-scoring tokens that overlap with higher-ranking competitors, allowing direct Top-K subset selection.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Visual tokens $V$ & Instruction query $q$"] --> B["Semantic Guidance Fusion<br/>Text relevance + Visual saliency"]
    B --> C["Semantics-Guided Ridge Leverage Scoring<br/>Covariance inversion attenuates dominant energy"]
    C --> D["Ranking-Based Directional Masking<br/>Asymmetric suppression of redundant lower-rank tokens"]
    D --> E["Single-Pass Top-K Selection<br/>Pruned token subset fed into LLM backbone"]

Key Designs

1. Semantic Guidance Fusion: Bridging Instruction Relevance and Intrinsic Saliency Purely visual compression often drops task-specific cues, while pure text-attention scoring amplifies dominance bias. This module establishes a grounded prior score for each visual token. Let \(V \in \mathbb{R}^{N \times D}\) denote the \(L_2\)-normalized visual token embedding matrix and \(q \in \mathbb{R}^{1 \times D}\) be the projected text embedding. Textual relevance \(S_{\mathrm{rel},i}\) is computed via temperature-scaled cosine similarity: $\(S_{\mathrm{rel},i} = \frac{\exp(v_i q^\top / \tau)}{\sum_{j=1}^N \exp(v_j q^\top / \tau)}\)$ Concurrently, intrinsic visual saliency \(S_{\mathrm{attn},i}\) is gathered from the vision encoder self-attention weights (using the CLS-to-patch attention distribution, or an aggregated spatial attention map when a canonical CLS token is absent, as in Qwen2.5-VL). The two signals are linearly combined via balancing factor \(\alpha\) and normalized via Min-Max scaling to produce the stable base prior \(\tilde{S}_{\mathrm{guide},i}\).

2. Semantics-Guided Ridge Leverage Scoring: Global Covariance-Level Attenuation Sequential subset optimization suppresses redundancy by repeatedly projecting out chosen vectors from the residual space. To eliminate this step-by-step loop, SFPruner leverages the Ridge Leverage Score (RLS) from Randomized Numerical Linear Algebra (RandNLA) to quantify geometric uniqueness globally. Defining the feature covariance matrix as \(C = V^\top V \in \mathbb{R}^{D \times D}\), the ridge leverage score for token \(i\) is formulated as: $\(\ell_i = v_i (C + \lambda I_D)^{-1} v_i^\top\)$ where \(\lambda\) represents a small ridge regularization hyperparameter. The inverse regularized covariance \((C + \lambda I_D)^{-1}\) penalizes the dominant eigendirections of the feature space. Pervasive background tokens that cluster along high-energy eigenspaces receive low leverage scores, while structurally distinctive tokens obtain elevated scores. Multiplying this score with the semantic prior produces \(S_{\mathrm{SG\text{-}RLS},i} = \ell_i \tilde{S}_{\mathrm{guide},i}\), operating as a soft-AND gate: high scores require both task relevance and structural uniqueness, cleanly suppressing task-irrelevant geometric outliers and redundant salient patterns simultaneously. By applying the Woodbury matrix identity, when \(N > D\), the matrix inversion is executed in \(D \times D\) space via stable Cholesky decomposition (\(O(ND^2)\) complexity), bypassing the intractable \(N \times N\) inversion.

3. Ranking-Based Directional Masking: Asymmetric Pairwise Competition in Parallel While SG-RLS effectively handles covariance-level redundancy, spatial near-duplicate tokens can still survive. Instead of iteratively updating similarity states after each token pick, directional masking resolves local overlap via parallel tensor operations. Given the pairwise cosine similarity matrix \(C_{\mathrm{sim}} = V V^\top \in \mathbb{R}^{N \times N}\), an asymmetric ranking mask \(M\) is defined using the SG-RLS prior: $\(M_{ij} = \mathbb{I}(S_{\mathrm{SG\text{-}RLS},j} > S_{\mathrm{SG\text{-}RLS},i})\)$ This enforces strict unidirectional competition: token \(j\) can suppress token \(i\) only if token \(j\) holds a strictly higher priority. The suppression penalty factor for token \(i\) is calculated as: $\(P_i = 1 - \max_j (M_{ij} C_{\mathrm{sim},ij})\)$ Since \(M_{ii} = 0\), \(P_i \le 1\) holds naturally. If a lower-ranked token exhibits substantial semantic overlap with an already preferred competitor, its penalty factor decreases sharply, diminishing its final score \(S_{\mathrm{final},i} = S_{\mathrm{SG\text{-}RLS},i} \cdot P_i\). The entire operation consists of matrix multiplications, element-wise comparisons, and row-wise reductions, entirely devoid of Python-level loops, enabling immediate Top-K retrieval.

Key Experimental Results

Main Results

Evaluation spans LLaVA-NeXT-7B, Qwen2.5-VL-7B, and LLaVA-Video-7B. In the challenging continuous single-sequence regime of Qwen2.5-VL-7B, performance retention and selection latency (measured on a single NVIDIA RTX 4090 GPU with synchronized CUDA timing) are compared below:

Method / Setting TextVQA AI2D MME POPE MMStar Rel. (%) Selection Time (ms) โ†“
Qwen2.5-VL-7B (Vanilla 100%) 85.3 80.8 2316.0 86.4 56.7 100.0 -
Retain 40% (Avg. 512 Tokens)
PruMerge (ICCV'25) 82.1 79.8 2287.0 86.0 55.6 98.1 1.2
VisionZip (CVPR'25) 82.6 79.5 2306.5 86.1 55.2 98.0 1.1
DivPrune (CVPR'25) 82.3 78.6 2230.6 84.5 54.7 97.0 23.1
CDPruner (NeurIPS'25) 83.7 79.6 2302.2 86.1 55.9 98.8 112.4
D2Pruner (AAAI'26) 83.2 82.9 2305.3 85.3 54.9 98.7 31.2
SFPruner (Ours) 84.0 79.9 2319.3 86.5 55.9 99.0 2.5
Retain 20% (Avg. 256 Tokens)
PruMerge (ICCV'25) 74.7 77.2 2280.1 84.5 52.8 94.6 1.1
VisionZip (CVPR'25) 75.9 77.6 2276.5 84.5 53.5 95.1 1.1
DivPrune (CVPR'25) 76.5 76.6 2203.7 83.6 51.8 93.7 11.8
CDPruner (NeurIPS'25) 79.1 78.4 2264.5 84.9 53.4 96.2 56.7
D2Pruner (AAAI'26) 78.6 79.4 2280.5 83.0 52.1 95.3 18.9
SFPruner (Ours) 79.8 78.3 2295.7 85.7 53.7 96.5 2.5

Under an extreme scaling stress test on Qwen2.5-VL with \(N = 9216\) tokens (\(D = 3584\), 30% retention), SFPruner executes token selection in just 28 ms (prefill latency: 1114 ms, peak memory: 18432 MB), compared to 576 ms for CDPruner and 458 ms for DivPrune. Across the full MME benchmark, SFPruner reduces end-to-end inference runtime from 6101 seconds to 3601 seconds.

Ablation Study

On LLaVA-NeXT-7B retaining 320 tokens per image, key scoring metrics and selection strategies were ablated:

Scoring Metric Selection Strategy GQA TextVQA POPE Rel. (%) Selection Time (ms) โ†“
SG (\(S_{\mathrm{guide}}\)) Naive Top-K 60.3 57.8 85.5 96.9 0.7
SG-RLS (Ours) Naive Top-K 60.5 58.1 86.1 97.6 2.2
SG (\(S_{\mathrm{guide}}\)) Sequential Search 61.2 58.2 86.3 98.0 18.6
SG (\(S_{\mathrm{guide}}\)) Directional Masking 61.2 58.3 86.4 97.9 0.8
SG-RLS (Ours) Directional Masking 61.5 58.3 86.7 98.3 2.3

Key Findings

  • Eliminating iterative loops unblocks actual speedups: While CDPruner exhibits a linear latency scaling with retained token count (112.4 ms at 512 tokens down to 56.7 ms at 256 tokens), SFPruner maintains an invariant 2.5 ms selection latency, regardless of budget \(K\).
  • Synergistic duality of covariance and pairwise suppression: Upgrading naive Top-K from SG to SG-RLS increases relative retention from 96.9% to 97.6%. Integrating directional masking achieves 98.3% retention, demonstrating that global energy attenuation and asymmetric pairwise suppression address distinct, non-overlapping dimensions of token redundancy.
  • Superiority in extreme sequence lengths: In multi-frame video inference (LLaVA-Video) and 9K-token single-image regimes, sequential optimization overheads balloon into seconds, cancelling FLOPs savings. SFPruner completes pruning in 28 ms, translating theoretical pruning into tangible acceleration.

Highlights & Insights

  • Algebraic reformulation of combinatorial subset search: Demonstrates that the diversity-relevance trade-off in token selection can be addressed globally via regularized covariance matrix inversion rather than through autoregressive greedy selection.
  • Asymmetric directional masking: Replaces iterative neighborhood suppression with a vectorized upper-triangular masking tensor, preserving diverse token coverage without introducing kernel launch bottlenecks.
  • Woodbury duality for scalable matrix operations: By solving the inversion in the smaller of sequence length \(N\) and channel dimension \(D\), the computational footprint is bounded at \(O(ND^2)\), preserving stable memory usage even at 16K visual tokens.

Limitations & Future Work

  • Author-admitted limitations: When deployed on architectures lacking dedicated text encoders or global tokens (such as Qwen2.5-VL), semantic guidance relies solely on pooled spatial attention as a proxy, slightly dampening sensitivity to complex textual queries.
  • Underlying model assumptions: The feature covariance matrix assumes linear dependency in the feature manifold; non-linear geometric clustering may not be fully captured.
  • Future directions: Adapting directional masking to local neighborhood windows to lower the \(O(N^2)\) pairwise similarity complexity to linear time, scaling to million-token video streams.
  • vs CDPruner (NeurIPS 2025): CDPruner enforces conditional diversity via greedy iterations (112.4 ms selection at 512 tokens on Qwen2.5-VL); SFPruner matches its accuracy while cutting latency to 2.5 ms (a 45x speedup).
  • vs DivPrune (CVPR 2025): DivPrune incurs up to 131 ms overhead in video token pruning due to iterative distance maximization; SFPruner processes video frames in 3.1 ms with superior retention (94.9% vs 94.5%).
  • vs Heuristic Methods (VisionZip, PruMerge+): While heuristic methods are fast (~1 ms), their lack of structural redundancy modeling leads to severe accuracy degradation at 20% or lower token budgets; SFPruner preserves near-lossless reasoning without compromising speed.

Rating

  • Novelty: โญโญโญโญโ˜† Elegantly bridges randomized numerical linear algebra (RLS) and directional tensor masking to bypass iterative combinatorial token pruning.
  • Experimental Thoroughness: โญโญโญโญโญ Rigorously tested across multi-patch architectures, dense continuous sequences, multi-frame video benchmarks, and 9K-token stress tests.
  • Writing Quality: โญโญโญโญโญ Clear mathematical formulations, insightful micro-architectural profiling, and coherent narrative structure.
  • Value: โญโญโญโญโญ Delivers practical wall-clock acceleration for high-resolution MLLMs, closing the gap between theoretical FLOPs reduction and real deployment latency.