Skip to content

GTR: Guide-Then-Refine Token Compression for Training-Free Acceleration of Video-LLMs

Conference: ECCV 2026
Paper: CVF Open Access
Area: Model Compression
Keywords: Video-LLMs, token pruning, training-free acceleration, dynamic budget allocation, multimodal efficiency

TL;DR

Proposes Guide-Then-Refine (GTR), a training-free two-stage visual token compression framework that dynamically allocates frame-level budgets using cross-frame global uniqueness and text-query relevance, followed by local spatial refinement within each frame, maintaining 95.8% of full performance with only 15% visual tokens.

Background & Motivation

Video Large Language Models (Video-LLMs) have achieved remarkable progress across video question answering, long-form action understanding, and embodied decision-making. However, as input video resolution and frame counts scale up, visual token volumes grow exponentially. The quadratic computational and memory complexity of the Transformer's self-attention mechanism incurs prohibitive inference latency and memory footprints, severely bottlenecking the practical deployment of Video-LLMs in latency-critical and resource-constrained environments.

Existing visual token compression approaches suffer from three prominent limitations: First, many efficient pruning techniques (e.g., STTM, DynTok, FLOC) rely exclusively on visual heuristics such as inter-frame similarity or spatial variance, ignoring downstream text query guidance. In long-video QA, subtle visual tokens critical to the query are frequently pruned. Second, conditional methods that do incorporate text guidance (e.g., HICom, CrossLMM, HoliTom) often require expensive cross-modal pre-training or intrusive in-LLM modifications (such as custom cross-attention layers or runtime attention filtering), breaking compatibility with optimized kernels like FlashAttention. Third, static compression strategies apply uniform retention ratios across all frames, disregarding temporal variations in information density and treating static background frames identically to information-dense action frames.

This paper tackles the challenge by moving all compression operations entirely before the LLM backbone, injecting lightweight multimodal guidance without modifying model weights. The core idea is to introduce Guide-Then-Refine (GTR), a two-stage training-free compression paradigm that uses cross-frame uniqueness and query alignment to allocate dynamic frame budgets in the guide stage, and leverages 2D spatial neighborhood variance to preserve fine-grained foreground tokens in the refine stage, providing seamless plug-and-play acceleration fully compatible with FlashAttention.

Method

Overall Architecture

GTR operates strictly between the visual projector and the LLM input embedding layer, requiring no retraining and keeping LLM attention and KV-caches untouched. For an input video of \(I\) frames where each frame yields \(N\) spatial patch tokens, GTR proceeds in two complementary stages: Stage 1 ("Guide") measures global cross-frame uniqueness against spatial mean features and computes semantic alignment with the query embedding, dynamically determining each frame's retention budget via range-normalized informativeness scores. Stage 2 ("Refine") operates within each frame under its assigned budget, calculating local Chebyshev neighborhood contrast to capture fine-grained boundaries and foreground objects, fusing the scores to select top-\(K\) tokens that are temporally concatenated and fed into the frozen LLM.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input video frames and text query<br/>Vision encoder extracts patch tokens"] --> B["Global & text guidance scoring<br/>Cross-frame mean contrast and text similarity"]
    B --> C["Dynamic frame-level budget allocation<br/>Linear mapping of frame score to retention ratio"]
    C --> D["Local spatial structure scoring<br/>Adaptive Chebyshev neighborhood contrast"]
    D --> E["Multi-score fusion and frame-wise pruning<br/>Select Top-Ki visual tokens"]
    E --> F["Concatenate retained tokens into LLM<br/>Preserve standard attention & fast kernel support"]

Key Designs

1. Global & text guidance scoring: Decoupling static backgrounds and query-critical targets

To eliminate static temporal redundancy while preserving query-relevant details, the guide stage calculates a cross-frame global uniqueness score \(S_{\text{glob}}(i, j)\) and a text relevance score \(S_{\text{text}}(i, j)\) for each token \(\mathcal{F}_{i,j}\) (the \(j\)-th spatial token in frame \(i\)). To quantify uniqueness, the average feature across all \(I\) frames at position \(j\) is calculated as \(\bar{\mathcal{F}}_j = \frac{1}{I} \sum_{i=1}^I \mathcal{F}_{i,j}\), representing the persistent spatial baseline across the video. The global score is defined as the complement of the cosine similarity between \(\mathcal{F}_{i,j}\) and \(\bar{\mathcal{F}}_j\):

\[S_{\text{glob}}(i, j) = 1 - \frac{\mathcal{F}_{i,j} \cdot \bar{\mathcal{F}}_j}{\|\mathcal{F}_{i,j}\|_2 \|\bar{\mathcal{F}}_j\|_2}\]

Dynamic foreground elements diverge from the cross-frame mean and yield high scores, whereas static backgrounds remain close to the average and receive low scores. For text relevance, a lightweight text encoder aligned with the vision backbone (e.g., SigLIP) encodes the question \(Q\) into a single global vector \(T \in \mathbb{R}^D\), and the cosine similarity between \(\mathcal{F}_{i,j}\) and \(T\) serves as \(S_{\text{text}}(i, j) = \frac{\mathcal{F}_{i,j} \cdot T}{\|\mathcal{F}_{i,j}\|_2 \|T\|_2}\). The combined guidance score is \(S_{\text{guide}}(i, j) = \alpha S_{\text{glob}}(i, j) + \beta S_{\text{text}}(i, j)\), with \(\alpha = 0.375\) and \(\beta = 0.625\).

2. Dynamic frame-level budget allocation: Tailoring capacity to temporal information density

To replace uniform frame compression with content-aware budgeting, GTR aggregates the comprehensive guidance scores across all \(N\) tokens in frame \(i\) as a proxy for frame-level information density: \(S_{\text{guide}}(i) = \sum_{j=1}^N S_{\text{guide}}(i, j)\). Frames with rich motion or strong query relevance obtain higher aggregate scores. These scores are range-normalized across all frames to compute a dynamic retention ratio \(r_i\):

\[r_i = r_{\text{base}} + (r_{\text{max}} - r_{\text{base}}) \times \frac{S_{\text{guide}}(i) - S_{\text{min}}}{S_{\text{max}} - S_{\text{min}}}\]

where \(S_{\text{min}}\) and \(S_{\text{max}}\) are the minimum and maximum frame scores across the video, \(r_{\text{base}} = 10\%\) safeguards low-value frames from catastrophic context loss, and \(r_{\text{max}} = 40\%\) caps token retention for computational efficiency. Frame \(i\) is assigned a budget of \(K_i = \lceil N \cdot r_i \rceil\) tokens.

3. Local spatial structure scoring: Preserving fine-grained spatial details via grid neighborhoods

Relying solely on global cross-frame deviation risks discarding subtle but essential local objects whose appearance does not shift dramatically across frames. The refine stage compensates for this by evaluating the local spatial distinctiveness of token \(\mathcal{F}_{i,j}\) within its 2D grid neighborhood. Based on Chebyshev distance, the immediate spatial neighborhood is defined as \(\mathcal{N}_{i,j} = \{ \mathcal{F}_{i,k} \mid \max(|u - u'|, |v - v'|) \le 1, (u', v') \in [1, H] \times [1, W] \}\), naturally adapting to image boundaries without synthetic padding and remaining independent of subsequent positional encodings. The local score measures the average dissimilarity against its valid neighbors:

\[S_{\text{local}}(i, j) = 1 - \frac{1}{|\mathcal{N}_{i,j}|} \sum_{\mathcal{F}_{i,k} \in \mathcal{N}_{i,j}} \frac{\mathcal{F}_{i,j} \cdot \mathcal{F}_{i,k}}{\|\mathcal{F}_{i,j}\|_2 \|\mathcal{F}_{i,k}\|_2}\]

Tokens belonging to homogeneous backgrounds (e.g., walls or clear sky) exhibit high similarity with neighbors and receive \(S_{\text{local}} \approx 0\), while boundary tokens, small tools, and textured objects stand out with elevated local scores.

4. Multi-score fusion and frame-wise pruning: Tri-perspective importance selection

In the final step, global uniqueness, query relevance, and local structure are integrated into a unified importance metric:

\[S_{\text{final}}(i, j) = \lambda_g S_{\text{glob}}(i, j) + \lambda_t S_{\text{text}}(i, j) + \lambda_l S_{\text{local}}(i, j)\]

with weights \(\lambda_g = 0.3, \lambda_t = 0.5, \lambda_l = 0.2\). Prioritizing text alignment ensures task fidelity, cross-frame scoring suppresses static redundancy, and local scoring preserves fine-grained spatial structures. Tokens within frame \(i\) are sorted in descending order of \(S_{\text{final}}\), the top \(K_i\) tokens are retained, and the selected tokens from all frames are concatenated in temporal order into the final sequence for the LLM.

Key Experimental Results

Main Results

Evaluations are conducted on LLaVA-OneVision-7B (LLaVA-OV-7B) and LLaVA-Video-7B across MVBench, LongVideoBench, MLVU, and VideoMME (Overall, Short, Medium, Long subsets).

Model / Method Retention Ratio MVBench LongVideoBench MLVU VideoMME (Overall) VideoMME (Long) Relative Avg (%)
LLaVA-OV-7B (Full) 100% 56.9 56.4 63.0 58.6 48.8 100.0
FastV 25% 55.5 53.3 59.6 55.3 47.0 94.9
SparseVLM 25% 56.4 53.9 60.7 57.3 48.1 97.5
VidCom2 25% 57.2 54.9 62.5 58.6 49.4 99.6
GTR (Ours) 25% 57.9 55.1 61.9 58.7 49.9 99.8
FastV 15% 51.6 48.3 55.0 48.1 43.3 85.0
SparseVLM 15% 52.9 49.7 57.4 53.4 47.0 91.2
VidCom2 15% 54.3 52.0 58.9 56.2 48.1 95.1
GTR (Ours) 15% 55.7 52.6 59.3 56.4 48.7 95.8

At a 25% retention ratio, GTR outperforms the full-token baseline on MVBench by 1.0% (57.9% vs. 56.9%) and reaches 49.9% on VideoMME-Long. Under aggressive 15% retention, GTR maintains an average of 95.8% of full-token capability, outperforming FastV by 10.8 percentage points.

Ablation Study

Ablations on LLaVA-OV-7B examine the individual contributions of the scoring components on MLVU and long-video understanding (VideoMME-Long):

Config MLVU Score Long Video Score Relative to Vanilla (%) Note
Full Model (Ours) 61.9 49.6 99.7 Full combination of \(S_{\text{text}} + S_{\text{glob}} + S_{\text{local}}\)
\(S_{\text{text}} + S_{\text{glob}}\) 61.7 49.3 99.4 Lacks local refinement, leading to minor drops on fine detail
\(S_{\text{text}} + S_{\text{local}}\) 61.2 48.6 98.6 Lacks cross-frame uniqueness, retaining redundant background
\(S_{\text{glob}} + S_{\text{local}}\) 60.5 48.2 97.5 Omits text guidance, failing to prioritize query targets
Only \(S_{\text{text}}\) 61.5 49.0 98.3 Strongest single signal, but noisy without visual context
Only \(S_{\text{glob}}\) 60.8 48.7 97.1 Relies purely on temporal changes, query-agnostic
Only \(S_{\text{local}}\) 60.0 47.5 96.0 Pure spatial contrast, lacks temporal and query awareness

In end-to-end efficiency testing on LLaVA-OV-7B, full-token inference requires 26 min 03 s with 17.7 GB GPU memory. At 25% retention, GTR reduces total runtime to 18 min 44 s (~1.4ร— acceleration, increasing throughput from 0.64 to 0.88 samples/s) and cuts memory to 16.1 GB. The guide and refine scoring overhead adds only 8.5 ms and 2.3 ms respectively, avoiding the deep in-LLM evaluation costs of FastV (LLM latency 179.8 ms vs. 260.9 ms, FastV memory 24.7 GB).

Key Findings

  • Text guidance is critical for long videos: Ablations confirm that \(S_{\text{text}}\) is the single most impactful component; omitting it causes long-video understanding to drop from 49.6% to 48.2%, highlighting that query-agnostic pruning is prone to discarding subtle evidence in lengthy footage.
  • Global and local scores are complementary: Global uniqueness identifies active temporal frames and dynamic regions, while local scores preserve fine spatial contours and small objects against backgrounds.
  • Orthogonal plug-and-play enhancement: Integrating GTR on top of existing pruning frameworks (FastV and SparseVLM) delivers consistent gains of +0.7% to +0.9% across benchmarks, proving that its pre-LLM budget allocation is fully orthogonal to intra-model token pruning.

Highlights & Insights

  • Elegant cross-frame mean contrast: Using the arithmetic mean of spatial patches across all frames as an anchor establishes an effective background suppression baseline via simple cosine dissimilarity, bypassing heavy graph constructions or spatio-temporal clustering.
  • FlashAttention-friendly decoupling: By executing all pruning strictly before LLM layers, GTR avoids attention mask fragmentation and memory overhead inside the Transformer, maintaining full compatibility with FlashAttention and standard KV-caches.
  • Clear spatiotemporal division of labor: The guide stage decides when to allocate more tokens across the timeline, and the refine stage determines where to retain tokens within the 2D frame, creating an intuitive, interpretable compression pipeline.

Limitations & Future Work

  • Constraint on streaming and ultra-long videos: Computing the cross-frame mean \(\bar{\mathcal{F}}_j\) requires access to all frames in the video, rendering the offline formulation unsuitable for open-ended streaming or real-time camera feeds without sliding-window adaptations.
  • Embedding alignment discrepancy: For models without a natively paired text encoder (e.g., Qwen2-VL), relying on an external SigLIP encoder may introduce subtle multimodal semantic misalignment on complex instructions.
  • Future directions: Extending GTR with sliding-window online mean tracking and coupling token pruning with latency-aware scheduling under dynamic compute budgets.
  • vs VidCom2: While VidCom2 introduces frame uniqueness for adaptive allocation, it remains purely visual; GTR demonstrates that integrating text relevance \(S_{\text{text}}\) yields substantial accuracy margins at high compression ratios (55.7% vs. 54.3% at 15% retention).
  • vs FastV / PDrop: FastV prunes tokens after initial LLM layers, destroying FlashAttention compatibility and demanding 24.7 GB GPU memory; GTR performs compression entirely upstream, operating at 16.1 GB memory with higher throughput.
  • vs HoliTom / CrossLMM: HoliTom requires dual pre-LLM and in-LLM modifications, and CrossLMM alters model architecture with cross-attention layers; GTR retains full zero-shot plug-and-play capability across heterogeneous Video-LLM backbones.

Rating

  • Novelty: โญโญโญโญโ˜† [Elegant integration of cross-frame mean divergence, text guidance, and local Chebyshev contrast in a two-stage pre-LLM workflow]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive evaluations across 4 benchmarks, 2 backbones, efficiency and GPU memory profiling, and orthogonal combination studies]
  • Writing Quality: โญโญโญโญโญ [Clear motivation, self-contained equations, well-structured prose, and coherent visualizations]
  • Value: โญโญโญโญโญ [High practical utility for accelerating long Video-LLMs in resource-constrained environments without model retraining]