Skip to content

Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/cau-hai-lab/PORTA.git
Area: Multimodal VLM / Model Compression
Keywords: vision-language model pruning / retraining-free / task-agnostic / activation variance / adaptive sparsity allocation

TL;DR

Addressing modality bias and calibration sensitivity in vision-language model pruning caused by LLM-oriented activation outlier assumptions, this paper proposes PORTA, which uses cross-modal stable activation variance to quantify representation utility and couples it with adaptive layer-wise sparsity allocation driven by output feature variability, enabling a single pruning pass on generic data to transfer zero-shot across classification, retrieval, and VQA without retraining.

Background & Motivation

Vision-language models (VLMs, such as CLIP, BLIP, and Qwen2-VL) have established a unified multimodal representation space through large-scale pretraining on web-scale image-text pairs, demonstrating remarkable generalization across zero-shot classification, cross-modal retrieval, and visual question answering. However, their burgeoning parameter counts and computational footprints present severe bottlenecks for deployment in resource-constrained environments such as edge devices and real-time systems. Post-training one-shot pruning has emerged as a promising avenue to eliminate parameter redundancy without changing network topology. In multimodal scenarios, an ideal pruning paradigm should be "task-agnostic": the model should be pruned only once using a small set of generic calibration samples and immediately deployed across diverse, previously unseen downstream tasks without task-specific recalibration, re-pruning, or retraining.

Existing VLM pruning approaches suffer from two fundamental tensions and technical limitations. First, mainstream pruning criteria (e.g., Wanda, SparseGPT) borrow heavily from large language models, using activation magnitude norms or extreme outliers to measure weight importance. However, the pronounced activation outlier patterns typical of text Transformers do not manifest identically in vision Transformers (ViTs); their activation scales and cross-layer dispersions differ substantially (for instance, the max-min range of text layers reaches 7.11 while vision layers sit at 3.73). Relying heavily on modality-specific activation distribution properties introduces severe modality bias, causing pruning to concentrate disproportionately on either the vision or language branch. Second, existing importance estimation methods are highly sensitive to calibration data choices, leading to sharp performance drops when downstream task-specific samples are absent; meanwhile, VLM-specific methods such as Multiflow still necessitate task-specific fine-tuning, failing to achieve true retraining-free plug-and-play utility.

This work revisits the connection between representation utility and activation statistics. Empirical findings reveal that, unlike magnitude metrics that oscillate dramatically, activation variance along the token/patch dimension exhibits stable dispersion across modalities and layers, reliably capturing a feature's global responsiveness and contribution to representational capacity across diverse inputs. Core idea: this paper introduces PORTA, a prune-once retraining-free VLM pruning framework that builds a modality-agnostic importance formulation based on activation variance and derives an adaptive layer-wise sparsity allocation strategy from output feature covariance spaces, enabling robust zero-shot transfer across heterogeneous multimodal tasks with generic calibration data.

Method

Overall Architecture

PORTA establishes an end-to-end compression pipeline tailored for multimodal architectures: "statistic estimation \(\to\) modality-agnostic weight scoring \(\to\) output variability layer-wise sparsity allocation \(\to\) one-shot retraining-free mask pruning." The framework requires no downstream task supervision or fine-tuning. Given generic image-text calibration samples, it first computes the input activation variance along the token/patch dimension for each layer. It then constructs a modality-agnostic weight importance score matrix by element-wise scaling. Concurrently, it projects the input covariance onto the weight space to quantify each layer's output representation capacity, dynamically distributing layer-wise sparsity quotas under a global target constraint before applying binary pruning masks.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Generic Image-Text Calibration Input"] --> B["Activation Variance Feature Utility<br/>Token/patch dimension variance estimation"]
    B --> C["Modality-Agnostic Weight Scoring<br/>Variance broadcast and weight magnitude product"]
    B --> D["Output Space Variability Estimation<br/>Normalized input covariance trace projection"]
    D --> E["Adaptive Layer-Wise Sparsity Allocation<br/>Global target constrained quota reweighting"]
    C --> F["One-Shot Retraining-Free Mask Pruning"]
    E --> F
    F --> G["Zero-Shot Plug-and-Play Multi-Task Deployment"]

Key Designs

1. Activation Variance Feature Utility: Eliminating Cross-Modal Outlier Bias Conventional magnitude-based importance criteria (such as Wanda) depend heavily on the raw amplitude of activation values. In VLMs, the text branch exhibits extreme outlier spikes from attention aggregation, whereas the vision branch maintains relatively moderate activation scales; directly applying magnitude criteria erroneously treats visual features as less critical, pruning them excessively. To resolve this pain point, PORTA employs input activation variance to gauge representation utility. Given layer input activations \(X \in \mathbb{R}^{B \times T \times D_{\text{in}}}\) (with batch size \(B\), token/patch length \(T\), and channel dimension \(D_{\text{in}}\)), the variability of the \(j\)-th feature dimension is defined as: $\(v_j = \mathrm{Var}_{b,t}\left(X_{b,t,j}\right), \quad j = 1, \dots, D_{\text{in}}\)$ The underlying insight is intuitive: low-variance features respond narrowly to localized patterns or remain near-constant, providing redundant information; conversely, high-variance features fluctuate widely across tokens and diverse multimodal contexts, reflecting broader engagement and high representational value. Empirical analysis demonstrates that variance statistics yield uniform dispersion ranges across both vision and text encoders, naturally neutralizing modality scale discrepancies.

2. Modality-Agnostic Weight Scoring: Combining Dynamic Fluctuations with Static Topology After deriving the representation utility vector across input dimensions, it must be paired with parameter topology to preserve critical connections. PORTA sidesteps the steep computational cost of second-order Hessian inversions by broadcasting the feature variance vector \(v \in \mathbb{R}^{D_{\text{in}}}\) across output dimensions and multiplying it element-wise with the absolute weight matrix \(W \in \mathbb{R}^{D_{\text{out}} \times D_{\text{in}}}\): $\(S_{i, j} = v_j \cdot \left| W_{i, j} \right|, \quad i = 1, \dots, D_{\text{out}}, \; j = 1, \dots, D_{\text{in}}\)$ Here, \(\left| W_{i, j} \right|\) measures static parameter magnitude, while \(v_j\) captures the dynamic representation utility of the input feature dimension. Their product reflects both parameter scale and representation capacity. This formulation avoids back-propagation or iterative optimization, ensuring high computational throughput while preventing overfitting to narrow downstream data distributions.

3. Output Space Variability Estimation: Quantifying Layer-Wise Representational Capacity Different layers in deep networks occupy distinct semantic hierarchies and multimodal fusion roles. Uniform layer-wise pruning risks over-pruning critical semantic stages and under-pruning early feature layers, causing performance cliffs. PORTA models the output feature space variability analytically. For a linear layer \(Y = X W^\top\), assuming input tokens \(x_t\) follow a zero-mean distribution with input covariance \(\Sigma_X = \mathbb{E}[x_t^\top x_t]\), the sequence-level variance of the \(o\)-th output channel \(y_{:, o} = X w_o^\top\) expands to \(w_o \Sigma_X w_o^\top\). To eliminate scale discrepancies and ensure parity across layers, PORTA normalizes the covariance by its trace, \(\widehat{\Sigma}_X = \frac{\Sigma_X}{\operatorname{Tr}(\Sigma_X)}\), and defines the layer representation capacity score as: $\(s_{\text{layer}} = \frac{1}{D_{\text{out}}} \sum_{o=1}^{D_{\text{out}}} w_o \widehat{\Sigma}_X w_o^\top = \frac{1}{D_{\text{out}}} \operatorname{Tr}\left(W^\top W \widehat{\Sigma}_X\right)\)$ To avoid maintaining full covariance matrices, PORTA adopts a diagonal covariance approximation \(\Sigma_X \approx \operatorname{diag}(v_1, \dots, v_{D_{\text{in}}})\). This allows direct reuse of the activation variances \(v_j\) calculated during the scoring stage, simplifying the layer score to: $\(s_{\text{layer}} = \frac{\sum_{j=1}^{D_{\text{in}}} v_j \| W_{:, j} \|_2^2}{D_{\text{out}} \sum_{k=1}^{D_{\text{in}}} v_k}\)$ This score essentially represents a variance-weighted average of squared column norms \(\| W_{:, j} \|_2^2\), quantifying how strongly high-utility input dimensions project energy into the layer's output space.

4. Adaptive Layer-Wise Sparsity Allocation: Global Target Constrained Smooth Quota Scheduling After computing raw layer scores \(\{s_i\}_{i \in \mathcal{L}}\) across all candidate layers \(\mathcal{L}\), PORTA executes an adaptive quota allocation algorithm. It first computes the normalized layer importance proportion \(p_i = s_i / \sum_{k \in \mathcal{L}} s_k\), then derives an unscaled pruning score \(u_i = (1 - p_i)^\alpha\), where \(\alpha > 0\) controls allocation sensitivity. It next computes the parameter-weighted average pruning score \(\bar{u} = (\sum_i w_i u_i) / \sum_i w_i\) using layer parameter counts \(w_i\). Finally, scaling by \(C = S / \bar{u}\) under target global sparsity \(S\) yields the exact layer pruning ratio \(r_i = C \cdot u_i\). This mechanism guarantees that the total pruned parameters strictly match the target sparsity while assigning lower sparsity to deep layers with higher output variability, preserving essential semantic representations.

Key Experimental Results

Main Results

Evaluated on OpenAI CLIP-ViT-bigG, PORTA is benchmarked against four representative pruning baselines (Wanda, SparseGPT, ECoFLaP, and Multiflow) across zero-shot image-text retrieval (MSCOCO) and zero-shot image classification (CIFAR-10, CIFAR-100, ImageNet-1K, and Flowers-102).

Sparsity Method MSCOCO TR@1 MSCOCO TR@5 MSCOCO IR@1 MSCOCO IR@5 CIFAR-10 Acc CIFAR-100 Acc ImageNet Acc Flowers102 Acc
0% Dense Baseline 68.56 87.70 52.28 76.09 97.04 87.50 78.45 80.19
55% Wanda 66.30 86.44 48.62 73.05 96.12 79.92 69.50 67.72
55% SparseGPT 65.78 86.44 48.57 73.29 95.07 79.95 68.61 65.16
55% ECoFLaP 51.14 76.04 33.14 58.44 80.69 51.79 43.11 30.00
55% Multiflow (Zero-Shot) 0.02 0.12 0.02 0.12 9.25 1.23 0.16 1.50
55% PORTA (Ours) 66.28 87.00 48.70 73.02 95.58 80.27 70.29 68.73
60% Wanda 59.78 82.82 42.55 68.01 93.20 70.28 57.35 51.48
60% SparseGPT 57.52 81.88 41.77 66.97 95.16 74.23 55.26 41.37
60% ECoFLaP 35.44 60.56 19.12 39.46 57.98 32.60 25.73 14.21
60% Multiflow (Zero-Shot) 0.04 0.16 0.01 0.08 9.79 1.10 0.08 0.49
60% PORTA (Ours) 60.14 83.04 43.11 68.18 95.06 75.59 59.66 55.50
65% Wanda 27.30 51.04 16.96 35.61 69.81 28.86 30.06 17.30
65% SparseGPT 25.10 49.10 15.93 35.00 86.35 49.23 25.96 12.09
65% ECoFLaP 13.60 29.62 7.22 18.32 33.18 16.70 10.20 6.66
65% Multiflow (Zero-Shot) 0.02 0.12 0.02 0.10 8.90 0.96 0.12 0.81
65% PORTA (Ours) 28.14 51.26 17.31 37.32 90.49 62.29 31.38 18.27

On generative multimodal architectures, PORTA is evaluated using Qwen2-VL on ScienceQA (at 50% sparsity) in a zero-shot visual question answering protocol:

Sparsity Method Natural Sci. (NAT) Social Sci. (SOC) Language (LAN) Text Context (TXT) Image Context (IMG) No Context (NO) Grade 1-6 Grade 7-12 Overall Average
0% Dense Baseline 60.66 67.27 60.00 62.37 60.98 73.77 65.57 55.24 61.87
50% Wanda 55.01 47.35 55.09 56.03 49.97 75.40 55.80 49.17 53.43
50% SparseGPT 49.42 49.04 55.00 53.16 47.74 67.21 53.56 45.74 50.78
50% ECoFLaP 51.33 50.95 52.19 52.61 49.88 63.93 54.52 46.01 51.47
50% Multiflow 51.95 47.80 54.18 54.55 47.89 73.77 53.92 47.59 51.66
50% PORTA (Ours) 55.23 50.05 55.81 56.50 51.46 72.13 56.82 49.76 54.30

Ablation Study

To isolate the individual gains from the importance metric versus the sparsity allocation strategy, a cross-combination ablation is conducted on CLIP at 60% sparsity:

Sparsity Importance Score Sparsity Allocation Ratio CIFAR-10 Acc (%) CIFAR-100 Acc (%) ImageNet-1K Acc (%) Note
60% Wanda (Magnitude) Uniform 93.20 70.28 57.35 Standard magnitude baseline
60% PORTA (Variance) Uniform 94.91 74.45 59.35 Variance scoring yields substantial gains across tasks
60% Wanda (Magnitude) Multiflow 10.03 1.58 0.11 Multiflow allocation collapses without fine-tuning
60% PORTA (Variance) Multiflow 10.07 1.40 0.11 Same collapse under Multiflow allocation
60% Wanda (Magnitude) ECoFLaP 93.52 71.16 61.14 Zeroth-order gradient allocation
60% PORTA (Variance) ECoFLaP 93.57 71.19 61.33 Minor improvement with variance scoring
60% Wanda (Magnitude) PORTA (Output Var.) 94.62 74.29 59.04 PORTA ratio boosts performance under magnitude scoring
60% PORTA (Variance) PORTA (Output Var.) 95.06 75.59 59.66 Optimal combination achieving peak performance

Ablation on modeling assumptions (zero-mean assumption and diagonal covariance approximation) on CLIP at 60% sparsity for zero-shot retrieval:

Configuration MSCOCO TR@1 MSCOCO TR@5 MSCOCO IR@1 MSCOCO IR@5 Average Recall
PORTA Full Model (Diagonal approx. + zero-mean) 60.14 83.04 43.11 68.18 63.62
w/o Zero-Mean Assumption (Explicit mean centering) 60.20 83.12 43.11 68.12 63.64
w/o Diagonal Approximation (Full covariance matrix) 60.70 82.80 43.13 68.27 63.73

Key Findings

  • High-Sparsity Resilience: At 65% sparsity, PORTA achieves a 21.5% and 12.6% relative improvement across 8 evaluation metrics over Wanda and SparseGPT, respectively. On ImageNet-1K, it retains 31.38% top-1 accuracy compared to 25.96% for SparseGPT and 30.06% for Wanda, effectively mitigating severe performance degradation.
  • Extreme Calibration Robustness: Shifting calibration distributions across Flickr30K, MSCOCO, and Visual Genome induces up to ~5% accuracy fluctuations in Wanda and SparseGPT, whereas PORTA remains stable within ~1%. Furthermore, scaling calibration sample counts from 24 to 211 shows near-flat accuracy, confirming that PORTA does not rely on massive calibration sets.
  • Superior Execution Efficiency: Total wall-clock pruning time on CLIP is 195.34 seconds for PORTA, making it 3.96\(\times\) faster than SparseGPT (773.35s) and 2.06\(\times\) faster than ECoFLaP (402.28s). This efficiency stems from the diagonal covariance formulation, which directly reuses the activation variance statistics computed during importance scoring.

Highlights & Insights

  • Activation Variance as Modality-Agnostic Proxy: Uncovers the fundamental mismatch between LLM activation outlier dynamics and vision feature distributions, replacing raw magnitude norms with statistical variance to achieve unbiased representation utility measurement across modalities.
  • Statistic-Reused Covariance Projection: Derives layer output variability via input covariance projection and trace normalization, showing that diagonal covariance approximation incurs negligible performance loss (\(<0.15\%\)) while eliminating full-matrix overhead and extra forward passes.
  • Truly Task-Agnostic Retraining-Free Paradigm: Demonstrates that multimodal compression does not require downstream label alignment or fine-tuning; the "variance utility + output capacity allocation" framework generalizes effectively to generative VLMs and diffusion backbones.

Limitations & Future Work

  • Author-Acknowledged Limitations: Under extreme sparsity (\(\ge 70\%\)), zero-shot accuracy still drops noticeably without post-pruning fine-tuning; the current study focuses on unstructured/weight pruning and does not explore unified pipelines with quantization or low-rank factorization.
  • Potential Technical Constraints: While diagonal covariance approximation is fast and produces minimal error in standard benchmarks, ignoring off-diagonal covariance cross-terms could potentially underestimate channel co-dependencies in dense specialized layers.
  • Promising Research Directions: Integrating lightweight retraining-free scaling compensation, and extending the variance formulation to dynamic token pruning and multi-head attention pruning.
  • vs Wanda: Wanda relies on weight magnitude multiplied by input \(\ell_2\) norm, which over-allocates importance to LLM outlier channels; PORTA leverages activation variance across sequence dimensions, eliminating modality scale imbalances.
  • vs SparseGPT: SparseGPT relies on iterative second-order Hessian inversions, which are computationally heavy and sensitive to calibration shifts; PORTA provides a closed-form statistical criterion that is 4\(\times\) faster with superior calibration stability.
  • vs ECoFLaP: ECoFLaP explores layer-wise ratio allocation via zeroth-order gradients but retains Wanda's magnitude metric; PORTA formulates a unified variance framework across both scoring and layer allocation, achieving clear margins at high sparsity.
  • vs Multiflow: Multiflow proposes task-agnostic principles but relies on downstream fine-tuning to recover accuracy; PORTA achieves genuine one-shot pruning that transfers zero-shot without any retraining.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Elegantly establishes activation variance as a modality-symmetric utility measure and develops an analytical output-variance layer allocation scheme.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive zero-shot validation across retrieval, classification, and VQA across multiple architectures (CLIP, BLIP, Qwen2-VL) and extensive sensitivity analyses.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical derivations, well-structured arguments, and self-contained experimental narratives.
  • Value: ⭐⭐⭐⭐⭐ Provides a fast, practical, and retraining-free pruning framework that directly addresses deployment bottlenecks in real-world multimodal systems.