Skip to content

CASS: Contribution-Aware Structured Sparsity for Model Merging

Conference: NeurIPS2026
arXiv: 2609.34184
Area: Model Compression
Keywords: model merging, structured sparsity, contribution score, task vector, gradient masking

TL;DR

CASS uses unlabeled task samples to identify attention heads and FFN neurons with prominent contributions, then filters task vectors or constrains fine-tuning gradients to reduce merging interference, improving Task Arithmetic's average accuracy across 20 tasks on ViT-B/16 from 65.02% to 69.99%.

Background & Motivation

Model merging aims to combine task experts derived from the same base model into one model without joint retraining. A common approach computes each expert's parameter difference from the base model, called a task vector, then sums or reweights these differences. The difficulty is that not every expert update is useful for its task: combining these updates with those from other experts can damage previously effective functionality. TIES handles conflicts through parameter magnitudes and signs, while DARE randomly drops updates, but neither directly identifies the functional components associated with an update.

A Transformer is not an unstructured collection of parameters. Each attention head adds a contribution to the residual stream through its output projection, and each intermediate FFN neuron adds a contribution through its corresponding down-projection direction. Consequently, individual weight magnitude need not indicate task importance: a channel with large weights may barely activate on the task, while an active channel's influence also depends on its output projection. On Qwen2.5-1.5B-Instruct, the paper observes that neuron contributions concentrate in a small subset of channels and that high-contribution sets from different tasks often have limited overlap; the appendix also stresses that this does not prove strict disjointness throughout the model.

This shifts the question from which parameters to delete to which functional components' task-specific updates should enter the merge. Attention heads require an additional distinction: some heads contribute strongly to every task, so large absolute contribution does not establish task specificity. Core idea: measure functional contribution using both activations and output projections, preserve updates to task-relevant components, and revert unselected components to base-model parameters rather than deleting network components.

Method

Overall Architecture

CASS is a structured update-filtering framework, not a new expert-routing network. Contribution estimation precedes structured mask construction, followed by an application branch determined by whether experts already exist; the final model retains the base architecture, and inference requires neither new probes nor dynamic expert selection.

CASS-Merging is the primary setting: given the base model, existing task experts, and unlabeled probes for each task, it estimates contributions on the experts and applies masks to their task vectors. CASS-Tuning instead estimates the same type of masks using the base model and unlabeled training samples before experts are trained, then restricts gradients during task-specific fine-tuning. The branches share a selection principle but differ in reference models, data sources, and masked objects; they are not the same post-hoc operation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Experts and unlabeled probes"] --> C["Contribution estimation"]
    B["Base model and<br/>unlabeled training samples"] --> C
    C --> D["Structured masks"]
    D -->|Existing experts: task vectors| E["Post-hoc task-vector filtering"]
    D -->|New fine-tuning: training gradients| F["Training-time gradient constraints"]
    E --> G["Existing merge operator"]
    F --> G
    G --> H["Single multi-task model"]

The samples in the diagram estimate component statistics rather than provide training supervision. Only the training-time branch additionally requires the original task's fine-tuning objective; the post-hoc branch does not retrain experts. The two inputs represent alternative settings, not a requirement to use expert and base-model statistics together in one run.

Key Designs

1. Contribution estimation: replace weight magnitude with input-conditioned functional contribution

Under a fixed reference model, an FFN neuron's residual contribution is its intermediate activation multiplied by the corresponding down-projection vector. CASS measures the magnitude of this contribution as a proxy for task relevance, rather than using activation frequency or weight magnitude alone. The single-token score is:

\[ \mathcal{I}_{\mathrm{neuron}}^{(j,t)}=\mathbb{E}_{\mathbf{x}\sim\mathcal{D}_{t}}\left[|m_j(\mathbf{x})|\,\|\mathbf{w}_{down}^{(j)}\|_2\right]. \]

Here, \(m_j\) is the neuron's intermediate activation, \(\mathbf{w}_{down}^{(j)}\) its down-projection vector, and \(\mathcal{D}_t\) the task probe set. The appendix also gives a sequence formulation using the Frobenius norm of the contribution matrix across the sequence. The score captures both how active a channel is on the inputs and how strongly it contributes to the residual stream, so a channel with large static weights but little activity does not automatically rank highly.

Attention heads cannot simply reuse this absolute-magnitude criterion. A general-purpose routing head may contribute strongly to every task, and selecting it need not reduce cross-task overlap. The paper first computes each head's absolute contribution, then measures its relative deviation on the current task from its average contribution across all tasks to be merged:

\[ \mathcal{A}_{h,t}=\mathbb{E}_{\mathbf{X}\sim\mathcal{D}_{t}}\left[\|\mathbf{H}_{h}(\mathbf{X})\mathbf{W}_{O}^{(h)}\|_F\right],\qquad \mu_{h,\mathrm{global}}=\frac{1}{K}\sum_{t'=1}^{K}\mathcal{A}_{h,t'},\qquad \mathcal{S}_{h,t}=\frac{\mathcal{A}_{h,t}-\mu_{h,\mathrm{global}}}{\mu_{h,\mathrm{global}}+\epsilon}. \]

\(\mathbf{H}_h\) is the head output, \(\mathbf{W}_O^{(h)}\) the corresponding output-projection block, \(K\) the number of tasks, and \(\epsilon\) a numerical stabilizer. Selection uses positive task-selectivity scores, not their absolute values: a strongly negative score means the head is suppressed on the current task rather than particularly active. Since the mean is computed across tasks, head contribution statistics must first be collected for all tasks; changing the task set can also change the head masks.

2. Structured masks: rank heads across layers and neurons within each layer

For each task, attention heads are ranked across the entire model by task selectivity, whereas FFN neurons are ranked within each layer by contribution. The experiments use a 20% retention ratio. Heads are relatively few, so global selection avoids requiring every layer to retain a few unimportant heads; neurons are numerous, and layer-wise allocation prevents differences in contribution scale from concentrating the entire budget in a few layers. This 20% is a component retention ratio, not deletion of 80% of the model's parameters, and not all parameters are masked.

Component selections are then broadcast to parameter slices. A head mask controls its Q, K, V projection blocks and output-projection block; an FFN neuron mask controls a column of the up projection and a row of the down projection. For Qwen's gated FFNs, the corresponding slices of gate, up, and down projections are handled together, rather than splitting a functional unit by filtering only one matrix.

Bias handling also has boundaries: head-specific Q, K, V bias slices and FFN up-projection bias channels can follow their component masks, but the shared attention output bias and FFN down-projection bias are not decomposed by component and remain unmasked in the appendix's implementation. These row and column directions follow the paper's matrix convention and must be aligned with actual tensor layouts in an implementation.

3. Post-hoc task-vector filtering: undo only irrelevant expert updates

In CASS-Merging, the reference model for each task is its fine-tuned expert, and probes require no labels. The expert's task vector is its parameter difference from the common base model; after filtering, it is passed to the original TA, TIES, TSV, WUDI, or other operator. The core operation is:

\[ \tau_t=\theta_t-\theta_{base},\qquad \tilde{\tau}_t=\mathbf{M}_t\odot\tau_t,\qquad \theta_{merged}=\theta_{base}+\mathrm{MergeOp}\left(\{\tilde{\tau}_t\}_{t=1}^{K}\right). \]

\(\odot\) denotes element-wise multiplication after structured broadcasting. A zero mask means that this expert contributes no task-specific update to the component; the component itself still exists. If another expert selects the same component, the final merged parameters can still receive that expert's updates, so a head unselected for one task need not remain entirely at its base state in the final model.

This is also not multiplication of an expert's forward output by zero. Attention's Q and K projections and FFN nonlinear activations change with the parameters, and the appendix explicitly notes that parameter-update filtering generally differs from deleting an expert's output contributions. The additive decomposition supports contribution measurement and component identification, but does not prove that the filtered network's output equals a linear removal of terms from the original expert output.

4. Training-time gradient constraints: restrict update locations before task vectors form

CASS-Tuning has no experts to measure yet, so the base model first processes unlabeled training samples from each task to generate fixed task masks, after which each expert is fine-tuned from the base model. The rationale is that task adaptation amplifies pathways already present in the base model; this is a plausible assumption, not a guarantee for every task. Training still uses the original task loss, changing which locations can update:

\[ \mathbf{g}'_t=\mathbf{M}_t\odot\nabla_{\theta}\mathcal{L}_t. \]

Unlike removing updates after training, this branch concentrates task vectors in selected structural subspaces as they form. Task masks typically overlap only partially, so the method encourages structural separation without enforcing complete disjointness or adding a cross-task exclusivity loss. After fine-tuning, the task vectors still use an existing merge operator.

With LoRA, the mask must align with channels of the effective low-rank update; it can act on the composed update matrix or a dimensionally aligned LoRA factor. Zeroing gradients alone does not automatically handle optimizer momentum or decoupled weight decay: the appendix requires inactive channels' effective updates to remain zero to preserve base-model behavior.

Loss & Training

CASS-Merging introduces no new training loss, only preprocessing to collect activation statistics and filter task vectors. Masks can be reused across merge operators; unlabeled does not mean data-free, since task selectivity still depends on task probes and the current task set.

The Qwen2.5 experiments use LoRA with rank 32, alpha 64, dropout 0.1, learning rate \(2\times10^{-4}\), a cosine schedule, warmup ratio 0.05, 1 epoch, and batch size 16. Randomized procedures use fixed seed 42; the paper does not report multi-seed means or confidence intervals.

The probe-size sensitivity experiment uses 8 to 1024 samples per task, with average accuracy varying by less than 0.5 percentage points for TA+CASS-M on both ViT architectures. The authors report that estimating masks with 8 samples per task takes only a few seconds on one A100; this is a one-time cost under specific experimental conditions, not a measured guarantee for every architecture and scale.

Key Experimental Results

Main Results

The vision experiments merge 20 CLIP ViT image-classification experts; RoBERTa uses 8 GLUE tasks and averages their standard task metrics rather than a uniform classification accuracy. Qwen2.5-0.5B/1.5B-Instruct covers coding, mathematics, instruction following, and safety; Instruct is omitted below.

Qwen scores are first normalized against the corresponding single-task expert: higher-is-better metrics divide the model's score by the expert's score, while the lower-is-better safety Micro Harm Rate is complemented before division. The four normalized scores are averaged arithmetically, so 1.000 represents the expert reference rather than a theoretical upper bound, and is not directly comparable to vision accuracy.

Model and evaluation Merge operator Without CASS CASS-Merging
ViT-B/32, average accuracy % across 20 tasks TA 61.37 65.29
ViT-B/16, average accuracy % across 20 tasks TA 65.02 69.99
ViT-B/16, average accuracy % across 20 tasks DARE 65.03 69.90
ViT-B/32, average accuracy % across 20 tasks TSV 73.67 74.91
RoBERTa, average performance across 8 tasks TA 66.36 70.12
RoBERTa, average performance across 8 tasks Iso-C 75.72 75.98
Qwen2.5-0.5B, average normalized performance across 4 tasks TA 0.743 0.776
Qwen2.5-0.5B, average normalized performance across 4 tasks TIES 0.695 0.773
Qwen2.5-0.5B, average normalized performance across 4 tasks Iso-C 0.738 0.729
Qwen2.5-1.5B, average normalized performance across 4 tasks WUDI 0.862 0.873

These data come from Tables 1 and 2. All eight vision baselines improve on both architectures, and all seven RoBERTa baselines improve; most Qwen averages improve, but Iso-C at 0.5B is an explicit counterexample. Table 2 labels TA's 0.743โ†’0.776 change as +0.034, whereas subtraction of the displayed values gives 0.033; the original endpoints are retained here without rewriting the two quantities as consistent.

Ablation Study

The numerical discussion of Figure 7 supports the component ablation below. Figure 6 uses random structured masks with the same sparsity budget, but the cached text does not provide all exact readings, so no random-mask values are invented.

Configuration on ViT-B/32 Average accuracy gain over TA, percentage points Note
TA 0.00 Baseline average accuracy 61.37%
Attention-head masks only +1.24 Reported in the Figure 7 discussion
FFN neuron masks only +2.20 Reported in the Figure 7 discussion
Attention-head and FFN masks +3.92 Table 1: 61.37%โ†’65.29%

FFN-only filtering provides the larger gain, while combining both is best, supporting complementary information in the two component types. Random structured masks do not reliably improve TA and reduce performance on RoBERTa; the gains therefore cannot be explained solely by merging fewer parameters.

The following supplementary comparison covers the training-time branch. All entries are Qwen average normalized performance across four tasks from Table 3; this branch requires new fine-tuning rather than free postprocessing under identical access conditions.

Model Operator or diagnostic setting Conventional fine-tuning/merging CASS-Tuning
Qwen2.5-0.5B Single-task expert average, Masked FT 1.000 0.954
Qwen2.5-1.5B Single-task expert average, Masked FT 1.000 0.924
Qwen2.5-0.5B TIES merging 0.695 0.816
Qwen2.5-1.5B WUDI merging 0.862 0.883

Key Findings

  • Structured filtering is particularly effective for simple merging, but can also complement weight-space methods: TSV on ViT-B/32 improves from 73.67% to 74.91%, so the benefit is not limited to TA.
  • High Post-Mask scores diagnose preserved expert capability, not multi-task merging performance: they are 0.910 and 0.975 for the two Qwen scales. Masked FT instead evaluates single-task experts trained under constraints; the diagnostics are not interchangeable.
  • Better averages do not guarantee improvement on every capability. In appendix Table 7, TA's normalized safety score on Qwen2.5-0.5B changes from 0.687 to 0.684, despite its average improving from 0.743 to 0.776.
  • OOD performance also improves in one setting: when only 12 vision tasks are merged, TA's average on the other 8 tasks changes from 0.570 to 0.588 on ViT-B/32. This is not the main experiment that merges 20 experts.

Highlights & Insights

  • Contribution statistics connect model merging to functional components. The method selects updates that an expert relies on for its task, rather than finding a permanently pruned subnetwork suitable for every task.
  • Different metrics for heads and neurons are an important detail. Absolute contribution suits sparsely activated neurons, while relative cross-task contribution avoids treating generally active heads as task-specific structure.
  • The same structured mask can support post-hoc filtering or training-time constraints. The transferable element is update-channel selection, not interchangeability of the two settings' reference models or costs.

Limitations & Future Work

  • The authors acknowledge that fixed 20% retention may not suit every task, layer, and architecture. Adaptive allocation based on contribution distributions or task complexity is a natural extension beyond uniform top-k selection.
  • Decoder-only experiments stop at 1.5B; models of 7B or larger, more experts, and more divergent task sets remain untested. Current results do not establish equivalent effectiveness at larger scales.
  • Contribution magnitude is a proxy for task relevance, not causal importance. Partial mask overlap, nonlinear component interactions, and changes in head scores when the task set changes warrant further investigation.
  • Fixed seed 42 limits conclusions about stability; robustness to probe size does not establish robustness to probe distribution. Multi-seed, distribution-shift, and component-selection stability experiments are needed.
  • The Iso-C counterexample and individual capability regressions motivate checking downstream operators and capability boundaries. Improvements in normalized safety averages in particular cannot replace separate safety evaluation before deployment.
  • vs TIES / DARE: These methods process individual updates using magnitude, signs, or random rules, whereas CASS selects complete components by functional contribution. It can filter first and then use these operators without replacing their conflict-resolution mechanisms.
  • vs TSV / Iso-C: Subspace methods process task-matrix geometry or singular-value spectra; CASS adds input-conditioned functional statistics. They can complement one another, but the Iso-C counterexample on Qwen2.5-0.5B shows that compatibility does not guarantee improvement.
  • vs TALL-Masks: Both localize task information, but TALL-Masks uses weight differences to construct parameter masks, whereas CASS uses probe activations and projection contributions to build head/neuron masks, requiring additional task data.
  • Research direction: Explicitly distinguish shared from task-specific heads under a fixed component budget, then test whether preserving shared routing improves capabilities prone to regression, such as instruction following. This is an extension, not a result established by the paper.

Rating

  • Novelty: 4/5. Unifies contribution-aware structural selection for merging and fine-tuning constraints, with distinct criteria for heads and neurons.
  • Experimental Thoroughness: 4/5. Covers vision, encoders, decoders, and multiple operator families, but lacks larger models, multiple seeds, and distribution-robustness evaluation.
  • Writing Quality: 4/5. The appendix clearly distinguishes update filtering from output deletion, though some gain annotations differ from the displayed endpoints because of rounding.
  • Value: 4/5. A practical lightweight pre-filter for existing merging pipelines, subject to operator compatibility and per-capability validation.