Skip to content

๐Ÿ“ฆ Model Compression

๐Ÿง  NeurIPS2026 ยท 5 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (89) ยท ๐Ÿ“ท CVPR2026 (108) ยท ๐Ÿ”ฌ ICLR2026 (240) ยท ๐Ÿ’ฌ ACL2026 (59) ยท ๐Ÿงช ICML2026 (117) ยท ๐Ÿค– AAAI2026 (60)

๐Ÿ”ฅ Top topics: Model Compression ร—2

CASS: Contribution-Aware Structured Sparsity for Model Merging

CASS uses unlabeled task samples to identify attention heads and FFN neurons with prominent contributions, then filters task vectors or constrains fine-tuning gradients to reduce merging interference, improving Task Arithmetic's average accuracy across 20 tasks on ViT-B/16 from 65.02% to 69.99%.

Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models

CaRE-KD selects Forward or Reverse KL per token using relative teacherโ€“student confidence and uses MC-dropout BALD to reject supervision when the teacher is uncertain but the student is comparatively certain, improving distillation quality in the tested settings and reaching 62.4 on MBPP and 73.9 on GSM8K, although gains are not universal and training becomes substantially more expensive.

Importance-Aware OBS Pruning for Diffusion Models

The paper injects prompt-related spatial importance into the layer-wise OBS reconstruction objective to improve subject fidelity in highly sparse diffusion models without fine-tuning, but its global mask-coefficient ablation has an unexplained inconsistency with the stated exact equations.

Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry Transformer

FGQ weights channel reconstruction errors in block-wise post-training quantization of VGGT using normalized squared gradients from its geometric tasks, prioritizing sensitive features during calibration; under W4A4, mean completeness error on 7-Scenes falls from QuantVGGT's 0.085 to 0.059, although not every metric recovers FP16 performance.

Stabilizing the Dynamic Low-Rank Training

SDLRT retains the leading singular directions discarded in the previous iteration and combines compensated basis augmentation with negative feedback on truncation tolerance to mitigate rank collapse under aggressive compression, achieving 85.15% average accuracy on six SuperGLUE validation tasks.