๐ฆ Model Compression¶
๐ง NeurIPS2026 ยท 5 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (89) ยท ๐ท CVPR2026 (108) ยท ๐ฌ ICLR2026 (240) ยท ๐ฌ ACL2026 (59) ยท ๐งช ICML2026 (117) ยท ๐ค AAAI2026 (60)
๐ฅ Top topics: Model Compression ร2
- CASS: Contribution-Aware Structured Sparsity for Model Merging
-
CASS uses unlabeled task samples to identify attention heads and FFN neurons with prominent contributions, then filters task vectors or constrains fine-tuning gradients to reduce merging interference, improving Task Arithmetic's average accuracy across 20 tasks on ViT-B/16 from 65.02% to 69.99%.
- Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models
-
CaRE-KD selects Forward or Reverse KL per token using relative teacherโstudent confidence and uses MC-dropout BALD to reject supervision when the teacher is uncertain but the student is comparatively certain, improving distillation quality in the tested settings and reaching 62.4 on MBPP and 73.9 on GSM8K, although gains are not universal and training becomes substantially more expensive.
- Importance-Aware OBS Pruning for Diffusion Models
-
The paper injects prompt-related spatial importance into the layer-wise OBS reconstruction objective to improve subject fidelity in highly sparse diffusion models without fine-tuning, but its global mask-coefficient ablation has an unexplained inconsistency with the stated exact equations.
- Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry Transformer
-
FGQ weights channel reconstruction errors in block-wise post-training quantization of VGGT using normalized squared gradients from its geometric tasks, prioritizing sensitive features during calibration; under W4A4, mean completeness error on 7-Scenes falls from QuantVGGT's 0.085 to 0.059, although not every metric recovers FP16 performance.
- Stabilizing the Dynamic Low-Rank Training
-
SDLRT retains the leading singular directions discarded in the previous iteration and combines compensated basis augmentation with negative feedback on truncation tolerance to mitigate rank collapse under aggressive compression, achieving 85.15% average accuracy on six SuperGLUE validation tasks.