โก LLM Efficiency¶
๐ง NeurIPS2026 ยท 6 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (1) ยท ๐ท CVPR2026 (8) ยท ๐ฌ ICLR2026 (171) ยท ๐ฌ ACL2026 (23) ยท ๐งช ICML2026 (48) ยท ๐ค AAAI2026 (9)
- Adaptive Mass-Segmented KV Compression for Long-Form Reasoning
-
AMS leaves token-importance scorers unchanged and instead uses attention-derived quality mass to form adaptive segments and allocate retention quotas before local selection, mitigating contiguous context loss in long-form reasoning; on Math500 with a 256-token cache budget, AMS-Expected improves over AdaKV-ExpE2 by 16.0 percentage points.
- Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
-
Without training the model or changing the base sampler's stopping rule, ADAS dynamically discounts confidence using a candidate's attention to positions already selected in the current step and their uncertainty, improving parallel decoding at low numbers of denoiser evaluations; task-and-method-average gains are 9.11 and 10.46 percentage points for the two models.
- Block Sparse Flash Attention
-
BSFA first computes all causally visible QK scores exactly inside FlashAttention-2, then uses offline-calibrated block-maximum thresholds to skip V loads, PV, and softmax-state updates for low-scoring blocks, achieving a LongBench score of 39.78% versus the dense baseline's 40.24% on Llama-3.1-8B and a 1.13ร end-to-end prefill speedup on the longest 10 samples.
- Fractional State Space Transition for Long Sequence Modeling
-
Frac approximates fractional long memory with finite exponential modes on a shared geometric timescale bank, then adds token-wise controls and independent read/write routing, preserving bounded recurrent state while raising the 1.3B model's average LongBench score from GDN's 16.0 to 17.9, without leading on every task or throughput measure.
- It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs
-
Synth expands sourced encyclopedic seeds into heterogeneous instruction and reasoning data, enabling the final Baguettotron student to acquire instruction-following and factual capabilities from random initialization in one training process; the 600M variant reaches 79.3% precision on verifiable facts about seed entities, but this does not establish broad knowledge coverage or assume that teachers and data generation are free.
- SEED: Self-Speculative Decoding via Implicit EncoderโDecoder
-
SEED trains the last two layers of an existing decoder-only model as a drafter that reuses deep KV caches from the previous full-model verification, achieving average decoding speedups of 2.6ร/2.7ร over AR on Qwen3-1.7B/4B while maintaining or improving task quality under task-specific fine-tuning and adaptive tree verification.