Skip to content

โšก LLM Efficiency

๐Ÿง  NeurIPS2026 ยท 6 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (1) ยท ๐Ÿ“ท CVPR2026 (8) ยท ๐Ÿ”ฌ ICLR2026 (171) ยท ๐Ÿ’ฌ ACL2026 (23) ยท ๐Ÿงช ICML2026 (48) ยท ๐Ÿค– AAAI2026 (9)

Adaptive Mass-Segmented KV Compression for Long-Form Reasoning

AMS leaves token-importance scorers unchanged and instead uses attention-derived quality mass to form adaptive segments and allocate retention quotas before local selection, mitigating contiguous context loss in long-form reasoning; on Math500 with a 256-token cache budget, AMS-Expected improves over AdaKV-ExpE2 by 16.0 percentage points.

Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

Without training the model or changing the base sampler's stopping rule, ADAS dynamically discounts confidence using a candidate's attention to positions already selected in the current step and their uncertainty, improving parallel decoding at low numbers of denoiser evaluations; task-and-method-average gains are 9.11 and 10.46 percentage points for the two models.

Block Sparse Flash Attention

BSFA first computes all causally visible QK scores exactly inside FlashAttention-2, then uses offline-calibrated block-maximum thresholds to skip V loads, PV, and softmax-state updates for low-scoring blocks, achieving a LongBench score of 39.78% versus the dense baseline's 40.24% on Llama-3.1-8B and a 1.13ร— end-to-end prefill speedup on the longest 10 samples.

Fractional State Space Transition for Long Sequence Modeling

Frac approximates fractional long memory with finite exponential modes on a shared geometric timescale bank, then adds token-wise controls and independent read/write routing, preserving bounded recurrent state while raising the 1.3B model's average LongBench score from GDN's 16.0 to 17.9, without leading on every task or throughput measure.

It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

Synth expands sourced encyclopedic seeds into heterogeneous instruction and reasoning data, enabling the final Baguettotron student to acquire instruction-following and factual capabilities from random initialization in one training process; the 600M variant reaches 79.3% precision on verifiable facts about seed entities, but this does not establish broad knowledge coverage or assume that teachers and data generation are free.

SEED: Self-Speculative Decoding via Implicit Encoderโ€“Decoder

SEED trains the last two layers of an existing decoder-only model as a drafter that reuses deep KV caches from the previous full-model verification, achieving average decoding speedups of 2.6ร—/2.7ร— over AR on Qwen3-1.7B/4B while maintaining or improving task quality under task-specific fine-tuning and adaptive tree verification.