Skip to content

โšก VLM Efficiency

๐Ÿง  NeurIPS2026 ยท 2 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (57) ยท ๐Ÿ“ท CVPR2026 (63) ยท ๐Ÿ”ฌ ICLR2026 (18) ยท ๐Ÿ’ฌ ACL2026 (6) ยท ๐Ÿงช ICML2026 (4) ยท ๐Ÿค– AAAI2026 (5)

Beyond Selection: Token Parameterization for Extreme Visual Token Compression

Rather than selecting a few original patches, Braco constructs a compact visual interface using a DCT low-frequency backbone, basis-coordinate embeddings, budget-dependent coordinate organization, and sparse-pooled spatial residuals; at 576โ†’16 tokens, it retains a Vanilla-normalized aggregate score of 94.0 while reducing full-pipeline prefill latency to 40.59 ms.

G2TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

G2TR uses VAE latent anchors to select and merge understanding-side visual tokens before LLM prefill in separate-encoder unified multimodal models; retaining 50% of the tokens yields 99.0% relative-average understanding performance and 98.0% editing performance on BAGEL, with approximately 1.94ร— lower prefill FLOPs, but not lossless performance on every task.