Skip to content

๐Ÿง  VLM Reasoning

๐Ÿง  NeurIPS2026 ยท 2 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (164) ยท ๐Ÿ“ท CVPR2026 (150) ยท ๐Ÿ”ฌ ICLR2026 (112) ยท ๐Ÿ’ฌ ACL2026 (32) ยท ๐Ÿงช ICML2026 (31) ยท ๐Ÿค– AAAI2026 (10)

Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?

LIFT extracts the hidden-state difference at identical answer tokens with and without a reasoning trace from a base large language model, then injects it into the corresponding vision-language model's language layers without updating the backbone, raising six-benchmark average accuracy from 66.6/67.8 to 68.3/69.0 with static intervention and 69.1/69.4 with vector adaptation.

OneCanvas: 3D Scene Understanding via Panoramic Reprojection

OneCanvas lifts multi-view image patch features into 3D and places them as separate continuous tokens in a shared panoramic coordinate system, using native RoPE for angles and frame order and an additive embedding for metric position, then combines synthetic spatial pretraining with real-scene QA adaptation to score 65.3, 71.3, and 72.1 on SQA3D, VSI-Bench, and SPBench, respectively.