Skip to content

๐Ÿงฉ Multimodal VLM

๐Ÿง  NeurIPS2026 ยท 9 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (225) ยท ๐Ÿ“ท CVPR2026 (419) ยท ๐Ÿ”ฌ ICLR2026 (211) ยท ๐Ÿ’ฌ ACL2026 (82) ยท ๐Ÿงช ICML2026 (89) ยท ๐Ÿค– AAAI2026 (75)

๐Ÿ”ฅ Top topics: Multimodal/VLM ร—7

Beyond Prediction: Steering VLM Agents with Retrospective World Modeling

RWM makes VLM agents infer the previous action from textual belief states derived from consecutive observations and converts action-transition consistency into training feedback; the full configuration achieves the paper's reported 0.81 Overall success rate across four task families, but its gains jointly involve retrospective reasoning, Bi-Level GAE, external judging, and SCR rather than SCR alone.

Binding Multiple Modalities via Multimodal Wasserstein Barycenter

BaryBind replaces a fixed modality anchor with a learnable Wasserstein barycenter map and applies volumetric contrastive learning to modality gap vectors around that anchor, improving zero-shot retrieval and classification on a VAST backbone, although its published dual derivation, potential cancellation, and some experimental numbers require clarification.

Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding

Exemplar2VQA compiles spatial QA exemplars into reusable simulator capture and geometric annotation programs, using four-role collaboration and execution feedback to reduce generation errors; approximately 10K synthetic examples raise the author-reported average score of Qwen2.5-VL-3B on the multiple-choice subset of VSI-Bench from 35.3 to 42.9.

MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

MulTaBench constructs a benchmark of 20 image-tabular and 20 text-tabular tasks by requiring both complementary multimodal signal and gains from target-aware representations over frozen embeddings, finding that adaptation gains generalize to new learners but are not consistently significant on every selected dataset.

SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data

SemMSA feeds incomplete language, visual, and acoustic evidence into a frozen LLM, recursively appends continuous hidden states as auxiliary semantics, and learns robust representations through within-instance kernel spectral alignment and cross-instance separation, achieving average Acc-2 scores of 74.36/73.91 on MOSI, 79.61/79.38 on MOSEI, and 75.46 on SIMS across ten intra-modal missing rates.

Spherical Interpolation for Backward-Compatible Multimodal Representations

Without rebuilding the old gallery index, Procrustes first aligns new-model queries to the old space, then spherical interpolation combines them with old-model queries for the same input; endpoint complementarity accounts for most average retrieval gains, while support-set weight selection further improves compatibility success rates.

The Alignment Illusion in Multimodal Large Language Models

Using norm-matched noise and irrelevant images in 13 multimodal large language models, this paper shows that shared MLP weights can produce high similarity and introduces the principal-angle gap to distinguish one-directional collapse from multi-directional structure, while emphasizing that geometry alone cannot establish the use of question-relevant visual content.

Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets

Argus compares uncertainty methods using unified single-step GUI click records under different observable interfaces, finding stronger cross-dataset method-ranking transfer at a fixed model than across models or from open weights to API-only systems, while error discrimination cannot replace probability calibration or spatial coverage checks.

Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs

Through representational analysis and activation interventions on controlled successful trajectories, this paper explains how audio-visual LLMs bind โ€œwho says whatโ€ through temporal and position IDs, then uses an existing active speaker detector for visual prompting and optional lightweight fine-tuning to improve conversation understanding, with training-free gains depending on visual-marker grounding capabilities.