๐งฉ Multimodal VLM¶
๐ง NeurIPS2026 ยท 9 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (225) ยท ๐ท CVPR2026 (419) ยท ๐ฌ ICLR2026 (211) ยท ๐ฌ ACL2026 (82) ยท ๐งช ICML2026 (89) ยท ๐ค AAAI2026 (75)
๐ฅ Top topics: Multimodal/VLM ร7
- Beyond Prediction: Steering VLM Agents with Retrospective World Modeling
-
RWM makes VLM agents infer the previous action from textual belief states derived from consecutive observations and converts action-transition consistency into training feedback; the full configuration achieves the paper's reported 0.81 Overall success rate across four task families, but its gains jointly involve retrospective reasoning, Bi-Level GAE, external judging, and SCR rather than SCR alone.
- Binding Multiple Modalities via Multimodal Wasserstein Barycenter
-
BaryBind replaces a fixed modality anchor with a learnable Wasserstein barycenter map and applies volumetric contrastive learning to modality gap vectors around that anchor, improving zero-shot retrieval and classification on a VAST backbone, although its published dual derivation, potential cancellation, and some experimental numbers require clarification.
- Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding
-
Exemplar2VQA compiles spatial QA exemplars into reusable simulator capture and geometric annotation programs, using four-role collaboration and execution feedback to reduce generation errors; approximately 10K synthetic examples raise the author-reported average score of Qwen2.5-VL-3B on the multiple-choice subset of VSI-Bench from 35.3 to 42.9.
- MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
-
MulTaBench constructs a benchmark of 20 image-tabular and 20 text-tabular tasks by requiring both complementary multimodal signal and gains from target-aware representations over frozen embeddings, finding that adaptation gains generalize to new learners but are not consistently significant on every selected dataset.
- SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data
-
SemMSA feeds incomplete language, visual, and acoustic evidence into a frozen LLM, recursively appends continuous hidden states as auxiliary semantics, and learns robust representations through within-instance kernel spectral alignment and cross-instance separation, achieving average Acc-2 scores of 74.36/73.91 on MOSI, 79.61/79.38 on MOSEI, and 75.46 on SIMS across ten intra-modal missing rates.
- Spherical Interpolation for Backward-Compatible Multimodal Representations
-
Without rebuilding the old gallery index, Procrustes first aligns new-model queries to the old space, then spherical interpolation combines them with old-model queries for the same input; endpoint complementarity accounts for most average retrieval gains, while support-set weight selection further improves compatibility success rates.
- The Alignment Illusion in Multimodal Large Language Models
-
Using norm-matched noise and irrelevant images in 13 multimodal large language models, this paper shows that shared MLP weights can produce high similarity and introduces the principal-angle gap to distinguish one-directional collapse from multi-directional structure, while emphasizing that geometry alone cannot establish the use of question-relevant visual content.
- Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets
-
Argus compares uncertainty methods using unified single-step GUI click records under different observable interfaces, finding stronger cross-dataset method-ranking transfer at a fixed model than across models or from open weights to API-only systems, while error discrimination cannot replace probability calibration or spatial coverage checks.
- Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs
-
Through representational analysis and activation interventions on controlled successful trajectories, this paper explains how audio-visual LLMs bind โwho says whatโ through temporal and position IDs, then uses an existing active speaker detector for visual prompting and optional lightweight fine-tuning to improve conversation understanding, with training-free gains depending on visual-marker grounding capabilities.