Skip to content

๐ŸŽฎ Reinforcement Learning

๐ŸŽž๏ธ ECCV2026 ยท 1 paper notes

๐Ÿ“Œ Same area in other venues: ๐Ÿ“ท CVPR2026 (25) ยท ๐Ÿ”ฌ ICLR2026 (400) ยท ๐Ÿ’ฌ ACL2026 (46) ยท ๐Ÿงช ICML2026 (110) ยท ๐Ÿค– AAAI2026 (58) ยท ๐Ÿง  NeurIPS2025 (143)

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval

ELVA proposes a rule-based verifiable reward RL framework that co-optimizes the retrieval-ranking capability of MLLMs through ranking rewards and margin rewards. This addresses the "grain blindness" issue that arises when adapting contrastive learning to retrieval tasks. ELVA achieves SOTA performance on both M-BEIR and a self-built multi-grained benchmark MRBench, with a 13.1% improvement on MRBench.