๐ฎ Reinforcement Learning¶
๐๏ธ ECCV2026 ยท 1 paper notes
๐ Same area in other venues: ๐ท CVPR2026 (25) ยท ๐ฌ ICLR2026 (400) ยท ๐ฌ ACL2026 (46) ยท ๐งช ICML2026 (110) ยท ๐ค AAAI2026 (58) ยท ๐ง NeurIPS2025 (143)
- ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval
-
ELVA proposes a rule-based verifiable reward RL framework that co-optimizes the retrieval-ranking capability of MLLMs through ranking rewards and margin rewards. This addresses the "grain blindness" issue that arises when adapting contrastive learning to retrieval tasks. ELVA achieves SOTA performance on both M-BEIR and a self-built multi-grained benchmark MRBench, with a 13.1% improvement on MRBench.