Skip to content

โš–๏ธ Alignment & RLHF

๐ŸŽž๏ธ ECCV2026 ยท 1 paper notes

๐Ÿ“Œ Same area in other venues: ๐Ÿ“ท CVPR2026 (12) ยท ๐Ÿ”ฌ ICLR2026 (102) ยท ๐Ÿ’ฌ ACL2026 (38) ยท ๐Ÿงช ICML2026 (37) ยท ๐Ÿค– AAAI2026 (17) ยท ๐Ÿง  NeurIPS2025 (36)

Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

Omni-RRM synthesizes preference data with five-criterion justifications through dual-teacher consensus, then uses SFTโ†’GRPO to learn pairwise response discrimination across images, video, and audio, raising the 7B backbone's five-benchmark average accuracy from 60.2% to 70.4% and improving response selection without updating the generator's parameters.