โ๏ธ Alignment & RLHF¶
๐๏ธ ECCV2026 ยท 1 paper notes
๐ Same area in other venues: ๐ท CVPR2026 (12) ยท ๐ฌ ICLR2026 (102) ยท ๐ฌ ACL2026 (38) ยท ๐งช ICML2026 (37) ยท ๐ค AAAI2026 (17) ยท ๐ง NeurIPS2025 (36)
- Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
-
Omni-RRM synthesizes preference data with five-criterion justifications through dual-teacher consensus, then uses SFTโGRPO to learn pairwise response discrimination across images, video, and audio, raising the 7B backbone's five-benchmark average accuracy from 60.2% to 70.4% and improving response selection without updating the generator's parameters.