Skip to content

โš–๏ธ Alignment & RLHF

๐Ÿ“น ICCV2025 ยท 2 paper notes

๐Ÿ“Œ Same area in other venues: ๐Ÿ“ท CVPR2026 (12) ยท ๐Ÿ”ฌ ICLR2026 (102) ยท ๐Ÿ’ฌ ACL2026 (38) ยท ๐Ÿงช ICML2026 (37) ยท ๐Ÿค– AAAI2026 (17) ยท ๐Ÿง  NeurIPS2025 (36)

Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models

This paper proposes HIMRD, a black-box multimodal jailbreak attack method that bypasses unimodal safety mechanisms by distributing malicious semantics across multiple modalities. A heuristic search strategy is employed to identify optimal understanding-enhancing prompts and inducing prompts, achieving average attack success rates of approximately 90% and 68% on open-source and closed-source multimodal large language models, respectively.

MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization

This paper proposes MagicID, a framework that constructs hybrid video pair data capturing identity and dynamic preferences, and designs a two-stage Hybrid Preference Optimization (HPO) training strategy. MagicID is the first work to apply DPO to identity-customized video generation, simultaneously addressing identity degradation and motion weakening caused by conventional self-reconstruction training.