Skip to content

โš–๏ธ Alignment & RLHF

๐Ÿง  NeurIPS2026 ยท 2 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (6) ยท ๐Ÿ“ท CVPR2026 (12) ยท ๐Ÿ”ฌ ICLR2026 (102) ยท ๐Ÿ’ฌ ACL2026 (38) ยท ๐Ÿงช ICML2026 (37) ยท ๐Ÿค– AAAI2026 (17)

Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable

Holding the wrong answer fixed and changing its endorser, the paper uses direction removal and an independently fitted attribution patch to show selectively intervenable differences between wrong-source deference and user agreement in three open-weight families, so one sycophancy evaluation cannot substitute for the other.

Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment

GAPO constructs an ephemeral geometric anchor near the current policy along the direction that decreases the batch-average preference margin, then reduces brittle pairs' update weights according to their margin degradation, improving multiple alignment evaluations and controlled noise experiments without winning every modelโ€“metric combination.