โ๏ธ Alignment & RLHF¶
๐ง NeurIPS2026 ยท 2 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (6) ยท ๐ท CVPR2026 (12) ยท ๐ฌ ICLR2026 (102) ยท ๐ฌ ACL2026 (38) ยท ๐งช ICML2026 (37) ยท ๐ค AAAI2026 (17)
- Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable
-
Holding the wrong answer fixed and changing its endorser, the paper uses direction removal and an independently fitted attribution patch to show selectively intervenable differences between wrong-source deference and user agreement in three open-weight families, so one sycophancy evaluation cannot substitute for the other.
- Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment
-
GAPO constructs an ephemeral geometric anchor near the current policy along the direction that decreases the batch-average preference margin, then reduces brittle pairs' update weights according to their margin degradation, improving multiple alignment evaluations and controlled noise experiments without winning every modelโmetric combination.