๐ LLM Safety¶
๐ง NeurIPS2026 ยท 5 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (20) ยท ๐ท CVPR2026 (12) ยท ๐ฌ ICLR2026 (185) ยท ๐ฌ ACL2026 (115) ยท ๐ค AAAI2026 (41) ยท ๐ง NeurIPS2025 (81)
๐ฅ Top topics: LLM ร3 ยท Alignment/RLHF ร2
- Contrastive Representation Shaping for LLM Unlearning
-
CLReg uses augmented views of the same forget example as positives and retain examples as negatives to shape hidden representations alongside a base unlearning objective, improving aggregate scores in multiple settings without equating representation separation with knowledge removal or providing privacy guarantees.
- LLM AlignmentโUtility Asymmetry under Semantic-Preserving Transformations
-
Through paired evaluation of original and semantic-preserving representations, this paper finds that task capability transfer does not guarantee corresponding transfer of safety behavior, and qualifies this empirical finding through benign-data conditions, protocol controls, and defensive supervision analysis.
- Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems
-
Using 213 selected cases in SkillCascade-Bench, the paper shows that passing individual skill reviews does not imply safe joint execution; sandbox stress tests across three agent systems and eight models yield an author-reported average success rate of 89.4%, while a behavior-composition defense achieves only partial improvement.
- Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
-
The paper measures inconsistent safety judgments between reasoning traces and final answers with DSAR and introduces SARA, an on-policy reinforcement learning method combining early safety awareness, full-trace safety, and answer safety; it improves consistency under perturbed reasoning on two DeepSeek-based models, but does not outperform answer-only rewards in every setting or safety metric.
- UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?
-
UnlearningSoup replaces some repeated tuning in training-based large language model unlearning with evaluation-guided weight interpolation: EfficientSoup searches using the original model and two trained unlearned models, while PerformanceSoup merges existing candidates, improving empirical forgettingโretention quality and selection costs across benchmarks without eliminating base unlearning training or providing certified forgetting.