Skip to content

๐Ÿ”’ LLM Safety

๐Ÿง  NeurIPS2026 ยท 5 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (20) ยท ๐Ÿ“ท CVPR2026 (12) ยท ๐Ÿ”ฌ ICLR2026 (185) ยท ๐Ÿ’ฌ ACL2026 (115) ยท ๐Ÿค– AAAI2026 (41) ยท ๐Ÿง  NeurIPS2025 (81)

๐Ÿ”ฅ Top topics: LLM ร—3 ยท Alignment/RLHF ร—2

Contrastive Representation Shaping for LLM Unlearning

CLReg uses augmented views of the same forget example as positives and retain examples as negatives to shape hidden representations alongside a base unlearning objective, improving aggregate scores in multiple settings without equating representation separation with knowledge removal or providing privacy guarantees.

LLM Alignmentโ€“Utility Asymmetry under Semantic-Preserving Transformations

Through paired evaluation of original and semantic-preserving representations, this paper finds that task capability transfer does not guarantee corresponding transfer of safety behavior, and qualifies this empirical finding through benign-data conditions, protocol controls, and defensive supervision analysis.

Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

Using 213 selected cases in SkillCascade-Bench, the paper shows that passing individual skill reviews does not imply safe joint execution; sandbox stress tests across three agent systems and eight models yield an author-reported average success rate of 89.4%, while a behavior-composition defense achieves only partial improvement.

Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models

The paper measures inconsistent safety judgments between reasoning traces and final answers with DSAR and introduces SARA, an on-policy reinforcement learning method combining early safety awareness, full-trace safety, and answer safety; it improves consistency under perturbed reasoning on two DeepSeek-based models, but does not outperform answer-only rewards in every setting or safety metric.

UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?

UnlearningSoup replaces some repeated tuning in training-based large language model unlearning with evaluation-guided weight interpolation: EfficientSoup searches using the original model and two trained unlearned models, while PerformanceSoup merges existing candidates, improving empirical forgettingโ€“retention quality and selection costs across benchmarks without eliminating base unlearning training or providing certified forgetting.