Skip to content

๐Ÿ›ก๏ธ AI Safety

๐Ÿง  NeurIPS2026 ยท 10 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (44) ยท ๐Ÿ“ท CVPR2026 (145) ยท ๐Ÿ”ฌ ICLR2026 (141) ยท ๐Ÿ’ฌ ACL2026 (5) ยท ๐Ÿงช ICML2026 (114) ยท ๐Ÿค– AAAI2026 (45)

๐Ÿ”ฅ Top topics: Adversarial Robustness ร—2

A Unified Uncertainty Representation for Graph Neural Networks via Doubly-Spectral Stochastic Expansion

DSS-GNN expands node representations along graph-frequency and stochastic chaos-order axes and propagates uncertainty through deterministic quadrature; its standalone mode obtains the lowest mean Brier among compared methods on 14 calibration benchmarks, while its hybrid mode outperforms the listed baselines on most node-OOD settings and all 7 GOOD concept-shift settings.

Anchoring Adversarial Trajectories to Data Manifolds: A Bilevel Transfer Optimization Framework

MABT combines empirical manifold anchoring with bilevel initialization learning under local surrogate uncertainty, increasing error transfer in controlled cross-model ImageNet evaluation, while its geometric interpretation, unknown-model distribution approximation, and convergence claims have explicit limits.

ASIRF: An Agentic Framework for Context-Dependent Sensitive Information Redaction

ASIRF moves sensitive-attribute definitions from fixed model labels into an inference-time knowledge base and adapts through context classification, definition retrieval, and value extraction; the authors report that at least one architecture exceeds OPF recall in 68 of 80 modelโ€“domain combinations, but the multi-agent architecture does not consistently outperform the single agent with the same retrieval tool.

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

GT-HarmBench organizes 1,535 strategic-interaction instances into six two-player, two-action game families, reports average utilitarian accuracy of 0.62 across 15 models, and observes aggregate improvements of +0.13 to +0.18 from institutional descriptions in a separate eight-model experiment, but measures textual advice rather than institutional execution or deployment safety.

How Much Must a Private Mempool Hide? Exact Leakage Thresholds for Sandwich Attacks

For a single-order model of a fee-free constant-product automated market maker, the paper proves that interval-leakage security boundaries are governed primarily by the lower endpoint of order size rather than interval width, while separating all-type guarantees, posterior expected incentives, and residual post-trade risks.

Less Supervision, Better Generalization: Weakly Supervised Fake Region Localization in Diffusion-Edited Images

ReGFLoW trains a dual-encoder, VAE-reconstruction-guided regional scoring network with image-level real/fake labels, improving average cross-generator F1 in mixed partial-edit/full-synthesis evaluation at a substantial in-domain localization cost; it does not beat the strongest mask-supervised baseline's OOD average on partial edits alone.

Rank-Constrained Adaptation for Reliable Real-World Performance

MARLA uses error probabilities from a frozen ERM model on a held-out adaptation set with task labels to construct a weighted feature subspace, learns a low-rank logit correction only within that subspace, and improves worst-group accuracy without subgroup labels while separating model-selection conditions with no, partial, and complete subgroup knowledge.

Social Choice Foundations for Simulation-Augmented Generation

This paper formalizes representative routing in simulation-augmented generation as proportional clustering: predict simulation-response viewpoint embeddings, then select a small set of simulators with SEAR; it establishes approximate representation guarantees under explicit reward-factorization and error assumptions and outperforms clustering and random-routing baselines on two simulation pools, without studying final-answer synthesis.

Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning

TRACE treats gradient leakage in embodied reinforcement learning as a temporally correlated privacy problem, evaluating observation and action leakage through temporal learning under a strict ordered per-step gradient visibility assumption, reaching approximately 18.8 dB PSNR in the main experiment and showing that lightweight compression cannot replace sequence-level privacy protection.

Watermarking Should Be Treated as a Monitoring Primitive

Through an observer-based threat model and controlled text and image experiments, this paper argues that persistent entity bindings and reliable inference can make watermarking support repeated attribution, motivating governance of both internal attribution access and design-dependent external source identification beyond per-sample robustness.