Skip to content

๐Ÿ“ Learning Theory

๐ŸŽž๏ธ ECCV2026 ยท 1 paper notes

๐Ÿ“Œ Same area in other venues: ๐Ÿ”ฌ ICLR2026 (293) ยท ๐Ÿงช ICML2026 (45) ยท ๐Ÿค– AAAI2026 (3) ยท ๐Ÿง  NeurIPS2025 (25) ยท ๐Ÿงช ICML2025 (16)

Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers

The paper incorporates attention probability distributions into a local bound on the self-attention Jacobian, uses tighter softmax spectral estimates to construct JaSMin regularization, and improves ViT-B accuracy under some adversarial attacks while exposing trade-offs among robustness, clean accuracy, and gradient propagation.