Breaking Rigidity in Adversarial Patch Attacks¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/visheshrajput2408/SmuDPatch.git
Area: AI Safety
Keywords: adversarial patch attack, physical adversarial defense, vision-language models, deformable shape optimization, robustness benchmarking
TL;DR¶
Addressing the pervasive reliance on rigid geometric shapes (such as squares and circles) in existing adversarial patch attacks and defenses, this paper proposes SmuDPatch—a universal smooth deformable adversarial patch generation framework utilizing centripetal Catmull-Rom splines and a two-stage evolutionary search—which demonstrates strong cross-architecture transferability across CNNs, ViTs, and VLMs, while exposing critical blind spots in state-of-the-art defense detectors.
Background & Motivation¶
Deep Neural Networks (DNNs) and Vision-Language Models (VLMs) have been extensively deployed in security-critical domains such as autonomous driving and biometric authentication, yet they remain vulnerable to adversarial perturbations. Among physical-world attack paradigms, adversarial patches represent one of the most practical threat vectors. Pioneered by early studies and standardized by benchmarks such as ImageNet-Patch, the overwhelming majority of existing patch generation frameworks inherently assume fixed, rigid geometries—predominantly squares, rectangles, or circles. This long-standing convention has inadvertently conditioned patch detectors (e.g., NAPGuard, AdvPatchXAI) and modern multimodal foundation models (e.g., LLaVA-NeXT, SAM-3) with strong shape inductive biases, generating a misleading sense of physical security.
The fundamental challenge in breaking shape rigidity lies in generating smooth, physically realizable, and universal deformable patches that generalize across unseen images and diverse model families under an identical perturbation budget. The few existing deformable patch methods (notably DAPatch) optimize shape and texture jointly on a per-image basis via radial polygon rays; this practice inevitably yields jagged, fragmented boundaries that are impractical to fabricate, while almost completely failing to generalize in universal attack settings. Furthermore, non-differentiable operations in shape rasterization and contour generation impede direct end-to-end gradient-based shape evolution.
This paper addresses the problem through a decoupled, two-stage methodology: first searching for an optimal smooth geometric boundary over a low-dimensional radial parameter space via centripetal Catmull-Rom splines under a fixed-area constraint, and subsequently optimizing universal adversarial textures across an ensemble of heterogeneous architectures with affine augmentations. Core idea: decouple universal adversarial patch generation into a two-stage sequential pipeline featuring centripetal Catmull-Rom spline evolutionary shape search and ensemble-driven robust texture optimization, yielding physically deployable, smooth, and highly transferable non-rigid patches (SmuDPatch).
Method¶
Overall Architecture¶
The SmuDPatch framework proceeds through two sequential stages: evolutionary shape optimization and robust physical texture optimization. Given a target class \(y_t\), training images sampled from ImageNet, and a model ensemble, the first stage parameterizes the patch boundary into a 25-dimensional radial vector, interpolates closed smooth contours using centripetal Catmull-Rom splines, normalizes candidate shapes to a constant pixel area budget, and identifies the best geometry using an evolutionary algorithm guided by inner-loop texture fitness. The second stage freezes the discovered optimal mask, applies stochastic physical affine transformations (rotation and translation), and iteratively refines the adversarial texture using gradient descent over a heterogeneous model ensemble (ViT, ResNet50, and MobileNetV2), producing the final universal SmuDPatch.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Target Class & Training Image Pool"] --> B["Low-Dimensional Radial Parameterization & Spline Smoothing<br/>N=25 Anchors + Catmull-Rom Curves"]
B --> C["Area-Matched Normalization Baseline<br/>Constant ~5% Image Area Constraint"]
C --> D["Two-Stage Evolutionary Shape Search<br/>Short Inner Texture Optimization Fitness"]
D --> E["Ensemble & Affine-Augmented Texture Optimization<br/>Multi-Architecture Gradients + Physical Robustness"]
E --> F["Universal Smooth Deformable Adversarial Patch (SmuDPatch)"]
Key Designs¶
1. Low-Dimensional Radial Parameterization & Spline Smoothing: Eliminating Self-Intersections and Cusps
Unlike prior deformable methods that optimize unconstrained dense binary masks or radial rays prone to irregular, serrated edges, this paper constrains the patch topology to a star-convex closed contour centered at \(c = (c_a, c_b)\). The angular space \([0, 2\pi)\) is uniformly sampled across \(N=25\) fixed directions \(\theta_i = \frac{2\pi i}{N}\), reducing the geometric degrees of freedom entirely to a radial distance vector \(\mathbf{r} = [r_0, r_1, \dots, r_{N-1}]\). The corresponding anchor coordinates are defined by \(P_i = [c_a + r_i \cos\theta_i, c_b + r_i \sin\theta_i]^T\). To prevent the harsh angular cusps produced by standard polygonal vertex connections, the method interpolates the anchors using a centripetal Catmull-Rom spline (CCRS), where each segment is determined by four neighboring anchor points:
Connecting all segments yields the closed smooth boundary \(\mu_{con} = \bigcup_{i=0}^{N-1} C_i\). Mathematically, centripetal parameterization prevents self-intersecting loops and cusps, producing smooth, organic silhouettes that can be readily printed, cut, and adhered to physical objects without tearing or edge degradation.
2. Area-Matched Normalization Baseline: Decoupling Geometry from Perturbation Budget
In adversarial patch research, the perturbed pixel area strongly correlates with attack success; any uncontrolled expansion in patch footprint compromises fair comparison. To isolate the intrinsic effectiveness of shape deformation from area advantages, every candidate spline boundary \(\mu_{con}\) is rasterized to compute its polygonal area \(A(\mu_{con})\) and strictly normalized to an invariant target area budget \(A_{\text{target}}\) (standardized to \(5\%\) of a \(224 \times 224\) image, i.e., 2,508 pixels, with \(8\%\) and \(10\%\) examined in supplementary studies). Each contour point is scaled radially from the center by:
This rigorous normalization guarantees that all attack gains stem solely from the deceptive geometric properties of the contour and its interaction with deep receptive fields, rather than an inflated perturbation budget.
3. Two-Stage Evolutionary Shape Search: Guiding Non-Differentiable Topology via Inner-Loop Texture Loss
Because spline curve generation, polygon rasterization, and binary masking involve non-differentiable operations, standard backpropagation cannot directly optimize the radial vector \(\mathbf{r}\). The framework adopts a population-based evolutionary algorithm (population size 30, evolved across 20 generations), where each individual corresponds to a candidate vector \(\mathbf{r}^{(k)}\). To quantify the attack potential of a specific geometry, a temporary random texture is initialized within the candidate mask and optimized for a minimal number of gradient steps while keeping the shape fixed. The final cross-entropy loss \(L_{\text{shape}}(\mathbf{r})\) achieved during this short inner optimization defines the shape's fitness score:
Higher fitness denotes a shape geometry that more readily fosters deceptive textures. At each generation, individuals are ranked; the top \(50\%\) are preserved as parents, while one-point crossover and Gaussian mutations generate children, striking an effective balance between global geometric exploration and compute efficiency.
4. Ensemble & Affine-Augmented Texture Optimization: Ensuring Physical and Cross-Architecture Transferability
Once the optimal binary mask \(\mu\) is selected, the second stage focuses on universal texture optimization. To ensure that the physical patch reliably misclassifies diverse images to target class \(y_t\) while enduring real-world viewing variations, the training objective integrates a stochastic distribution of affine transformations \(\mathcal{T}\) (rotations within \([-\pi/8, \pi/8]\) and translations within \(\pm 68\) pixels to avoid boundary clipping) alongside an ensemble of three architecturally distinct models: ResNet50 (deep CNN), MobileNetV2 (compact CNN), and Vision Transformer (ViT, self-attention without convolutional inductive bias). The optimization minimizes:
Gradients are accumulated across all ensemble models and transformation passes at each epoch to update patch parameters \(\delta\), ensuring the resulting texture disrupts both convolutional receptive fields and multi-head self-attention maps.
Loss & Training¶
During the texture optimization stage, the patch is trained using Stochastic Gradient Descent (SGD) with a learning rate of 1.0 for 2,000 epochs under cross-entropy loss. For each of the 10 target classes, an optimization set is constructed by randomly sampling 20 non-target images from the ImageNet validation set. Optimization runs on a single NVIDIA RTX A6000 GPU (48 GB VRAM) within a standard PyTorch environment.
Key Experimental Results¶
Main Results¶
Evaluation is performed on the ImageNet 5,000-image test subset (aligned with the RobustBench protocol) targeting 10 classes, covering 15 deep models categorized into Ensemble (ENS), Standard, Adversarially Robust (Adv), Augmentation-trained (Augm.), and Billion-scale Pretrained (More-Data).
Table 1 details Clean Accuracy (Clean Acc), Robust Accuracy (Robust Acc, lower is better), and Attack Success Rate (ASR, higher is better) comparing DAPatch, fixed circular patches (Circle), and the proposed SmuDPatch.
| Metric | MobileNetV2 (ENS) | ResNet50 (ENS) | ViT (ENS) | ResNet18 (Standard) | GoogLeNet (Standard) | Wong et al. [53] (Adv) | Hendrycks [17] (Augm.) | Yalniz [58]-b (More-Data) |
|---|---|---|---|---|---|---|---|---|
| Clean Acc. (%) | 71.9 | 76.7 | 80.9 | 69.7 | 69.8 | 53.5 | 76.8 | 82.1 |
| DAPatch Robust ↓ (%) | 60.5 | 66.9 | 79.0 | 59.6 | 60.1 | 47.0 | 73.6 | 77.3 |
| DAPatch ASR ↑ (%) | 1.4 | 3.5 | 0.8 | 0.6 | 0.8 | 0.1 | 0.3 | 1.7 |
| Circle Robust ↓ (%) | 46.3 | 49.7 | 74.2 | 53.7 | 55.3 | 41.1 | 64.1 | 69.3 |
| Circle ASR ↑ (%) | 17.9 | 27.9 | 7.4 | 2.3 | 4.0 | 0.2 | 5.8 | 8.7 |
| SmuDPatch Robust ↓ (%) | 31.6 | 28.5 | 65.5 | 45.5 | 46.0 | 38.4 | 52.2 | 56.8 |
| SmuDPatch ASR ↑ (%) | 44.2 | 62.1 | 20.4 | 12.0 | 18.4 | 0.2 | 24.2 | 26.9 |
Ablation & Defense Evasion Experiments¶
To evaluate how breaking shape rigidity affects SOTA physical patch detectors and promptable segmentation foundation models, Table 2 summarizes performance when attacking the promptable segmentation model SAM-3 (8,000 images per shape family prompted with "adversarial patch") and standard object detectors (5 SOTA detectors evaluated on clean vs. SmuDPatch images).
| Target System / Task | Evaluation Metric | Rigid Square Patch / Clean | SmuDPatch (Deformable) | Performance Degradation / Analysis |
|---|---|---|---|---|
| SAM-3 Segmentation [6] | mIoU (%) | 39.54 | 9.41 | Massive relative collapse of -76.2% |
| SAM-3 Segmentation [6] | mDice (%) | 39.84 | 9.56 | Drop of -30.28% in mask alignment |
| SAM-3 Segmentation [6] | Successful Detections | 3212 / 8000 | 778 / 8000 | Localization rate falls from 40.2% to 9.7% |
| Object Detection: Faster R-CNN | mAP (%) | 36.88 (Clean) | 32.50 | Cross-task drop of -4.38 mAP |
| Object Detection: YOLO-v11 | mAP (%) | 45.66 (Clean) | 35.09 | Classification patch causes -10.57 mAP loss |
| Object Detection: Average Across 5 | Average mAP (%) | 41.86 (Clean) | 36.63 | Consistent transfer degradation of -5.23 mAP |
Key Findings¶
- Generalization of Deformable Patches: In universal attack settings, per-image optimization methods like DAPatch fail completely (achieving only 3.5% ASR on ResNet50). In contrast, SmuDPatch achieves a 62.1% targeted ASR on ResNet50, reducing robust accuracy to 28.5%, proving that smooth non-rigid patches can attain universal effectiveness.
- Vulnerabilities in SOTA Defense and Foundation Models: Promptable segmentation models such as SAM-3 fail to locate SmuDPatch, dropping from 39.54% mIoU on squares to 9.41%, missing 7,222 out of 8,000 instances. Similarly, LLaVA-NeXT detects 1,704 square patches but only 765 SmuDPatch instances, while the dedicated defense NAPGuard drops from 64.07% mAP to 45.97%, proving that current defenses overfit to rigid rectangular boundaries.
- Physical Feasibility and Cross-Task Transfer: In physical experiments involving four everyday objects (orange, knife, computer mouse, remote control) across 120 images, ResNet50 accuracy collapsed from 66.6% to 2.5%, and ViT dropped from 74.17% to 31.67%. Additionally, without any detector-specific training, SmuDPatch caused a 10.57 mAP degradation on YOLO-v11.
Highlights & Insights¶
- Centripetal Splines Overcome Geometric Jaggedness: Constraining 25-dimensional radial variables with centripetal Catmull-Rom splines guarantees cusp-free and loop-free smooth contours, transforming theoretical deformable patches into physically cuttable and deployable artifacts.
- Decoupled Two-Stage Search Outperforms Joint Optimization: The framework demonstrates that guiding non-differentiable geometric evolution via short inner-loop texture loss circumvents gradient truncation in polygon rasterization, offering far higher stability than unconstrained joint relaxation.
- Large-Scale Non-Rigid Benchmarks for the Defense Community: The authors contribute two extensive benchmarks—comprising 100,000 ImageNet-Patch images and 40,000 COCO out-of-distribution patch images—providing the research community with the foundational data necessary to train shape-agnostic patch defenses.
Limitations & Future Work¶
- Acknowledged Limitations: While superior to prior attacks, the targeted attack success rate across unseen adversarially robust architectures and certain vision transformers leaves room for improvement compared to white-box attacks tuned to a single model.
- Experimental Constraints: Physical tests on 120 images established feasibility, but broader evaluations covering non-planar cloth wrinkles, extreme camera angles, and dynamic illumination remain to be fully benchmarked. Moreover, evolutionary search requires multiple population evaluations, making patch generation slower than fixed-shape baselines.
- Improvement Directions: Integrating the proposed differentiable variant (using Laplacian pyramid blending and soft masks) into differentiable rendering engines could enhance environmental camouflage; incorporating multimodal reward feedback could further guide geometric search to evade vision-language models.
Related Work & Insights¶
- vs. Brown et al. (2017) & Pintor et al. (2023) (ImageNet-Patch): Prior works relied strictly on rigid square or circular patches, fostering false confidence in existing defenses; SmuDPatch demonstrates that non-rigid shapes dramatically erode defense detection rates.
- vs. DAPatch (Chen et al., ECCV 2022): DAPatch optimizes per-image polygonal rays, leading to jagged contours that fail in universal attack settings (ASR of 0.6%–3.5%); SmuDPatch uses smooth spline modeling and evolutionary search to achieve high universal attack efficacy (up to 62.1% ASR).
- vs. NAPGuard [55] & AdvPatchXAI [24]: Modern patch defense methods exhibit substantial performance drops (14.1%–29.9% mAP drop for NAPGuard) against smooth deformable patches, highlighting the urgency of adopting shape-agnostic defense paradigms.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First systematic framework breaking the rigidity assumption in universal physical patch attacks via centripetal splines and evolutionary search.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive coverage spanning 15 classifiers, 5 detectors, SAM-3, LLaVA-NeXT, 140,000 benchmark images, and 120 physical real-world test images.
- Writing Quality: ⭐⭐⭐⭐☆ Clear geometric formulations and thorough discussions connecting attack design to defense vulnerabilities.
- Value: ⭐⭐⭐⭐⭐ High impact in exposing blind spots across computer vision models and providing essential datasets for developing universal, shape-agnostic defenses.