Deformable and Multi-view Gradient-Aligned Physical Adversarial Camouflage¶
Conference: ECCV 2026
Paper: ECCV 2026 Official Page
Area: AI Safety
Keywords: Physical Adversarial Attack, 3D Adversarial Camouflage, Object Detection, Deformable Texture Field, First-Order Meta-Optimization
TL;DR¶
Addressing spatial gradient discontinuity and destructive cross-view gradient conflicts in 3D physical adversarial camouflage, this paper proposes a continuous Deformable Texture Field (DTF) decoupling appearance from geometric deformation, coupled with a Gradient-Aligned Meta-Optimization (GAMO) strategy that implicitly maximizes cross-view gradient alignment via multi-step exploration trajectories to achieve robust, perspective-invariant vehicle evasion.
Background & Motivation¶
Deep neural networks have become foundational in modern object detection systems deployed across autonomous driving and surveillance infrastructures. However, their vulnerability to adversarial perturbations poses severe real-world security risks. Unlike localized 2D adversarial patches that remain strictly viewpoint-dependent and constrained to planar regions, full-coverage 3D physical adversarial camouflage seeks to optimize the entire 3D mesh surface to reliably evade detection across omni-directional viewing angles. Nonetheless, synthesizing robust and physically realizable 3D camouflage textures remains fundamentally difficult due to the severe structural mismatch between static 2D texture parameterizations and dynamic 3D physical observations.
This optimization difficulty manifests as two critical hurdles within the differentiable rendering pipeline. The first is spatial gradient inconsistency: geometric self-occlusion restricts gradient back-propagation strictly to visible surfaces under each specific camera view. This asynchronous and disjoint update across iterations disrupts local spatial correlations among adjacent texels, generating fragile high-frequency noise and topological tearing that fail to survive physical printing and real-world environmental distortions. The second hurdle is cross-view gradient conflict: joint optimization over omni-directional viewpoint distributions frequently yields contradictory gradient directions (e.g., frontal versus rear viewpoints displaying negative cosine similarities). Standard joint or sequential updates merely average these opposing signals, resulting in destructive interference that traps parameters in suboptimal oscillations or causes optimization divergence rather than reaching a shared evasion manifold.
To overcome these two core bottlenecks, this work reframes the 3D texture optimization paradigm without resorting to rigid fixed grids or compute-heavy explicit gradient projection solvers. Core idea: parameterize the texture as a continuous Deformable Texture Field (DTF) driven by learnable content and geometric flow grids, and employ Gradient-Aligned Meta-Optimization (GAMO) via multi-step local exploration trajectories to implicitly maximize cross-view gradient inner products, steering convergence toward a robust, mutually reinforcing parameter manifold.
Method¶
Overall Architecture¶
The proposed framework addresses both spatial and multi-view hurdles through two unified components: the forward Deformable Texture Field (DTF) and the backward Gradient-Aligned Meta-Optimization (GAMO). Rather than optimizing discrete, unconstrained high-resolution pixels directly, the model decouples the texture into a pair of low-resolution control grids: a Content Grid capturing underlying color patterns and a Flow Grid encoding localized coordinate offsets. Dense continuous deformation fields are generated through bilinear upsampling and used to resample the content grid differentiably, inherently imposing spatial smoothness while enabling flexible geometric warping. For optimization, viewpoint clusters are treated as distinct tasks within a dual-phase meta-learning paradigm: an inner loop explores view-specific curvature through multi-step trajectories with reset optimizer momentum, while an outer loop aggregates trajectory displacements to implicitly align gradient directions across conflicting viewpoints.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Initialize Global Parameters<br/>ฮ = {Content Grid C, Flow Grid F}"] --> B["Deformable Texture Field Synthesis<br/>Bilinear Dense Flow + Inverse Resampling"]
B --> C["Differentiable Physical Rendering<br/>Mesh M Projected to Multi-view Images"]
C --> D["Multi-step Local Exploration Trajectory<br/>Reset Momentum & Update k Steps on View v"]
D --> E["Gradient-Aligned Global Meta-Update<br/>Trajectory Displacement Aggregation & Implicit Alignment"]
E --> F["Output Robust Physical Camouflage Texture T"]
Key Designs¶
1. Deformable Texture Field: Unifying Spatial Smoothness Priors with Adaptive Geometric Representation To eliminate the high-frequency artifacts and topological tearing caused by sparse, visibility-gated pixel updatesโas well as the rigid representational constraints of fixed UV gridsโthe proposed framework decouples the target high-resolution texture map \(T \in \mathbb{R}^{H \times W \times 3}\) into two low-resolution learnable control grids: a Content Grid \(C \in \mathbb{R}^{H' \times W' \times 3}\) and a Flow Grid \(F \in \mathbb{R}^{H' \times W' \times 2}\) (where \(H', W' \ll H, W\), e.g., \(64 \times 64\) control grids synthesizing a \(2048 \times 2048\) texture map). The framework first upsamples the sparse Flow Grid \(F\) to the target texture resolution using a bilinear interpolation operator \(\Phi\), yielding a dense deformation field \(\Delta P \in \mathbb{R}^{H \times W \times 2}\): $\(\Delta P(u) = \Phi(F)(u), \quad \forall u \in \Omega_{tex} = [0, 1]^2\)$ For every target texel coordinate \(u\), its warped sampling location \(p' = u + \Delta P(u)\) is computed within the continuous Content Grid space. The final color value is obtained via differentiable bilinear inverse resampling over the four nearest integer neighbor nodes in \(C\). This formulation naturally bakes spatial coherence into the representation, ensuring that adjacent pixels receive smooth, mutually supportive updates even under sparse gradient feedback. Simultaneously, the learnable geometric flow allows the pattern to adaptively warp and concentrate representational capacity on salient adversarial object boundaries without expanding the parameter budget. An \(L_2\) regularizer \(\mathcal{L}_{reg}(F) = \|F\|_2^2\) is applied to penalize excessive deformation and preserve physical printability.
2. Gradient-Aligned Meta-Optimization: Trajectory-Driven Conflict Resolution Across Viewpoints Standard joint optimization over wide viewpoint distributions leads to severe destructive interference because gradients from disparate perspectives often exhibit negative inner products. Instead of relying on explicit multi-objective gradient projections, this work reformulates multi-view adversarial learning into a dual-phase first-order meta-optimization strategy operating on global parameters \(\Theta = \{C, F\}\) and local exploration parameters \(\phi\). At the start of each episode, local parameters are initialized as \(\phi_0 \leftarrow \Theta\) and the inner optimizer state is strictly reset to flush accumulated historical momentum. The inner loop then executes \(k\) sequential update steps under currently sampled viewpoint \(v_t\): $\(\phi_{t+1} \leftarrow \phi_t - \eta \nabla_{\phi} \mathcal{L}(\phi_t, v_t)\)$ This multi-step local exploration gathers curvature information along the trajectory, producing a net displacement \((\phi_k - \phi_0)\). The outer loop subsequently updates global parameters via \(\Theta \leftarrow \Theta + \epsilon (\phi_k - \Theta)\). Applying a first-order meta-learning Taylor expansion reveals that this update direction mathematically approximates a modified objective: $\(\Delta \Theta \approx -\mathbb{E}[g_i] + \frac{\eta(k - 1)}{4} \nabla_{\Theta} \mathbb{E}[g_i \cdot g_j]\)$ While the first term enforces primary adversarial loss minimization, the emergent second term acts as an intrinsic regularizer that maximizes the inner product \(g_i \cdot g_j\) between gradients of conflicting viewpoints. This implicitly drives global parameters toward a shared, collaborative loss basin where view-specific gradients align, bypassing the prohibitive computational overhead of explicit Hessian calculations or pairwise projection quadratic programs.
Loss & Training¶
The framework is optimized in an end-to-end manner. The overall loss evaluated at rendered image \(I = \mathcal{R}(M, T, v)\) combines an adversarial objective with flow grid regularization: $\(\mathcal{L} = \mathcal{L}_{adv}(f(\mathcal{R}(M, T, v)), y_{gt}) + \lambda \mathcal{L}_{reg}(F)\)$ where \(\mathcal{L}_{adv}\) suppresses the detector's classification confidence on ground-truth foreground category \(y_{gt}\), and \(\mathcal{R}\) denotes PyTorch3D differentiable rendering equipped with binary visibility masks. Key hyper-parameters include \(64 \times 64\) grid resolution, flow regularization weight \(\lambda = 0.01\), inner exploration steps \(k = 16\), inner learning rate \(\eta = 0.1\) under SGD, and outer learning rate \(\epsilon = 1.0\). The entire model converges stably within 3 epochs on a single NVIDIA A100 GPU.
Key Experimental Results¶
Main Results¶
Experiments were conducted using the CARLA driving simulator (Audi E-Tron target vehicle, 20,000 multi-view images sampled from 5m to 20m) and 1:24 scale physical printed models. Attack performance is quantified using target vehicle Average Precision at IoU 0.5 ([email protected], %); lower values denote stronger camouflage evasion.
The table below reports black-box transferability across five advanced detector architectures (spanning modern ConvNets, Vision Transformers, and Vision-Language models) when attacked by camouflage trained against white-box YOLOv3:
| Methods | RTMDet | ConvNeXt | GLIP | DINO | Grounding DINO |
|---|---|---|---|---|---|
| Normal (Clean) | 94.49 | 95.36 | 96.24 | 85.95 | 91.24 |
| DAS | 92.61 | 83.92 | 95.54 | 79.11 | 83.69 |
| FCA | 71.56 | 56.95 | 82.58 | 60.26 | 72.12 |
| DTA | 72.18 | 61.02 | 80.84 | 35.12 | 62.60 |
| ACTIVE | 67.69 | 55.13 | 81.99 | 41.42 | 59.92 |
| RAUCA | 51.03 | 37.68 | 60.12 | 28.55 | 40.96 |
| GRAC | 52.90 | 32.66 | 41.42 | 23.03 | 47.39 |
| Ours | 45.44 | 27.84 | 43.72 | 13.36 | 35.56 |
The table below presents physical evaluation results on 1:24 scale vehicle models across diverse camera elevation angles (0ยฐ, 20ยฐ, 45ยฐ) evaluated on white-box YOLOv3 and black-box Faster R-CNN:
| Methods | YOLOv3 (0ยฐ) | YOLOv3 (20ยฐ) | YOLOv3 (45ยฐ) | Faster R-CNN (0ยฐ) | Faster R-CNN (20ยฐ) | Faster R-CNN (45ยฐ) |
|---|---|---|---|---|---|---|
| Normal | 98.01 | 94.70 | 88.08 | 99.45 | 97.42 | 94.70 |
| DAS | 85.43 | 68.21 | 73.51 | 91.39 | 84.77 | 80.79 |
| FCA | 72.85 | 52.32 | 54.30 | 98.01 | 84.77 | 68.87 |
| DTA | 68.87 | 43.71 | 35.10 | 89.40 | 80.79 | 60.93 |
| ACTIVE | 45.77 | 42.02 | 29.30 | 79.47 | 64.90 | 41.72 |
| RAUCA | 39.74 | 29.30 | 23.39 | 78.15 | 60.26 | 30.45 |
| GRAC | 41.72 | 26.44 | 20.53 | 74.17 | 54.30 | 27.08 |
| Ours | 34.53 | 23.39 | 18.53 | 70.20 | 41.15 | 23.39 |
Ablation Study¶
To isolate the individual and synergistic effects of DTF and GAMO, the table below compares modular configurations across three detector architectures and multiple elevation angles:
| DTF | GAMO | YOLOv3 (0ยฐ) | YOLOv3 (20ยฐ) | YOLOv3 (45ยฐ) | Faster R-CNN (0ยฐ) | Faster R-CNN (20ยฐ) | Faster R-CNN (45ยฐ) | PVT (0ยฐ) | PVT (20ยฐ) | PVT (45ยฐ) |
|---|---|---|---|---|---|---|---|---|---|---|
| ร | ร | 61.70 | 42.02 | 17.06 | 85.17 | 41.15 | 20.70 | 80.12 | 46.65 | 25.82 |
| โ | ร | 33.06 | 18.95 | 9.26 | 57.92 | 20.96 | 15.79 | 56.42 | 24.52 | 11.67 |
| ร | โ | 40.13 | 12.08 | 7.71 | 66.53 | 22.92 | 8.32 | 62.22 | 18.06 | 10.13 |
| โ | โ | 19.05 | 2.19 | 1.80 | 42.92 | 9.01 | 1.04 | 41.30 | 11.67 | 1.41 |
To evaluate gradient-conflict resolution specifically, different conflict-handling algorithms were benchmarked while keeping the DTF texture parameterization identical:
| Method | [email protected] (%) โ | Gradient Cosine Similarity โ | Conflict Ratio โ |
|---|---|---|---|
| DTF (Standard Average) | 40.26 | 0.0048 | 0.4871 |
| DTF + PCGrad | 20.25 | 0.0051 | 0.4709 |
| DTF + MGDA | 13.42 | 0.0063 | 0.4507 |
| DTF + GRAC | 12.34 | 0.0055 | 0.4494 |
| DTF + GAMO (Ours) | 7.68 | 0.0070 | 0.4345 |
Key Findings¶
- Adding DTF alone reduces YOLOv3 0ยฐ [email protected] from 61.70% to 33.06%, demonstrating that enforcing continuous spatial interpolation removes fragile high-frequency noise. Adding GAMO alone decreases AP to 40.13%, confirming the necessity of mitigating cross-view gradient conflicts. Their combination exhibits a powerful orthogonal synergy, dropping [email protected] to 19.05% at 0ยฐ and 1.80% at 45ยฐ.
- In controlled gradient-conflict benchmark experiments with DTF fixed, GAMO achieves the lowest AP (7.68%), highest gradient cosine similarity (0.0070), and lowest conflict ratio (0.4345), significantly outperforming pairwise projection approaches like PCGrad and MGDA.
- In computational efficiency, GAMO requires only 140.84 GFLOPs and 135.17 ms per view, markedly faster than GRAC (8505.60 GFLOPs, 430.10 ms) and RAUCA (973.21 GFLOPs, 265.67 ms), since outer loop meta-updates avoid computing costly second-order derivatives or projection matrices.
Highlights & Insights¶
- First-order meta-learning enables implicit gradient alignment: By leveraging the mathematical emergence of \(\nabla_\Theta (g_i \cdot g_j)\) through multi-step inner loop trajectories, GAMO achieves cross-view gradient alignment without the prohibitive memory and runtime costs of explicit gradient surgery.
- Continuous deformable field resolves physical tear-up: Decoupling appearance from deformation via a low-resolution Content Grid and Flow Grid injects a robust smoothness prior, preserving continuous spatial coherence across discrete visibility changes while maintaining rich adversarial representational capacity.
- Broad transferability across architectures: The generated camouflage exhibits strong black-box transferability against modern Transformers and vision-language models, dropping DINO [email protected] to 13.36% and Grounding DINO to 35.56%.
Limitations & Future Work¶
- Simplified lighting and reflectance modeling: The differentiable renderer focuses primarily on geometric projection under standard diffuse lighting, leaving specular highlights, non-Lambertian surfaces, and severe weather scattering (e.g., dense fog, rain streaks) for future exploration.
- Sensitivity to inner loop trajectory length: The number of inner exploration steps \(k\) and learning rate \(\eta\) govern the scale of implicit gradient alignment; an excessively large \(k\) risks inner loop divergence or overfitting to specific viewpoints.
- Future directions: Integrating neural rendering primitives such as 3D Gaussian Splatting (3DGS) to handle view-dependent specular effects, and exploring multi-modal physical camouflage targeting open-vocabulary vision-language foundational detectors.
Related Work & Insights¶
- vs FCA / DAS: Early techniques optimize 2D UV maps directly at pixel granularity, resulting in severe spatial tearing and overfitting to specific digital camera views; DTF guarantees intrinsic smoothness and flexible geometric deformation.
- vs GRAC / PCGrad: Prior multi-view conflict mitigations rely on explicit pairwise gradient projections or dynamic loss reweighting, causing computational complexity to escalate rapidly with viewpoint count; GAMO achieves superior alignment implicitly through trajectory dynamics at a fraction of the computational budget.
Rating¶
- Novelty: โญโญโญโญโญ Elegant integration of deformable fields with first-order meta-optimization for multi-view physical adversarial camouflage.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive digital evaluation across five diverse detector families, rigorous ablations on gradient conflicts, and full-scale 1:24 physical model validation.
- Writing Quality: โญโญโญโญโญ Lucid formulation of the dual optimization hurdles, mathematically grounded derivations, and well-structured empirical validation.
- Value: โญโญโญโญโญ Exposes critical physical vulnerabilities in state-of-the-art vision systems, providing a solid benchmark for physically grounded defense research.