Skip to content

FaceArmor: A Universal Facial Image Protection Against Diffusion-Based Manipulations

Conference: ECCV 2026
Paper: ECCV Official
Full-Text Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-5785.txt
Area: Image Generation
Keywords: Diffusion Defense, Adversarial Perturbations, Face Privacy Protection, Cross-Architecture Transferability, Spatial Gradient Decoupling

TL;DR

FaceArmor departs from model-specific white-box defenses by attacking the intrinsic representation feature space of facial images across decoupled structural and semantic tiers, resolving spatial gradient conflicts via surrogate attention heatmaps and PCGrad to achieve universal zero-knowledge protection across unseen diffusion architectures and manipulation paradigms.

Background & Motivation

The rapid advancement of diffusion models has significantly lowered the technical barrier for photorealistic facial synthesis and personalized manipulation, giving rise to acute privacy vulnerabilities, non-consensual impersonation, and disinformation campaigns. Malicious actors routinely harvest facial portraits from public web platforms and deploy diverse generative manipulation paradigms against them. The modern threat landscape encompasses four major paradigms: localized region reconstruction via Image Inpainting, holistic stylistic and attribute modifications via Image Editing (e.g., InstructPix2Pix), subject likeness cloning via Face Mimicry (e.g., DreamBooth, LoRA), and visual-prompt-guided Identity-Preserving Generation (e.g., InstantID). Establishing proactive defense mechanisms that safeguard facial images prior to online dissemination has thus become an urgent research imperative.

Existing proactive defense methodologies overwhelmingly operate under restrictive white-box assumptions, optimizing adversarial perturbations against specific internal components of known generative pipelines. Representative defenses like PhotoGuard and EditShield focus heavily on disrupting the variational autoencoder (VAE) encoder, AdvDM and SDST target intermediate representations within the denoising U-Net, while FaceLock and IDProtector concentrate on biometric recognition or cross-attention layers. Because distinct manipulation paradigms exploit disparate internal computational pathways, perturbations overfitted to isolated modules collapse when confronted with unseen generative architectures or alternative manipulation pipelines. Consequently, existing defenses fail to maintain consistent efficacy across the broader spectrum of real-world threats.

This paper is motivated by a fundamental observation: regardless of specific network topologies, all diffusion-based manipulation frameworks fundamentally rely on extracting and aligning two universal dimensions of an image—physical spatial structure and high-level conceptual semantics. Moving defense optimization from model-dependent modules to the intrinsic representation feature space offers an architecture-invariant attack surface. Core idea: decouple facial representations into a structural tier (disrupting latent reconstruction and fine textures) and a semantic tier (severing cross-modal alignment and biometric identities), and employ continuous attention heatmaps extracted directly from surrogate extractors as adaptive soft weights alongside PCGrad to resolve multi-objective spatial conflicts for universal zero-knowledge transferability.

Method

Overall Architecture

FaceArmor generates an imperceptible adversarial perturbation \(\delta\) bounded by \(\|\delta\|_\infty \le \eta\) on a clean facial photograph \(x\) to yield a protected image \(x + \delta\). The end-to-end framework operates across two coupled stages. In the first stage, the facial representation is decomposed into four feature-level defense objectives spanning two orthogonal tiers: a structural tier that corrupts macroscopic latent reconstruction and dense spatial textures, and a conceptual tier that breaks vision-language alignment and obscures biometric identity. In the second stage, acknowledging that distinct feature gradients naturally dominate divergent spatial regions, FaceArmor utilizes intrinsic continuous attention heatmaps extracted during forward passes of surrogate networks as soft spatial weights. It couples this with a smooth-transition PCGrad projection to resolve optimization conflicts along semantic boundaries, optimizing the final perturbation under an Expectation over Transformation (EOT) framework using Projected Gradient Descent (PGD).

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Clean Face Image x"] --> B["Multi-Tier Intrinsic Feature Decoupling<br/>Structural Tier (VAE/ReferenceNet) + Conceptual Tier (CLIP/ArcFace)"]
    B --> C["Self-Guided Attention Soft Weighting<br/>Extract Face/Foreground/Background Spatial Heatmaps"]
    C --> D["Smooth-Transition Gradient Projection<br/>Hadamard Soft Masking + PCGrad Gradient Surgery"]
    D --> E["Expectation over Transformation & PGD Update<br/>EOT Transformation Robustness Augmentation"]
    E --> F["Protected Output Image x + δ<br/>Thwarts Inpainting/Editing/Mimicry/ID Preservation"]

Key Designs

1. Multi-Tier Intrinsic Feature Decoupling: Architecture-Agnostic Representation Attack

Conventional defenses overfit to specific components of a known diffusion backbone, rendering them brittle under cross-architecture transfer. FaceArmor instead establishes an intrinsic representation attack surface comprising four complementary dimensions structured across two tiers. Within the structural tier, to corrupt macroscopic physical layout reconstruction, FaceArmor targets an ensemble of pre-trained VAE encoders \(\mathcal{E} = \{E_i\}_{i=1}^N\) drawn from distinct generative families (e.g., SD v1.5 and SDXL), maximizing the average latent Euclidean distance: $$ \mathcal{L}E(x, x+\delta) = \frac{1}{N} \sum^N |E_i(x) - E_i(x+\delta)|2 $$ Concurrently, a pre-trained ReferenceNet \(R\) is leveraged to target dense spatial feature representations via \(\mathcal{L}_R(x, x+\delta) = \|R(x) - R(x+\delta)\|_2\), inducing catastrophic textural and physical consistency collapse in localized non-facial structures. Within the conceptual tier, cross-modal correspondence is disrupted via a pre-trained CLIP image encoder \(C\) using both an untargeted divergence loss and a targeted misalignment loss that steers visual features toward an orthogonal concept \(t_{\text{target}}\): $$ \mathcal{L}_C^t = \frac{(C(x) - C(x+\delta)) \cdot (C(t $$ This is complemented by an ArcFace biometric feature extractor }}) - C(t_{\text{target}}))}{|C(x) - C(x+\delta)| \cdot |C(t_{\text{orig}}) - C(t_{\text{target}})|\(F\) maximizing the identity distance \(\mathcal{L}_F(x, x+\delta) = \|F(x) - F(x+\delta)\|_2\), ensuring that any downstream personalized synthesis inevitably yields an identity-incongruent face.

2. Self-Guided Attention Soft Weighting: Intrinsic Semantic Spatial Decoupling

Simultaneously optimizing multi-objective adversarial perturbations introduces severe spatial interference: identity loss gradients \(\nabla_\delta \mathcal{L}_F\) are meaningful exclusively within core facial landmarks, whereas fine-grained texture loss gradients \(\nabla_\delta \mathcal{L}_R\) primarily impact non-facial areas such as hair, apparel, and background. Naive linear aggregation triggers destructive gradient cancellation and produces conspicuous visual boundary artifacts. Rather than relying on external segmentation models, FaceArmor extracts intrinsic continuous attention heatmaps directly from the surrogate models during the forward pass. Specifically, the spatial attention map from the final convolutional layer of ArcFace serves as the facial weight \(W_{\text{face}}\); Grad-CAM activations from the CLIP visual encoder conditioned on target text prompts form the foreground weight \(W_{\text{fg}}\), with background weight \(W_{\text{bg}} = 1 - W_{\text{fg}}\); ReferenceNet loss is restricted to non-facial zones via \(W_{\text{ref}} = 1 - W_{\text{face}}\); and global latent disruption operates uniformly with \(W_{\text{latent}} = 1\). This soft-masking mechanism ensures natural spatial continuity and preserves visual imperceptibility.

3. Smooth-Transition Gradient Projection: Conflict Elimination and Real-World Robustness

While soft spatial weighting substantially dampens broad cross-region collisions, overlapping transition zones along facial contours still encounter conflicting gradient directions among weighted objectives \(\hat{g}_j = W_j \odot \nabla_\delta \mathcal{L}_j(x, x+\delta)\) for \(j \in \{F, C, R, E\}\). FaceArmor integrates continuous soft weights with Projecting Conflicting Gradients (PCGrad). Whenever two weighted gradients exhibit negative inner products (\(\hat{g}_i \cdot \hat{g}_j < 0\)), \(\hat{g}_i\) is projected orthogonally onto the normal plane of \(\hat{g}_j\). The continuous soft weights establish an overlapping buffer zone that smooths numerical transitions and eliminates sharp boundary artifacts. Furthermore, to survive social platform image re-encoding, transmission noise, and resizing, the optimization incorporates an Expectation over Transformation (EOT) formulation across transformation set \(T\) (Gaussian blur, JPEG compression, Gaussian noise, random resizing): $$ g_t^{\text{update}} = \mathbb{E}{t \sim T} \left[ \text{PCGrad}\left( { W_j \odot \nabla\delta \mathcal{L}j(t(x), t(x+\delta_t)) } \right) \right] $$ The perturbation is iteratively updated along the sign of \(g_t^{\text{update}}\) within the \(L_\infty\) ball of radius \(\eta\), yielding robust and imperceptible universal protection.

Loss & Training

Adversarial optimization is executed via sign-based Projected Gradient Descent (PGD) with EOT over \(N = 100\) iterations with a perturbation budget of \(\epsilon = \eta = 8/255\). At each iteration \(t\): $$ \delta_{t+1} = \Pi_\eta \left( \delta_t + \alpha \cdot \text{sign}(g_t^{\text{update}}) \right) $$ All surrogate models (VAE ensemble, CLIP, ArcFace, ReferenceNet) remain frozen throughout optimization. Defenses are generated once per user image offline prior to public sharing, without requiring gradient access or backpropagation through downstream generative networks.

Key Experimental Results

Main Results

Experiments were conducted on 800 facial images (200 identities) sampled from CelebA-HQ and VGG-Face2 across four manipulation tasks: Image Inpainting, Image Editing (InstructPix2Pix), Face Mimicry (DreamBooth), and Identity-Preserving Generation (InstantID). Surrogate defenses crafted on SD v1.5 were evaluated both on SD v1.5 and transferred to SD XL. Performance metrics include generation degradation metrics (FID, LPIPS; higher indicates stronger disruption) and source similarity metrics (SSIM, CLIP score, Face Recognition match rate FR; lower indicates superior defense).

Task Target Model Defense Method FID ↑ LPIPS ↑ SSIM ↓ CLIP ↓ FR ↓
Image Inpainting SD v1.5 PhotoGuard 117.41 0.3721 0.5754 0.8548 0.8784
AdvDM 140.89 0.3733 0.5439 0.8689 0.8988
FaceLock 119.04 0.3512 0.5492 0.8866 0.6752
FaceArmor (Ours) 139.15 0.4772 0.4385 0.8177 0.6162
SD XL PhotoGuard 106.39 0.3416 0.6013 0.8772 0.8960
AdvDM 127.30 0.3483 0.5691 0.8857 0.9059
FaceLock 108.83 0.3274 0.5712 0.9045 0.7109
FaceArmor (Ours) 129.30 0.4562 0.4617 0.8338 0.6355
Image Editing SD v1.5 AdvDM 131.59 0.5235 0.3748 0.6187 0.4950
FaceLock 132.80 0.4932 0.3968 0.6635 0.3002
FaceArmor (Ours) 163.43 0.5727 0.3121 0.6247 0.3425
SD XL AdvDM 119.00 0.4876 0.4002 0.6475 0.5203
FaceLock 120.90 0.4620 0.4225 0.6855 0.3302
FaceArmor (Ours) 150.10 0.5466 0.3355 0.6380 0.3610
Face Mimicry SD v1.5 Mist 167.09 0.5933 0.2762 0.6721 0.3956
AdvDM 136.91 0.6338 0.2147 0.6751 0.3716
FaceArmor (Ours) 160.70 0.6634 0.2252 0.6745 0.3657
SD XL Mist 152.55 0.5605 0.3029 0.6925 0.4275
FaceArmor (Ours) 149.65 0.6419 0.2484 0.6887 0.3857
ID Preservation SD v1.5 IDProtector 161.05 0.5344 0.3502 0.7591 0.6006
FaceArmor (Ours) 168.94 0.5533 0.3484 0.7647 0.5528
SD XL IDProtector 147.62 0.5125 0.3732 0.7765 0.6257
FaceArmor (Ours) 159.57 0.5336 0.3786 0.7834 0.5780

In zero-knowledge cross-architecture black-box transfer tests on modern Diffusion Transformer (DiT) models (SD 3 and Flux) under Image Inpainting, prior defenses exhibited severe performance degradation, whereas FaceArmor maintained robust protection across all five metrics: - Stable Diffusion 3: FaceArmor attained FID 156.44, LPIPS 0.3131, SSIM 0.6839, CLIP 0.7826, and FR 0.9479, clearly outperforming the strongest baseline AdvDM (FID 143.88, LPIPS 0.2830, SSIM 0.7048, CLIP 0.8480); - Flux: FaceArmor achieved FID 115.23, LPIPS 0.2906, SSIM 0.7218, CLIP 0.7332, and FR 0.9464, substantially outstripping Mist (FID 108.17, CLIP 0.9148) and AdvDM (FID 105.92, CLIP 0.8946).

Ablation Study

To evaluate the contribution of each individual feature objective across all four manipulation paradigms, the component-wise ablation is summarized below:

Configuration \(\mathcal{L}_E\) (Latent) \(\mathcal{L}_C\) (Semantic) \(\mathcal{L}_F\) (Biometric ID) \(\mathcal{L}_R\) (Texture) FID ↑ LPIPS ↑ SSIM ↓ CLIP ↓ FR ↓ Note
Baseline ✓ 132.15 0.5215 0.3830 0.7604 0.7540 Baseline VAE disruption
+ Semantic ✓ ✓ 145.40 0.5329 0.3523 0.6928 0.7615 Markedly reduces CLIP alignment
+ Face ID ✓ ✓ 138.63 0.5086 0.3614 0.7533 0.4308 Slashes identity match to 0.4308
+ Texture ✓ ✓ ✓ 162.01 0.5510 0.3405 0.7010 0.4796 High FID through texture collapse
Full Model ✓ ✓ ✓ ✓ 158.32 0.5781 0.3205 0.6846 0.4522 Best overall balanced defense

In the ablation on optimization mechanisms: - Utilizing baseline loss \(\mathcal{L}_E\) alone achieved a mean four-task LPIPS of 0.5215; - Aggregating all four losses via naive summation without region-aware soft weights or PCGrad improved LPIPS to 0.5510, but induced severe visual boundary artifacts; - Integrating attention-guided soft weighting with smooth-transition PCGrad projection achieved an optimal LPIPS of 0.5781 and SSIM of 0.3205 while eliminating unnatural boundary traces.

In real-world transformation robustness testing, FaceArmor scored LPIPS 0.557 / SSIM 0.325 without transformation. Under JPEG compression (Quality=75), it maintained LPIPS 0.431 / SSIM 0.449, far surpassing EditShield (0.265 / 0.640) and AdvDM (0.310 / 0.595). Under Gaussian noise (\(\sigma=0.05\)), random rotation (\(\pm 10^\circ\)), and resizing (\(0.9\times\)), FaceArmor consistently maintained LPIPS above 0.40, demonstrating resilience against channel post-processing.

Key Findings

  • Critical Role of Biometric Disruption: Disputing latent or semantic features alone leaves facial identification largely intact (FR 0.7540–0.7615), leaving subjects exposed to identity mimicry; incorporating \(\mathcal{L}_F\) immediately cuts the biometric match rate to 0.4308, validating explicit identity-space perturbation as essential.
  • Gradient Conflict Elimination: Naive multi-objective aggregation creates severe gradient cancellation around facial perimeters; combining intrinsic attention soft weighting with PCGrad resolves gradient collisions, boosting global perceptual disruption (LPIPS) by +0.027 while keeping the user's source image visually pristine.
  • Seamless DiT Cross-Architecture Generalization: On modern DiT models (SD 3 and Flux), baselines that overfit to U-Net denoising dynamics experience severe defensive drop-offs, whereas FaceArmor's intrinsic representation attack transfers smoothly without architectural assumptions.

Highlights & Insights

  • Architecture-Agnostic Intrinsic Attack Manifold: By grounding the defense on intrinsic physical layout and semantic feature tiers rather than model-specific internal blocks, FaceArmor avoids representation overfitting and achieves unprecedented zero-knowledge black-box transferability.
  • Segmentation-Free Intrinsic Attention Decoupling: Repurposing ArcFace convolutional attention maps and CLIP Grad-CAM activations as spatial soft masks bypasses the need for auxiliary segmentation models, naturally aligning loss functions with their semantic domains while establishing continuous transition buffers.
  • Practical Deployment Resilience via EOT: Seamlessly integrating physical transformation expectations directly into the projected multi-objective gradient flow ensures that defensive perturbations survive social media re-compression and downsampling pipelines.

Limitations & Future Work

  • Optimization Overhead: Evaluating forward and backward passes across multiple surrogate networks (VAE ensemble, ReferenceNet, CLIP, ArcFace) alongside PCGrad orthogonalization results in an average optimization time of 90.2 seconds per image on an NVIDIA A800 GPU (compared to 28.5s for PhotoGuard). While acceptable for one-time offline protection before web upload, lighter surrogate distillation is desirable for real-time mobile execution.
  • Advanced Purification Defenses: Highly aggressive diffusion purification or continuous-time stochastic smoothing filters may partially attenuate high-frequency adversarial perturbations; exploring low-frequency structural perturbations warrants future investigation.
  • vs PhotoGuard / EditShield: Early defenses perturb specific components (e.g., VAE encoders or cross-attention) in a white-box SD v1.5 setup. FaceArmor demonstrates that such targeted attacks fail under cross-architecture transfer and multi-task manipulation, resolving this via decoupled structural and conceptual representation attacks.
  • vs FaceLock / IDProtector: Identity-centric methods confine perturbations exclusively to the inner facial region, leaving background manipulation and semantic editing unhindered. FaceArmor achieves a balanced defense across both facial identity and non-facial context via region-aware attention soft weighting.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Decouples defense from generative network topologies into an intrinsic dual-tier feature space with self-guided attention soft weighting.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 4 manipulation tasks, 4 diffusion architectures (including DiT-based SD 3 and Flux), 7 baselines, transformation robustness, and ablation studies.
  • Writing Quality: ⭐⭐⭐⭐⭐ Coherent structure, clear technical problem formulation, sound mathematical grounding, and readable visual flowcharts.
  • Value: ⭐⭐⭐⭐⭐ Delivers an actionable, highly transferable, and imperceptible proactive defense safeguarding personal facial photographs against malicious diffusion manipulations.