FaceArmor: A Universal Facial Image Protection Against Diffusion-Based Manipulations¶
Conference: ECCV 2026
Paper: ECCV Official
Full-Text Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-5785.txt
Area: Image Generation
Keywords: Diffusion Defense, Adversarial Perturbations, Face Privacy Protection, Cross-Architecture Transferability, Spatial Gradient Decoupling
TL;DR¶
FaceArmor departs from model-specific white-box defenses by attacking the intrinsic representation feature space of facial images across decoupled structural and semantic tiers, resolving spatial gradient conflicts via surrogate attention heatmaps and PCGrad to achieve universal zero-knowledge protection across unseen diffusion architectures and manipulation paradigms.
Background & Motivation¶
The rapid advancement of diffusion models has significantly lowered the technical barrier for photorealistic facial synthesis and personalized manipulation, giving rise to acute privacy vulnerabilities, non-consensual impersonation, and disinformation campaigns. Malicious actors routinely harvest facial portraits from public web platforms and deploy diverse generative manipulation paradigms against them. The modern threat landscape encompasses four major paradigms: localized region reconstruction via Image Inpainting, holistic stylistic and attribute modifications via Image Editing (e.g., InstructPix2Pix), subject likeness cloning via Face Mimicry (e.g., DreamBooth, LoRA), and visual-prompt-guided Identity-Preserving Generation (e.g., InstantID). Establishing proactive defense mechanisms that safeguard facial images prior to online dissemination has thus become an urgent research imperative.
Existing proactive defense methodologies overwhelmingly operate under restrictive white-box assumptions, optimizing adversarial perturbations against specific internal components of known generative pipelines. Representative defenses like PhotoGuard and EditShield focus heavily on disrupting the variational autoencoder (VAE) encoder, AdvDM and SDST target intermediate representations within the denoising U-Net, while FaceLock and IDProtector concentrate on biometric recognition or cross-attention layers. Because distinct manipulation paradigms exploit disparate internal computational pathways, perturbations overfitted to isolated modules collapse when confronted with unseen generative architectures or alternative manipulation pipelines. Consequently, existing defenses fail to maintain consistent efficacy across the broader spectrum of real-world threats.
This paper is motivated by a fundamental observation: regardless of specific network topologies, all diffusion-based manipulation frameworks fundamentally rely on extracting and aligning two universal dimensions of an image—physical spatial structure and high-level conceptual semantics. Moving defense optimization from model-dependent modules to the intrinsic representation feature space offers an architecture-invariant attack surface. Core idea: decouple facial representations into a structural tier (disrupting latent reconstruction and fine textures) and a semantic tier (severing cross-modal alignment and biometric identities), and employ continuous attention heatmaps extracted directly from surrogate extractors as adaptive soft weights alongside PCGrad to resolve multi-objective spatial conflicts for universal zero-knowledge transferability.
Method¶
Overall Architecture¶
FaceArmor generates an imperceptible adversarial perturbation \(\delta\) bounded by \(\|\delta\|_\infty \le \eta\) on a clean facial photograph \(x\) to yield a protected image \(x + \delta\). The end-to-end framework operates across two coupled stages. In the first stage, the facial representation is decomposed into four feature-level defense objectives spanning two orthogonal tiers: a structural tier that corrupts macroscopic latent reconstruction and dense spatial textures, and a conceptual tier that breaks vision-language alignment and obscures biometric identity. In the second stage, acknowledging that distinct feature gradients naturally dominate divergent spatial regions, FaceArmor utilizes intrinsic continuous attention heatmaps extracted during forward passes of surrogate networks as soft spatial weights. It couples this with a smooth-transition PCGrad projection to resolve optimization conflicts along semantic boundaries, optimizing the final perturbation under an Expectation over Transformation (EOT) framework using Projected Gradient Descent (PGD).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Clean Face Image x"] --> B["Multi-Tier Intrinsic Feature Decoupling<br/>Structural Tier (VAE/ReferenceNet) + Conceptual Tier (CLIP/ArcFace)"]
B --> C["Self-Guided Attention Soft Weighting<br/>Extract Face/Foreground/Background Spatial Heatmaps"]
C --> D["Smooth-Transition Gradient Projection<br/>Hadamard Soft Masking + PCGrad Gradient Surgery"]
D --> E["Expectation over Transformation & PGD Update<br/>EOT Transformation Robustness Augmentation"]
E --> F["Protected Output Image x + δ<br/>Thwarts Inpainting/Editing/Mimicry/ID Preservation"]
Key Designs¶
1. Multi-Tier Intrinsic Feature Decoupling: Architecture-Agnostic Representation Attack
Conventional defenses overfit to specific components of a known diffusion backbone, rendering them brittle under cross-architecture transfer. FaceArmor instead establishes an intrinsic representation attack surface comprising four complementary dimensions structured across two tiers. Within the structural tier, to corrupt macroscopic physical layout reconstruction, FaceArmor targets an ensemble of pre-trained VAE encoders \(\mathcal{E} = \{E_i\}_{i=1}^N\) drawn from distinct generative families (e.g., SD v1.5 and SDXL), maximizing the average latent Euclidean distance: $$ \mathcal{L}E(x, x+\delta) = \frac{1}{N} \sum^N |E_i(x) - E_i(x+\delta)|2 $$ Concurrently, a pre-trained ReferenceNet \(R\) is leveraged to target dense spatial feature representations via \(\mathcal{L}_R(x, x+\delta) = \|R(x) - R(x+\delta)\|_2\), inducing catastrophic textural and physical consistency collapse in localized non-facial structures. Within the conceptual tier, cross-modal correspondence is disrupted via a pre-trained CLIP image encoder \(C\) using both an untargeted divergence loss and a targeted misalignment loss that steers visual features toward an orthogonal concept \(t_{\text{target}}\): $$ \mathcal{L}_C^t = \frac{(C(x) - C(x+\delta)) \cdot (C(t $$ This is complemented by an ArcFace biometric feature extractor }}) - C(t_{\text{target}}))}{|C(x) - C(x+\delta)| \cdot |C(t_{\text{orig}}) - C(t_{\text{target}})|\(F\) maximizing the identity distance \(\mathcal{L}_F(x, x+\delta) = \|F(x) - F(x+\delta)\|_2\), ensuring that any downstream personalized synthesis inevitably yields an identity-incongruent face.
2. Self-Guided Attention Soft Weighting: Intrinsic Semantic Spatial Decoupling
Simultaneously optimizing multi-objective adversarial perturbations introduces severe spatial interference: identity loss gradients \(\nabla_\delta \mathcal{L}_F\) are meaningful exclusively within core facial landmarks, whereas fine-grained texture loss gradients \(\nabla_\delta \mathcal{L}_R\) primarily impact non-facial areas such as hair, apparel, and background. Naive linear aggregation triggers destructive gradient cancellation and produces conspicuous visual boundary artifacts. Rather than relying on external segmentation models, FaceArmor extracts intrinsic continuous attention heatmaps directly from the surrogate models during the forward pass. Specifically, the spatial attention map from the final convolutional layer of ArcFace serves as the facial weight \(W_{\text{face}}\); Grad-CAM activations from the CLIP visual encoder conditioned on target text prompts form the foreground weight \(W_{\text{fg}}\), with background weight \(W_{\text{bg}} = 1 - W_{\text{fg}}\); ReferenceNet loss is restricted to non-facial zones via \(W_{\text{ref}} = 1 - W_{\text{face}}\); and global latent disruption operates uniformly with \(W_{\text{latent}} = 1\). This soft-masking mechanism ensures natural spatial continuity and preserves visual imperceptibility.
3. Smooth-Transition Gradient Projection: Conflict Elimination and Real-World Robustness
While soft spatial weighting substantially dampens broad cross-region collisions, overlapping transition zones along facial contours still encounter conflicting gradient directions among weighted objectives \(\hat{g}_j = W_j \odot \nabla_\delta \mathcal{L}_j(x, x+\delta)\) for \(j \in \{F, C, R, E\}\). FaceArmor integrates continuous soft weights with Projecting Conflicting Gradients (PCGrad). Whenever two weighted gradients exhibit negative inner products (\(\hat{g}_i \cdot \hat{g}_j < 0\)), \(\hat{g}_i\) is projected orthogonally onto the normal plane of \(\hat{g}_j\). The continuous soft weights establish an overlapping buffer zone that smooths numerical transitions and eliminates sharp boundary artifacts. Furthermore, to survive social platform image re-encoding, transmission noise, and resizing, the optimization incorporates an Expectation over Transformation (EOT) formulation across transformation set \(T\) (Gaussian blur, JPEG compression, Gaussian noise, random resizing): $$ g_t^{\text{update}} = \mathbb{E}{t \sim T} \left[ \text{PCGrad}\left( { W_j \odot \nabla\delta \mathcal{L}j(t(x), t(x+\delta_t)) } \right) \right] $$ The perturbation is iteratively updated along the sign of \(g_t^{\text{update}}\) within the \(L_\infty\) ball of radius \(\eta\), yielding robust and imperceptible universal protection.
Loss & Training¶
Adversarial optimization is executed via sign-based Projected Gradient Descent (PGD) with EOT over \(N = 100\) iterations with a perturbation budget of \(\epsilon = \eta = 8/255\). At each iteration \(t\): $$ \delta_{t+1} = \Pi_\eta \left( \delta_t + \alpha \cdot \text{sign}(g_t^{\text{update}}) \right) $$ All surrogate models (VAE ensemble, CLIP, ArcFace, ReferenceNet) remain frozen throughout optimization. Defenses are generated once per user image offline prior to public sharing, without requiring gradient access or backpropagation through downstream generative networks.
Key Experimental Results¶
Main Results¶
Experiments were conducted on 800 facial images (200 identities) sampled from CelebA-HQ and VGG-Face2 across four manipulation tasks: Image Inpainting, Image Editing (InstructPix2Pix), Face Mimicry (DreamBooth), and Identity-Preserving Generation (InstantID). Surrogate defenses crafted on SD v1.5 were evaluated both on SD v1.5 and transferred to SD XL. Performance metrics include generation degradation metrics (FID, LPIPS; higher indicates stronger disruption) and source similarity metrics (SSIM, CLIP score, Face Recognition match rate FR; lower indicates superior defense).
| Task | Target Model | Defense Method | FID ↑ | LPIPS ↑ | SSIM ↓ | CLIP ↓ | FR ↓ |
|---|---|---|---|---|---|---|---|
| Image Inpainting | SD v1.5 | PhotoGuard | 117.41 | 0.3721 | 0.5754 | 0.8548 | 0.8784 |
| AdvDM | 140.89 | 0.3733 | 0.5439 | 0.8689 | 0.8988 | ||
| FaceLock | 119.04 | 0.3512 | 0.5492 | 0.8866 | 0.6752 | ||
| FaceArmor (Ours) | 139.15 | 0.4772 | 0.4385 | 0.8177 | 0.6162 | ||
| SD XL | PhotoGuard | 106.39 | 0.3416 | 0.6013 | 0.8772 | 0.8960 | |
| AdvDM | 127.30 | 0.3483 | 0.5691 | 0.8857 | 0.9059 | ||
| FaceLock | 108.83 | 0.3274 | 0.5712 | 0.9045 | 0.7109 | ||
| FaceArmor (Ours) | 129.30 | 0.4562 | 0.4617 | 0.8338 | 0.6355 | ||
| Image Editing | SD v1.5 | AdvDM | 131.59 | 0.5235 | 0.3748 | 0.6187 | 0.4950 |
| FaceLock | 132.80 | 0.4932 | 0.3968 | 0.6635 | 0.3002 | ||
| FaceArmor (Ours) | 163.43 | 0.5727 | 0.3121 | 0.6247 | 0.3425 | ||
| SD XL | AdvDM | 119.00 | 0.4876 | 0.4002 | 0.6475 | 0.5203 | |
| FaceLock | 120.90 | 0.4620 | 0.4225 | 0.6855 | 0.3302 | ||
| FaceArmor (Ours) | 150.10 | 0.5466 | 0.3355 | 0.6380 | 0.3610 | ||
| Face Mimicry | SD v1.5 | Mist | 167.09 | 0.5933 | 0.2762 | 0.6721 | 0.3956 |
| AdvDM | 136.91 | 0.6338 | 0.2147 | 0.6751 | 0.3716 | ||
| FaceArmor (Ours) | 160.70 | 0.6634 | 0.2252 | 0.6745 | 0.3657 | ||
| SD XL | Mist | 152.55 | 0.5605 | 0.3029 | 0.6925 | 0.4275 | |
| FaceArmor (Ours) | 149.65 | 0.6419 | 0.2484 | 0.6887 | 0.3857 | ||
| ID Preservation | SD v1.5 | IDProtector | 161.05 | 0.5344 | 0.3502 | 0.7591 | 0.6006 |
| FaceArmor (Ours) | 168.94 | 0.5533 | 0.3484 | 0.7647 | 0.5528 | ||
| SD XL | IDProtector | 147.62 | 0.5125 | 0.3732 | 0.7765 | 0.6257 | |
| FaceArmor (Ours) | 159.57 | 0.5336 | 0.3786 | 0.7834 | 0.5780 |
In zero-knowledge cross-architecture black-box transfer tests on modern Diffusion Transformer (DiT) models (SD 3 and Flux) under Image Inpainting, prior defenses exhibited severe performance degradation, whereas FaceArmor maintained robust protection across all five metrics: - Stable Diffusion 3: FaceArmor attained FID 156.44, LPIPS 0.3131, SSIM 0.6839, CLIP 0.7826, and FR 0.9479, clearly outperforming the strongest baseline AdvDM (FID 143.88, LPIPS 0.2830, SSIM 0.7048, CLIP 0.8480); - Flux: FaceArmor achieved FID 115.23, LPIPS 0.2906, SSIM 0.7218, CLIP 0.7332, and FR 0.9464, substantially outstripping Mist (FID 108.17, CLIP 0.9148) and AdvDM (FID 105.92, CLIP 0.8946).
Ablation Study¶
To evaluate the contribution of each individual feature objective across all four manipulation paradigms, the component-wise ablation is summarized below:
| Configuration | \(\mathcal{L}_E\) (Latent) | \(\mathcal{L}_C\) (Semantic) | \(\mathcal{L}_F\) (Biometric ID) | \(\mathcal{L}_R\) (Texture) | FID ↑ | LPIPS ↑ | SSIM ↓ | CLIP ↓ | FR ↓ | Note |
|---|---|---|---|---|---|---|---|---|---|---|
| Baseline | ✓ | 132.15 | 0.5215 | 0.3830 | 0.7604 | 0.7540 | Baseline VAE disruption | |||
| + Semantic | ✓ | ✓ | 145.40 | 0.5329 | 0.3523 | 0.6928 | 0.7615 | Markedly reduces CLIP alignment | ||
| + Face ID | ✓ | ✓ | 138.63 | 0.5086 | 0.3614 | 0.7533 | 0.4308 | Slashes identity match to 0.4308 | ||
| + Texture | ✓ | ✓ | ✓ | 162.01 | 0.5510 | 0.3405 | 0.7010 | 0.4796 | High FID through texture collapse | |
| Full Model | ✓ | ✓ | ✓ | ✓ | 158.32 | 0.5781 | 0.3205 | 0.6846 | 0.4522 | Best overall balanced defense |
In the ablation on optimization mechanisms: - Utilizing baseline loss \(\mathcal{L}_E\) alone achieved a mean four-task LPIPS of 0.5215; - Aggregating all four losses via naive summation without region-aware soft weights or PCGrad improved LPIPS to 0.5510, but induced severe visual boundary artifacts; - Integrating attention-guided soft weighting with smooth-transition PCGrad projection achieved an optimal LPIPS of 0.5781 and SSIM of 0.3205 while eliminating unnatural boundary traces.
In real-world transformation robustness testing, FaceArmor scored LPIPS 0.557 / SSIM 0.325 without transformation. Under JPEG compression (Quality=75), it maintained LPIPS 0.431 / SSIM 0.449, far surpassing EditShield (0.265 / 0.640) and AdvDM (0.310 / 0.595). Under Gaussian noise (\(\sigma=0.05\)), random rotation (\(\pm 10^\circ\)), and resizing (\(0.9\times\)), FaceArmor consistently maintained LPIPS above 0.40, demonstrating resilience against channel post-processing.
Key Findings¶
- Critical Role of Biometric Disruption: Disputing latent or semantic features alone leaves facial identification largely intact (FR 0.7540–0.7615), leaving subjects exposed to identity mimicry; incorporating \(\mathcal{L}_F\) immediately cuts the biometric match rate to 0.4308, validating explicit identity-space perturbation as essential.
- Gradient Conflict Elimination: Naive multi-objective aggregation creates severe gradient cancellation around facial perimeters; combining intrinsic attention soft weighting with PCGrad resolves gradient collisions, boosting global perceptual disruption (LPIPS) by +0.027 while keeping the user's source image visually pristine.
- Seamless DiT Cross-Architecture Generalization: On modern DiT models (SD 3 and Flux), baselines that overfit to U-Net denoising dynamics experience severe defensive drop-offs, whereas FaceArmor's intrinsic representation attack transfers smoothly without architectural assumptions.
Highlights & Insights¶
- Architecture-Agnostic Intrinsic Attack Manifold: By grounding the defense on intrinsic physical layout and semantic feature tiers rather than model-specific internal blocks, FaceArmor avoids representation overfitting and achieves unprecedented zero-knowledge black-box transferability.
- Segmentation-Free Intrinsic Attention Decoupling: Repurposing ArcFace convolutional attention maps and CLIP Grad-CAM activations as spatial soft masks bypasses the need for auxiliary segmentation models, naturally aligning loss functions with their semantic domains while establishing continuous transition buffers.
- Practical Deployment Resilience via EOT: Seamlessly integrating physical transformation expectations directly into the projected multi-objective gradient flow ensures that defensive perturbations survive social media re-compression and downsampling pipelines.
Limitations & Future Work¶
- Optimization Overhead: Evaluating forward and backward passes across multiple surrogate networks (VAE ensemble, ReferenceNet, CLIP, ArcFace) alongside PCGrad orthogonalization results in an average optimization time of 90.2 seconds per image on an NVIDIA A800 GPU (compared to 28.5s for PhotoGuard). While acceptable for one-time offline protection before web upload, lighter surrogate distillation is desirable for real-time mobile execution.
- Advanced Purification Defenses: Highly aggressive diffusion purification or continuous-time stochastic smoothing filters may partially attenuate high-frequency adversarial perturbations; exploring low-frequency structural perturbations warrants future investigation.
Related Work & Insights¶
- vs PhotoGuard / EditShield: Early defenses perturb specific components (e.g., VAE encoders or cross-attention) in a white-box SD v1.5 setup. FaceArmor demonstrates that such targeted attacks fail under cross-architecture transfer and multi-task manipulation, resolving this via decoupled structural and conceptual representation attacks.
- vs FaceLock / IDProtector: Identity-centric methods confine perturbations exclusively to the inner facial region, leaving background manipulation and semantic editing unhindered. FaceArmor achieves a balanced defense across both facial identity and non-facial context via region-aware attention soft weighting.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Decouples defense from generative network topologies into an intrinsic dual-tier feature space with self-guided attention soft weighting.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 4 manipulation tasks, 4 diffusion architectures (including DiT-based SD 3 and Flux), 7 baselines, transformation robustness, and ablation studies.
- Writing Quality: ⭐⭐⭐⭐⭐ Coherent structure, clear technical problem formulation, sound mathematical grounding, and readable visual flowcharts.
- Value: ⭐⭐⭐⭐⭐ Delivers an actionable, highly transferable, and imperceptible proactive defense safeguarding personal facial photographs against malicious diffusion manipulations.