Visible Yet Unrecognizable: Frequency-Selective Facial Privacy via Attention¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/atulkr05/MIRAGE.git
Area: Human Understanding
Keywords: Facial Privacy Protection, Frequency-Selective Manipulation, Attention Mechanism, Deep Face Recognition, Visual Fidelity
TL;DR¶
Addressing the core privacy-utility tradeoff where existing de-identification either ruins visual quality or produces completely unfamiliar faces, MIRAGE leverages Discrete Wavelet Transform to separate face images into low- and high-frequency components, perturbing identity via AdaIN cross-attention on low frequencies while preserving visual attributes via CBAM self-attention on high frequencies, backed by a frequency-aware consistency loss to fool nine deep face recognition models while retaining realistic visual appearance.
Background & Motivation¶
With the proliferation of deep learning and unauthorized face scraping, public online portraits are routinely collected without consent to train commercial surveillance systems and biometric classifiers, epitomized by large-scale scraping controversies. Protecting facial privacy requires modifying images so that automated deep face recognition (DFR) systems fail to match or identify the individual, while the visual appearance remains natural and consistent to human observers, ensuring the image retains its social utility for everyday sharing.
However, existing privacy-preserving methodologies suffer from a severe dichotomy. Traditional spatial obfuscations such as blurring, pixelation, or synthetic masking introduce conspicuous artifacts that render images unusable. Conversely, contemporary generative and latent-space de-identification models (e.g., A3GAN, FALCO, FAS) synthesize completely different facial identities; while they successfully disrupt face recognition, the resulting outputs no longer look like the original person, destroying their social utility. On the other hand, imperceptible adversarial noise attacks often exhibit weak cross-model transferability and are easily purified by off-the-shelf diffusion models.
Frequency analysis of deep face recognition models reveals an indispensable insight: face recognition networks exhibit asymmetric sensitivity across frequency bands. Under 2D Discrete Wavelet Transform (DWT), matching using only the low-frequency approximation subband (LL) yields recognition accuracy remarkably close to that of full-spectrum images, demonstrating that biometric identity signals predominantly concentrate in low-frequency bands. Conversely, high-frequency detail subbands (LH, HL, HH) primarily encode visual appearance attributes, textures, and edge details. Core Idea: Decouple facial images into frequency subbands via discrete wavelet transform, perturb the identity-discriminative low-frequency subband via AdaIN-based cross-attention against a target identity embedding while recalibrating high-frequency appearance details via unconditional self-attention, jointly optimized with a Frequency-Aware Consistency (FAC) loss to achieve robust machine de-identification alongside human visual plausibility.
Method¶
Overall Architecture¶
MIRAGE takes a source face image and a target pseudo-identity embedding to synthesize a privacy-protected face image. The framework consists of three sequential stages: dual-feature encoding with single-level DWT decomposition, attention-guided frequency manipulation, and inverse DWT reconstruction followed by spatial refinement and adaptive face masking. In the low-frequency branch, source LL spatial features attend to the target identity embedding to displace biometric signatures. In the high-frequency branch, independent CBAM modules refine textures without external identity conditioning. Finally, IDWT reconstruction is refined by a U-Net and blended with the input using a learned soft face mask.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Face Image x"] --> B["Dual Encoding & 1-level DWT Decomposition"]
B --> C["Low-Frequency Identity Perturbation<br/>AdaIN Cross-Attention with target z"]
B --> D["High-Frequency Appearance Preservation<br/>CBAM Self-Attention Modules"]
C --> E["Inverse Discrete Wavelet Transform IDWT"]
D --> E
E --> F["Spatial Refinement & Adaptive Blending<br/>U-Net Refinement + Face Mask M"]
F --> G["Protected Face Output x'"]
Key Designs¶
1. Low-Frequency Identity Perturbation: AdaIN Cross-Attention Identity Redirection To overcome the fragility and poor transferability of heuristic spatial perturbations, this module operates directly on the low-frequency approximation subband \(x^{LL}\), which carries the core facial topology and identity features. The module integrates Adaptive Instance Normalization (AdaIN) with a cross-attention mechanism: spatial features from \(x^{LL}\) serve as queries, while a randomly sampled target identity embedding \(z\) serves as keys and values: $\(x'^{LL} = \text{AdaIN-CrossAttn}(x^{LL}, z)\)$ Cross-attention allows source spatial structures to selectively align with target identity vectors, while AdaIN shifts low-frequency feature statistics to match the target identity distribution. This forces downstream DFR models to extract identity representations corresponding to the target identity rather than the original subject, without introducing high-frequency visual artifacts.
2. High-Frequency Appearance Preservation: Unconditioned CBAM Texture Recalibration To preserve fine visual detailsโsuch as facial wrinkles, gaze expression, hair strands, and lighting ambianceโeach high-frequency subband (\(x^{LH}, x^{HL}, x^{HH}\)) is processed independently by a Convolutional Block Attention Module (CBAM) without any external identity conditioning: $\(x'^{HF} = \text{CBAM}(x^{HF}), \quad HF \in \{LH, HL, HH\}\)$ The channel attention submodule adaptively recalibrates frequency channel responses, while the spatial attention submodule highlights salient perceptual structures. Operating strictly on the subband's intrinsic representations eliminates identity leakage and ensures that fine-grained visual details remain authentic and faithful to the source subject.
3. Spatial Refinement and Adaptive Blending: U-Net Refinement with Soft Facial Mask Recombining altered frequency subbands via single-level IDWT produces a coarse protected image \(\tilde{x}\) that may exhibit minor boundary artifacts. To eliminate these wavelet reconstruction artifacts, \(\tilde{x}\) is concatenated with the original image \(x\) to form a 6-channel input for a lightweight U-Net refinement network \(\mathcal{R}\). Furthermore, a lightweight mask subnet predicts a continuous facial soft mask \(M \in [0, 1]^{H \times W}\) to blend the refined face with the original image: $\(x' = M \odot x_{\text{refined}} + (1 - M) \odot x\)$ This ensures that non-facial regions (e.g., ears, hair periphery, and background) are preserved pixel-for-pixel, constraining all perturbations strictly within the facial region.
Loss & Training¶
The framework is optimized end-to-end via a composite objective: $\(\mathcal{L}_{\text{total}} = \lambda_{\text{FAC}}\mathcal{L}_{\text{FAC}} + \lambda_{\text{attr}}\mathcal{L}_{\text{attr}} + \lambda_{\text{recon}}\mathcal{L}_{\text{recon}} + \lambda_{\text{perc}}\mathcal{L}_{\text{perc}} + \lambda_{\text{adv}}\mathcal{L}_{\text{adv}}\)$ The core component is the Frequency-Aware Consistency (FAC) Loss, \(\mathcal{L}_{\text{FAC}} = \mathcal{L}_{\text{f-margin}} + \mathcal{L}_{\text{f-focal}}\):
- Frequency-Aware Margin Loss (\(\mathcal{L}_{\text{f-margin}}\)): Enforces identity separation across frequency subbands by penalizing cosine similarities that fall below band-specific margins, avoiding negative triplet mining: $\(\mathcal{L}_{\text{f-margin}} = \sum_{b \in \{LL, LH, HL, HH\}} w_b \cdot \text{ReLU}\big(m_b - d_{\text{cos}}(e_i^b, e_i'^b)\big)\)$ with band weights \(w_b \in \{1.5, 0.5, 0.5, 0.3\}\) and margins \(m_b \in \{0.7, 0.3, 0.3, 0.2\}\), heavily prioritizing the identity-critical \(LL\) subband.
- Frequency-Consistency Focal Loss (\(\mathcal{L}_{\text{f-focal}}\)): Focuses high-frequency texture reconstruction on challenging regions using dynamic SSIM weighting: $\(\mathcal{L}_{\text{f-focal}} = \sum_{b \in \{LH, HL, HH\}} \big(1 - \text{SSIM}(x^b, x'^b)\big)^\gamma \cdot \|x^b - x'^b\|_2\)$ These terms are complemented by attribute preservation loss \(\mathcal{L}_{\text{attr}} = \|e_a - e'_a\|_2\), combined \(L_1/L_2\) reconstruction loss \(\mathcal{L}_{\text{recon}}\), multi-scale VGG perceptual loss \(\mathcal{L}_{\text{perc}}\), and PatchGAN adversarial loss \(\mathcal{L}_{\text{adv}}\).
Key Experimental Results¶
Main Results¶
Following standard biometric verification protocols (1 gallery image, 3 probe images per identity), the authors evaluate nine deep face recognition models (AdaFace, MagFace, ArcFace, Dlib, VGG-Face, SFace, FaceNet, FaceNet-512, DeepFace) on three benchmark datasets (CelebA-HQ test set, 105-Pins, LFW) under both Proactive (gallery protection) and Reactive (probe protection) modes.
The table below summarizes recognition accuracy under the gender-agnostic setting comparing MIRAGE against four SOTA methods across AdaFace and MagFace (lower accuracy indicates stronger privacy protection):
| DFR Model | Method | CelebA-HQ (Proactive) โ | 105-Pins (Proactive) โ | LFW (Proactive) โ | CelebA-HQ (Reactive) โ | 105-Pins (Reactive) โ | LFW (Reactive) โ |
|---|---|---|---|---|---|---|---|
| AdaFace | FPP | 96.19% | 94.45% | 90.52% | 96.18% | 94.05% | 89.54% |
| AdaFace | DiffAM | 96.39% | 96.16% | 88.05% | 96.45% | 96.25% | 87.78% |
| AdaFace | FAS | 68.35% | 70.22% | 71.70% | 68.45% | 71.11% | 72.03% |
| AdaFace | FALCO | 56.77% | 54.46% | 60.05% | 59.70% | 56.22% | 56.92% |
| AdaFace | MIRAGE (Ours) | 70.24% | 71.63% | 65.42% | 68.56% | 69.48% | 65.63% |
| MagFace | FPP | 77.77% | 84.09% | 60.80% | 77.82% | 83.69% | 61.01% |
| MagFace | DiffAM | 83.49% | 88.25% | 58.67% | 82.25% | 90.94% | 58.42% |
| MagFace | FAS | 57.06% | 61.76% | 56.37% | 58.02% | 59.68% | 56.89% |
| MagFace | FALCO | 56.83% | 53.42% | 52.61% | 56.39% | 55.15% | 52.08% |
| MagFace | MIRAGE (Ours) | 57.87% | 61.75% | 44.13% | 56.67% | 64.21% | 44.89% |
While FALCO achieves lower recognition accuracy on several subsets, it drastically alters facial appearance, creating unrecognizable strangers. MIRAGE achieves competitive protection (e.g., dropping clean AdaFace accuracy on CelebA-HQ from 96.16% to 70.24% and achieving 65.42% on LFW) while maintaining high visual quality (average PSNR of 32.85 dB and SSIM of 0.89).
Ablation Study¶
The ablation analysis evaluates the individual contributions of low-frequency cross-attention, high-frequency self-attention, and the FAC loss:
| DFR Model | Config | CelebA-HQ (Proactive) โ | 105-Pins (Proactive) โ | LFW (Proactive) โ | CelebA-HQ (Reactive) โ | 105-Pins (Reactive) โ | LFW (Reactive) โ | Note |
|---|---|---|---|---|---|---|---|---|
| AdaFace | w/o cross-attention | 96.30% | 94.66% | 87.61% | 95.77% | 95.84% | 88.51% | Identity disruption fails; near clean baseline |
| AdaFace | w/o self-attention | 92.20% | 94.05% | 86.53% | 93.86% | 95.56% | 86.88% | Visual degradation destabilizes protection |
| AdaFace | w/o FAC Loss | 92.98% | 94.17% | 87.02% | 92.74% | 95.58% | 86.64% | Relaxed frequency margins degrade privacy |
| AdaFace | Full MIRAGE | 70.24% | 71.63% | 65.42% | 68.56% | 69.48% | 65.63% | Optimal protection across all scenarios |
| MagFace | w/o cross-attention | 80.20% | 93.61% | 58.35% | 79.81% | 85.19% | 58.86% | Drastic loss in privacy disruption |
| MagFace | w/o self-attention | 77.14% | 86.38% | 55.20% | 77.66% | 84.21% | 56.54% | Texture instability impairs overall quality |
| MagFace | w/o FAC Loss | 78.01% | 86.55% | 57.41% | 77.68% | 83.55% | 56.66% | Loss of frequency-band separation |
| MagFace | Full MIRAGE | 57.87% | 61.75% | 44.13% | 56.67% | 64.21% | 44.89% | Best overall privacy protection |
Crucially, naive DWT sub-band swapping yields only 95.47% on AdaFace (CelebA-HQ Proactive), showing that learned attention and FAC supervision are essential.
Key Findings¶
- Cross-attention on LL subband is the primary privacy driver: Removing cross-attention raises AdaFace accuracy from 70.24% to 96.30%, confirming that identity perturbation must be applied directly to the low-frequency carrier.
- Robustness against compression and resizing: Evaluating protected images under JPEG compression (QF=50 to 95) and spatial downsampling (\(112 \times 112\)) shows that recognition accuracy drops even further (e.g., AdaFace accuracy drops to 64.02% at QF=50), because lossy re-encoding further disrupts the fragile perturbed identity signal.
- Resilience to adaptive fine-tuning: When attackers fine-tune AdaFace on MIRAGE-protected faces, accuracy improves by only +1.84%, demonstrating strong resistance against adaptive defenses.
Highlights & Insights¶
- Grounded Physical Decomposition: Rather than treating disentanglement as an unconstrained latent space problem, MIRAGE grounds identity-appearance separation in the physics of 2D wavelet transforms, exploiting the fact that deep recognition models attend primarily to low frequencies.
- Asymmetric Attention Mechanism: The decoupled architecture applies AdaIN cross-attention to inject external target identities into low-frequency subbands, while employing unconditioned CBAM self-attention on high frequencies to insulate textures from identity interference.
- Triplet-Free Frequency Supervision: The FAC margin loss applies band-specific thresholds (\(LL > LH, HL > HH\)) directly on cosine distances, achieving efficient multi-scale identity separation without costly negative mining.
Limitations & Future Work¶
- Residual Accuracy under Modern Margin-Loss Models: While MIRAGE reduces AdaFace accuracy by ~26%, the absolute accuracy remains around ~70% on some sets (unlike older models like DeepFace and FaceNet512, which drop to 20%~25%), indicating that advanced margin-based architectures retain residual robustness against subtle perturbations.
- Single-Scale Haar Transform Limits: The fixed single-level Haar wavelet may cause frequency leakage under extreme facial poses, complex lighting, or severe partial occlusions.
- Future Directions: Exploring learnable multi-scale wavelet representations and extending frequency-selective privacy protection to video streams and multimodal vision-language models.
Related Work & Insights¶
- vs FALCO (CVPR 2023) / FAS (WACV 2025): FALCO and FAS rely on GAN inversion and latent space optimization, producing unrecognizable surrogate faces that destroy personal identity and social utility. MIRAGE preserves original human perception (0.89 SSIM, 32.85 dB PSNR) while obscuring identity from automated models.
- vs DiffAM (CVPR 2024) / FPP (CVPR 2025): DiffAM and FPP maintain appearance via adversarial makeup transfer or diffusion purification weakening, but their protection against modern DFR models remains weak (>90% accuracy). MIRAGE reliably suppresses recognition rates into the 60%~70% range.
Rating¶
- Novelty: โญโญโญโญโ Principled combination of wavelet decomposition and asymmetric attention provides intuitive, physically grounded de-identification.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive validation across 9 DFR models, 3 datasets, proactive/reactive protocols, compression, resizing, and adaptive retraining.
- Writing Quality: โญโญโญโญโญ Clear motivation, well-formulated methodology, and cohesive narrative.
- Value: โญโญโญโญโ Provides a practical and visually realistic solution for social media privacy protection against unauthorized web-scale scraping.