Breaking High Confidence: Practical Face Impersonation under High-Security Thresholds¶
Conference: ECCV 2026
Paper: ECCV Official
Area: AI Safety
Keywords: Face Recognition Systems, Score-based Black-box Attack, High-Security Thresholds, Orthogonal Face Set, Template Inversion
TL;DR¶
Addressing critical high-security scenarios such as banking eKYC and airport identity checks (FMR = \(10^{-6}\) or commercial API confidence score 99), this paper systematically analyzes error accumulation in score-based black-box impersonation across metric distortion, subspace projection, and template inversion, proposing a correction matrix, PCA-based orthogonal face set, and SPNet to achieve over 92% attack success rate against AWS Rekognition at a confidence threshold of 99 within 100 non-adaptive queries.
Background & Motivation¶
The widespread deployment of face recognition systems (FRSs) across critical infrastructure has established exceptionally stringent security requirements. From remote financial onboarding under electronic Know-Your-Customer (eKYC) mandates to border control and airport identity checks, even marginal false acceptance rates can lead to severe identity theft and unauthorized access. To defend against impersonation attacks where non-matching faces are erroneously accepted, real-world deployments enforce conservative decision boundaries. For instance, the National Institute of Standards and Technology (NIST FRTE 1:1 evaluation) evaluates systems at a false match rate of \(\text{FMR} = 10^{-6}\), while commercial APIs such as AWS Rekognition recommend setting confidence score thresholds at 99% or higher in public-safety and law enforcement contexts. Under these regimes, the margin for representation discrepancy is vanishingly small.
However, existing adversarial and impersonation attacks on FRSs have predominantly focused on medium-security settings (such as \(\text{FMR} = 10^{-3}\) or the default threshold of 80 in commercial APIs). In the practical black-box score-based threat model, an adversary cannot inspect internal model weights, network architectures, training datasets, or scoring formulas, but can only observe scalar similarity scores or confidence values. Concurrently, commercial deployments enforce tight rate limits and financial costs per API call. Conventional zero-order optimization, hill-climbing, or evolutionary algorithms require tens or hundreds of thousands of adaptive queries, which not only trigger pattern-based query filtering defenses but also completely collapse under low query budgets (e.g., at most 100 queries). While recent non-adaptive geometric attacks significantly reduce query counts, their attack success rate drastically diminishes to near 0% under high-security thresholds, exposing an unaddressed security blind spot.
A rigorous theoretical examination reveals that this failure is driven by compounded approximation errors: metric distortion between the surrogate model and the target black-box embedding space, subspace truncation caused by projecting high-dimensional features onto a limited number of query directions, and synthesis fidelity degradation during template-to-pixel reconstruction. Under elevated decision thresholds, these errors push the recovered faces outside the narrow acceptance hyper-ball. Core idea: by mathematically modeling the template space geometry and its multi-stage error propagation, the paper introduces a metric correction matrix to rectify manifold distortion, a PCA-based orthogonal face set to minimize subspace projection gaps, and SPNet—an enhanced inversion network fusing style adaptation with skip connections—to eliminate the error bottleneck and execute successful black-box impersonation under high-security constraints without extra query overhead.
Method¶
Overall Architecture¶
The proposed black-box score-based impersonation attack operates under strict rate limits, where the adversary possesses only a locally trained surrogate FRS and an offline-trained template inversion network. The attack pipeline consists of four coordinated phases: offline construction of an optimal orthogonal query set, online black-box querying and score conversion, metric-corrected target feature projection, and high-fidelity facial image reconstruction. Offline, the adversary computes principal component analysis over large-scale public face embeddings to construct a PCA-based orthogonal face set (OFS) of size \(Q\). Online, the adversary submits these \(Q\) faces to the target black-box FRS to collect confidence scores with respect to the enrolled victim. The scalar scores are inverted into cosine similarities, calibrated via a precomputed correction matrix that accounts for inter-manifold distortion, and projected to yield an accurate target template estimate. Finally, the estimated template is passed through SPNet to synthesize the spoofed face image for target verification.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Offline Data Stream: Surrogate Embedding Space<br/>Large-scale Face Dataset MS1MV3"] --> B["PCA-based Orthogonal Face Set<br/>Principal components minimize projection gap"]
B --> C["Black-box Query Stage: Submit Q Face Queries<br/>Collect victim confidence scores from target FRS"]
C --> D["Numerical Transformation: Confidence to Similarity<br/>Invert black-box monotone outputs to distances"]
D --> E["Correction Matrix: Metric Distortion & Target Projection<br/>Weighted least squares under covariance approximation"]
E --> F["Style-Aware Skip-Connected Inversion: High-Fidelity Synthesis<br/>SPNet decodes estimated template into spoofed face"]
F --> G["Authentication Success on Target FRS<br/>Surpasses FMR=10⁻⁶ or confidence threshold 99"]
Key Designs¶
1. Correction Matrix: Taming Metric Distortion and Target Space Projection
Prior low-query attacks formulate the black-box score response as \(Ay = s - \epsilon\), where \(A \in \mathbb{R}^{Q \times d}\) is the template matrix of \(Q\) query faces in the surrogate model, \(s\) is the target similarity vector, and \(\epsilon\) denotes the unknown error arising from surrogate-target misalignment. The baseline method directly computes the pseudoinverse \(\hat{y} = A^\dagger s\), essentially solving a standard least-squares problem in the surrogate Euclidean space. This formulation neglects the metric deformation between heterogeneous feature extractors, amplifying the error term \(\|A^\dagger \epsilon\|_2\) by the inverse of the smallest singular value \(\sigma_{\min}(A)^{-1}\).
To correct this cross-model metric distortion, the authors introduce a first-order local distortion approximation via a positive definite symmetric matrix \(\Sigma \in \mathbb{R}^{d \times d}\), modeling target inner products as weighted inner products in the surrogate space, \(x^T \Sigma y\). Under this formulation, the pairwise target similarity matrix among the query set becomes \(S = A \Sigma A^T\), while the victim similarity vector satisfies \(s = A \Sigma y\). The adversary computes the corrected estimate:
Mathematically, \(\hat{y}\) precisely equals the orthogonal projection of the victim template \(y\) onto the row space \(R(A)\) under the inner product metric induced by \(\Sigma\):
By finding the optimal projection directly under the target system's induced metric, the correction matrix mitigates manifold deformation bias and substantially shrinks the initial metric translation error \(E_1\).
2. PCA-based Orthogonal Face Set: Minimizing Subspace Projection Gaps
Even in the absence of cross-model metric distortion, reconstructing a \(d\)-dimensional template (\(d = 512\)) from only \(Q\) scalar queries (\(Q \ll d\), typically \(Q = 100\)) is an underdetermined inverse problem that inevitably incurs a subspace truncation error \(E_2\). The baseline approach solely enforces mutual orthogonality among query templates to minimize matrix condition numbers, ignoring the fact that real-world facial representations concentrate on low-dimensional sub-manifolds. Consequently, arbitrary orthogonal bases capture irrelevant directions of low variance.
The authors leverage the empirical distribution of facial embeddings by deriving the orthogonal query set via Principal Component Analysis (PCA). Given a large-scale public face dataset \(X \subset \mathbb{R}^d\) encoded by the surrogate model, PCA computes an optimal row-orthogonal matrix \(W \in \mathbb{R}^{Q \times d}\) that minimizes the average reconstruction error:
Selecting the facial images corresponding to the top \(Q\) principal directions ensures that the observation subspace aligns with the directions of maximum feature variance while preserving strict orthogonality (condition number near 1). Empirical evaluation demonstrates that substituting baseline orthogonal queries with PCA-derived queries shifts the average reconstruction similarity from 0.5289 to 0.6392, allowing over 97% of ideal reconstruction samples to surpass the stringent \(10^{-6}\) FMR threshold.
3. Style-Aware Skip-Connected Inversion: High-Fidelity Feature Reconstruction
The transformation of the estimated template \(\hat{y}\) into a pixel-level face image relies on a black-box template inversion model. Prior architectures like NbNet suffer from structural blurring, loss of high-frequency details, and identity drift, resulting in substantial reconstruction error \(E_3\) when the generated face is re-embedded by the target FRS.
To bridge this gap, the authors design SPNet. Building upon the base architecture of NbNet, SPNet integrates StyleGAN-inspired style mapping by treating the 512-dimensional template as a global style vector and modulating intermediate feature maps through Adaptive Instance Normalization (AdaIN). To preserve high-frequency facial topography, SPNet incorporates DSCasConv-style dense skip connections from shallow spatial layers directly into deep deconvolutional blocks. Furthermore, several micro-architectural refinements are implemented: the final \(\tanh\) activation is removed to prevent dynamic range saturation, pixel-wise normalization (PNorm) is applied across deconvolution layers, and standard ReLUs are replaced with smooth GELU activations. SPNet is trained using a composite loss function:
where the identity loss \(\mathcal{L}_{\text{ID}}\) combines multiple open-source FRS extractors (including the surrogate model) to enforce multi-space identity consistency. These architectural innovations boost inversion precision from 0.8809 (baseline NbNet) to 0.9782, providing the final precision required to conquer high-security thresholds.
Key Experimental Results¶
Main Results¶
The attack is evaluated on three benchmark datasets—LFW, CFP-FP, and AgeDB—against both the commercial AWS Rekognition CompareFace API and representative open-source face recognition models. On AWS Rekognition, the attack is tested under default (80), elevated (90), and public-safety recommended (99) confidence score thresholds. For open-source models, evaluation is conducted at the NIST-recommended extreme threshold of \(\text{FMR} = 10^{-6}\).
| Method / Configuration | Target System | Query Budget | LFW@90 | LFW@99 | CFP-FP@99 | AgeDB@99 |
|---|---|---|---|---|---|---|
| Baseline [36] Reproduced | AWS Rekognition | 100 | 12.54% | 0.60% | 0.57% | 0.37% |
| NbNet + Correction Matrix (C) | AWS Rekognition | 100 | 48.98% | 2.83% | 2.46% | 2.10% |
| NbNet + PCA-OFS (P) | AWS Rekognition | 100 | 97.37% | 46.68% | 51.26% | 55.67% |
| NbNet + C + P | AWS Rekognition | 100 | 98.20% | 56.10% | 60.60% | 63.50% |
| Arc2Face (Diffusion) + C + P | AWS Rekognition | 100 | 97.83% | 53.92% | 53.03% | 44.83% |
| Ours (SPNet + C + P) | AWS Rekognition | 100 | 99.83% | 92.80% | 95.24% | 96.87% |
In black-box evaluations against open-source FRS models at \(\text{FMR} = 10^{-6}\) (using ArcFace ResNet-100 trained on Glint360K as the surrogate model \(F_S\)), the proposed attack demonstrates consistent efficacy across diverse architectures and loss formulations:
| Target Model Identifier | Backbone Architecture | Loss Function & Mechanism | Training Dataset | FMR=\(10^{-6}\) Threshold | Attack ASR (%) |
|---|---|---|---|---|---|
| \(F_S\) (White-box Baseline) | ResNet-100 | ArcFace | Glint360K | 0.5422 | 76.83% |
| \(F_1\) (Cross-architecture & Loss) | ViT-Base | AdaFace + KPRPE | WebFace12M | 0.5195 | 32.63% |
| \(F_2\) (Topology Alignment) | ResNet-100 | CosFace + TopoFR | MS1MV2 | 0.5216 | 89.27% |
| \(F_3\) (Hyperspherical Revived) | ResNet-100 | SphereFace-R | MS1MV2 | 0.5537 | 86.40% |
Ablation Study¶
Ablation experiments analyze the individual contributions of architectural modifications in SPNet toward inversion precision and downstream ASR on AWS Rekognition.
| Model Variant | Core Modification | Inversion Precision | AWS@99 ASR (%) | Note |
|---|---|---|---|---|
| NbNet [47] | Style mapping + AdaIN baseline | 0.8809 | 56.10% | Starting baseline |
| + Micro-refinements | Remove tanh, add PNorm & GELU | 0.9687 | \(91.41 \pm 1.75\)% | Eliminates blur and saturation |
| + DSCasConv [58] (Full SPNet) | Dense cross-layer skip connections | 0.9782 | \(91.88 \pm 1.60\)% | Preserves fine spatial topology |
Key Findings¶
- High-security threshold defense is fundamentally bypassed: Prior score-based methods suffered a catastrophic drop at the AWS 99 threshold (yielding only 0.60% ASR). Integrating the PCA-based OFS and correction matrix elevates ASR to 56.10%, and deploying SPNet drives performance to 92.80%, demonstrating that elevated thresholds alone do not provide adversarial security.
- Fast query budget saturation: Under moderate decision thresholds (80 and 90), ASR reaches near-saturation with 40-50 queries. Under the stringent threshold of 99, ASR exhibits a sharp transition around 50 queries and stabilizes above 90% at 100 queries, remaining well within practical API rate limits.
- Architectural divergence as an empirical barrier: The Vision Transformer target (\(F_1\)) exhibits greater empirical robustness (32.63% ASR) against the CNN-based surrogate model than fellow CNN architectures (\(F_2\) and \(F_3\), exceeding 86% ASR). Furthermore, enforcing uniform hyperspherical feature distribution (UniformFace) disperses PCA variance, reducing ASR by over 43% and pointing toward promising defense paradigms.
Highlights & Insights¶
- Formulating manifold mismatch as a weighted metric projection: Rather than relying on unrealistic assumptions of universal linear basis completeness, the paper models cross-network embedding discrepancies through a positive-definite metric tensor, solving the estimation as a generalized projection without requiring extra query steps.
- Variance-guided observation subspace design: The work identifies the inefficiency of arbitrary orthogonal bases in low-query regimes and demonstrates that aligning observation directions with the principal components of the natural face distribution drastically diminishes subspace truncation error.
- Disparity between human perception and deep embedding metrics: Reconstructed faces frequently appear visually dissimilar to human observers while yielding over 99% confidence match scores in deep feature extractors, underscoring the discrepancy between perceptual identity and high-dimensional metric boundaries.
Limitations & Future Work¶
- Physical presentation attacks and liveness detection: The current evaluation is restricted to the digital domain with direct image submission to APIs. Real-world execution requires physical presentation (e.g., printed artifacts or 3D masks), which faces presentation attack detection (PAD) and liveness verification countermeasures.
- Output image resolution: SPNet currently synthesizes \(128 \times 128\) pixel images. While these images consistently pass commercial face detection and match verification, incorporating state-of-the-art high-resolution generative models (such as 512-resolution diffusion backbones) while maintaining feature alignment represents a valuable avenue for refinement.
- Surrogate-target alignment under domain shift: When target systems utilize proprietary, long-tailed training distributions or unknown loss formulations, the transferability of the surrogate model can degrade, warranting further theoretical exploration of cross-architecture bounds.
Related Work & Insights¶
- vs. Iterative Optimization Attacks (Razzhigaev et al., Park et al.): Prior approaches rely on genetic algorithms or zero-order gradient estimation over 4,000 to 300,000 queries, leaving identifiable query traces that trigger API defenses. In contrast, this work uses at most 100 non-adaptive, deterministic queries.
- vs. Low-Query Geometric Impersonation (Kim et al., S&P 2024): The baseline attack works well under moderate thresholds but collapses to 0.6% under strict thresholds due to unmitigated error accumulation. This paper establishes a comprehensive error bound and introduces three targeted remedies to break the high-confidence barrier.
- vs. White-box Template Inversion (NbNet, Arc2Face): Classical inversion assumes direct access to enrolled deep feature vectors via data breaches. This work seamlessly links black-box score feedback to high-precision inversion, lowering the practical barrier to entry.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First systematic and successful black-box score-based attack against FRSs under extreme security thresholds (FMR = \(10^{-6}\) and AWS confidence 99).
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous evaluation across LFW, CFP-FP, and AgeDB against commercial AWS APIs and diverse open-source targets, supported by sound ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear progression from error quantification to mathematical formulations, intuitive visualizations, and disciplined reporting.
- Value: ⭐⭐⭐⭐⭐ Discloses critical security vulnerabilities in high-assurance biometrics and provides foundational benchmarks for robust FRS design.