Noise-Robust Face Recognition via Non-target Similarity Distribution Guided Sample Selection¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/wfl95/DDLN
Area: Human Understanding
Keywords: Face Recognition, Label Noise Detection, Sample Selection, Robust Learning, Cosine Similarity
TL;DR¶
By identifying that the cosine-similarity distribution of discarded noisy targets matches that of clean non-target pairs, DDLN achieves noise-rate-agnostic sample selection with only 0.2%~0.3% computational overhead while establishing state-of-the-art recognition accuracy under extreme label noise.
Background & Motivation¶
Large-scale face recognition systems increasingly rely on web-crawled datasets comprising millions of facial images. However, open-web data inevitably introduces substantial label noise, which corrupts training dynamics and significantly impairs feature discrimination and generalizationโparticularly under strict low false-acceptance regimes (such as FAR=\(10^{-5}\)). Existing strategies primarily fall into two categories: soft reweighting and multi-center or multi-model cross-validation (e.g., Sub-Center ArcFace and Co-mining), which incur prohibitive memory footprints and computational overhead; or dynamic loss/cosine thresholding (such as OTSU-based RVFace), which relies heavily on the assumption of a bimodal loss or similarity distribution. Unfortunately, this bimodal structure disappears during early training or collapses under severe noise rates, rendering estimated boundaries brittle and unstable.
From a backpropagation perspective, deep networks exhibit a fundamental property: sampleโincorrect-label pairs and sampleโnon-target-class pairs both reflect a mismatch between the sample and an assigned identity. Both operate in the long tail of the softmax distribution, where the gradient magnitude of non-target similarities is proportional to their near-zero posterior probabilities, leading to saturated, weak gradient coupling. Similarly, once suspected noisy pairs are excluded from backpropagation, the gradient-driven attraction toward their incorrect class prototypes vanishes. As a result, clean non-target cosine similarities and unfitted noisy target cosine similarities exhibit an almost identical distribution profile across training steps.
This paper exploits the observable clean non-target similarity distribution as an uncorrupted, data-driven anchor, bypassing the fragile bimodality assumption on target similarities and avoiding prior noise-rate estimation. Core idea: track the statistical upper bound of high-confidence clean non-target cosine similarities to establish a dynamic noise boundary, combined with an early progressive relaxation schedule to discard mislabeled instances while preserving hard but clean samples at negligible computational overhead.
Method¶
Overall Architecture¶
DDLN performs noise detection and sample selection directly within the cosine-similarity space of standard margin-based softmax face recognition objectives. Given a mini-batch of normalized feature embeddings and labels, DDLN assigns a binary selection mask \(q_i \in \{0, 1\}\) to each sample, permitting only verified clean pairs to contribute to loss computation and backpropagation. The selection mask is governed by a dynamic threshold \(T_{n-1}\) updated from the preceding step. Within each mini-batch, the retained clean set is used to estimate the statistical upper and lower bounds of non-target similarities, which are then interpolated via a progressive schedule and smoothed across steps using an exponential moving average (EMA).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Mini-batch Features and Assigned Labels"] --> B["Binary Masking and Noise Discarding<br/>Compare target cosine similarity with threshold Tn-1"]
B --> C["Non-target Statistical Boundary Estimation<br/>Compute retained top-k non-target upper and lower bounds"]
C --> D["Progressive Scheduling and EMA Smoothing<br/>Linear interpolation relaxation + EMA global update"]
D --> E["Forward Loss Computation and Backpropagation Update"]
Key Designs¶
1. Binary Masking and Noise Discarding: Eliminating Corrupted Gradient Updates
To overcome error accumulation inherent in soft confidence reweighting, DDLN enforces strict sample selection. For each sample \(x_i\) and assigned label \(y_i\) in the current mini-batch, its target cosine similarity \(\cos \theta_{i, y_i}\) is evaluated against the global dynamic threshold \(T_{n-1}\). Samples meeting or exceeding the threshold receive \(q_i = 1\) and form the retained clean subset \(\mathcal{Q}_n\), whereas those falling below are masked out (\(q_i = 0\)): $$ \mathcal{L}{\text{DDLN}} = -\sum $$ where } q_i \log \frac{e^{s f(\theta_{i, y_i})}}{e^{s f(\theta_{i, y_i})} + \sum_{j \neq y_i} e^{s f(\theta_{i, j})}\(f(\theta)\) represents the margin penalty function (such as \(\cos \theta - m\) in CosFace or \(\cos(\theta + m)\) in ArcFace). By zeroing out the loss for suspected noisy samples, gradient backpropagation pulling the feature representation toward an incorrect class prototype is immediately severed.
2. Non-target Statistical Boundary Estimation: A Prior-Free Data-Driven Noise Upper Bound
Conventional thresholding strategies attempt to fit a separation boundary directly over the target similarity distribution, which easily collapses when high noise rates dismantle the bimodal structure. DDLN exploits the empirical fact that clean non-target similarities share an almost indistinguishable distribution profile with discarded noisy target similarities (with the distribution range discrepancy bounded around 0.005). Within the retained clean set \(\mathcal{Q}_n\), DDLN computes the \(k\)-th largest non-target similarity \(\cos \theta_{i, (k)}\) per sample (default \(k=10\)), adds a slack margin \(\alpha = 0.02\), and simultaneously records the batch-level minimum non-target similarity: $$ T'{\max} = \frac{1}{|\mathcal{Q}_n|} \sumn} \cos \theta + \alpha, \quad T'{\min} = \frac{1}{|\mathcal{Q}_n|} \sumn} \min $$ Adopting a moderate order statistic } \cos \theta_{i, j\(k=10\) provides a faithful estimate of the non-target upper tail while avoiding susceptibility to extreme single negative outliers.
3. Progressive Scheduling and EMA Smoothing: Protecting Hard Clean Pairs During Early Stages
Deep networks fit clean, generic patterns before memorizing noisy samples. Applying a strict filtering boundary \(T'_{\max}\) at the very beginning of training would prematurely eliminate hard but clean samples whose initial target similarities are modest. DDLN resolves this dilemma through a progressive linear transition from the conservative lower bound \(T'_{\min}\) to the tightened upper bound \(T'_{\max}\): $$ T'n = T' + \beta_n (T'{\max} - T' $$ where }), \quad \beta_n = \frac{\min(n, n_{\max})}{n_{\max}\(n_{\max}\) completes before the first learning-rate decay (epochs 8~10). Early on, the threshold hovers near \(T'_{\min}\), allowing unhindered representation learning on hard-clean pairs. As class representations consolidate, the threshold tightens toward \(T'_{\max}\) to cleanly filter noise. To suppress mini-batch stochastic fluctuations, the local estimate \(T'_n\) is smoothed via an exponential moving average (EMA, \(\lambda=0.01\)): \(T_n = (1 - \lambda) T_{n-1} + \lambda T'_n\). Once noisy pairs drop below this boundary, they lose direct prototype attraction and rarely cross back, yielding stable, monotonic data purification.
Key Experimental Results¶
Main Results¶
Experiments were conducted on simulated synthetic datasets (using cleaned CASIA-WebFace with injected outliers and label flips) and real-world noisy benchmarks (raw CASIA-WebFace, MS-Celeb-1M, and WebFace2M-Noise). The table below reports noise detection and verification results on simulated benchmarks using SEResNet50-IR (BLUFR Avg represents the average across SLLFW and BLUFR TAR at FAR=\(10^{-3}, 10^{-4}, 10^{-5}\)):
| Setting (Outlier-Flip %) | Method | SLLFW (%) | BLUFR Avg (%) | Precision (%) | Recall (%) | F1-score (%) |
|---|---|---|---|---|---|---|
| 10-5-5 (CosFace) | None (Noisy baseline) | 96.86 | 92.61 | โ | โ | โ |
| 10-5-5 (CosFace) | Clean (Oracle) | 97.20 | 93.53 | โ | โ | โ |
| 10-5-5 (CosFace) | OTSU (RVFace) | 96.92 | 93.91 | 90.51 | 96.41 | 93.36 |
| 10-5-5 (CosFace) | DDLN (Ours) | 97.15 | 93.46 | 97.04 | 99.64 | 98.32 |
| 40-20-20 (CosFace) | None (Noisy baseline) | 92.62 | 83.42 | โ | โ | โ |
| 40-20-20 (CosFace) | Clean (Oracle) | 96.43 | 92.33 | โ | โ | โ |
| 40-20-20 (CosFace) | OTSU (RVFace) | 95.70 | 91.44 | 94.74 | 98.58 | 96.62 |
| 40-20-20 (CosFace) | DDLN (Ours) | 96.45 | 93.38 | 99.14 | 99.70 | 99.42 |
| 60-30-30 (CosFace) | None (Noisy baseline) | 83.10 | 58.76 | โ | โ | โ |
| 60-30-30 (CosFace) | Clean (Oracle) | 95.33 | 91.81 | โ | โ | โ |
| 60-30-30 (CosFace) | OTSU (RVFace) | 93.35 | 85.12 | 95.05 | 98.00 | 96.50 |
| 60-30-30 (CosFace) | DDLN (Ours) | 95.15 | 91.12 | 99.17 | 99.14 | 99.16 |
| 80-40-40 (CosFace) | None (Noisy baseline) | 70.40 | 30.45 | โ | โ | โ |
| 80-40-40 (CosFace) | Clean (Oracle) | 92.75 | 86.14 | โ | โ | โ |
| 80-40-40 (CosFace) | OTSU (RVFace) | 83.92 | 60.06 | 92.65 | 98.30 | 95.39 |
| 80-40-40 (CosFace) | DDLN (Ours) | 91.70 | 78.11 | 95.16 | 98.13 | 96.63 |
On the challenging WebFace2M-Noise dataset (sampled from WebFace260M with ~80% severe web noise), evaluated with ResNet-50 across standard benchmarks and the low-quality TinyFace dataset:
| Backbone & Denoising Strategy | LFW (%) | CALFW (%) | CPLFW (%) | AgeDB (%) | CFP (%) | TinyFace R1 (%) | 8-Benchmark Avg (%) |
|---|---|---|---|---|---|---|---|
| CosFace baseline | 98.58 | 90.50 | 82.37 | 85.97 | 86.07 | 50.40 | 77.02 |
| CosFace + Sub-center | 97.95 | 89.63 | 80.50 | 84.96 | 82.89 | 50.99 | 75.86 |
| CosFace + SubRe (Retrain) | 99.05 | 92.52 | 85.07 | 90.43 | 89.21 | 54.29 | 79.57 |
| CosFace + RVFace (OTSU) | 99.46 | 93.83 | 88.32 | 93.90 | 93.49 | 61.72 | 83.52 |
| CosFace + DDLN (Ours) | 99.43 | 94.46 | 89.25 | 94.47 | 94.33 | 63.52 | 84.43 |
| ArcFace baseline | 98.52 | 90.65 | 82.60 | 86.13 | 86.61 | 50.86 | 77.02 |
| ArcFace + RVFace (OTSU) | 99.32 | 93.98 | 88.88 | 93.13 | 93.34 | 61.86 | 83.53 |
| ArcFace + RobustFace | 98.42 | 90.72 | 82.32 | 85.67 | 85.11 | 53.14 | 77.55 |
| ArcFace + DDLN (Ours) | 99.53 | 94.57 | 90.18 | 95.10 | 95.10 | 63.25 | 84.70 |
Ablation Study¶
On the simulated #40-20-20 dataset with ResNet-50 and CosFace, ablating threshold scheduling, rank order \(k\), slack \(\alpha\), and EMA factor \(\lambda\):
| Configuration | Epoch \(n_{\max}\) | Slack \(\alpha\) | EMA \(\lambda\) | Precision (%) | Recall (%) | F1-score (%) | Note |
|---|---|---|---|---|---|---|---|
| Fixed Threshold \(T'_{\max}(k=10)\) | โ | โ | โ | 67.72 | 99.86 | 80.70 | Without scheduling, many hard-clean samples are discarded early |
| Fixed Threshold 0.20 | โ | โ | โ | 99.82 | 93.87 | 96.75 | Heuristic threshold, lower recall |
| Fixed Threshold 0.45 | โ | โ | โ | 96.19 | 99.96 | 98.04 | High recall but lower precision |
| Progressive (\(k=1\)) | 9 | 0.00 | 0.01 | 97.97 | 99.91 | 98.93 | Sensitive to extreme negative outliers |
| Progressive (\(k=10\)) | 9 | 0.00 | 0.01 | 99.38 | 98.90 | 99.14 | Standard order statistic |
| Progressive (\(k=50\)) | 9 | 0.00 | 0.01 | 99.65 | 97.50 | 98.56 | Over-smooths boundary, dropping recall |
| Full Model (Default) | 9 | 0.02 | 0.01 | 99.14 | 99.70 | 99.42 | Incorporating slack margin \(\alpha\) yields best F1 |
| Schedule Variation (\(n_{\max}=8\)) | 8 | 0.02 | 0.01 | 99.05 | 99.72 | 99.38 | Minimal sensitivity to schedule length |
| Schedule Variation (\(n_{\max}=10\)) | 10 | 0.02 | 0.01 | 99.15 | 99.62 | 99.39 | Robust across convergence schedules |
Key Findings¶
- Progressive scheduling is indispensable: Directly applying the upper threshold without progressive scheduling drops precision to 67.72% (F1 falls from 99.42% to 80.70%). Early hard-clean samples exhibit low initial similarities, and a premature hard cutoff permanently removes valuable informative pairs.
- Superiority under extreme noise: Under 80% synthetic noise, the collapse of bimodal target distributions forces OTSU to discard 31.19% of clean samples (clean FNR), delivering only 60.06% recognition accuracy. DDLN reduces clean sample loss to 19.96% and improves recognition accuracy to 78.11% (+18.05% gain).
- Negligible computational overhead: Because non-target similarities are inherently evaluated in the forward pass of standard softmax-based losses, DDLN adds only in-batch sorting and EMA tracking, incurring merely ~1.07 ms extra latency per batch (~0.2%~0.3% overhead), compared to multi-network or sub-center alternatives that double resource consumption.
Highlights & Insights¶
- Root-cause insight via backpropagation analysis: Instead of imposing heuristic parametric assumptions on loss or target similarity distributions, DDLN grounds boundary estimation on the mathematical isomorphism between unfitted label noise and negative pairs in saturated gradient regimes.
- Tuning-free universal deployment: Operates without prior noise rates or validation sets. The single set of default hyperparameters (\(k=10, \alpha=0.02, \lambda=0.01\)) generalizes seamlessly from 10% to 80% noise rates across millions of facial images.
- Direct transferability to open-set metric learning: The principle of anchoring target boundaries using retained non-target statistics is broadly applicable to person re-identification (ReID), vehicle retrieval, and fine-grained classification models utilizing margin softmax losses.
Limitations & Future Work¶
- Boundary jitter under extreme long-tail sparsity: Identities with minimal sample counts combined with severe noise may fail to form stable non-target distributions, introducing residual boundary noise.
- Irreversible sample exclusion: Once filtered out after threshold convergence, samples are excluded from subsequent backpropagation entirely, lacking an adaptive mechanism to recover potential pseudo-noise as representations mature.
- Future directions: Integrating momentum prototype banks or semi-supervised consistency regularization to reuse excluded high-confidence outliers as unlabeled data for self-supervised contrastive pretraining.
Related Work & Insights¶
- vs RVFace (OTSU): RVFace relies on a bimodal target similarity assumption that dissolves under high noise; DDLN anchors to non-target statistics, outperforming RVFace by over 18% under 80% noise.
- vs Sub-Center ArcFace: Sub-Center ArcFace incurs substantial parameter and memory overhead with multi-stage retraining; DDLN accomplishes single-stage end-to-end cleaning with an average 5% performance advantage.
- vs Co-mining / Co-teaching: Co-mining requires dual synchronized networks, doubling computation and GPU memory; DDLN operates on a single network with ~0.2% overhead and achieves superior robustness.
Rating¶
- Novelty: โญโญโญโญโญ Elegant discovery of distributional alignment between clean non-target pairs and unfitted noisy targets, circumventing target bimodality assumptions.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive validation spanning 8 synthetic noise levels and 3 multi-million web benchmarks, complete with fine-grained FNR/FPR metrics.
- Writing Quality: โญโญโญโญโญ Rigorous backpropagation gradient derivation, intuitive motivation, and high-quality visualizations.
- Value: โญโญโญโญโญ Plug-and-play formulation with minimal computational overhead, highly practical for industrial-scale face recognition pipelines.