Robustness Meets Uncertainty: Evidential Adversarial Training for Robust Selective Classification¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/NicolasSournac/Robustness_Meets_Uncertainty.EV-AT
Area: AI Safety
Keywords: Adversarial Robustness, Uncertainty Estimation, Selective Classification, Evidential Deep Learning, Adversarial Training
TL;DR¶
To address the pervasive issue where standard adversarial training degrades uncertainty ranking and induces confident misclassifications, this paper proposes Evidential Adversarial Training (EV-AT), which constrains Dirichlet predictive representations in log-concentration space to shift the Pareto frontier of robustness and selective risk.
Background & Motivation¶
In safety-critical applications such as autonomous driving and medical diagnosis, dependable deep classifiers must not only resist malicious input perturbations, but also accurately gauge their own predictive confidence. Selective classification offers an essential safety mechanism by enabling classifiers to abstain when confidence is low, deferring uncertain cases to human operators or downstream fallback protocols. For such selective systems to operate reliably under adversarial attacks, input perturbations that cause misclassifications must be assigned high uncertainty so they can be filtered out, rather than slipping through as confident errors.
However, existing adversarial robustness evaluations predominantly focus on robust accuracy, overlooking the fragile integrity of predictive uncertainty under perturbation. Standard state-of-the-art adversarial training defenses frequently disrupt uncertainty ranking while defending against label flips. Adversarial samples can easily trick the network into making incorrect predictions with inflated confidence, bypassing rejection thresholds and manifesting as hazardous silent failures at practical operating points. Furthermore, the absence of a standardized benchmark across architectures, augmentations, and threat models has historically hindered systematic exploration of the trade-off between robustness and uncertainty.
To reconcile the core tension between increasing adversarial accuracy and maintaining faithful risk-coverage dynamics, this paper departs from conventional logit-space consistency regularization by leveraging evidential deep learning to untangle class predictions from evidential strength. Core Idea: Parameterize predictions as a Dirichlet posterior distribution and enforce robust evidence alignment directly in log-concentration space via an evidence-targeted adversary, thereby stabilizing both class probabilities and epistemic confidence to achieve robust selective classification.
Method¶
Overall Architecture¶
EV-AT maps network outputs to Dirichlet distribution concentration parameters, allowing simultaneous estimation of predictive class probabilities and total evidence strength (serving as an epistemic uncertainty proxy). During training, clean samples are supervised with an evidential marginal likelihood loss, while an evidence-targeted adversary generates worst-case perturbations in log-concentration space via multi-step PGD. The framework minimizes both clean evidence loss and robust evidence alignment discrepancies between clean and perturbed representations, augmented with adversarial weight perturbation (AWP).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Sample x"] --> B["Dirichlet Parameterization<br/>Evidence Vector and Concentration α"]
B --> C["Evidential Learning<br/>Marginal Likelihood and Prior Regularization"]
B --> D["Evidence-Targeted Adversary<br/>PGD Perturbation in Log Space x_adv"]
D --> E["Robust Evidence Alignment<br/>Log-Space IKL Divergence Consistency"]
C --> F["Joint Objective Optimization<br/>Clean Loss + Alignment Penalty + AWP"]
E --> F
F --> G["Entropy-Based Selective Classification<br/>Abstain on High-Uncertainty Inputs"]
Key Designs¶
1. Dirichlet Parameterization: Decoupling Categorical Probabilities from Evidential Strength
Standard neural networks rely on softmax normalization over unconstrained logits, which cannot distinguish whether a prediction is backed by rich evidence or results from a vacuous lack of discriminatory features. EV-AT employs evidential deep learning by having the backbone predict a non-negative evidence vector \(e_\theta(x)\), which defines Dirichlet concentration parameters \(\alpha = e_\theta(x) + 1\). The resulting predictive mean \(\bar{\pi} = \alpha / S\) (where total concentration \(S = \sum_{c=1}^C \alpha_c\)) acts as the categorical classification vector, while total strength \(S\) inversely quantifies epistemic uncertainty. At inference time, the selective classifier evaluates predictive entropy \(u(x) = H[\text{Cat}(\bar{\pi})]\) as the selection score, accepting predictions only when \(u(x) \le \tau\).
2. Log-Concentration Space Mapping: Multiplicative Evidence Stability and Variance Bounding
Directly penalizing deviations in the raw concentration space \(\alpha\) or predictive mean space \(\bar{\pi}\) leads to asymmetric gradients and numerical instability across varying evidence magnitudes. EV-AT maps concentrations into log-concentration space \(\eta = \log \alpha\) (acting as a pseudo-logit representation), where additive perturbations in \(\eta\) correspond directly to multiplicative scalings in \(\alpha\). Theoretical analysis shows that when the log-concentration drift is bounded by \(\|\eta' - \eta\|_\infty \le \rho\), the total strength ratio is bounded by \(e^{-\rho} \le S'/S \le e^\rho\), and the predictive mean ratio satisfies \(e^{-2\rho} \le \bar{\pi}'_c / \bar{\pi}_c \le e^{2\rho}\). Furthermore, since the Dirichlet posterior variance is strictly bounded by \(\text{Var}[\pi_c] \le \frac{1}{4(S+1)}\), regularizing \(\eta\) simultaneously bounds predictive mean drift and posterior variance dilation, theoretically guaranteeing entropy continuity and decision stability under attack.
3. Evidence-Targeted Adversary: Stress-Testing Posterior Evidential Integrity
Standard adversarial attacks construct adversarial perturbations by maximizing label cross-entropy loss, which merely pushes representations across the decision boundary and often manufactures overconfident misclassifications. In contrast, EV-AT constructs perturbations that explicitly target evidential structures. Within an \(\ell_p\)-norm perturbation ball, the adversary uses multi-step PGD to maximize the log-parameter discrepancy \(D(\eta, \eta_{\text{adv}})\). This adversary directly destabilizes the Dirichlet posterior, actively probing for worst-case evidence degradation and unmasking overconfident blind spots.
4. Robust Evidence Alignment: Posterior Consistency via Improved KL Divergence
In the outer minimization loop, EV-AT balances clean data fitting and adversarial posterior invariance. The clean branch optimizes the negative log marginal likelihood combined with a KL penalty toward a flat Dirichlet prior on non-target classes to curb spurious evidence. The adversarial branch minimizes the robust evidence alignment loss \(\mathcal{L}_{\text{REA}} = D(\eta, \eta_{\text{adv}})\) to penalize attack-induced posterior drift. EV-AT instantiates the discrepancy metric \(D\) using the Decoupled Improved Kullback-Leibler (IKL) divergence, ensuring well-behaved and stable gradients for both adversarial generation and alignment optimization.
Loss & Training¶
The overall training objective combines the clean evidential loss \(\mathcal{L}_{\text{EV}}\) with the weighted robust evidence alignment loss \(\mathcal{L}_{\text{REA}}\):
where \(\mathcal{L}_{\text{EV}}\) incorporates Type-II maximum likelihood marginal integration alongside uniform prior regularization, and \(\beta\) balances clean accuracy against adversarial alignment (empirically optimal around 60~100). Additionally, Adversarial Weight Perturbation (AWP) is incorporated during optimization to flatten the loss landscape with respect to network weights, suppressing robust overfitting and improving selective generalization.
Key Experimental Results¶
Main Results¶
The authors establish a standardized benchmark across Clean, AutoAttack (AA, \(\ell_\infty, \varepsilon=8/255\)), and common corruptions (CIFAR-10-C). On WideResNet-34-10, EV-AT is compared against top adversarial defenses across multiple data augmentations using classification accuracy and the Area Under the Generalized Risk-Coverage curve (AUGRC, lower is better):
| Augmentation | Method | Venue | Acc.: Clean (%) | Acc.: AA (%) | Acc.: Clean/AA (%) | AUGRC: Clean (↓) | AUGRC: AA (↓) | AUGRC: Clean/AA (↓) |
|---|---|---|---|---|---|---|---|---|
| Basic | AT | ICML'18 | 85.11 | 51.63 | 68.37 | 2.79 | 12.32 | 7.56 |
| Basic | AT-AWP | NeurIPS'20 | 85.28 | 53.57 | 69.42 | 2.96 | 11.64 | 7.30 |
| Basic | TRADES | ICML'19 | 84.53 | 52.92 | 68.73 | 3.83 | 12.82 | 8.32 |
| Basic | TRADES-AWP | NeurIPS'20 | 84.89 | 55.85 | 70.37 | 3.58 | 11.26 | 7.42 |
| Basic | IKL-AT | NeurIPS'24 | 85.04 | 56.18 | 70.61 | 3.48 | 11.27 | 7.38 |
| Basic | TRADES-EMFF | TPAMI'25 | 84.97 | 50.35 | 67.66 | 3.67 | 14.03 | 8.85 |
| Basic | EV-AT (Ours) | - | 88.16 | 55.38 | 71.77 | 2.20 | 10.70 | 6.45 |
| AugMix | AT | ICML'18 | 83.38 | 52.17 | 67.78 | 3.51 | 12.35 | 7.93 |
| AugMix | AT-AWP | NeurIPS'20 | 81.41 | 52.76 | 67.09 | 4.37 | 12.45 | 8.41 |
| AugMix | TRADES | ICML'19 | 84.06 | 47.71 | 65.88 | 4.06 | 15.64 | 9.85 |
| AugMix | TRADES-AWP | NeurIPS'20 | 84.64 | 52.37 | 68.50 | 3.90 | 13.00 | 8.45 |
| AugMix | IKL-AT | NeurIPS'24 | 82.85 | 55.38 | 69.11 | 4.65 | 12.27 | 8.46 |
| AugMix | TRADES-EMFF | TPAMI'25 | 84.23 | 49.90 | 67.07 | 4.25 | 14.52 | 9.38 |
| AugMix | EV-AT (Ours) | - | 87.10 | 55.93 | 71.52 | 2.78 | 10.73 | 6.75 |
Ablation Study¶
Ablations on CIFAR-10 / WRN-34-10 (AugMix) investigate objective components and alignment design choices under AutoAttack (\(\ell_\infty, \varepsilon=8/255\)):
| Variant | Investigated Aspect | Acc.: Clean/AA (%) | AUGRC: Clean/AA (↓) | Note |
|---|---|---|---|---|
| EV-AT w/ CE | Clean Loss Replacement | 69.56 | 7.88 | Replacing evidential clean loss with cross-entropy impairs uncertainty ranking |
| EV-AT w/o REA (\(\beta=0\)) | Removal of Alignment | 20.11 | 37.67 | Without REA, evidential posterior collapses under adversarial perturbation |
| EV-AT w/o AWP | Removal of Weight Perturbation | 68.46 | 8.21 | Weight smoothing absence degrades both robust generalization and selective risk |
| EV-AT (Full model) | Complete Objective | 71.51 | 6.67 | Achieves superior joint robustness and uncertainty quality |
| Alignment Space: \(\alpha\) | Space Selection | 67.32 | 10.26 | Raw concentration space is scale-sensitive with unstable optimization |
| Alignment Space: \(\bar{\pi}\) | Space Selection | 64.64 | 10.43 | Aligning only categorical mean ignores the total evidence strength \(S\) |
| Alignment Space: \(\eta = \log \alpha\) | Space Selection | 71.51 | 6.67 | Multiplicatively scales evidence and controls both mean and precision |
| Divergence: \(L_2\) | Divergence Metric | 69.81 | 7.64 | Euclidean metric fails to reflect distributional geometry |
| Divergence: KL | Divergence Metric | 70.61 | 7.60 | Standard KL is slightly less stable than decoupled formulation |
| Divergence: IKL | Divergence Metric | 71.51 | 6.67 | Provides optimal gradient signal and balanced trade-off |
Key Findings¶
- State-of-the-art defenses fall into uncertainty ranking traps: Top-performing baselines such as IKL-AT and TRADES-AWP achieve high robust accuracy under AutoAttack, but exhibit disproportionately elevated AUGRC (e.g., IKL-AT has an adversarial AUGRC of 11.27 under Basic augmentation). This reveals that narrowing the classification margin without posterior constraints produces severely overconfident misclassifications.
- Log-concentration space is critical for dual stabilization: Shifting the REA constraint from \(\eta = \log \alpha\) to mean space \(\bar{\pi}\) or raw concentration space \(\alpha\) inflates AUGRC above 10. Log-parameter alignment provides the necessary mathematical coupling to constrain both predictive direction and posterior concentration.
- Coverage under strict risk constraints improves dramatically: At a stringent 5% adversarial risk budget (Coverage @ 5% Risk), TRADES-AWP achieves 26% coverage and IKL-AT achieves 23%, whereas EV-AT significantly expands retained sample volume, effectively mitigating silent failures at practical operational thresholds.
Highlights & Insights¶
- Decoupled view of selective classification and adversarial defenses: Points out the fundamental pitfall of relying exclusively on robust accuracy, uncovering the real-world safety hazard where defenses improve adversarial accuracy while degrading uncertainty-based rejection.
- Theoretical guarantees via log-Dirichlet pseudo-logits: Formalizes how bounding log-concentration perturbations yields rigorous multiplicative bounds on total evidence strength \(S\) and predictive mean \(\bar{\pi}\) (Lemma 1~3), providing end-to-end continuity guarantees for entropy-based selective rejection.
- Consistent multi-regime Pareto frontier expansion: Delivers robust improvements across multiple data augmentations, model capacities (WRN-34-10 and PreActResNet-18), attack budgets (\(\varepsilon \in [0, 8/255]\)), and corruptions (CIFAR-C), consistently setting a new state-of-the-art Pareto frontier.
Limitations & Future Work¶
- Training computational overhead: While inference latency remains identical to standard classifiers, generating evidence-targeted PGD adversaries requires taking higher-order gradients through log-Dirichlet parameters, which adds memory and compute overhead during training.
- Scope confined to single-label classification: Experiments focus on CIFAR-10/100 image classification benchmarks. Extending evidential adversarial training to complex vision tasks such as dense semantic segmentation, 3D perception, or open-vocabulary detection remains to be validated.
- Sensitivity to alignment weight \(\beta\): The alignment coefficient \(\beta\) exhibits a clear optimal band (around 60~100); suboptimal tuning can either lead to insufficient evidential regularization or over-regularization that harms clean accuracy.
Related Work & Insights¶
- vs TRADES / TRADES-AWP: TRADES balances clean accuracy and adversarial prediction consistency in logit space but treats output confidence implicitly; EV-AT explicitly models posterior uncertainty via a Dirichlet distribution, regularizing both classification direction and confidence magnitude.
- vs IKL-AT: IKL-AT introduces decoupled KL divergence for traditional logits; EV-AT adapts this divergence to log-Dirichlet parameter space to govern epistemic evidence drift.
- vs Evidential Deep Learning (EDL): Vanilla EDL methods yield well-ranked uncertainty on clean distributions but collapse under adversarial attack; EV-AT fills this critical gap by introducing an evidence-targeted adversary and robust evidence alignment.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Unifies evidential deep learning with adversarial training in log-concentration space, offering a fresh angle on selective classification safety.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation covering 4 data augmentations, 2 network architectures, multiple threat models, and thorough ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Exemplary mathematical rigor and structural clarity, bridging motivation, theoretical lemmas, and empirical findings seamlessly.
- Value: ⭐⭐⭐⭐☆ Highly impactful for deploying reliable vision models in high-consequence domains where both attack resistance and self-awareness of failure are mandatory.