Skip to content

Improving Adversarial Robustness by Mitigating Instability through Relearning

Conference: ECCV 2026
Paper: ECCV Official
Area: AI Safety
Keywords: Adversarial Training, Robust Overfitting, Robust Instability, Relearning, Decision Boundary Distortion

TL;DR

Addressing severe robust overfitting where adversarial training degrades sharply after learning rate decay, this paper reveals the counterintuitive phenomenon of robust instability across adjacent training epochs, formalizes it via the Robust Stability Rate (RSR), and proposes Relearning Adversarial Training (RAT) to establish a new min-max dynamic with temporal ensembling soft labels that eliminates overfitting without extra models.

Background & Motivation

Adversarial training is the most effective approach to defend against adversarial attacks by casting empirical risk minimization into a min-max optimization problem. However, adversarial training routinely suffers from severe robust overfitting: shortly after the first learning rate decay, the model robust accuracy on test data deteriorates continuously as training epochs increase. Conventional regularization techniques such as weight decay and data augmentation fail to halt this decline and often underperform early stopping, while generative methods that synthesize massive external data incur prohibitive computational costs and sample complexity.

Investigating the micro-level dynamics across training epochs, this paper identifies a counterintuitive phenomenon: after an entire training epoch, the model actually performs worse on the exact same adversarial examples than it did in the previous epoch. The authors formalize this optimization challenge as Robust Instability and quantify it via the Robust Stability Rate (RSR). Theoretical derivations and empirical evidence demonstrate that standard adversarial training causes a harmful distortion of the decision boundary within the perturbation ball—improving robustness along the previous epoch's perturbation direction while degrading it along the current one. The persistent accumulation of this instability is shown to be a primary driving mechanism behind robust overfitting.

Core Idea: Reformulate adversarial training from memory-based fitting into a smooth trajectory by proposing Relearning Adversarial Training (RAT), which establishes a new min-max game where adversarial perturbations actively induce instability while model optimization enforces robust stability guided by historical momentum soft labels.

Method

Overall Architecture

Relearning Adversarial Training (RAT) reconstructs the standard min-max optimization paradigm by introducing a sample-wise relearning label \(p_i\) that aggregates temporal historical predictions. Instead of forcing the model to strictly fit hard ground-truth labels on adversarial points, RAT establishes a cooperative-competitive dynamic: the attack phase generates adversarial samples that maximize both classification loss and instability divergence; the model training phase optimizes classification while penalizing divergence from historical soft knowledge; and the epoch conclusion phase recursively updates the relearning labels via momentum temporal ensembling.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Pair (x, y) & Current Weights θ"] --> B["Stage 1: Instability-Inducing Adversarial Generation<br/>Joint gradient ascent maximizing L_ce and L_mse"]
    B --> C["Stage 2: Relearning Model Weight Optimization<br/>Minimize joint relearning loss L_re to smooth boundary"]
    C --> D["Stage 3: Momentum Temporal Ensembling Update<br/>Aggregate clean and adversarial predictions into p"]
    D --> E["Output Robust Model & Smooth History"]

Key Designs

1. Instability-Inducing Adversarial Generation: Injecting instability destruction into the attack phase Standard adversarial example generation relies solely on maximizing cross-entropy loss, ignoring the turbulent boundary shifts across training epochs. To expose the most vulnerable directions under the min-max game, the inner maximization is augmented with a mean squared error (MSE) penalty that measures deviation from the historical relearning labels. At the \((k+1)\)-th projected gradient step, the perturbation updates as:

\[ \widehat{x}_i^{k+1} = x_i + \text{clip}\left( \widehat{x}_i^k + \alpha \cdot \text{sign}\left( \nabla_{\widehat{x}_i^k} \left\{ \mathcal{L}_{ce}(f_\theta(\widehat{x}_i^k), y_i) + \omega \cdot \mathcal{L}_{mse}(f_\theta(\widehat{x}_i^k), p_i) \right\} \right) - x_i, -\epsilon, \epsilon \right) \]

This formulation forces adversarial examples to actively seek perturbations that both induce misclassification and shatter the prediction stability accumulated from past models, presenting a more rigorous defense target for the outer training step.

2. Relearning Model Weight Optimization: Suppressing destructive decision boundary reconstruction To resolve the conflicting boundary movements where loss decreases on \(\widehat{x}_{i-1}\) but increases on \(\widehat{x}_i\), the outer model training replaces pure hard-label supervision with the composite relearning loss \(\mathcal{L}_{re}\). It combines cross-entropy on ground-truth targets with an MSE regularizer aligning predictions with the historical relearning label \(p_i\):

\[ \mathcal{L}_{re}(\theta; \widehat{x}_i; y_i; p_i; \omega) = \mathcal{L}_{ce}(f_\theta(\widehat{x}_i), y_i) + \omega \cdot \mathcal{L}_{mse}(f_\theta(\widehat{x}_i), p_i) \]

By penalizing erratic shifts in the model output distribution, this quadratic constraint implicitly restrains drastic parameter updates across training epochs, preventing catastrophic forgetting of previously acquired robust features.

3. Momentum Temporal Ensembling Update: Zero-overhead multi-epoch historical aggregation Relying on a single preceding checkpoint risks propagating poor supervisory signals if that checkpoint experiences an optimization dip; however, keeping an explicit ensemble of historical models in GPU memory is computationally infeasible. RAT resolves this by maintaining a sample-level relearning label \(p_i\) via temporal ensembling. In each epoch, an auxiliary term combining clean and adversarial predictions \(q_i = \gamma \cdot f_\theta(x_i) + (1-\gamma) \cdot f_\theta(\widehat{x}_i)\) is calculated, followed by an exponential moving average update with momentum \(\eta\):

\[ p_i \leftarrow \eta \cdot p_i + (1 - \eta) \cdot q_i \]

Unrolling this recursion demonstrates that \(p_i\) acts as an exponential decay ensemble over all preceding historical epochs initialized at the one-hot ground-truth label. This provides a remarkably stable and smooth teaching signal at negligible computational and memory cost, ensuring continuous decision boundary evolution rather than disruptive shifts.

Loss & Training

The overall training operates in two stages: prior to the first learning rate decay, standard adversarial training is executed so the model learns fundamental representations and baseline robustness; shortly before the first decay, the Relearning Adversarial Training framework is initiated under the unified min-max formulation:

\[ \min_\theta \sum_i \max_{\widehat{x}_i \in \mathcal{B}_\epsilon(x_i)} \left( \mathcal{L}_{ce}(f_\theta(\widehat{x}_i), y_i) + \omega \cdot \mathcal{L}_{mse}(f_\theta(\widehat{x}_i), p_i) \right) \]

Models are trained using SGD with momentum 0.9, weight decay \(5 \times 10^{-4}\), initial learning rate 0.1, decayed by a factor of 10 at epochs 100 and 150 over 200 total epochs. Default hyperparameters are set to momentum \(\eta = 0.9\), balance factor \(\gamma = 0.5\), and relearning weight \(\omega = 30\).

Key Experimental Results

Main Results

On CIFAR-10 and CIFAR-100 under the standard \(L_\infty\) perturbation threat model (\(\epsilon = 8/255\)) using ResNet-18, the proposed relearning strategy (+RE) is integrated into four baseline adversarial training methods: PGD-AT, TRADES, MART, and AT-AWP. Evaluation metrics include clean Natural accuracy, PGD-20, Square Attack, CW, Auto Attack (AA), and RSR (where values closer to 0 indicate superior epoch-to-epoch stability).

Dataset Method Natural (%) PGD-20 (%) Square (%) CW (%) Auto Attack (%) RSR (%)
CIFAR-10 PGD-AT 82.40 41.40 59.97 40.98 40.74 -88.61
CIFAR-10 PGD-AT+RE 82.50 55.40 64.29 52.75 50.83 -22.80
CIFAR-10 TRADES 82.42 50.10 62.37 48.88 47.29 -100.90
CIFAR-10 TRADES+RE 82.58 54.66 63.61 51.29 49.76 -32.09
CIFAR-10 MART 82.18 48.45 60.13 45.81 43.61 -104.30
CIFAR-10 MART+RE 82.30 54.66 62.10 50.75 49.30 -29.42
CIFAR-10 AT-AWP 81.15 55.03 63.93 51.67 49.45 -45.36
CIFAR-10 AT-AWP+RE 81.18 57.23 63.19 52.74 51.23 -31.75
CIFAR-100 PGD-AT 58.08 20.81 31.14 21.15 19.69 -44.17
CIFAR-100 PGD-AT+RE 58.36 31.39 35.45 27.99 26.00 -20.66
CIFAR-100 TRADES 56.35 27.19 33.02 24.61 23.73 -56.95
CIFAR-100 TRADES+RE 58.35 30.92 35.10 26.68 25.42 -44.53
CIFAR-100 MART 55.19 24.13 30.93 22.37 21.00 -56.33
CIFAR-100 MART+RE 55.65 31.05 34.12 26.76 25.65 -29.37
CIFAR-100 AT-AWP 53.89 30.21 34.80 27.25 25.30 -15.79
CIFAR-100 AT-AWP+RE 54.88 32.11 34.61 27.75 25.87 -12.40

Ablation Study

On CIFAR-10 using ResNet-18, the impact of incorporating the relearning loss \(\mathcal{L}_{re}\) across the two distinct optimization phases (Generation vs Training) is examined under PGD-20 evaluation (reporting Best accuracy, Final accuracy, and the robust degradation Diff):

Generation Loss Training Loss Best (%) Final (%) Degradation Diff (%) Note
\(\mathcal{L}_{ce}\) \(\mathcal{L}_{ce}\) 50.70 41.40 -9.30 Standard PGD-AT baseline with severe overfitting
\(\mathcal{L}_{ce}\) \(\mathcal{L}_{re}\) 54.49 51.58 -2.91 Relearning applied only during weight updates
\(\mathcal{L}_{re}\) \(\mathcal{L}_{ce}\) 53.70 49.08 -4.62 Instability induction applied only during attack
\(\mathcal{L}_{re}\) \(\mathcal{L}_{re}\) 56.02 55.40 -0.62 Full RAT framework, robust overfitting virtually eliminated

Key Findings

  • Synergistic min-max dynamics are essential: Applying \(\mathcal{L}_{re}\) to only generation or only training yields robust degradations of -4.62% and -2.91%, respectively. Operating both phases jointly narrows the drop to a mere -0.62% while boosting final robust accuracy to 55.40% (a +14.00% gain over the baseline's 41.40%), confirming the vital complementarity between destabilizing attacks and stabilizing defenses.
  • RSR directly mirrors robust overfitting: Baseline PGD-AT experiences severe negative stability on CIFAR-10 with an RSR of -88.61%. With relearning, RSR improves to -22.80%. Similarly, TRADES+RE and MART+RE cut negative RSR magnitudes by more than half, validating that mitigating step-wise decision instability arrests cumulative overfitting.
  • Plug-and-play generality across models and baselines: RAT consistently boosts robust accuracy and maintains natural accuracy across architectures (ResNet-18, WideResNet-34-10, ViT). Combined with AT-AWP on WideResNet-34-10, it attains 61.21% final robust accuracy with a negligible -0.11% drop.

Highlights & Insights

  • Microscopic stability lens on macroscopic overfitting: By identifying that adjacent checkpoints perform worse on the exact same adversarial examples, the paper introduces the concept of Robust Instability and the RSR metric, shedding new light on the long-standing mystery of robust overfitting.
  • Teacher-free historical ensembling: Rather than relying on heavyweight pre-trained teachers or offline distillation, RAT utilizes the online trajectory via recursive exponential moving averages, securing a high-quality soft supervisory signal with zero extra model weights and flattening the loss landscape.
  • Dual-phase instability dynamic: Traditional adversarial attacks only maximize misclassification error; RAT turns perturbation generation into an active search for stability breakdown, yielding richer adversarial samples for robust boundary formation.

Limitations & Future Work

  • Heuristic phase activation: The method currently activates shortly prior to the scheduled learning rate drop; designing an automated trigger based on real-time RSR variance could remove this manual scheduling.
  • Extension to large vision-language models: While validated on ResNets, WideResNets, and ViT, evaluating whether instability-driven relearning scales to parameter-efficient fine-tuning (PEFT) on large multimodal foundation models remains an open question.
  • Hyperparameter co-dependence: The balance factor \(\gamma\) and loss weight \(\omega\) require mild calibration across different dataset scales, leaving room for adaptive loss weighting schemes.
  • vs PGD-AT (Madry et al., 2017): PGD-AT relies exclusively on hard targets in static min-max updates, leading to sharp boundary oscillations and overfitting. RAT stabilizes boundary movements via historical soft-label relearning.
  • vs FOMO (Ramkumar et al., 2024) / KD-SWA (Chen et al., 2020): These methods rely on stochastic forgetting or post-hoc weight averaging, which are cumbersome and slower to adapt to dynamic perturbations. RAT updates smooth guidance online at the sample level.
  • vs AT-AWP (Wu et al., 2020): AWP explicitly perturbs weights to flatten the loss landscape. RAT acts on prediction distributions, functioning orthogonally and synergistically with AWP to reach 51.23% Auto Attack accuracy on CIFAR-10.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates the microscopic concept of robust instability and constructs a novel instability-stability min-max dynamic.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Thorough evaluation across benchmarks, architectures, multiple attack protocols, loss landscape visualizations, and ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Cohesive theoretical analysis, intuitive mathematical formulations, and clear experimental presentation.
  • Value: ⭐⭐⭐⭐⭐ Plug-and-play enhancement that decisively mitigates robust overfitting without external models or extra data.