Skip to content

Indelible Backdoors: On the Limits of Post-Training Defenses

Conference: ECCV 2026
Paper: ECCV 2026
Area: AI Safety
Keywords: Backdoor Attack / Gradient Alignment / Post-Training Defense / Fisher Information / Continual Learning Security

TL;DR

By identifying task-critical neurons via Fisher Information and explicitly maximizing the gradient cosine similarity between backdoor and clean samples on these parameters, this paper presents an indelible backdoor attack that forces fine-tuning and pruning defenses to destroy clean accuracy whenever they attempt to suppress backdoor effects.

Background & Motivation

With deep neural networks widely deployed across mission-critical domains, backdoor attacks have emerged as one of the most perilous security threats in artificial intelligence. An adversary embeds hidden malicious triggers during the training phase, ensuring the model maintains competitive accuracy on benign inputs while coercing any trigger-stamped input into a predefined adversary-chosen target class. In modern Machine Learning as a Service (MLaaS) workflows and foundation model ecosystems, downstream users rarely possess full governance over pretraining data. Consequently, practitioners predominantly rely on post-training defense paradigms—such as standard Fine-Tuning (FT), Neural Attention Distillation (NAD), Reconstructive Neuron Pruning (RNP), Feature Shift Tuning (FST), and Two-Stage Backdoor Defense (TSBD)—to purge latent trojans using a modest set of clean validation samples via catastrophic forgetting.

However, existing post-training defense mechanisms inherently rely on an untested core assumption: benign and malicious behaviors occupy separable parameter subspaces or distinct representation manifolds, implying that optimizing over small clean datasets can selectively overwrite or prune backdoor pathways without impairing clean task utility. This fundamental premise of latent separability exposes an overlooked vulnerability against adaptive, defense-aware adversaries. If the optimization trajectory of the backdoor task is deliberately entangled with that of the benign task, the defender's dual objective—eliminating the backdoor while preserving clean performance—collapses into an unresolvable structural conflict.

To investigate this theoretical and empirical vulnerability, this paper introduces an adaptive attack that couples the backdoor optimization direction directly with clean decision boundaries. By probing parameter sensitivity on a surrogate model, the attacker aligns backdoor gradients with clean gradients on the most critical neurons. Core idea: leverage the diagonal Fisher Information Matrix (FIM) to construct binary masks over task-critical neurons, and train an input-dependent trigger generator that explicitly maximizes the cosine similarity between backdoor and benign gradients on these masked parameters, embedding the backdoor directly into the clean task backbone so that any post-training unlearning inevitably degrades clean classification accuracy.

Method

Overall Architecture

The indelible backdoor framework operates across three sequential stages: Stage 1 estimates the diagonal Fisher Information Matrix on a clean subset to identify parameters essential to clean task performance and extracts a top-k% binary mask; Stage 2 trains an input-dependent trigger generator using a joint objective with an explicit gradient alignment loss, forcing backdoor gradients to mimic clean gradients on masked neurons; Stage 3 freezes the generator, crafts poisoned training samples with sample-specific perturbations, and trains the victim model from scratch under standard supervision.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Clean Subset + Surrogate Model"] --> B["Stage 1: Fisher Information Computation<br/>Diagonal FIM estimation & Top-k% mask"]
    B --> C["Stage 2: Gradient-Aligned Trigger Generation<br/>Maximize cosine similarity on masked neurons"]
    C --> D["Stage 3: Victim Model Training<br/>Inject sample-specific triggers & standard training"]
    D --> E["Delivered Victim Model<br/>Resilient to post-training fine-tuning & pruning"]

Key Designs

1. Fisher Information Computation: Pinpointing Task-Critical Decision Parameters

To maximize the destructive trade-off imposed on post-training defenses, the attacker must align backdoor gradients specifically with parameters that govern benign predictions, rather than arbitrary or noisy weights. Given a surrogate model \(f_{\text{sur},\theta}\) trained on clean data, the diagonal empirical Fisher Information Matrix is computed across \(N\) clean mini-batches:

\[F_i = \frac{1}{N} \sum_{n=1}^N \left( \nabla_{\theta_i} \mathcal{L}_{\text{CE}}(f_{\text{sur},\theta}(x_n), y_n) \right)^2\]

where \(\mathcal{L}_{\text{CE}}\) denotes the cross-entropy classification loss. Parameters exhibiting high \(F_i\) values receive large, directionally stable gradients across clean samples, forming the backbone of the model's decision boundaries; parameters with near-zero \(F_i\) experience noisy, self-canceling updates that contribute negligibly to benign classification. By sorting \(\{F_i\}\) and extracting the top-\(k\%\) (default \(k=10\%\)) parameters, the attacker builds a layer-wise binary importance mask \(M_i\), precisely isolating the sensitive parameter subspace that downstream fine-tuning defenses rely upon to preserve clean accuracy.

2. Gradient-Aligned Trigger Generation: Entangling Optimization Manifolds

Unlike conventional dynamic backdoors that optimize trigger generators solely for visual imperceptibility and classification bias, the proposed generator \(T_\phi(x)\) maps an input \(x\) into a bounded perturbation, producing poisoned samples \(\tilde{x} = \text{clip}(x + \epsilon \cdot T_\phi(x), 0, 1)\) while optimizing for post-training persistence. The generator is trained under a four-component composite loss:

\[\mathcal{L}_T = \lambda_{\text{bd}} \mathcal{L}_{\text{bd}} + \lambda_{\text{align}} \mathcal{L}_{\text{align}} + \lambda_{\text{clean}} \mathcal{L}_{\text{clean}} + \mathcal{L}_{L_2}\]

where \(\mathcal{L}_{\text{bd}}\) enforces misclassification to the target class \(c\), \(\mathcal{L}_{\text{clean}}\) preserves benign classification performance, and \(\mathcal{L}_{L_2}\) penalizes perturbation norm. The central mechanism is the alignment loss \(\mathcal{L}_{\text{align}}\). For each trainable layer \(i\), the masked clean gradient \(g_i^c = \text{vec}(M_i \odot \nabla_{\theta_i} \mathcal{L}_{\text{clean}})\) and masked backdoor gradient \(g_i^b = \text{vec}(M_i \odot \nabla_{\theta_i} \mathcal{L}_{\text{bd}})\) are extracted, and their negative cosine similarity is minimized across all \(L\) masked layers:

\[\mathcal{L}_{\text{align}} = - \frac{1}{L} \sum_{i=1}^L \frac{g_i^c \cdot g_i^b}{\|g_i^c\|_2 \|g_i^b\|_2}\]

By driving \(g_i^b\) to align with \(g_i^c\), any gradient step taken by downstream defenders using clean data simultaneously pushes model parameters in directions favorable to both benign accuracy and the backdoor objective, effectively turning clean fine-tuning into a backdoor reinforcer.

3. Victim Model Training: Embedding Indelible Representations

Once generator training converges, \(T_\phi\) is frozen. The attacker poisons a small fraction \(\rho\) of the training set by generating sample-conditioned perturbations and reassigning their labels to target class \(c\). Because each input receives an individualized, image-dependent perturbation pattern, the poisoned dataset easily bypasses static pattern detection or trigger reverse-engineering tools (e.g., Neural Cleanse, STRIP). The victim model \(f_\theta\) is then trained from scratch under conventional supervised learning. The embedded gradient alignment ensures that the backdoor functionality fuses into the primary feature extraction circuits, leaving no isolated "backdoor neurons" for post-training defenses to excise.

Loss & Training

The generator is trained for 20 epochs with hyperparameters set to \(\lambda_{\text{align}} = 1.0\), \(\lambda_{\text{bd}} = 2.0\), and \(\lambda_{\text{clean}} = 0.1\) with perturbation budget \(\epsilon = 0.1\). The victim model is trained for 100 epochs with an initial learning rate of 0.01. Fisher estimation is performed over 1,000 clean samples to select the top 10% parameters. Defense resistance is comprehensively quantified using the Defense Effectiveness Rate (DER):

\[\text{DER} = \frac{\max(0, \Delta \text{ASR}) - \max(0, \Delta \text{ACC}) + 1}{2}\]

which rewards ASR suppression while penalizing clean accuracy degradation. A defense that cuts ASR only by collapsing model utility achieves a low DER, revealing its practical ineffectiveness.

Key Experimental Results

Main Results

The attack is evaluated on CIFAR-10 and Tiny-ImageNet against 7 baseline attacks (BadNets, Blended, WaNet, LC, Input-Aware, Narcissus, COMBAT) and 8 post-training defenses (FT, ANP, NAD, I-BAU, RNP, Unit, FST, TSBD). The quantitative evaluation on CIFAR-10 is summarized below:

Defense Method Metric BadNets WaNet Narcissus COMBAT Ours Empirical Observation
Pretrained ACC (%)
ASR (%)
91.44
94.41
92.67
99.54
93.09
94.64
93.94
94.47
93.27
100.00
All baseline and proposed models achieve strong initial utility and high ASR
FT (Fine-Tuning) ACC (%)
ASR (%)
DER (%)
90.56
1.47
96.03
92.50
13.91
92.73
92.35
89.81
52.05
93.46
72.83
60.58
92.20
99.99
49.47
Baseline ASR collapses to 1%-14%, while Ours persists at 99.99% without drop
NAD (Attention Distill) ACC (%)
ASR (%)
DER (%)
89.33
2.08
95.11
91.88
9.98
94.39
91.27
88.06
52.38
93.48
70.96
61.52
91.54
99.97
49.15
Attention distillation fails to separate aligned features, leaving ASR at 99.97%
I-BAU (Hypergradient) ACC (%)
ASR (%)
DER (%)
88.13
7.91
91.60
86.52
20.04
86.68
89.27
33.01
78.91
91.01
1.98
94.78
91.14
88.47
54.70
COMBAT is cleared to 1.98%, whereas Ours retains 88.47% ASR
RNP (Reconstructive) ACC (%)
ASR (%)
DER (%)
87.63
3.76
93.42
90.34
0.17
98.52
92.97
91.52
51.50
91.40
24.23
83.85
86.03
96.51
48.13
Pruning Fisher-important neurons degrades clean accuracy while ASR stays 96.51%
FST (Feature Shift) ACC (%)
ASR (%)
DER (%)
87.06
2.08
93.98
92.40
0.58
99.35
92.18
93.91
49.91
91.25
30.65
80.57
89.68
99.58
48.42
Advanced feature shift tuning is completely bypassed by aligned gradients
TSBD (Weight Change) ACC (%)
ASR (%)
DER (%)
90.13
1.78
95.66
92.48
1.08
99.14
92.85
82.16
56.12
92.91
35.57
78.94
93.11
99.84
50.00
Reinitializing correlated weights fails as clean updates re-activate the backdoor

On the more complex Tiny-ImageNet benchmark, recent defenses suffer catastrophic failure: Unit reduces ASR to 66.92% only by destroying clean accuracy (dropping ACC to 3.09%); FST suppresses ASR to 21.38% but sacrifices clean ACC down to 32.74%; and TSBD achieves 53.90% ACC but leaves an intact 91.38% ASR against the proposed attack.

Ablation Study

The ablation experiments investigate the sensitivity of the attack against varying perturbation budgets \(\epsilon\) and poisoning ratios \(\rho\). The table below reports resilience across 5 defenses under different perturbation bounds \(\epsilon\):

Perturbation Budget \(\epsilon\) Pretrained (ACC/ASR) NAD (ACC/ASR) I-BAU (ACC/ASR) Unit (ACC/ASR) FST (ACC/ASR) TSBD (ACC/ASR)
\(\epsilon = 0.100\) 93.27% / 100.00% 91.54% / 99.97% 91.14% / 88.47% 88.15% / 92.07% 89.68% / 99.58% 93.11% / 99.84%
\(\epsilon = 0.075\) 92.82% / 100.00% 90.21% / 98.88% 88.62% / 94.97% 87.61% / 96.62% 89.67% / 99.12% 91.44% / 97.44%
\(\epsilon = 0.050\) 92.96% / 99.99% 90.43% / 87.97% 90.45% / 72.43% 88.25% / 90.09% 89.70% / 87.02% 91.92% / 81.02%
\(\epsilon = 0.030\) 92.91% / 99.92% 89.98% / 57.06% 89.20% / 51.84% 87.77% / 69.06% 89.19% / 71.99% 91.92% / 47.76%

Furthermore, when evaluating attack budgets with minimal poisoning ratios (Table 4), the proposed attack maintains \(\ge 99.29\%\) ASR under all defenses (FT, NAD, I-BAU, Unit, FST) even at \(\rho = 0.01\). In stark contrast, Narcissus drops precipitously from 94.64% to 38.26% initial ASR and further down to 33.56% under NAD at \(\rho = 0.01\), demonstrating the exceptional data efficiency of gradient-aligned poisoning.

Key Findings

  • Clean fine-tuning serves as a backdoor booster: For all optimization-based defenses (FT, NAD, FST, TSBD), high gradient alignment on Fisher-critical weights causes benign gradient steps to inadvertently reinforce backdoor pathways, maintaining near-100% ASR across extended fine-tuning epochs.
  • Neuron pruning incurs severe utility collapse: Pruning defenses (ANP, RNP) depend on the assumption that backdoor-related weights are dormant or redundant for clean tasks. Forcing backdoors into top-10% Fisher neurons breaks this premise, making neuron excision directly cannibalize benign utility.
  • Universal architectural applicability: Across PreAct-ResNet18, VGG19_BN, MobileNetV3-Large, EfficientNet-B3, and ViT-B/16, defenses consistently fail to enter the "Sweet Spot" (\(\Delta \text{ACC} \le 5\%\) and \(\text{ASR} \le 20\%\)), proving that gradient alignment targets fundamental optimization dynamics rather than model-specific inductive biases.

Highlights & Insights

  • From latent separability to gradient entanglement: Shifts the paradigm of backdoor persistence from static heuristic masking or clean-label clustering to direct optimization of gradient directional alignment, invalidating the foundational assumptions of post-training unlearning.
  • Fisher Information as a surgical entanglement guide: Rather than penalizing global weight divergence, the method exploits EWC-inspired parameter sensitivity to align gradients solely on the top-10% variance-bearing neurons, achieving maximal defense resistance with minimal optimization overhead.
  • Demarcating the empirical boundary of post-training defense: Provides conclusive evidence that post-training defenses operating solely on small clean datasets cannot reliably distinguish deeply entangled backdoors from benign features, underscoring the urgent need for pretraining auditing or causal intervention.

Limitations & Future Work

  • Surrogate model access requirement: Computing the Fisher Information Matrix and training the alignment generator requires access to a representative architecture and clean dataset; extending this approach to black-box, zero-knowledge, or cross-modality foundation model transfer warrants further study.
  • Dirty-label reliance: While sample-specific triggers evade static reverse-engineering, the poisoning stage still flips labels to target class \(c\), leaving the dataset vulnerable to human inspection or sophisticated data sanitization filters.
  • Defense co-evolution: The findings incentivize defensive researchers to develop higher-order curvature analysis, Hessian asymmetry audits, or causal counterfactual testing that go beyond first-order clean fine-tuning.
  • vs Input-Aware & WaNet: Early sample-dependent and warping backdoors bypassed pattern reverse-engineering but remained vulnerable to feature shift tuning and neuron pruning; this paper provides resilience via direct parameter gradient alignment.
  • vs Narcissus & COMBAT: SOTA clean-label attacks exhibit severe performance collapse under low poisoning rates (\(\rho = 0.01\)) or complex datasets (Tiny-ImageNet); the proposed method sustains near-saturated ASR across all budget regimes.
  • vs NAD & TSBD: Existing defense literature assumes attention maps or unlearning trajectories diverge between clean and trojan tasks; this work proves that intentional gradient co-directionality shatters such divergence.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates an innovative gradient-alignment objective anchored on Fisher-critical parameters to expose the structural vulnerability of post-training defenses.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Exhaustively benchmarks 7 attacks and 8 defenses across multiple vision architectures, data scales, poisoning ratios, and perturbation budgets.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulation, lucid problem motivation, and comprehensive analysis of empirical failure modes.
  • Value: ⭐⭐⭐⭐⭐ Establishes a critical benchmark demonstrating the limitations of post-training backdoor sanitization in real-world MLaaS pipelines.