Skip to content

Mechanistic interventions for explainable digital pathology uncovers adversarial vulnerabilities

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Medical Imaging
Keywords: Digital Pathology, Mechanistic Interpretability, Adversarial Robustness, Sparse Autoencoders, Activation Patching

TL;DR

Addressing the black-box opacity and vulnerability of digital pathology models to adversarial perturbations, this work presents a mechanistic framework combining sparse autoencoders and causal activation patching to uncover spurious artifact dependencies (e.g., staining variations and scanner noise), and introduces Mechanism-Informed Adversarial Training (MIAT) to surgically regularize vulnerable neural circuits while maintaining clean diagnostic accuracy.

Background & Motivation

Deep learning has demonstrated remarkable success in automated cancer grading across whole-slide images (WSIs) in digital pathology, offering transformative potential for clinical diagnostics. However, high-stakes medical settings require absolute trust and reliability, whereas modern deep neural networks remain opaque black boxes. Critically, these networks are exceptionally susceptible to adversarial perturbations—imperceptible input noise that drastically alters diagnostic outputs such as Gleason or Nottingham cancer grades. Such failure modes pose severe risks to patient health and violate emerging regulatory standards for Software as a Medical Device (SaMD).

The fundamental vulnerability stems from shortcut learning: deep pathology models frequently exploit non-biological artifacts, such as laboratory-specific staining variations, digitization sensor noise, and compression artifacts, rather than generalizable morphology. Meanwhile, conventional adversarial defenses treat robustness as a monolithic optimization problem, applying indiscriminate global regularization across the entire parameter space. This causes severe feature over-smoothing, significantly degrades diagnostic accuracy on clean images, and fails to diagnose the root mechanistic causes of vulnerability within internal representations.

This paper tackles the challenge by adopting mechanistic interpretability tools to uncover feature disentanglement and causal circuit pathways in pathology networks. Core idea: decompose high-level representations into morphological concepts versus non-biological artifacts using sparse autoencoders (SAEs), isolate the neural circuits causally responsible for vulnerability via activation patching, and propose Mechanism-Informed Adversarial Training (MIAT) to surgically regularize vulnerable circuits for Pareto-optimal robustness and clean accuracy.

Method

Overall Architecture

The proposed MIAT pipeline operates across three foundational stages: first, an overcomplete sparse autoencoder (SAE) decomposes penultimate-layer activations into monosemantic morphological patterns versus non-biological artifacts; second, causal activation patching and path patching systematically intervene across layers to identify the minimal set of neural circuits sufficient and necessary for adversarial vulnerability; finally, circuit-targeted adversarial perturbations are generated and paired with a circuit-specific consistency regularization loss during training, with dynamic circuit discovery iteratively updating the target circuits as the model evolves.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Pathology Images<br/>(Clean WSI patches / Adversarial inputs)"] --> B["Feature Disentanglement & Circuit Discovery<br/>Penultimate-layer SAE & Causal Patching"]
    B --> C["Targeted Adversarial Example Generation<br/>Maximizing vulnerable circuit activation shift"]
    C --> D["Circuit-Specific Regularization<br/>Penalizing activation deviation on vulnerable circuits"]
    D --> E["Dynamic Circuit Discovery<br/>Periodic re-identification and reinforcement"]
    E --> F["Robust Pathology Model<br/>Pareto-optimal clean and adversarial accuracy"]

Key Designs

1. Feature Disentanglement & Circuit Discovery: Uncovering Non-Biological Artifacts and Causal Pathways

Standard linear probing only tests whether linear representations correlate with ground-truth grades, failing to distinguish clinically genuine glandular structures from digitization artifacts. In this work, an overcomplete SAE with dictionary \(D \in \mathbb{R}^{k \times d}\) (\(k=4096\), \(4\times\) overcomplete relative to activation dimension \(d\)) is trained on penultimate activations \(h \in \mathbb{R}^d\) via: $\(\mathcal{L}_{\text{SAE}} = \|h - D z\|_2^2 + \lambda \|z\|_1\)$ Analyzing maximally activating exemplars reveals that approximately 60% of feature directions encode true morphology (epithelial density, glandular formation), while 40% encode non-morphological artifacts (H&E staining inconsistency, scanner sensor noise, JPEG compression blocks). Subsequently, activation patching intervenes causally: substituting clean activations into an adversarial forward pass checks if the upstream circuit is sufficient for robustness; corrupting clean activations with adversarial states checks if the downstream circuit is necessary for vulnerability, isolating the vulnerable circuit set \(\mathcal{C}\).

2. Targeted Adversarial Example Generation: Amplifying Activation Shifts in Vulnerable Pathways

Standard PGD attacks optimize purely against classification cross-entropy, dispersing adversarial perturbations across unconstrained feature dimensions. MIAT directly crafts perturbations designed to exploit the discovered vulnerable circuits \(\mathcal{C}\) by maximizing the activation divergence within those specific pathways: $\(\delta^* = \arg\max_{\|\delta\|_p \le \epsilon} \sum_{a \in \mathcal{C}} \|a(x + \delta) - a(x)\|_2^2\)$ where \(a(x)\) denotes the activation of a neuron or attention head within circuit \(\mathcal{C}\) under input \(x\). This forces adversarial examples to aggressively trigger artifact-sensitive features, providing maximally informative adversarial training signals.

3. Circuit-Specific Regularization: Precision Surgical Fortification of Vulnerable Subnetworks

Conventional adversarial training imposes global regularization, dampening clean predictive features. MIAT introduces a targeted consistency regularizer acting exclusively on the identified vulnerable circuits \(\mathcal{C}\): $\(\mathcal{L}_{\text{reg}} = \sum_{a \in \mathcal{C}} \|a(x') - a(x)\|_2^2\)$ The total training objective balances classification performance with circuit stability: $\(\mathcal{L}_{\text{MIAT}} = \mathcal{L}_{\text{CE}}(f_\theta(x'), y) + \alpha \mathcal{L}_{\text{reg}}\)$ where \(\alpha\) controls regularization strength (\(\alpha=0.01\)). This constraint enforces representation invariance specifically across vulnerable pathways without over-regularizing healthy morphological circuits, effectively preventing adversarial noise from propagating into diagnostic decisions.

4. Dynamic Circuit Discovery: Adapting to Evolving Neural Representations

As model parameters update during adversarial training, the suppression of primary vulnerable circuits may cause the network to route information through alternative pathways or develop secondary vulnerabilities. To maintain targeted defense fidelity, MIAT dynamically updates circuit set \(\mathcal{C}\) by periodically re-running lightweight SAE decomposition and causal patching throughout training iterations. This closed-loop audit guarantees that the regularizer continuously adapts to the model's current failure modes.

Key Experimental Results

Main Results

Evaluated on TCGA-Prostate (Gleason grading, ~500,000 patches) and BreastPathQ (Nottingham grading, ~300,000 patches) across Vision Transformer (ViT-B/16) and ResNet-50 backbones under Clean, PGD (\(\epsilon=8/255\)), Carlini-Wagner (CW), and DeepFool attacks.

Dataset Model Training Method Clean Acc (%) PGD Acc (%) CW Acc (%) DeepFool Acc (%)
TCGA-Prostate ViT-B/16 Standard Clean 88.2 12.5 18.3 22.1
TCGA-Prostate ViT-B/16 Standard AT 85.1 68.4 55.2 48.7
TCGA-Prostate ViT-B/16 MIAT (Ours) 86.8 75.3 62.8 55.6
TCGA-Prostate ResNet-50 Standard Clean 86.5 10.8 15.7 19.4
TCGA-Prostate ResNet-50 Standard AT 83.9 62.1 50.3 44.2
TCGA-Prostate ResNet-50 MIAT (Ours) 85.2 70.5 58.9 51.8
BreastPathQ ViT-B/16 Standard Clean 87.1 11.9 17.6 21.3
BreastPathQ ViT-B/16 Standard AT 84.3 65.8 53.1 47.2
BreastPathQ ViT-B/16 MIAT (Ours) 85.9 73.7 60.4 54.1
BreastPathQ ResNet-50 Standard Clean 85.4 10.2 14.9 18.7
BreastPathQ ResNet-50 Standard AT 82.8 60.5 48.9 43.0
BreastPathQ ResNet-50 MIAT (Ours) 84.6 69.2 57.3 50.5

Ablation Study & Benchmark Comparisons

Evaluated on TCGA-Prostate against state-of-the-art defenses (TRADES, MART, LAT) and ablated configurations.

Config / Defense Method Model Clean Acc (%) PGD Acc (%) CW Acc (%) Note
TRADES ViT-B/16 84.3 65.2 52.1 Global trade-off loss, excessive smoothing
MART ViT-B/16 83.8 67.5 54.8 Focuses on misclassified adversarial samples
LAT (Latent AT) ViT-B/16 85.0 70.1 57.3 Layerwise latent adversarial perturbations
MIAT (Full, \(\alpha=0.01\)) ViT-B/16 86.8 75.3 62.8 Full model: mechanism-informed circuit regularization
Ablation: \(\alpha=0\) (w/o circuit reg) ViT-B/16 85.1 68.4 55.2 Degrades entirely to standard AT
Ablation: Random circuit selection ViT-B/16 84.9 68.6 55.0 Regularizing arbitrary circuits yields no gain
TRADES ResNet-50 83.1 63.5 50.8 Baseline CNN defense
MART ResNet-50 82.6 65.2 52.4 Baseline CNN defense
LAT ResNet-50 84.0 68.3 55.7 Layerwise CNN defense
MIAT (Full, \(\alpha=0.01\)) ResNet-50 85.2 70.5 58.9 Full model: optimal trade-off on CNN

Key Findings

  • Architectural Distribution of Vulnerabilities: In Vision Transformers, adversarial vulnerabilities are sharply localized in intermediate layers (layers 6–9), where specific attention heads (e.g., layer 7) disproportionately attend to scanner artifacts. In ResNet-50, vulnerable circuits are more broadly dispersed across later convolutional blocks (blocks 3 and 4), responding to high-frequency noise and staining artifacts.
  • Causal Intervention Efficacy: In ViT, patching clean activations at layer 6 into adversarial forward passes rescues 78.3% of misclassified cases (sufficient for robustness); conversely, injecting adversarial activations at layer 9 into clean passes triggers misclassifications in 81.7% of cases (necessary for vulnerability).
  • Representation Distance Reduction: MIAT decreases the mean \(L_2\) activation distance between clean and adversarial inputs within vulnerable circuits by 42.1%, directly demonstrating its ability to suppress adversarial noise propagation while preserving clean accuracy within 1.4% of the clean baseline.

Highlights & Insights

  • Bridging Mechanistic Interpretability and Active Defense: Unlike passive interpretability tools, MIAT converts post-hoc circuit discovery via SAEs and activation patching into a closed-loop training regularizer, transforming mechanistic insights into actionable model defense.
  • Unmasking Artifact Exploitation in Pathology: Demonstrating that 40% of learned feature directions in diagnostic models attend to staining inconsistencies and digitization artifacts provides vital empirical evidence for why clinical pathology models fail to generalize out-of-distribution.
  • Regulatory Alignment for Medical AI: Directly addresses FDA SaMD and EU MDR Class III medical device compliance mandates by offering auditable, circuit-level explanations and risk mitigation documentation for life-critical diagnostics.

Limitations & Future Work

  • Computational Overhead in Discovery: Training overcomplete SAEs (\(k=4096\)) and executing iterative activation patching incurs notable GPU overhead during training. Although inference latency remains unchanged, resource-efficient approximations (e.g., randomized or quantized SAEs) are needed for lightweight deployment.
  • Potential Exposure to Higher-Order Adaptive Attacks: The evaluation primarily benchmarks standard first-order attacks (PGD, CW, DeepFool); assessing resilience against sophisticated white-box adaptive attackers that dynamically circumvent patched circuits warrants deeper exploration.
  • Multi-Modal and Multi-Center Expansion: Future validation should extend beyond H&E histology to multi-modal foundation models integrating genomics, spatial transcriptomics, and rare histological variants.
  • vs TRADES / MART: Standard defenses enforce global smoothing across all layers, leading to noticeable clean accuracy degradation (3–5% drop); MIAT isolates the specific vulnerable minority of circuits, achieving Pareto-optimal accuracy and robustness.
  • vs Saliency & Heatmap Explanations (Grad-CAM): Pixel-level attribution maps fail to uncover which internal computational mechanisms or attention heads are responsible for shortcut reliance; MIAT provides causal, circuit-level isolation and actionable intervention targets.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneering integration of sparse autoencoders and causal activation patching for adversarial defense in digital pathology]
  • Experimental Thoroughness: ⭐⭐⭐⭐ [Evaluated across two major clinical cohorts, ViT and ResNet architectures, and multiple attack methods with rigorous ablations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, mathematically rigorous formulation, cohesive narrative spanning theory, clinical application, and regulatory requirements]
  • Value: ⭐⭐⭐⭐⭐ [Provides an essential blueprint for building explainable, robust, and regulation-ready AI systems in safety-critical medical diagnostics]