Mechanistic interventions for explainable digital pathology uncovers adversarial vulnerabilities¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Medical Imaging
Keywords: Digital Pathology, Mechanistic Interpretability, Adversarial Robustness, Sparse Autoencoders, Activation Patching
TL;DR¶
Addressing the black-box opacity and vulnerability of digital pathology models to adversarial perturbations, this work presents a mechanistic framework combining sparse autoencoders and causal activation patching to uncover spurious artifact dependencies (e.g., staining variations and scanner noise), and introduces Mechanism-Informed Adversarial Training (MIAT) to surgically regularize vulnerable neural circuits while maintaining clean diagnostic accuracy.
Background & Motivation¶
Deep learning has demonstrated remarkable success in automated cancer grading across whole-slide images (WSIs) in digital pathology, offering transformative potential for clinical diagnostics. However, high-stakes medical settings require absolute trust and reliability, whereas modern deep neural networks remain opaque black boxes. Critically, these networks are exceptionally susceptible to adversarial perturbations—imperceptible input noise that drastically alters diagnostic outputs such as Gleason or Nottingham cancer grades. Such failure modes pose severe risks to patient health and violate emerging regulatory standards for Software as a Medical Device (SaMD).
The fundamental vulnerability stems from shortcut learning: deep pathology models frequently exploit non-biological artifacts, such as laboratory-specific staining variations, digitization sensor noise, and compression artifacts, rather than generalizable morphology. Meanwhile, conventional adversarial defenses treat robustness as a monolithic optimization problem, applying indiscriminate global regularization across the entire parameter space. This causes severe feature over-smoothing, significantly degrades diagnostic accuracy on clean images, and fails to diagnose the root mechanistic causes of vulnerability within internal representations.
This paper tackles the challenge by adopting mechanistic interpretability tools to uncover feature disentanglement and causal circuit pathways in pathology networks. Core idea: decompose high-level representations into morphological concepts versus non-biological artifacts using sparse autoencoders (SAEs), isolate the neural circuits causally responsible for vulnerability via activation patching, and propose Mechanism-Informed Adversarial Training (MIAT) to surgically regularize vulnerable circuits for Pareto-optimal robustness and clean accuracy.
Method¶
Overall Architecture¶
The proposed MIAT pipeline operates across three foundational stages: first, an overcomplete sparse autoencoder (SAE) decomposes penultimate-layer activations into monosemantic morphological patterns versus non-biological artifacts; second, causal activation patching and path patching systematically intervene across layers to identify the minimal set of neural circuits sufficient and necessary for adversarial vulnerability; finally, circuit-targeted adversarial perturbations are generated and paired with a circuit-specific consistency regularization loss during training, with dynamic circuit discovery iteratively updating the target circuits as the model evolves.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Pathology Images<br/>(Clean WSI patches / Adversarial inputs)"] --> B["Feature Disentanglement & Circuit Discovery<br/>Penultimate-layer SAE & Causal Patching"]
B --> C["Targeted Adversarial Example Generation<br/>Maximizing vulnerable circuit activation shift"]
C --> D["Circuit-Specific Regularization<br/>Penalizing activation deviation on vulnerable circuits"]
D --> E["Dynamic Circuit Discovery<br/>Periodic re-identification and reinforcement"]
E --> F["Robust Pathology Model<br/>Pareto-optimal clean and adversarial accuracy"]
Key Designs¶
1. Feature Disentanglement & Circuit Discovery: Uncovering Non-Biological Artifacts and Causal Pathways
Standard linear probing only tests whether linear representations correlate with ground-truth grades, failing to distinguish clinically genuine glandular structures from digitization artifacts. In this work, an overcomplete SAE with dictionary \(D \in \mathbb{R}^{k \times d}\) (\(k=4096\), \(4\times\) overcomplete relative to activation dimension \(d\)) is trained on penultimate activations \(h \in \mathbb{R}^d\) via: $\(\mathcal{L}_{\text{SAE}} = \|h - D z\|_2^2 + \lambda \|z\|_1\)$ Analyzing maximally activating exemplars reveals that approximately 60% of feature directions encode true morphology (epithelial density, glandular formation), while 40% encode non-morphological artifacts (H&E staining inconsistency, scanner sensor noise, JPEG compression blocks). Subsequently, activation patching intervenes causally: substituting clean activations into an adversarial forward pass checks if the upstream circuit is sufficient for robustness; corrupting clean activations with adversarial states checks if the downstream circuit is necessary for vulnerability, isolating the vulnerable circuit set \(\mathcal{C}\).
2. Targeted Adversarial Example Generation: Amplifying Activation Shifts in Vulnerable Pathways
Standard PGD attacks optimize purely against classification cross-entropy, dispersing adversarial perturbations across unconstrained feature dimensions. MIAT directly crafts perturbations designed to exploit the discovered vulnerable circuits \(\mathcal{C}\) by maximizing the activation divergence within those specific pathways: $\(\delta^* = \arg\max_{\|\delta\|_p \le \epsilon} \sum_{a \in \mathcal{C}} \|a(x + \delta) - a(x)\|_2^2\)$ where \(a(x)\) denotes the activation of a neuron or attention head within circuit \(\mathcal{C}\) under input \(x\). This forces adversarial examples to aggressively trigger artifact-sensitive features, providing maximally informative adversarial training signals.
3. Circuit-Specific Regularization: Precision Surgical Fortification of Vulnerable Subnetworks
Conventional adversarial training imposes global regularization, dampening clean predictive features. MIAT introduces a targeted consistency regularizer acting exclusively on the identified vulnerable circuits \(\mathcal{C}\): $\(\mathcal{L}_{\text{reg}} = \sum_{a \in \mathcal{C}} \|a(x') - a(x)\|_2^2\)$ The total training objective balances classification performance with circuit stability: $\(\mathcal{L}_{\text{MIAT}} = \mathcal{L}_{\text{CE}}(f_\theta(x'), y) + \alpha \mathcal{L}_{\text{reg}}\)$ where \(\alpha\) controls regularization strength (\(\alpha=0.01\)). This constraint enforces representation invariance specifically across vulnerable pathways without over-regularizing healthy morphological circuits, effectively preventing adversarial noise from propagating into diagnostic decisions.
4. Dynamic Circuit Discovery: Adapting to Evolving Neural Representations
As model parameters update during adversarial training, the suppression of primary vulnerable circuits may cause the network to route information through alternative pathways or develop secondary vulnerabilities. To maintain targeted defense fidelity, MIAT dynamically updates circuit set \(\mathcal{C}\) by periodically re-running lightweight SAE decomposition and causal patching throughout training iterations. This closed-loop audit guarantees that the regularizer continuously adapts to the model's current failure modes.
Key Experimental Results¶
Main Results¶
Evaluated on TCGA-Prostate (Gleason grading, ~500,000 patches) and BreastPathQ (Nottingham grading, ~300,000 patches) across Vision Transformer (ViT-B/16) and ResNet-50 backbones under Clean, PGD (\(\epsilon=8/255\)), Carlini-Wagner (CW), and DeepFool attacks.
| Dataset | Model | Training Method | Clean Acc (%) | PGD Acc (%) | CW Acc (%) | DeepFool Acc (%) |
|---|---|---|---|---|---|---|
| TCGA-Prostate | ViT-B/16 | Standard Clean | 88.2 | 12.5 | 18.3 | 22.1 |
| TCGA-Prostate | ViT-B/16 | Standard AT | 85.1 | 68.4 | 55.2 | 48.7 |
| TCGA-Prostate | ViT-B/16 | MIAT (Ours) | 86.8 | 75.3 | 62.8 | 55.6 |
| TCGA-Prostate | ResNet-50 | Standard Clean | 86.5 | 10.8 | 15.7 | 19.4 |
| TCGA-Prostate | ResNet-50 | Standard AT | 83.9 | 62.1 | 50.3 | 44.2 |
| TCGA-Prostate | ResNet-50 | MIAT (Ours) | 85.2 | 70.5 | 58.9 | 51.8 |
| BreastPathQ | ViT-B/16 | Standard Clean | 87.1 | 11.9 | 17.6 | 21.3 |
| BreastPathQ | ViT-B/16 | Standard AT | 84.3 | 65.8 | 53.1 | 47.2 |
| BreastPathQ | ViT-B/16 | MIAT (Ours) | 85.9 | 73.7 | 60.4 | 54.1 |
| BreastPathQ | ResNet-50 | Standard Clean | 85.4 | 10.2 | 14.9 | 18.7 |
| BreastPathQ | ResNet-50 | Standard AT | 82.8 | 60.5 | 48.9 | 43.0 |
| BreastPathQ | ResNet-50 | MIAT (Ours) | 84.6 | 69.2 | 57.3 | 50.5 |
Ablation Study & Benchmark Comparisons¶
Evaluated on TCGA-Prostate against state-of-the-art defenses (TRADES, MART, LAT) and ablated configurations.
| Config / Defense Method | Model | Clean Acc (%) | PGD Acc (%) | CW Acc (%) | Note |
|---|---|---|---|---|---|
| TRADES | ViT-B/16 | 84.3 | 65.2 | 52.1 | Global trade-off loss, excessive smoothing |
| MART | ViT-B/16 | 83.8 | 67.5 | 54.8 | Focuses on misclassified adversarial samples |
| LAT (Latent AT) | ViT-B/16 | 85.0 | 70.1 | 57.3 | Layerwise latent adversarial perturbations |
| MIAT (Full, \(\alpha=0.01\)) | ViT-B/16 | 86.8 | 75.3 | 62.8 | Full model: mechanism-informed circuit regularization |
| Ablation: \(\alpha=0\) (w/o circuit reg) | ViT-B/16 | 85.1 | 68.4 | 55.2 | Degrades entirely to standard AT |
| Ablation: Random circuit selection | ViT-B/16 | 84.9 | 68.6 | 55.0 | Regularizing arbitrary circuits yields no gain |
| TRADES | ResNet-50 | 83.1 | 63.5 | 50.8 | Baseline CNN defense |
| MART | ResNet-50 | 82.6 | 65.2 | 52.4 | Baseline CNN defense |
| LAT | ResNet-50 | 84.0 | 68.3 | 55.7 | Layerwise CNN defense |
| MIAT (Full, \(\alpha=0.01\)) | ResNet-50 | 85.2 | 70.5 | 58.9 | Full model: optimal trade-off on CNN |
Key Findings¶
- Architectural Distribution of Vulnerabilities: In Vision Transformers, adversarial vulnerabilities are sharply localized in intermediate layers (layers 6–9), where specific attention heads (e.g., layer 7) disproportionately attend to scanner artifacts. In ResNet-50, vulnerable circuits are more broadly dispersed across later convolutional blocks (blocks 3 and 4), responding to high-frequency noise and staining artifacts.
- Causal Intervention Efficacy: In ViT, patching clean activations at layer 6 into adversarial forward passes rescues 78.3% of misclassified cases (sufficient for robustness); conversely, injecting adversarial activations at layer 9 into clean passes triggers misclassifications in 81.7% of cases (necessary for vulnerability).
- Representation Distance Reduction: MIAT decreases the mean \(L_2\) activation distance between clean and adversarial inputs within vulnerable circuits by 42.1%, directly demonstrating its ability to suppress adversarial noise propagation while preserving clean accuracy within 1.4% of the clean baseline.
Highlights & Insights¶
- Bridging Mechanistic Interpretability and Active Defense: Unlike passive interpretability tools, MIAT converts post-hoc circuit discovery via SAEs and activation patching into a closed-loop training regularizer, transforming mechanistic insights into actionable model defense.
- Unmasking Artifact Exploitation in Pathology: Demonstrating that 40% of learned feature directions in diagnostic models attend to staining inconsistencies and digitization artifacts provides vital empirical evidence for why clinical pathology models fail to generalize out-of-distribution.
- Regulatory Alignment for Medical AI: Directly addresses FDA SaMD and EU MDR Class III medical device compliance mandates by offering auditable, circuit-level explanations and risk mitigation documentation for life-critical diagnostics.
Limitations & Future Work¶
- Computational Overhead in Discovery: Training overcomplete SAEs (\(k=4096\)) and executing iterative activation patching incurs notable GPU overhead during training. Although inference latency remains unchanged, resource-efficient approximations (e.g., randomized or quantized SAEs) are needed for lightweight deployment.
- Potential Exposure to Higher-Order Adaptive Attacks: The evaluation primarily benchmarks standard first-order attacks (PGD, CW, DeepFool); assessing resilience against sophisticated white-box adaptive attackers that dynamically circumvent patched circuits warrants deeper exploration.
- Multi-Modal and Multi-Center Expansion: Future validation should extend beyond H&E histology to multi-modal foundation models integrating genomics, spatial transcriptomics, and rare histological variants.
Related Work & Insights¶
- vs TRADES / MART: Standard defenses enforce global smoothing across all layers, leading to noticeable clean accuracy degradation (3–5% drop); MIAT isolates the specific vulnerable minority of circuits, achieving Pareto-optimal accuracy and robustness.
- vs Saliency & Heatmap Explanations (Grad-CAM): Pixel-level attribution maps fail to uncover which internal computational mechanisms or attention heads are responsible for shortcut reliance; MIAT provides causal, circuit-level isolation and actionable intervention targets.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneering integration of sparse autoencoders and causal activation patching for adversarial defense in digital pathology]
- Experimental Thoroughness: ⭐⭐⭐⭐ [Evaluated across two major clinical cohorts, ViT and ResNet architectures, and multiple attack methods with rigorous ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, mathematically rigorous formulation, cohesive narrative spanning theory, clinical application, and regulatory requirements]
- Value: ⭐⭐⭐⭐⭐ [Provides an essential blueprint for building explainable, robust, and regulation-ready AI systems in safety-critical medical diagnostics]