SPICE: Simple Polysemantic feature Interpretation via Clustering-based Explanations¶
Conference: ECCV 2026
Paper: ECCV 2026
Area: Interpretability
Keywords: polysemanticity / concept disentanglement / attribution clustering / mechanistic interpretability / vision transformer
TL;DR¶
Addressing the challenge of polysemanticity where individual neurons activate for multiple unrelated concepts, SPICE introduces an architecture-agnostic framework that adaptively determines cluster count \(K\) via attribution footprints and kernel density estimation, enabling the first unified polysemanticity comparison across CNNs and Vision Transformers.
Background & Motivation¶
Mechanistic interpretability aims to reverse-engineer the internal representations and computational mechanisms learned by neural networks. In computer vision, foundational neuron-level interpretation methods historically operated under the monosemanticity assumption, presuming that each filter or latent dimension maps onto a single coherent human-understandable concept. Under this view, labeling a unit simply requires inspecting the shared visual traits of its top-activating dataset examples. However, recent theoretical and empirical studies have revealed that superposition and capacity constraints induce widespread polysemanticity: individual neurons routinely activate across multiple, mutually disjoint concepts. Treating these units as single-concept detectors distorts interpretation faithfulness and conceals the actual circuit dynamics governing the model.
To decompose mixed concepts within individual neurons, researchers have turned to Sparse Autoencoders (SAEs) and attribution clustering. Yet existing attribution-clustering approaches suffer from two fundamental bottlenecks. First, they lack architectural generality, being largely engineered around CNN-specific propagation rules (such as CRP grammars) and unable to accommodate global attention mechanisms in modern Vision Transformers. Second, they lack scalability because they rely on manual heuristics—predominantly requiring practitioners to specify a fixed number of concept clusters \(K\) (e.g., \(K=2\)) per neuron. Tuning or hard-coding \(K\) across tens of thousands of neurons across different layers is labor-intensive and subjective. Furthermore, as network depth increases, feature space anisotropy causes attribution vectors to collapse into high mutual similarity, severely undermining clustering separation.
This paper tackles the challenge from the perspective of upstream computational tracing: when a polysemantic neuron fires on disparate concepts, it necessarily draws upon distinct combinations of predecessor neurons and attribution gradients. By explicitly mitigating anisotropy-induced similarity inflation and dynamically estimating per-neuron cohesion thresholds, concept separation can be fully automated. Core idea: trace activation origins via predecessor Input×Gradient attribution footprints, mitigate representation anisotropy via dimension normalization, and dynamically disentangle concept clusters using an adaptive KDE-based cohesion threshold without preset cluster count \(K\) across CNNs and Transformers.
Method¶
Overall Architecture¶
Given a pretrained vision backbone and a target polysemantic neuron, SPICE partitions its highly activating image collection into distinct, semantically cohesive concept subsets. For the target neuron, the framework first computes Input×Gradient attribution footprints from the immediately preceding layer and normalizes them along the neuron dimension to neutralize anisotropy. Next, an adaptive cohesion threshold estimator analyzes the empirical pairwise similarity distribution to identify the conceptual cohesion boundary. Finally, a progressive iterative clustering procedure discovers and isolates high-cohesion clusters one by one, allowing residual outlier samples to settle as fine-grained individual concepts.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Target Neuron Activating Samples<br/>CNN channel / ViT hidden dimension"] --> B["Attribution Footprint & Normalization<br/>Predecessor Input×Gradient + Anisotropy mitigation"]
B --> C["Adaptive Cohesion Threshold Estimation<br/>95th percentile & Bimodal KDE peak selection"]
C --> D["Progressive Iterative Clustering<br/>Dynamic k adjustment & Cohesion verification"]
D --> E["Disentangled Concept Clusters<br/>Separated concept groups + VLM annotation"]
Key Designs¶
1. Attribution Footprint Representation & Anisotropy Normalization: Tracing upstream pathways and overcoming deep representation collapse
Standard interpretation evaluates only input samples, ignoring internal computational origins. SPICE grounds the concept semantics of target neuron \(u^l\) in layer \(l\) by quantifying contributions from all predecessor neurons in layer \(l-1\). In CNNs, each channel serves as a neuron; in Transformers, each hidden dimension acts as a neuron. Because activations span multiple spatial positions or patch tokens, they are summed into a scalar activation \(\bar{a}_u(x) = \sum_{h,w} a_u(x)_{h,w}\). For each top-activating sample \(x_i \in \mathcal{D}_{u^l}\), the Input×Gradient attribution vector is formed by multiplying the predecessor layer's aggregated activation vector \(\bar{\mathbf{a}}^{l-1}(x_i)\) with the gradient of the target objective: $\(\boldsymbol{\phi}(x_i) = \bar{\mathbf{a}}^{l-1}(x_i) \odot \frac{\partial \bar{a}_{u^l}(x_i)}{\partial \bar{\mathbf{a}}^{l-1}(x_i)}\)$ Stacking these attribution vectors across all activating samples produces the footprint matrix \(\mathcal{A}_{u^l}\). To counteract feature space anisotropy—where deep layer representations become artificially aligned and pairwise similarities cluster close to 1—SPICE normalizes attribution matrices along the neuron dimension prior to clustering. This centers the comparison on relative attribution patterns rather than absolute magnitude shifts, preventing deep-layer cluster degradation.
2. Adaptive Cohesion Threshold Estimation: Unifying shallow unimodal and deep bimodal concept emergence
To eliminate manual tuning of a fixed cluster count \(K\), SPICE defines cluster quality via \(\text{Cohesion}(C)\), the mean pairwise cosine similarity within cluster \(C\). A global static threshold \(\tau\) fails catastrophically: shallow layers exhibit compact, unimodal long-tailed similarity distributions, whereas deep layers develop pronounced bimodal distributions as distinct semantic categories crystallize. A rigid static threshold either yields near-zero neuron coverage or causes over-fragmentation. SPICE dynamically determines the per-neuron cohesion threshold \(\tau\) by taking the maximum between two data-driven candidates: $\(\tau = \max \bigl( Q_{0.95}(S),\, p_2(S) \bigr)\)$ where \(S\) is the pairwise cosine similarity set of current footprints, \(Q_{0.95}(S)\) denotes the 95th percentile, and \(p_2(S)\) represents the second local maximum of a Gaussian Kernel Density Estimate (KDE) whose probability density exceeds 0.1. When the distribution is strictly unimodal, the threshold gracefully falls back to \(Q_{0.95}(S)\), ensuring a contextually grounded cohesion standard across the entire depth hierarchy.
3. Progressive Iterative Clustering: Fully data-driven concept discovery without preset hyperparameters
Once \(\tau\) is established, SPICE executes an iterative search to identify cohesive groupings without predetermining the number of concepts. The algorithm initializes with candidate cluster count \(k=2\) on working set \(\mathcal{A}'\). At each iteration, it partitions \(\mathcal{A}'\) into \(k\) clusters and assesses whether any cluster satisfies \(\text{Cohesion}(C) > \tau\). All clusters meeting this condition are accepted as discovered cohesive concepts, archived into \(C^*\), and subtracted from \(\mathcal{A}'\). The search parameter then steps back according to the number of resolved groups: \(k \leftarrow \max(2, k - (|C_{\text{cohesive}}| - 1))\). If no candidate cluster satisfies the cohesion threshold, \(k\) is incremented by 1 (\(k \leftarrow k + 1\)). The loop terminates when the remaining working set is smaller than \(k\), with residual singletons treated as individual concepts. This guarantees clean cluster boundaries while remaining entirely parameter-free.
Loss & Training¶
SPICE is a post-hoc, zero-training interpretability framework. It requires no fine-tuning parameters, auxiliary classifiers, or objective modifications on the underlying vision backbone. Footprints are computed using standard Input×Gradient (which achieves parity with Integrated Gradients and GradientSHAP at substantially lower compute cost). Multi-modal VLMs (e.g., GPT-4V) can optionally be prompted with exemplar images to produce descriptive textual labels for each discovered cluster.
Key Experimental Results¶
Main Results¶
Quantitative evaluations were performed on the ImageNet validation set across ResNet-50, ViT-B/16, DenseNet, and CLIP ViT-B. Concept Separability is measured as the ratio of mean intra-cluster CLIP visual embedding similarity to mean inter-cluster similarity (higher indicates sharper concept boundaries).
| Model / Layer | PURE* (CVPR'24) | LE (ICML'24) | CPE (CVPR'25) | SPICE (Ours) |
|---|---|---|---|---|
| ResNet-50 L2.1 | 1.018 ± 0.031 | 1.112 ± 0.122 | 0.996 ± 0.014 | 1.092 ± 0.056 |
| ResNet-50 L3.3 | 0.999 ± 0.012 | 1.108 ± 0.153 | 0.995 ± 0.007 | 1.160 ± 0.081 |
| ResNet-50 L4.2 | 1.009 ± 0.071 | 1.187 ± 0.221 | 0.998 ± 0.011 | 1.291 ± 0.224 |
| ViT-B B2 | 1.002 ± 0.010 | 1.199 ± 0.181 | 0.994 ± 0.009 | 1.137 ± 0.062 |
| ViT-B B6 | 1.012 ± 0.013 | 1.185 ± 0.153 | 0.994 ± 0.009 | 1.240 ± 0.089 |
| ViT-B B11 | 1.035 ± 0.025 | 1.357 ± 0.181 | 0.998 ± 0.018 | 1.577 ± 0.089 |
| DenseNet L2.1 | 1.012 ± 0.028 | 1.090 ± 0.158 | 0.994 ± 0.007 | 1.071 ± 0.059 |
| DenseNet L3.12 | 1.053 ± 0.054 | 1.062 ± 0.139 | 0.994 ± 0.005 | 1.114 ± 0.060 |
| DenseNet L4.16 | 1.025 ± 0.035 | 1.191 ± 0.116 | 0.994 ± 0.009 | 1.033 ± 0.032 |
| CLIP ViT-B B2 | 1.007 ± 0.018 | 1.088 ± 0.082 | 0.994 ± 0.008 | 1.090 ± 0.051 |
| CLIP ViT-B B6 | 1.008 ± 0.016 | 1.165 ± 0.148 | 0.997 ± 0.009 | 1.212 ± 0.090 |
| CLIP ViT-B B11 | 1.021 ± 0.040 | 1.244 ± 0.192 | 0.998 ± 0.013 | 1.410 ± 0.102 |
Across extreme activation regimes (Top, Middle, Bottom 100 samples), text-prior methods degrade severely in deeper, low-activation zones (e.g., LE plunges to 1.04 in ViT-B B10 Bottom), while SPICE sustains a dominant separability score of 1.41.
Ablation Study¶
The ablation study validates the adaptive thresholding strategy against static alternatives across 100 randomly sampled neurons, reporting Separability and Coverage (the count of neurons yielding at least one valid cohesive cluster):
| Threshold Variant | ResNet-50 L3.3 Score | ResNet-50 L3.3 Coverage | ViT-B B6 Score | ViT-B B6 Coverage | Note |
|---|---|---|---|---|---|
| Fixed \(K=2\) (no \(\tau\)) | 1.00 | 100 / 100 | 1.03 | 100 / 100 | Forced partition merges heterogeneous concepts |
| Static \(\tau \ge 0.9\) | 1.56 | 6 / 100 | 2.18 | 7 / 100 | High separability but catastrophic coverage drop |
| Static \(\tau \ge 0.8\) | 1.56 | 9 / 100 | 2.18 | 10 / 100 | Coverage remains below 10% |
| Static \(\tau \ge 0.7\) | 1.39 | 27 / 100 | 1.88 | 21 / 100 | Moderate coverage gain but lacks robustness |
| Adaptive \(\tau\) (SPICE) | 1.15 | 100 / 100 | 1.25 | 100 / 100 | Full 100% coverage with balanced high separability |
Key Findings¶
- Macro Architectural Inductive Biases: CNNs sustain a high number of fine-grained concepts uniformly across most intermediate layers via localized receptive fields, collapsing sharply only at the final global pooling stages. In contrast, Vision Transformers exhibit a characteristic U-shaped cohesion trajectory, fostering extensive conceptual diversification across intermediate layers before re-converging toward task targets in late blocks.
- Dual Pathways of Polysemanticity: Tracing upstream attribution circuits reveals distinct origin mechanics. Neurons encoding related concepts (e.g., dial telephones and pay phones in unit #24) share 42% of their top-50 predecessor attribution connections. Conversely, neurons encoding totally unrelated visual classes (e.g., burritos and baseballs in unit #157) exhibit largely disjoint upstream circuits with only 18% connection overlap, corroborating genuine superposition.
Highlights & Insights¶
- Heuristic-Free Concept Discovery: Eliminating fixed cluster count assumptions via adaptive KDE peak tracking provides a principled, fully automated tool for network-wide polysemanticity auditing.
- Countering Anisotropy via Normalization: Spotting that representation collapse in deep layers creates artificial attribution alignment, SPICE employs simple channel-wise normalization to restore true relative semantic geometry.
- Circuit-Grounded Disentanglement: Grounding concept extraction directly in backward attribution paths bridges individual neuron analysis with upstream computational subgraphs.
Limitations & Future Work¶
- Reliance on External Verification Embeddings: Separability relies on CLIP image embeddings, which emphasize high-level semantics over low-level visual textures, understating performance differences in early layers like B2.
- Computational Overhead of Iterative Clustering: Compared to single-pass fixed-\(K\) clustering, dynamic iterative evaluation over thousands of neurons entails higher execution latency.
- Scaling to Full Multilayer Circuit Discovery: The current formulation focuses on single-neuron predecessor attribution pairs; extending this recursively into multi-layer end-to-end circuit discovery remains an exciting open avenue.
Related Work & Insights¶
- vs PURE (CVPR 2024): PURE depends on CNN-specific CRP rules and manual cluster count specification \(K\); SPICE removes architecture dependencies, working seamlessly on both CNNs and ViTs while deriving \(K\) adaptively.
- vs CPE / LE (CVPR 2025 / ICML 2024): CPE and LE rely heavily on text concepts or VLM generation that suffer from vision-language misalignment and fail in low-activation regimes; SPICE discovers concepts purely from internal activations and attribution gradients.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ [Novel adaptive cohesion clustering without preset K, bridging CNN and Transformer interpretability]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluated across 4 architectures, multiple layers, full activation ranges, generative simulation, and circuit tracing]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation chain, elegant mathematical framing, and well-structured empirical validation]
- Value: ⭐⭐⭐⭐⭐ [Provides a foundational, scalable diagnostic tool for mechanistic interpretability and representation research]