Skip to content

๐Ÿ”ฌ Interpretability

๐Ÿง  NeurIPS2026 ยท 10 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (30) ยท ๐Ÿ“ท CVPR2026 (34) ยท ๐Ÿ”ฌ ICLR2026 (195) ยท ๐Ÿ’ฌ ACL2026 (63) ยท ๐Ÿงช ICML2026 (91) ยท ๐Ÿค– AAAI2026 (37)

๐Ÿ”ฅ Top topics: Alignment/RLHF ร—2 ยท LLM ร—2

Can Circuit Alignment Predict OOD Generalization?

The paper defines Circuit Alignment Score (CAS) as same-class cross-domain circuit similarity minus cross-class circuit similarity, achieving a mean Spearman correlation of 0.88 on PACS for source-domain model selection without target data, but its OOD ranking consistency requires an additional monotonicity assumption and is not a fully data-free weight diagnostic.

Deep Minds and Shallow Probes

The paper derives a polynomial hierarchy of shallow probes from coordinate symmetries at the final readout, uses low-rank CP probes to read interaction concepts and probe-visible quotients to transfer concept readouts, and improves cross-token agreement AUROC by 16.8โ€“20.0 percentage points while explicitly separating transfer accuracy from concept coverage.

Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations

Using a shared generator to separate latent geometric changes from generator-concept changes, the paper compares prediction stability and smoothed-classifier certificates for standard and concept bottleneck models, finding that interpretability offers no uniform robustness advantage but changes where sensitivity appears; the conclusions remain conditional on generator-native inputs and specific task settings.

Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

After freezing an LLM, the method calibrates an attention head and selects answers using pre-RoPE dot products between the final prompt query and option-ending keys, exposing a gap between internal selection and final output; particular zero-shot configurations gain 27.4 and 49.8 percentage points on HellaSwag and HaluDialogue, respectively, but not every model benefits.

Parameter symmetries determine representational geometry in overparameterized nonlinear networks

For one-hidden-layer nonlinear networks, this paper reduces parameter symmetries preserving the global function to feature addition, duplication, and scaling, proves that they can substantially reshape representational geometry, and gives sufficient conditions for minimum-norm selection to restore geometric identifiability.

PersonaManifold: Revealing and Exploiting Curved Geometry in LLM Persona Representations

PersonaManifold models LLM persona activations as a curved, anisotropic low-dimensional manifold, using graph shortest paths weighted by local metrics to measure similarity and interpolate personas; combined BST triplet consistency improves over Euclidean distance by 5.1โ€“6.1 percentage points across three 7โ€“8B models, primarily benefiting high-deviation persona pairs.

PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders

PULSE identifies SAE features associated with demonstration utility from a small labeled discovery set, then uses them for complete-set ranking and cacheable per-example retrieval, improving selection across six text tasks while providing associational rather than causal-intervention evidence.

SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery

SMILE identifies decomposable structure in data, fits a network with fixed symbolic activations, and recovers compact expressions through pruning, constant refitting, and gradient-based rounding, achieving strong symbolic recovery under high noise in SRBench without universally leading noise-free recovery or prediction accuracy.

The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

The paper uses standard knowledge distillation as an analytical tool to transfer an event teacher's predictions into an RGB student, finding joint improvements in tolerance to color changes, shape preference, and high-frequency-band noise, but with lower clean accuracy and sensitivity to disrupted geometric continuity rather than comprehensive robustness.

Witness Overlap: Directional Provenance Inside Open-Weight Model Families

Within known same-family open-weight models with aligned parameters, Witness Overlap introduces a third checkpoint, compares each candidate endpoint's update overlap toward the target and witness, and predicts the lower-overlap endpoint as the parent; Frobenius cosine achieves 95.3% orientation accuracy over 1,542 single-witness triplets from 16 LLM families.