Skip to content

Deep Minds and Shallow Probes

Conference: NeurIPS2026
arXiv: 2605.11448
Area: Interpretability
Keywords: affine invariance, polynomial probes, low-rank tensors, probe-visible quotient, cross-model transfer

TL;DR

The paper derives a polynomial hierarchy of shallow probes from coordinate symmetries at the final readout, uses low-rank CP probes to read interaction concepts and probe-visible quotients to transfer concept readouts, and improves cross-token agreement AUROC by 16.8–20.0 percentage points while explicitly separating transfer accuracy from concept coverage.

Background & Motivation

A probe is a small predictor trained on frozen neural representations, whose success is often taken as evidence that a property is accessible in those representations. Linear probes are tractable and capacity-limited, but read only linearly separable structure; replacing them with powerful nonlinear networks can introduce task-solving capacity into the probe itself. The important question is not whether nonlinearity is stronger, but which nonlinear readouts still support credible claims about the original representation.

A second difficulty is that coordinates are not unique. Two models implementing the same output computation need not use the same hidden coordinates. If a probe restriction holds only in a particular basis, success or failure can reflect that basis rather than the presence of a concept. Starting from the final linear or affine readout, the authors characterize coordinate freedom under nondegeneracy assumptions and require the probe family to preserve realizable scores under coordinate changes. This is a requirement on the function family, not a requirement that trained parameters remain unchanged.

Cross-model transfer faces a related problem: full hidden states contain many directions unrelated to the monitored concepts, and reconstructing them need not be the necessary objective. The paper formalizes shallowness as a finite-dimensional linear score space and shared structure as equality of concept function families, rather than treating architectural similarity or output similarity on a finite test set as sufficient. Core idea: constrain probe families by representation symmetries and transfer objects by the quotient visible to the probes, thereby jointly addressing readout capacity, coordinate stability, and coverage boundaries.

Method

Overall Architecture

The paper develops two related but distinct arguments. The first derives affine coordinate changes from final-readout equivalence and classifies shallow score spaces satisfying exact affine closure. The second derives a probe-visible quotient from a shared linear concept family and realizes this abstract object as SVD coordinates of the probe weight matrix. The former explains polynomial and low-rank interactions; the latter explains what transfer should preserve.

For interaction readout, primitive linear probes are first trained with source concept labels, and a quadratic composition head is trained on their low-dimensional scores. For transfer, the source bank first determines visible directions, after which unlabeled paired activations from the two models on identical inputs fit a map from target hidden states to source quotient coordinates. Inference needs only target activations, this map, and the source readout parameters.

These theoretical objects are not a sequence of neural modules, so the theorem list is not drawn as a network architecture. Supervised concept-probe training must be distinguished from alignment fitting without target labels: the latter still requires paired inputs and activations from both models, and still depends on the source concept labels used by the former.

Key Designs

1. Affine-stable score spaces: turn coordinate freedom into a function-family constraint

For linear readouts, Theorem 2.1 assumes equal hidden dimensions, source representations spanning the entire space, both readout matrices having full column rank, and exactly identical output scores for every input in the comparison domain. It then obtains a unique invertible linear hidden-coordinate transformation. A biased readout can be handled by appending a constant coordinate, permitting hidden-space translation under the corresponding nondegeneracy assumptions. Equality of softmax probabilities additionally permits a shared logit shift across classes, which must not be confused with hidden-state translation. The result is anchored at the final readout interface; it does not automatically cover full hidden spaces of different dimensions or arbitrary intermediate layers.

After a coordinate change, probe parameters should change as well so that each input retains its score. The authors therefore require closure of the score family under precomposition by all invertible affine maps. The exact quantifiers of Theorem 2.3 are that a finite-dimensional linear subspace of continuous real-valued functions, exactly closed under the full affine group, must be either zero or a complete bounded-degree polynomial space.

\[ V\subset C(\mathbb{R}^{n}),\quad \dim V<\infty,\quad f\in V,\ g\in\operatorname{Aff}(n)\Rightarrow f\circ g\in V \quad\Longrightarrow\quad V=\{0\}\ \text{or}\ V=\mathcal{P}_{\leq\ell}(\mathbb{R}^{n}),\qquad \ell\geq 0. \]

Here, a linear space means that functions can be linearly combined, not that each function is linear. Degree 0 includes constants, and degree 1 includes linear scores with intercepts. The proof uses translation closure to obtain smoothness and derivative closure, dilations to eliminate nonzero exponential modes, and the general linear group to fill out the entire highest-degree homogeneous space; differentiation fills in lower degrees. It explains why the hierarchy is complete, rather than merely showing that polynomials are permissible.

The theorem does not imply that every finite-parameter MLP is polynomial: a nonlinear finite-parameter family generally is not a finite-dimensional linear function space. Nor does it prove that approximate closure on finite samples implies approximate polynomiality. The vector-valued corollary guarantees only a uniform degree bound on output coordinates, not all possible output directions. For classification, the classified mathematical object is the continuous score, not discontinuous thresholded labels or sigmoid probabilities.

2. Affine-completed CP probes: compress quadratic readouts with low-rank interactions

The coefficient count of complete polynomial spaces grows rapidly with dimension and degree, making full quadratic probes particularly expensive on language-model hidden states. Canonical Polyadic (CP) probes express higher-order terms as sums of a few products of linear forms. An invertible linear transformation changes the factor vectors and preserves the specified rank bound. In contrast, mixing coordinates can expand a few nonzero monomials into many terms, so monomial sparsity lacks the same geometric stability.

Homogeneous CP alone guarantees stability only under the general linear group. Translation creates lower-degree terms, so the authors introduce factor biases and all necessary lower-degree completion terms. The experimental quadratic instance is:

\[ f(h)=\sum_{r=1}^{R}\alpha_r\bigl(\langle u_r,h\rangle+a_r\bigr)\bigl(\langle v_r,h\rangle+b_r\bigr)+\langle w,h\rangle+c. \]

Products of two affine factors carry the interaction, while the final linear term and constant provide affine completion. This fixed-rank structured family is not the complete linear function space classified by Theorem 2.3; it is a coordinate-stable structured subfamily inside it. Adding distinct low-rank functions can require a higher rank. Complete-space classification and low-rank compression are therefore compatible, but CP rank 1 is not sufficient for every quadratic concept.

The cross-token experiment trains 30 primitive probes for part of speech, number, tense, punctuation, and dependency relations on Universal Dependencies English-EWT, concatenates subject and verb scores into 60 dimensions, and fits composition heads. Multiplication between number features expresses agreement where a linear sum struggles. Both classes use cross-sentence pairing to reduce sentence-identity confounds. Because the scores come from linear probes, quadratic score composition remains quadratic in the original concatenated hidden states, while the 60-dimensional composition is more manageable than operating on full activations.

3. Probe-visible quotients: align only directions actually read by the concept bank

Given a hidden space and a linear probe space, directions to which every probe is insensitive form the invisible subspace. Hidden vectors differing only along such directions produce identical scores for the entire bank and should count as the same readable state. This is not maximum-variance compression by PCA, but task-relevant compression defined by trained concept readouts.

\[ K(V)=\bigcap_{\ell\in V}\ker\ell,\qquad Z(V)=H/K(V),\qquad \mathcal{C}_i=\{\ell\circ h_i:\ell\in V_i\}. \]

Theorem 3.3 requires both models to realize the same finite-dimensional concept function family and their evaluation maps to assign different concepts to different probes: no nonzero probe vanishes on every represented input. Under these assumptions, both quotients are canonically isomorphic to the dual of the common concept family. Hidden dimensions may differ because the shared object consists of visible concepts, not full hidden states.

\[ \mathcal{C}_1=\mathcal{C}_2=\mathcal{C},\quad E_i:\ell\mapsto\ell\circ h_i\ \text{injective} \quad\Longrightarrow\quad H_1/K(V_1)\cong\mathcal{C}^{*}\cong H_2/K(V_2). \]

This sharing assumption is more specific than similar output behavior: behavioral agreement alone does not establish concept-family equality for an arbitrary chosen bank. The proof treats a hidden vector as an evaluator of concepts, with the invisible subspace exactly its kernel; removing that kernel gives the common evaluation space. Canonicality does not mean that the numerical SVD basis is unique.

In implementation, source linear-probe weights are stacked row-wise, sufficiently large right singular directions are retained, and their transpose defines quotient coordinates. Intercepts are retained separately, and real activations are centered using training means. The source weight-matrix kernel describes directions in centered representation space, rather than requiring zero probe intercepts.

\[ W=U\Sigma R^{\top},\qquad Q=R_k^{\top},\qquad k=\#\{j:\sigma_j>10^{-3}\sigma_1\},\qquad z=Qh. \]

The numerical threshold means that implementation retains a truncated quotient rather than every exactly nonzero direction. Target-side probes with the same labels need not be trained first: Appendix E.5 directly predicts source quotient coordinates from full target activations. This relates to the ideal theorem's two-sided quotient isomorphism, but the experiment must not be described as having trained both banks with target labels.

4. Coverage and conditioning diagnostics: separate transferability from a single accuracy score

Quotient transfer fits a Ridge map on training-split paired activations and pulls source concept readouts back to the target space. A new concept direction within the bank span is expressible through quotient coordinates. When much of its direction is projected away, transfer failure can reflect missing coverage rather than an alignment optimizer failing to find parameters. The in-span fraction (ISF) measures the fraction of a source concept weight's squared norm inside the source visible space.

\[ \operatorname{ISF}(w)=\frac{\|Qw\|_2^2}{\|w\|_2^2},\qquad w\neq 0. \]

ISF needs no target labels, but does require an existing source concept direction. It is geometric coverage in fixed coordinates with a Euclidean metric, not a certificate of semantic correctness. On real data, correlations with bank concepts can make low-ISF concepts predictable, while moderate ISF may carry little useful discriminative signal. Low ISF therefore must not be presented as necessarily precluding classification. Abstaining from using a transferred monitor with insufficient coverage is a conservative information-boundary decision, not proof that the original model entirely lacks that concept.

The Appendix C.1 error bound controls coordinate error by concept-family score mismatch divided by the smallest source-bank visible singular value. This assumes a given score-matching condition; it does not automatically guarantee finite-sample Ridge generalization or target AUROC. Appendix C.3 explains how both banks' conditioning affects the induced transport's stability, while classification agreement additionally needs sufficient source-score margin.

Probe count cannot substitute for coverage quality. Controlled redundancy experiments show that replacing independent directions shrinks the effective visible span and degrades transfer; merely appending near-duplicates lets SVD thresholding discard them without changing performance. Thresholding and Ridge absorb some numerical ill-conditioning, so effective dimension, singular spectrum, and concept diversity should be considered together rather than ranking banks solely by condition number.

A Worked Example

In the synthetic transfer experiment, shared latent concepts have dimension 8, source hidden states have dimension 64, and target states have dimension 128. Both sides additionally contain independent nuisance and observation noise. Five source primitive concepts train probes that typically produce a five-dimensional visible space, after which 5,000 activation pairs without target labels fit the map from target states to source quotient coordinates.

One held-out concept is a linear combination of the five primitive directions, whereas another is orthogonal to them in latent space. The former remains expressible through source quotient coordinates. For the latter, quotient projection removes the directions it needs, producing near-chance classification. Even if full-state OLS transfers the latter accurately, this establishes broader information preservation by the full activation map, not direct coverage by those five probes.

In original Table 2, increasing nuisance dimension from 0 to 56 changes quotient in-span balanced accuracy from 0.997 to 0.991 and out-of-span accuracy from 0.484 to 0.519. Full-state OLS remains at 0.995 and 0.989 on the out-of-span concept. The comparison illustrates a trade-off between selectivity and coverage transparency versus broader transfer, not an accuracy failure of full-state OLS.

Loss & Training

  • Primitive classification probes use logistic regression with an intercept, C=1.0, and max_iter=5000. Model activations are frozen; validation and test centering use the training mean.
  • CP classification heads optimize binary cross-entropy using L-BFGS with strong-Wolfe line search, with 30–50 random restarts depending on the experiment and validation-AUROC selection. Synthetic regression CP uses squared loss and Adam with learning rate \(10^{-3}\), up to 500 epochs, patience 30, and 2–3 restarts.
  • Quotient alignment uses multi-output Ridge with regularization \(10^{-4}\). Unlabeled paired alignment uses training splits; target labels are used only for testing on splits disjoint from alignment data. Zero-label means zero target labels.
  • Most experiments use five seeds. The cross-token main table explicitly reports means and standard deviations over 10 resamples, whereas monitor transfer reports bootstrap interval half-widths. Different tables' error bars must not be described as one statistical quantity.

Key Experimental Results

Main Results

Table 1: Cross-token subject–verb number agreement, AUROC. Original Table 1; a 60-dimensional primitive-probe score space, with means and standard deviations over 10 resamples. Gain is the CP-1 minus linear-head AUROC difference in percentage points, not an accuracy improvement.

Model Linear head Full quadratic head CP-1 CP-1 gain
Pythia-70m 0.733 ± 0.005 0.819 ± 0.013 0.901 ± 0.004 16.8 pp
Pythia-160m 0.740 ± 0.003 0.855 ± 0.007 0.930 ± 0.002 19.0 pp
Pythia-410m 0.755 ± 0.004 0.872 ± 0.011 0.955 ± 0.003 20.0 pp

Full quadratic outperforms linear but trails rank 1. Appendix Table 21 reports full-quadratic train–test gaps of 0.176, 0.145, and 0.128 across the three scales, compared with 0.026, 0.029, and 0.012 for CP-1. This supports capacity-driven overfitting, but is not a strict causal proof about optimization.

Table 2: Selected concepts in cross-model monitor transfer, AUROC. Selected from original Table 3, retaining only general sentiment and aggregate moderation tasks. The source is Qwen-2.5-7B-Instruct with an 11-probe bank and condition number approximately 5.8. Target ± values are 95% bootstrap interval half-widths; source values are in-model performance.

Concept Source Qwen-3B Qwen-14B Qwen-Coder Mistral Random initialization
Sentiment 0.975 0.966 ± 0.01 0.974 ± 0.01 0.967 ± 0.01 0.972 ± 0.01 0.511 ± 0.04
Moderation 0.890 0.836 ± 0.05 0.857 ± 0.04 0.867 ± 0.04 0.820 ± 0.05 0.512 ± 0.07

Cross-architecture results establish partial concept portability, not equal reliability across the entire bank. In the complete original Table 3, the five core concepts span 0.669–0.972 on Mistral, and the weakest direction still requires additional validation. The random-initialization control approaches chance on most concepts, but not every metric is exactly 0.5.

Ablation Study

Table 3: Effect of nuisance dimension on synthetic transfer, balanced accuracy. Original Table 2; shared latent dimension 8 and source/target dimensions 64/128, with means and standard deviations over five seeds. In-span and out-of-span refer to the concept span of source primitive probes.

Alignment method In-span, nuisance 0 In-span, nuisance 56 Out-of-span, nuisance 0 Out-of-span, nuisance 56
Quotient Ridge 0.997 ± 0.002 0.991 ± 0.004 0.484 ± 0.020 0.519 ± 0.023
Full-state OLS 0.996 ± 0.001 0.990 ± 0.003 0.995 ± 0.001 0.989 ± 0.006
PCA projection 0.996 ± 0.001 0.573 ± 0.036 0.996 ± 0.001 0.585 ± 0.022
Random projection 0.927 ± 0.029 0.574 ± 0.018 0.903 ± 0.065 0.567 ± 0.027

The quotient remains stable for in-span concepts, while PCA is affected by high-variance nuisance. Full-state OLS is accurate for both types of concept; what it lacks is a probe-bank coverage diagnostic. Near-chance out-of-span quotient performance is an expected consequence of projection, not evidence of a validated automatic rejection protocol.

Key Findings

  • Coordinate stability differs from learning-procedure stability. Appendix F.2 reports maximum score error \(1.14\times10^{-11}\) for analytical full-quadratic transport without retraining; under the same transforms, a sparse probe retaining its original 50 monomial slots has \(R^2=-18.3\pm3.5\). This isolates function-family closure, not an inability to retrain all sparse models.
  • Alignment text must activate the relevant concepts. In Appendix Table 27, SST-2-only alignment yields sentiment AUROC 0.966, while two aggregate safety concepts reach only 0.596 and 0.523. No target labels does not mean that arbitrary unlabeled text suffices.
  • In controlled redundancy replacement, Qwen-3B mean AUROC falls from 0.897 to 0.855; appending duplicate directions leaves it at 0.897. Replacing effective visible directions matters more than increasing probe count itself.
  • Polynomial degree, exact label-representation degree, and polynomial-threshold decision degree must not be conflated. Appendix F.1 explicitly assigns threshold decision degree 1 to AND, three-way AND, and majority; a high-degree product expression for labels does not establish the need for a high-degree classification head.

Highlights & Insights

  • Probe failure has more explanations than a model lacking a concept. Insufficient function-family capacity, coordinate-dependent restrictions, missing bank coverage, and transfer mismatch should be separated rather than collapsed into one accuracy score.
  • Low rank is not arbitrary parameter compression; it organizes interactions through linear forms. With lower-degree completion, coordinate changes can be absorbed by factor parameters, respecting readout symmetry better than fixed monomial slots.
  • A quotient preserves the information a bank commits to reading, not all information in a model. Coverage transparency and task accuracy should be evaluated as distinct objectives when comparing with full-state transfer.

Limitations & Future Work

  • Exact affine symmetry is derived at the final readout under nondegeneracy assumptions. Intermediate layers may have weaker or approximate symmetries; the paper does not prove that every layer admits the same full affine group.
  • Classification of probe spaces requires continuous scalar scores, a finite-dimensional linear space, and exact closure. Approximate closure, finite-data fitting, and particular neural probe architectures require additional analysis rather than direct application of the theorem.
  • ISF measures geometric overlap of source weights and depends on the chosen Euclidean coordinates; its numerical value need not be invariant under arbitrary nonorthogonal changes. It does not replace semantic coverage, out-of-distribution robustness, or target-side validation.
  • Cross-model transfer depends on source labels and unlabeled paired activations covering the relevant directions. Bank composition, centering, and alignment-pool deduplication affect results, so different appendix configurations must not be treated as repetitions of one identical experiment.
  • Statistical descriptions differ in the source: Table 1 specifies 10 resamples and Table 3 bootstrap half-widths, whereas the checklist summarizes Tables 1–4 as five-seed standard deviations. This note follows the specific table captions and retains that inconsistency.
  • High-level aggregate monitoring and behavioral evaluation establish only partial transfer within their experimental scope, not a deployment guarantee. Future work could study data-metric coverage diagnostics, uncertainty-calibrated abstention, and banks jointly optimized for coverage and conditioning.
  • vs linear probes and control tasks: Linear probes limit readout capacity, and control tasks test whether probes learn the task themselves. The paper adds coordinate stability as a complementary criterion, which does not replace those controls.
  • vs structural probes: Earlier syntax probes use quadratic geometry. This work explains a complete degree hierarchy through affine symmetry and controls interaction-readout cost using affine-completed low-rank structure.
  • vs model stitching and representation similarity: Full activation alignment asks whether hidden states correspond; this paper asks whether a specified concept family corresponds in a common visible space. PCA chooses directions by variance, while quotient construction chooses them by probe weights; they serve different objectives.
  • Research lead: Preserve existing source-concept coverage while selecting new probes by the visible singular spectrum and incremental directional gain, then test ISF stability across domains. This is a follow-up idea motivated by Appendix C and the redundancy ablation, not an algorithm already validated in the paper.

Rating

  • Novelty: 5/5 — Unifies probe-family classification and cross-model visible spaces as representation geometry problems.
  • Experimental Thoroughness: 4/5 — Synthetic, linguistic-interaction, and cross-model studies complement one another, but coverage diagnostics have real-data exceptions.
  • Writing Quality: 4/5 — The theory–experiment connection is clear; statistical conventions and configurations require careful separation.
  • Value: 5/5 — Provides reusable boundaries for choosing nonlinear probes and interpreting transfer failures.