Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models¶
Conference: NeurIPS 2026
arXiv: 2609.34348
Area: Optimization & Theory
Keywords: training order, Lie bracket, token attribution, path dependence, parametric memory
TL;DR¶
The paper decomposes the noncommutative residue of two training updates into a signed token readout, localizing and intervening on training-order differences under local SGD conditions; targeted interventions close the held-out loss gap by a median 32.0% on Qwen-3-4B, while paired-endpoint order assignment reaches 66/72 = 91.7%.
Background & Motivation¶
Language model post-training usually encounters several data sources sequentially. Even with identical data and total exposure, training on code and then news need not be equivalent to reversing the order. This is not merely a consequence of random batches: the first update changes the model's location, thereby changing the second source's gradient. Comparing final losses or benchmark scores establishes an ordering effect but does not identify which outputs carry it, or separate interpretable source interactions from numerical noise.
Prior work describes the noncommutativity of two gradient fields with a Lie bracket and projects it onto an evaluation-loss gradient to predict the better order. A scalar still cannot identify the support of that difference. Another line of training-history research detects recently learned information in activations, studying representational recency rather than the antisymmetric difference between equal-exposure training paths. This paper seeks an object that can be localized, intervened on, and read from paired parameter endpoints, without treating it as a memory store capable of retrieving individual examples.
A central difficulty is that visible sparsity need not originate in training order: random parameter directions passed through the same logit readout can also produce heavy-tailed distributions. The paper therefore examines not only whether a few tokens carry substantial readout mass, but also whether they align with actual endpoint differences, remain stable under batch resampling, and provide stronger intervention targets than frequency-matched tokens with near-zero scores. Core Idea: decompose the Lie bracket of two updates into vocabulary coordinates through error-weighted logit directional derivatives, then test localization, causal effects, and readability with support alignment, token interventions, and paired-endpoint assignment.
Method¶
Overall Architecture¶
The inputs are a base model, two specified training sources, and a held-out evaluation slice. The procedure first computes the source-pair Lie bracket at the base, then reads it into signed token scores at the shared first-order reference. Actual training paths subsequently test predicted support, targeted interventions, and the order trace in paired endpoints. This is a diagnostic workflow, not a new language model architecture or an inference-time memory retrieval system.
The paper separates forecasting an ordering difference from the base model and checking or correcting that difference after obtaining actual endpoints. The token map can be computed before running either order, but sign filtering in the main intervention experiment uses the measured baseline gap; paired-endpoint assignment explicitly requires both sets of weights. These experiments cannot be merged into a general training-history recovery system requiring only one final model.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
I["Base, two sources,<br/>evaluation slice"] --> A["Local Bracket Estimation"]
A --> B["Error-Weighted Token Readout"]
B --> C["Support Validation and Intervention"]
C --> D["Paired-Endpoint Readout"]
T["Fixed-batch training<br/>in both orders"] -.->|Measured gap and endpoints| C
T -.-> D
D --> O["Path-local diagnostic report"]
Solid arrows indicate diagnostic processing order; dashed arrows supply validation material produced by actual training. There is no generation-time storage, retrieval, or supervision branch.
Key Designs¶
1. Local Bracket Estimation: isolate the residue that reverses when order is swapped
The two sources' gradients and Hessians are evaluated at the base model. Taking one SGD step on source A followed by one on source B shares its first-order update with the reverse path; their difference starts at second order. The Lie bracket subtracts two cross-curvature responses: how A's update changes B's gradient, minus how B's update changes A's gradient. Swapping source names reverses the direction, so it describes an antisymmetric path residue rather than ordinary learning from either source.
This identity and scalar order prediction are inherited, not introduced here. The new emphasis is on decomposing the residue, validating its support, and intervening on it. Computation uses two exact Hessian–vector products without materializing a Hessian. These calculations use fp32; bf16 model storage does not make precision irrelevant to second-order measurement.
The shared reference is \(\theta_{\rm ref}=\theta_0-\eta(g_A+g_B)\). Evaluating the target-loss gradient there reduces contamination from shared first-order drift. The theory requires sufficiently smooth local losses and a step size in an empirically calibrated BCH-local window: the second-order signal must clear numerical resolution without higher-order terms becoming dominant. This single-step expansion is not an exact identity for long training runs, nor does it transfer directly to AdamW with moment buffers and a step counter.
2. Error-Weighted Token Readout: decompose an aggregate gap into signed vocabulary coordinates
A parameter direction is difficult to interpret directly. The paper computes the logit change along the bracket and multiplies it by cross-entropy error in each token coordinate: predicted probability minus the label indicator. A logit change becomes a loss contribution only in conjunction with the current prediction and label. A token can therefore contribute through its predicted probability even if it never appears as a label.
Here \(e=\operatorname{softmax}(z)-\operatorname{onehot}(y)\) and \(\delta z(x;v)=J_\theta z(x;\theta_{\rm ref})v\). A positive score means that this coordinate contributes positively to the predicted AB-minus-BA loss gap; a negative score means the opposite. This is attribution aggregated over an evaluation slice, not a label-token NLL difference computed separately for an arbitrary sentence, and not circuit localization.
The identity is exact for the true Jacobian–vector product, but the implemented logit directional derivative uses central finite differences. Exact HVPs and approximate logit JVPs must be distinguished: the former avoid an explicit Hessian, while the latter retain truncation error. SFT uses \(\epsilon=1.0\); corrected DPO identity checks use \(\epsilon=0.1\).
Experimental readouts therefore satisfy the analytic sum identity only approximately. Individual token scores also depend on native logit coordinates: a shared position-wise logit shift leaves the sum invariant, but not necessarily each attribution. This prevents interpreting the token map as a parameterization-independent absolute explanation.
3. Support Validation and Intervention: distinguish heavy-tailed readouts from order-specific support
The paper examines concentration by absolute score but does not treat high Gini as sufficient evidence of specificity. The substantive test passes measured endpoint displacements, brackets recomputed on independent batches, and norm-matched random directions through the same readout operator, then compares their high-mass support with the original bracket's support. Random directions may also be sparse but should not recover a particular source pair's support equally well. The controls thus test source specificity rather than merely the steepness of a concentration curve.
The causal test selects the ten largest absolute scores whose signs match the measured baseline gap. These “harmful tokens” are coordinates contributing in the same direction as a particular ordering gap; the term does not mean that the token content is harmful. During source A's update, Qwen-3-4B downweights cross-entropy at these label positions and compensates by upweighting other positions, preserving the mean position weight. This avoids mistaking a smaller overall learning rate for successful targeted intervention.
Here \(p\) is the fraction of training positions whose targets belong to the selected token set. The main experiment uses \(\alpha=0.1\), leaving selected label positions with 10% of their original weight. Controls use label-frequency-matched tokens with near-zero bracket scores to keep intervention doses comparable. The small-signal Qwen-2.5-1.5B experiment instead scales learning rates for corresponding output-layer rows to reduce noise amplified through hidden-state pathways; its effect size is not directly comparable with that of the other model.
Closure measures whether the gap between the orders decreases, not whether both models improve in absolute loss. Its upper bound is 1, with no lower bound; when the baseline gap approaches numerical noise, the denominator makes the ratio unstable.
4. Paired-Endpoint Readout: remove shared drift before assigning order
Projecting a single endpoint's displacement from the reference onto the bracket mixes in the second-order symmetric term shared by both paths. Even when the antisymmetric component exists theoretically, both single-endpoint scores can have the same sign. The paper therefore uses the difference between actual endpoints rather than generalizing a single-model sign test across scales.
When the first endpoint comes from AB, the leading term is positive; reversing endpoint presentation reverses the sign. Paired subtraction removes shared drift, allowing order assignment locally when the leading term dominates finite-step, sampling, and numerical errors. This requires a base model, the candidate source pair, and endpoints for both alternative orders. It does not recover a complete training history from an unknown single model.
A Worked Example¶
For the code–news source pair, diagnosis forms the difference between the two cross-HVPs at the base and generates a token map on the held-out slice. Different references and evaluation slices can yield different surface tokens: the main interpretability experiment emphasizes proper nouns and domain markers, whereas another base-reference readout surfaces code syntax markers. Token-by-token agreement should not be imposed across these protocols.
In the specified Qwen-3-4B intervention protocol, ten high-scoring tokens aligned with the baseline gap are used for reweighting, followed by retraining and held-out gap comparison, producing a median closure of 32.0%. This establishes a causal effect of the edit, not that these tokens explain exactly 32.0% of the original gap or that the remainder is an independently identified higher-order mechanism. The leading bracket after editing still predicts 71% of the original gap, but this statistic and median closure use different aggregation rules and cannot be subtracted to construct a component decomposition.
Loss & Training¶
The main experiments use SFT on fixed two-source paths from The Pile's code, news, math, legal, biomedical, and wikipedia partitions, with sequence length 512. Support experiments cover all 15 source pairs with three seeds each; main interventions use three fixed pairs with thirty seeds each. Paired order assignment uses one SGD update per source on six pairs with three seeds each, rather than arbitrary long training trajectories.
Matched-path parameter correction follows \(\theta_{AB}-\eta^2b_{AB}=\theta_{BA}+O(\eta^3)\) and aims to approach the reverse-order endpoint. Its iterative version recomputes the bracket and determines the projection step using the observed reverse endpoint, so approximately 90% iterative closure is not a prediction result obtained without a target endpoint. A separate bracket-only update constructed at the shared reference conditionally beats both orders: symmetric drift must increase target loss and the bracket projection sign must be estimated correctly. This is not an unconditional training-improvement guarantee.
DPO treats different preference-data sources as independent sources, not the chosen/rejected terms within one preference loss. The GRPO-style extension freezes rollouts shared by both orders and uses analytic rewards with a KL-regularized surrogate; clipping is inactive in the tests. This is not a general result for online GRPO. AdamW instead replays updates in an augmented state containing parameters, moments, and the counter. Under its fixed clock, order effects can start at first order in step size, so SGD's second-order bracket scaling cannot be reused.
Key Experimental Results¶
Main Results¶
The following table is based on paper Table 3. Closure and confidence intervals are expressed as percentages; the mean is 5/95 winsorized, and the intervals are not median intervals. Qwen-3-4B uses all 90 held-out trials; Qwen-2.5-1.5B retains only the predefined 28/90 fp32 trials above the \(10^{-4}\) denominator floor.
| Model and protocol | Token set | Median closure | Winsorized mean | Mean 95% CI | Trials with smaller gap |
|---|---|---|---|---|---|
| Qwen-3-4B, loss reweighting | High-score, gap-aligned | +32.0% | +31.2% | [+21.4%, +39.7%] | 68/90 |
| Qwen-3-4B, loss reweighting | Frequency-matched, near-zero score | +0.2% | +1.2% | [−5.8%, +8.7%] | 45/90 |
| Qwen-2.5-1.5B, output-row LR scaling | High-score, gap-aligned | +47.9% | +48.6% | [+40.5%, +57.1%] | 27/28 |
| Qwen-2.5-1.5B, output-row LR scaling | Frequency-matched, near-zero score | −0.003% | −0.006% | [−0.03%, +0.02%] | 11/28 |
Paired-endpoint assignment is a separate readability test. Paper Table 4 uses 18 pair–seed units per model, counting each model once; chance is 50%.
| Model and precision | Correct order assignment | Wilson 95% CI |
|---|---|---|
| Llama-3.2-1B, fp32 | 14/18 = 77.8% | [55%, 91%] |
| Qwen-2.5-1.5B, fp32 | 16/18 = 88.9% | [67%, 97%] |
| Qwen-3-4B, bf16 | 18/18 = 100.0% | [82%, 100%] |
| Llama-3.1-8B, bf16 | 18/18 = 100.0% | [82%, 100%] |
| Combined | 66/72 = 91.7% | [83%, 96%] |
Ablation Study¶
The following table uses the fp32 control subset from Table 2 and Appendix R: three pairs with three seeds each per model. Both support-overlap columns compare with the original bracket's top-20, not directly across models' vocabularies.
| Model | Vocabulary fraction carrying 80% of absolute readout mass, mean/median | Measured endpoint support overlap | Independent-batch bracket support overlap | Norm-matched random overlap |
|---|---|---|---|---|
| Llama-3.2-1B | 1.47% / 1.47% | 99% | 97% | 35% |
| Qwen-2.5-1.5B | 0.99% / 1.16% | 99% | 97% | 39–40% |
| Qwen-3-4B | 0.30% / 0.042% | 82% | 93% | 36–49% |
Both HVP terms contribute to useful attribution: selecting intervention tokens with the full bracket on Qwen-3-4B gives a paired Cohen's effect size of \(d=0.53\); using only \(H_Bg_A\) gives \(d=0.25\). This shows more directly than concentration alone that subtracting the two cross-responses is substantive, not notation that can be reduced arbitrarily to one term.
Key Findings¶
- Sparsity is not specificity. Random directions can be more concentrated than the bracket on smaller models; the key evidence is that endpoint and resampled-bracket support overlap substantially exceeds random controls sharing the same readout operator.
- Paired readout is stronger than single-endpoint readout. The single-endpoint method reaches 77.8% only on Qwen-3-4B and is near chance on the other three models; the headline 91.7% is not a general single-model order-identification rate.
- In the original 180 matched-batch DPO trials, median gap closure is 77.0%. Across original and fresh pairs, all 12/12 source pairs have positive median bracket-minus-random closure differences, and 11/12 pass the per-trial bootstrap strict-CI check. Shared source domains prevent treating the 316/360 aligned trials as independent samples for significance testing.
- The frozen matched-rollout surrogate has median single-step loss closure of 0.805, compared with −0.158 for a bracket from unmatched rollouts. Path matching separates the positive result from the negative control rather than being an incidental implementation detail.
- Under continued training on a third domain, aligned traces remain in 15 of 20 runs after four updates, 14 after eight, and 13 after sixteen. Among the 18 initially aligned runs, median retained fractions fall approximately from 27% to 14% and 7%, with sign reversals in some runs.
Highlights & Insights¶
- The method advances from predicting the better order to identifying where the difference appears. The token map is tied to the aggregate bracket forecast rather than a separately trained unconstrained explainer, although its coordinates depend on the model and evaluation slice.
- Controls address the actual confounds. Random directions share the logit readout, and frequency-matched interventions control common-token effects, making the support and causal evidence stronger than an isolated high Gini value.
- Paired subtraction is a reusable experimental design. Removing drift shared by both paths exposes the component that changes sign with order and is more robust than treating a single-endpoint projection sign as a history label.
Limitations & Future Work¶
- The theoretical object is a local two-source SGD residue, with limited update counts and parameter subspaces in the experiments. Multi-step support persistence does not imply accurate whole-vocabulary rankings or gap magnitudes over long training runs.
- Closure is not an absolute performance gain. The paper does not establish general improvements in GSM8K exact match or HumanEval pass@1 from these results. Small denominators, precision, and threshold selection must accompany effect-size reporting.
- Some interventions use a measured gap sign, iterative correction needs the reverse endpoint, DPO correction needs matched batches, and the reward surrogate needs frozen matched rollouts. The evidence supports controlled diagnosis, not a deployment-ready universal update tool.
- Sparse outputs do not automatically establish sparse full parameter vectors. The appendix's observable/null estimates cover a selected parameter subspace and contain damping, incompletely converged conjugate-gradient solves, and bf16 rounding error.
- Future work could test whether cheaper curvature approximations preserve order signs and high-mass support, and validate correction without an observed reverse endpoint. Cross-optimizer studies must include optimizer state and clocks rather than transplanting SGD equations.
Related Work & Insights¶
- vs scalar Lie-bracket order prediction: prior work supplies the second-order endpoint identity and target-gradient projection. This paper principally adds token decomposition, causal intervention, and paired-endpoint readout; inherited mechanisms and new contributions should be evaluated separately.
- vs activation recency probes: those identify which knowledge is more recent, whereas this paper detects the antisymmetric parameter residue of exchanging a specified source pair. Both study training history, but their access conditions and explanatory objects differ.
- vs sample influence and mechanistic attribution: influence functions and TRAK attribute effects to training examples, while circuit interpretation localizes internal structure. This paper attributes to output tokens only for training-path interactions and does not promise recovery of sample origins.
- vs EWC, GEM, PCGrad, and model editing: these methods address forgetting, gradient conflict, or knowledge modification; this paper first measures the noncommutativity of composed updates. Its paired controls can inform diagnosis of ordinary post-training interactions, but NLL-gap closure is not equivalent to preserving every capability.
Rating¶
- Novelty: 4/5. The second-order identity is inherited, but token readout, controls, and paired readability define a clear new problem.
- Experimental Thoroughness: 4/5. Multiple models, controls, and negative results provide substantial coverage, with remaining limits from local paths, floors, and small seed clusters.
- Writing Quality: 4/5. Inherited results, finite differences, and matching conditions are distinguished clearly, although the numerous appendix protocols increase reading effort.
- Value: 4/5. The framework provides interpretable diagnosis of post-training order interactions; deployment benefits and long-horizon validity remain unverified.