Skip to content

PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders

Conference: NeurIPS2026
arXiv: 2609.32469
Area: Interpretability
Keywords: in-context learning, sparse autoencoders, demonstration utility, feature localization, demonstration retrieval

TL;DR

PULSE identifies SAE features associated with demonstration utility from a small labeled discovery set, then uses them for complete-set ranking and cacheable per-example retrieval, improving selection across six text tasks while providing associational rather than causal-intervention evidence.

Background & Motivation

In-context learning (ICL) adapts a model without parameter updates by placing examples with their answers before a prediction query, making demonstration selection a central part of adaptation. Lex-Sim and SBERT rely on lexical or semantic similarity, while CEIL additionally considers set diversity; these methods can find material that resembles the query without necessarily finding material that helps the target model. Two similarly relevant examples may activate different internal representations in different models, and their benefits can also change with redundancy, label composition, or interactions within a context.

The gap is not merely the absence of a stronger retriever: the relationship between demonstration utility and internal representations has not been directly localized. Raw residual-stream dimensions suffer from feature superposition, mixing multiple semantic factors and potentially obscuring task-relevant changes. Sparse autoencoders (SAEs) provide coordinates better suited to separating these factors, but ordinary SAE-space cosine similarity still treats coordinates approximately equally and does not identify which ones track demonstration gains.

PULSE therefore begins with a small supervised discovery procedure within the training split rather than labeling test demonstrations as good or bad. For the same query, it compares utility differences between complete demonstration sets with their SAE activation differences and retains consistently associated coordinates. Core idea: localize utility-related sparse features from within-query demonstration-set pairs, validate set-level prediction with a signed vector, and use its magnitude to construct a feature-relevance metric for per-example retrieval.

Method

Overall Architecture

The inputs are a frozen target language model, a pretrained SAE for the selected layer, a support pool with supervised targets, and a small set of training-split discovery queries; the output is a demonstration set for the final ICL prompt. PULSE first performs โ€œUtility Feature Discovery,โ€ then branches into two uses: โ€œComplete-Set Rankingโ€ validates localization on queries excluded from discovery, while โ€œMagnitude-Mask Retrievalโ€ supports practical per-example retrieval. Ranking validation is not a mandatory retrieval-time stage and does not relearn features from test labels.

SAE representations come from the residual stream: the full prompt is passed through the model, encoded by the SAE at the selected layer, and mean-pooled over prompt tokens. Discovery prompts contain demonstrations with their targets followed by a query with its answer hidden; retrieval encodes only candidate inputs with their targets omitted, making candidate representations cacheable. A candidate target enters the final prompt as a demonstration answer only after that candidate is selected.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Training discovery queries<br/>and labeled support pool"] --> B["Utility Feature Discovery"]
    B -->|Signed vector; held-out queries| C["Complete-Set Ranking"]
    C --> V["Validation: select a complete set<br/>and evaluate downstream performance"]
    B -->|Frozen vector magnitude; no query labels| D["Magnitude-Mask Retrieval"]
    Q["Query and candidate inputs<br/>Zero-shot SAE encodings"] --> D
    C ~~~ D
    D --> O["Inference: assemble examples and answers<br/>Target model performs ICL"]

Key Designs

1. Utility Feature Discovery: localize demonstration-induced changes while controlling query difficulty

The discovery set comes from the training split and does not overlap with evaluation queries. Each discovery query receives multiple randomly sampled complete few-shot demonstration sets; the query itself is excluded if it appears in the support pool. For classification, the utility proxy is the gold-label log probability minus the largest incorrect-label log probability. For CommonGen and GSM8K, it is the length-normalized teacher-forced log likelihood of the supervised target sequence. The latter measures support for the target under correct prefixes, not actual generation quality or mathematical answer accuracy.

Zero-shot-relative utility is defined as:

\[ \mathcal{U}_{\mathcal{T}}(q_x,E)=\mathcal{G}_{\mathcal{T}}(q_x,E)-\mathcal{G}_{\mathcal{T}}(q_x,\varnothing). \]

Here, \(\mathcal{G}_{\mathcal{T}}\) is the task-specific proxy described above, and \(E\) is a demonstration set with supervised targets. This definition distinguishes improvements from degradation relative to the original prediction. Feature estimation then pairs sets within the same query, so the zero-shot term cancels and query difficulty is not mistaken for a cross-query utility difference.

For each set pair, PULSE computes a utility difference \(\Delta\mathcal{U}_{\mathcal{T}}\) and, for SAE coordinate \(j\), an activation difference \(\Delta A_{j,l}\). All discovery queries and their unordered set pairs contribute to the feature score:

\[ S_{j,l}=\frac{\sum_{(q_d,E_a,E_b)}\Delta\mathcal{U}_{\mathcal{T}}(q_d;E_a,E_b)\,\Delta A_{j,l}(q_d;E_a,E_b)}{n_{\mathrm{pairs}}\sqrt{\mathrm{Var}(\Delta A_{j,l})+\varepsilon}}. \]

The denominator uses activation-difference variance across all pairs, reducing the advantage of coordinates whose changes are large simply because of their scale. This is neither a full Pearson correlation coefficient nor an established causal effect. PULSE retains the largest positive and most negative scores as their original weights, zeroing the remaining coordinates to obtain the sparse utility-localization vector \(\mathbf{w}_l\). A negative weight means that higher activation is associated with lower discovery-time utility, not that the coordinate represents negative sentiment or a neuron that must be suppressed.

The default uses 64 discovery queries and 32 candidate sets per query, yielding \(64\binom{32}{2}=31{,}744\) set pairs. Those pairs come from 2,048 queryโ€“set configurations and share queries and sets; they are not 31,744 independent supervised examples. Discovery cost is dominated by full-prompt model forward passes rather than the closed-form weight calculation itself.

2. Complete-Set Ranking: test whether the signed vector predicts whole-context gains

Because discovery learns a direction of change for complete contexts, validation first operates at the same scale. For each held-out query, every method ranks exactly the same presampled complete demonstration sets. PULSE subtracts the zero-shot SAE activation from a candidate's complete-context activation and takes its inner product with the signed vector:

\[ s_{\mathrm{rank}}(q_x,E)=\mathbf{w}_l^{\top}\bigl[\mathbf{A}_l(q_x,E)-\mathbf{A}_l(q_x,\varnothing)\bigr]. \]

Downstream prediction is evaluated after selecting the highest-scoring set. Increasing positive-weight coordinates or decreasing negative-weight coordinates raises the score, preserving discovery-time directional information. Under this protocol, Lex-Sim, SBERT, and KNN-SAE average queryโ€“example similarities within each set, while CEIL uses its set-level diversity objective; they are not compared using different candidate pools.

This experiment asks whether the localization vector predicts the usefulness of complete sets for unseen queries, not whether practical retrieval exactly optimizes set utility. Predictive success still does not establish that these coordinates causally mediate ICL: the authors do not ablate, clamp, or intervene on feature activations.

3. Magnitude-Mask Retrieval: turn set-level information into a cacheable relevance metric

With a support pool of roughly 2,000 items, enumerating all few-shot subsets is infeasible. Even full-prompt scoring of labeled singletons would require a separate forward pass for every queryโ€“candidate pair and would still ignore interactions among multiple demonstrations. PULSE-Retriever therefore adopts a weaker but scalable proxy: it compares input-only, zero-shot SAE encodings of queries and candidates along utility-informative coordinates.

\[ \begin{aligned} s_{\mathrm{mask}}(q_x,x_i)&=\cos\!\bigl(|\mathbf{w}_l|\odot\mathbf{a}^{(q)},\;|\mathbf{w}_l|\odot\mathbf{a}^{(i)}\bigr),\\ s_{\mathrm{blend}}(q_x,x_i)&=(1-\beta)\hat{s}_{\mathrm{mask}}(q_x,x_i)+\beta\hat{s}_{\mathrm{cos}}(q_x,x_i). \end{aligned} \]

Here, \(s_{\mathrm{cos}}\) is ordinary unweighted SAE cosine similarity, and hats denote z-score normalization across candidates for the same query. With the default \(\beta=0.3\), scores are normalized before the 0.7/0.3 blend rather than directly mixed in their original scales. Applying the magnitude mask to both sides is equivalent to a diagonal metric with squared weights: it measures similarity along important coordinates, not a candidate's signed utility.

Magnitude is used because of the scale mismatch: a negative set-level weight has not been calibrated to mean that a candidate input should rank lower. Applying signed weights symmetrically before cosine similarity also removes their signs from the inner product and norms, producing exactly the same result as the magnitude mask. Appendix B therefore compares against a genuinely sign-sensitive bilinear score rather than treating signed-mask cosine as a distinct method.

Retrieval keeps the 50 highest-scoring candidates and greedily composes the final context. The first step chooses the most relevant candidate; later steps subtract the largest unweighted SAE cosine similarity between a candidate and already selected examples, using a penalty coefficient of 0.3. Selected items are excluded from subsequent steps. This discourages redundant neighbors from occupying the entire context, but it does not provide an exact additive decomposition of set-level utility.

Loss & Training

PULSE trains no new retrieval head, does not fine-tune the target language model, and does not retrain the SAE. It uses labels to compute utility proxies and estimates weights in closed form, so it is not unsupervised. Classification decisions use teacher-forced label likelihood; final CommonGen and GSM8K evaluation uses greedy generation with maximum new-token budgets of 32 and 220, respectively.

Gemma2-2B uses a layer-12 Gemma-Scope dictionary with 16,384 coordinates, retaining 512 positive and 512 negative features. Llama3.1-8B uses a 131,072-coordinate Llama-Scope dictionary, with layers 12, 16, and 24 for classification, CommonGen, and GSM8K, respectively, retaining 2,048 positive and 2,048 negative features. The variance stabilizer is \(10^{-6}\). Hyperparameters are calibrated on held-out validation data and fixed before testing; the test split is not used for feature discovery.

Key Experimental Results

Main Results

Experiments cover AGNews, REST14, LAP14, EMOC, CommonGen, and GSM8K. AGNews, REST14, and EMOC use fixed 512-query held-out evaluation subsets; LAP14 uses its full 463-query test split. Support pools generally contain 2,000 items, except LAP14 with 1,850. Classification results therefore should not be described as evaluations on every dataset's complete official test split.

The following table selects the four-dataset classification-average accuracies (%) from the paper's Table 2. Each cell lists 1/2/4/8-shot results. These are practical pool-scale retrieval results, not the controlled complete-set ranking results in Table 1.

Method Gemma2-2B: 1/2/4/8-shot Llama3.1-8B: 1/2/4/8-shot
Random 49.40 / 61.81 / 67.02 / 70.91 62.83 / 69.52 / 74.31 / 75.87
SBERT 68.58 / 71.01 / 73.45 / 77.37 72.99 / 77.56 / 78.46 / 79.68
CEIL 68.58 / 72.13 / 74.67 / 78.50 72.99 / 78.94 / 79.00 / 81.10
KNN-SAE 66.88 / 70.33 / 74.65 / 78.32 71.45 / 77.89 / 78.88 / 81.36
PULSE-Retriever 70.64 / 74.31 / 77.34 / 79.61 74.38 / 79.96 / 80.77 / 82.45

Generation and reasoning results from the paper's Table 3 are shown below as averages over 1/2/4/8-shot settings. The strongest baseline is selected by its average for the corresponding task and backbone.

Task and backbone Metric Strongest baseline PULSE-Retriever Gain
CommonGen, Gemma2-2B BLEU-4 CEIL: 9.09 10.02 +0.93
CommonGen, Llama3.1-8B BLEU-4 SBERT: 9.84 10.45 +0.61
GSM8K, Llama3.1-8B Exact match (%) CEIL: 48.26 51.42 +3.16 percentage points

Ablation Study

The following table summarizes Gemma2-2B 4-shot classification-average accuracies from the paper's Tables 4, 5, and 22. These controls answer different questions and should not all be interpreted as individual module removals.

Config Average accuracy (%) Note
PULSE-Retriever 77.34 SAE feature identification, blended scoring, and greedy redundancy control
PCA basis 73.87 Replace the internal basis while preserving paired identification and retrieval
Raw hidden-state basis 73.38 Mean-pooled raw hidden dimensions from the same layer
PULSE-Mask 73.46 Magnitude mask only, without ordinary SAE similarity
Anti-Retrieval 48.52 Same mask, but select the lowest-scoring items
Sparse-Random 72.14 Replace identified coordinates with random sparse features
Bottom-Feats 71.31 Use features with the smallest absolute identification scores
PULSE-NoTopK 73.36 Remove the top-feature restriction in the mask-only control
Pure SAE cosine, \(\beta=1\) 74.65 No utility mask; control for blended scoring

Key Findings

  • Controlled ranking achieves best or tied-best results in 30 of 32 datasetโ€“shotโ€“backbone cells; practical retrieval wins in 25/32. The former validates complete-set information, while the latter validates its usefulness after conversion to a per-example proxy. The former count is not a retrieval result.
  • The SAE-basis average of 77.34 exceeds PCA's 73.87 and Raw's 73.38 by 3.47 and 3.96 percentage points, respectively. Mask-only retrieval at 73.46 is also weaker than blended retrieval at 77.34, indicating complementarity between sparse utility information and broader SAE similarity.
  • Appendix F's 4-shot EPR-style control scores 77.27 and 80.22 on Gemma2-2B and Llama3.1-8B, versus PULSE's 77.34 and 80.77. The classification-average advantages are small, and EPR is stronger on AGNews and LAP14. Fewer discovery queries do not establish strictly matched compute or feedback budgets.
  • Gemma2-2B 4-shot leave-one-out union transfer averages 76.71, compared with 77.34 for target-specific discovery; REST14 and LAP14 have an active-feature Jaccard overlap of only 0.112. Partial transfer does not imply a single task-independent utility direction.
  • The paper has comparison-scope inconsistencies: the abstract describes a 2โ€“3-point classification gain over the strongest baseline, but the main text's 2.93/1.99 gains are against KNN-SAE. Averaging each method across shots in Table 2 makes CEIL the stronger baseline, with PULSE ahead by approximately 2.01/1.38 points. Appendix F also compares EPR's 4-shot GSM8K result of 44.53 with Random's across-shot average of 47.61; the matched 4-shot Random result is 49.20. Original values are retained with their scopes distinguished rather than rewritten as consistent evidence.

Highlights & Insights

  • The shift from queryโ€“example resemblance to internal coordinates associated with utility directly connects feature selection to the target model. The contribution is not simply using SAE representations, but selecting representation dimensions through complete-set utility contrasts.
  • Within-query pairing reduces confounding by query difficulty, and ranking validation preserves the set-level scale of directional information. This experimental organization clarifies what the localization vector captures beyond reporting a single downstream retrieval improvement.
  • The explicit distinction between set-level signs and per-example relevance avoids treating negative weights as automatically identifying bad demonstrations. The magnitude mask also enables cached candidate encodings, connecting interpretability analysis to a practical retrieval proxy.

Limitations & Future Work

  • The approach requires a usable target-model SAE, access to internal activations, and a labeled discovery set, so it does not directly apply to closed models exposed only through output APIs. Appendix H reports approximately 141โ€“277 seconds for a Gemma classification static cache and 448โ€“884 seconds for discovery forward passes; gradient-free does not mean compute-free.
  • Whole-prompt mean-pooled SAE activations may depend on demonstration length; within-query pairing does not automatically remove length differences between sets. The authors suggest query-segment pooling, length-matched sets, or length-residualized activation differences.
  • Feature interpretations come from highly activating training-pool snippets and post-hoc semantic analysis, without causal interventions. Future work should test whether changing specific features actually changes demonstration utility rather than treating interpretable descriptions as mechanism proofs.
  • GSM8K discovery relies on supervised solution sequences, not only final answers. Answer likelihood, execution-based rewards, or generated rationales under weaker supervision require separate validation.
  • The three-seed analysis varies only PULSE discovery and candidate sampling, without matched multi-seed significance tests for every baseline. Appendix I reports Llama3.1-8B 4-shot GSM8K at \(55.20\pm0.91\); this cannot be directly combined with the across-shot +3.16 gain to establish significance.
  • vs SBERT / KNN-SAE: SBERT uses external sentence embeddings, and KNN-SAE uses ordinary similarity in internal SAE space. PULSE additionally discovers important coordinates from target-model feedback, incurring extra supervision and offline compute.
  • vs CEIL: CEIL selects combinations through set diversity, while PULSE emphasizes utility-related internal representations. Practical PULSE retrieval still includes redundancy control, so the results do not show that semantic features remove the need for context-composition design.
  • vs EPR-style: EPR-style uses singleton feedback for gradient-based dense retriever training, whereas PULSE uses paired complete-set feedback for closed-form feature localization. The appendix supports complementary strengths, but does not isolate every difference in supervision scale, feedback unit, or model capacity.
  • Research direction: Query-segment SAE activations and length-controlled discovery could be combined with budget-matched feature interventions and multi-seed comparisons to distinguish utility prediction, retrieval gains, and causal mechanisms. This is an extension, not an experiment completed in the paper.

Rating

  • Novelty: 4/5 โ€” Combines paired set utility with SAE feature localization while distinguishing ranking and retrieval scales.
  • Experimental Thoroughness: 4/5 โ€” Six tasks, two backbones, and multiple controls, but budget matching and statistical testing remain incomplete.
  • Writing Quality: 4/5 โ€” Clearly explains proxy boundaries, although the abstract and some appendix comparisons require care about scope.
  • Value: 4/5 โ€” A reusable approach to internal-feature-guided demonstration retrieval, with deployment dependent on SAEs and extra offline resources.