Skip to content

Spectral Reversal: Counteracting Singular Value Bias for Graph Prompting

Conference: NeurIPS2026
arXiv: 2609.32143
Code: https://github.com/keris-yang/Spectral-Reversal
Area: Graph Learning
Keywords: graph prompt learning, singular value bias, spectral reversal, null space, few-shot adaptation

TL;DR

SRP preferentially reads weak directions through singular-value soft masks at each frozen GNN layer, supplements features discarded by the original weights through null-space PCA, and adds low-rank prompts after aggregation; it achieves 68.11% accuracy on 5-shot Cora with GraphCL pretraining, but does not outperform the strongest baseline on every dataset.

Background & Motivation

Graph prompt learning aims to adapt a graph neural network (GNN) with few additional parameters rather than retraining the entire model. GPF primarily modifies node features, EdgePrompt introduces edge-level prompts, and GraphPrompt adjusts readout, but these approaches generally do not explicitly exploit the internal geometry of frozen weights. This paper shifts attention to a different question: which input directions does a pretrained weight transform into strong outputs, and which directions barely affect the representation?

The authors argue that self-supervised pretraining concentrates capacity on frequently observed variations. Graph propagation further changes the covariance of training features, reinforcing this imbalance. However, features needed to distinguish downstream classes need not align with large singular values. Small-singular-value directions remain reachable, although the original transformation responds weakly to them; the exact null space is different, because inputs in it vanish under that frozen linear transformation. Training a classifier on frozen representations may therefore require large coefficients to compensate for weak directions, while being unable to recover directions that were genuinely discarded.

Rather than treating all weak directions as noise, the paper provides an alternative route for reading them. Here, “spectral” primarily refers to the singular spectrum of a weight matrix, not the Fourier frequencies of the graph Laplacian. Data and propagation can connect these spectra, but they are not the same coordinate axis. Core Idea: allocate prompt capacity according to the singular geometry of frozen weights, combining weak singular directions and null-space information in a post-aggregation prompt instead of continuing to prioritize the strongest pretrained directions.

Method

Overall Architecture

The inputs remain the graph structure and node features, and the outputs remain node or graph classification predictions. SRP preserves the original message-passing branch and adds a prompt branch conditioned on the current layer inputs: prepare frozen bases, perform spectral reversal, and then map spectral and null-space coordinates into a shared bottleneck before fusion and projection to the layer output dimension.

In the diagram, Frozen bases is an offline preparation stage, whereas Spectral reversal and Prompt fusion form the prompt path used during training and inference. Labels enter only the downstream classification loss, not SVD or null-space PCA. Training updates adapters and thresholds; inference uses the learned parameters.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    W["Frozen weights + static features"] -->|offline SVD and null-space PCA| B["Frozen bases"]
    H["Current layer features"] --> R["Spectral reversal"]
    B -->|singular basis and values| R
    B -->|null-space basis| F["Prompt fusion"]
    H -->|null coordinates or fallback input| F
    R -->|masked spectral coordinates| F
    H --> G["Frozen aggregation"]
    F --> O["Add prompt and activate<br/>node or graph prediction"]
    G --> O
    Y["Training labels"] -.-> L["Classification loss"]
    O -.-> L
    L -.->|train adapters and threshold only| R
    L -.-> F

Understanding SRP requires separating its diagnostic metrics from its training rules. The authors measure downstream demand using the weights of a linear classifier on raw features and define the Downstream–Pretraining Misalignment Score for each nonzero singular direction as \(\mathrm{DPMS}_{j}=\|W_{\mathrm{down}}^{\top}v_j\|_2^2/\sigma_j^2\). The numerator measures downstream classifier weight energy along that direction, while the denominator measures the response energy of the frozen transformation. A large ratio indicates a direction that downstream prediction needs but pretraining serves weakly.

The authors also partition nonzero singular values into five groups in descending order and define \(\Delta_q=\frac{\sum_{j\in Q_q}\|W_{\mathrm{down}}^{\top}v_j\|_2^2}{\sum_j\|W_{\mathrm{down}}^{\top}v_j\|_2^2}-\frac{\sum_{j\in Q_q}\sigma_j^2}{\sum_j\sigma_j^2}\). A positive value means that a group's share of demand exceeds its share of capacity. The first group contains the largest singular values, the fifth contains the weakest nonzero directions, and exact zero directions are handled separately. The frozen weights determine these groups, not the downstream classifier.

DPMS explains misalignment; running SRP does not require first training an additional downstream classifier. The actual mask reads singular values and learns its threshold through downstream supervision. A small singular value does not automatically imply high demand, because the numerator of DPMS also matters.

Key Designs

1. Frozen bases: prepare reachable and discarded directions separately

A compact singular value decomposition (SVD) of each frozen weight, \(W=U\Sigma V^{\top}\), exposes its reachable directions in the layer's input feature space through the right singular vectors. Coordinates along these directions enter the spectral channel. Their orthogonal complement is described by \(P_{\mathcal N}=I-VV^{\top}\), the exact right-null-space projector. Very small nonzero directions and exact zero directions must not be conflated: a linear head can recover the former, whereas the latter are inaccessible after a single frozen linear transformation.

SRP applies PCA to data projected into the null space, retaining a few directions with the largest variance and passing them to a separate adapter. This does not make the entire high-dimensional null space trainable. Instead, it first extracts a low-dimensional, data-dependent basis within the region ignored by the frozen transformation. The optimality in Appendix B.3 concerns retained projected-data variance, not class discrimination. Whether high-variance directions are useful still requires experiments and downstream supervision.

Both SVD and PCA bases are frozen. The concrete two-layer implementation in Appendix D computes null-space PCA offline from static raw features at the first layer and uses the projective fallback at later layers; PCA does not need to be recomputed at every training step. Basis projections of static first-layer inputs can be cached. However, earlier prompts change later hidden features during training, so feature projections at all layers cannot indiscriminately be described as permanently cacheable.

The source contains transpose inconsistencies in its null-space basis notation: Section 5.2 defines a row-stored basis but subsequently uses some products as if it were column-stored. Spectral adapter expressions also mix row and column conventions. This note preserves the mechanism—project into complementary subspaces and then map to a bottleneck—without treating the inconsistent expressions as exact implementation-ready dimension specifications.

2. Spectral reversal: assign larger relative reading weights to weak directions

The spectral channel first projects current inputs onto the right singular basis, then uses a learnable scalar to control the spectral cutoff. Its core rule is \(\tau=\operatorname{sigmoid}(\tau_0),\quad\mu_j=\operatorname{sigmoid}((\tau\sigma_{\max}-\sigma_j)\gamma)\). With a positive temperature coefficient, singular values below the cutoff receive larger masks and large singular values receive smaller masks. The temperature coefficient controls transition sharpness, while the threshold determines its location.

This is not direct inverse-singular-value multiplication, nor does it multiply every weak coordinate by a gain greater than 1. The mask stays within 0–1, so “amplifying weak directions” more precisely means relatively preserving them and encouraging the trainable adapter to use them. Adapters and the output projection jointly determine final prompt magnitude. Differentiability allows the threshold to be optimized through classification loss without specifying the retained nonzero directions in advance.

This also differs from directly changing the singular values of frozen weights. SRP selects coordinates read by the prompt branch, while the original message-passing branch retains its original response. The combined representation may exhibit compensation, reinforcement, or cancellation, but the mask alone does not imply that the original backbone's strong directions are necessarily attenuated.

3. Prompt fusion: generate post-aggregation low-rank prompts from complementary inputs

Each channel maps its coordinates into the same low-dimensional bottleneck. Their outputs are summed, projected to the output space, and combined with a prompt bias. Using the row-feature convention in Appendix C.6, the prompt is \(p(H)=[(HV\odot\mu)A_s+(HN)A_n]Q+b_p,\quad\operatorname{col}(N)\subseteq\ker(W)\). Although the input subspaces are complementary, the shared bottleneck and output projection can still compress information. Orthogonality of the input subspaces does not imply lossless fusion for arbitrary inputs.

The prompt is added after neighborhood aggregation and before the nonlinear activation, without being filtered again by the propagation operator in the current layer. It can therefore complement node-local variations not preserved by that layer's original message-passing path. Subsequent layers still propagate features, so this bypass does not guarantee immunity to smoothing throughout the network. Nor can a weak weight-singular direction be directly identified with a high graph-frequency direction.

When the input dimension does not exceed the reachable dimension plus the configured PCA dimension, the null space is already small or empty, and SRP uses a projective fallback. For row-vector inputs, its reweighted representation is \(h_{\mathrm{weak}}=h-(hV\odot(1-\mu))V^{\top}\), which then passes through a single low-rank adapter. This preserves the masked reachable component and the complete null-space component in the original input space rather than performing unnecessary PCA.

The algebraic equivalence in Appendix B.5 requires a basis covering the entire null space and the corresponding linear adapters. It establishes the same function space as that complete dual-channel parameterization, not full reconstruction of information compressed by a bottleneck. With small input dimensions, gains rely more on spectral selection than on discovering a large additional null space.

A Worked Example

Consider Cora's first layer: raw features have 1,433 dimensions, the appendix specifies a hidden dimension of 128, and the reported right-null-space dimension is 1,305. A node's input contains both a component readable by the frozen weights and a component in this 1,305-dimensional space. A classifier trained directly on frozen outputs cannot recover the latter.

SRP first selects a PCA basis offline from null-space-projected Cora features. During training, the same node's right-singular coordinates pass through the reverse mask and spectral adapter, while its PCA coordinates pass through the null-space adapter. The two outputs fuse in the shared bottleneck and produce a 128-dimensional prompt, which is added to frozen aggregation before activation. Supervision changes prompt parameters, not the original SVD bases or GNN weights.

Later layers receive hidden representations, often with smaller null spaces, and use the projective fallback. This example explains why high-dimensional node features provide substantial complementary capacity. It does not imply that all 1,305 null-space directions contain label information or that node count alone explains the gains.

Loss & Training

Downstream classification supervision optimizes prompt parameters while the pretrained GNN remains frozen. Trainable components include spectral and null-space adapters, shared output projections, prompt biases, and layer-specific thresholds. Fallback layers instead train their low-rank adapters. There is no additional DPMS loss or label-trained PCA stage.

The main experiments use GraphCL and SimGRACE pretraining, with 5-shot per class for node classification and 50-shot per class for graph classification, aggregated over five random seeds. Cross-backbone experiments in the appendix follow the main few-shot settings and report means and standard deviations over five runs. The cache does not fully specify the optimizer, learning rate, or pretraining epochs, so it cannot support an invented reproduction recipe.

Section 6.4 and Appendix C.1 give default bottleneck and PCA dimensions of 32 and 16, respectively, while Appendix C.2 states that the temperature coefficient is 10 for all datasets. Appendix D.2 instead lists default bottleneck and PCA dimensions of 8 and 64. This is a source-level configuration conflict; neither pair should be selected silently and presented as the configuration behind every table.

The theoretical support must be interpreted conditionally. The reachability residual concerns a frozen linear transformation followed by a linear head. Gradient concentration requires concentrated inputs and vanishing cross-step terms, while the appendix's expectation-based version adds an independence discussion. The PAC-Bayes derivation also depends on prior, posterior, and parameter-norm assumptions; it is not a proof that the practical nonlinear, multilayer SRP necessarily outperforms full fine-tuning.

Key Experimental Results

Main Results

The following selection comes from Tables 1–2. The metric is accuracy (%), reported as mean ± standard deviation over five seeds. “Strongest baseline” means the highest baseline mean for the same dataset and pretraining strategy. Differences are percentage points, not evidence of statistical significance.

Pretraining Task and dataset SRP Strongest baseline Mean difference
GraphCL Node, Cora, 5-shot 68.11 ± 2.29 GraphLoRA, 63.35 ± 4.32 +4.76
GraphCL Node, CiteSeer, 5-shot 49.02 ± 2.26 EdgePrompt+, 46.20 ± 0.99 +2.82
GraphCL Node, ogbn-arxiv, 5-shot 24.75 ± 1.79 GraphLoRA, 23.26 ± 1.89 +1.49
GraphCL Node, Flickr, 5-shot 26.39 ± 2.11 GraphPrompt, 26.08 ± 3.44 +0.31
SimGRACE Node, Cora, 5-shot 67.24 ± 7.31 GraphLoRA, 63.06 ± 3.64 +4.18
SimGRACE Node, Flickr, 5-shot 28.48 ± 3.26 EdgePrompt, 30.12 ± 5.04 −1.64
GraphCL Graph, ENZYMES, 50-shot 36.32 ± 2.65 GraphLoRA, 36.05 ± 1.89 +0.27
GraphCL Graph, NCI109, 50-shot 68.18 ± 0.55 EdgePrompt+, 66.52 ± 0.91 +1.66
SimGRACE Graph, NCI1, 50-shot 66.53 ± 2.94 EdgePrompt+, 67.07 ± 1.96 −0.54
SimGRACE Graph, Mutagenicity, 50-shot 68.92 ± 2.73 EdgePrompt+, 68.31 ± 1.36 +0.61

Under GraphCL, SRP has the highest mean on all ten main datasets. Under SimGRACE, Flickr and NCI1 are clear exceptions. The source's summary of “best or second best” on almost all datasets should not be treated as an exact ranking claim: both EdgePrompt and EdgePrompt+ exceed SRP on SimGRACE Flickr.

Ablation Study

Table 3 reports only Cora and CiteSeer, so its conclusions cannot directly be extended to all ten datasets. The table below retains its one-decimal precision and the authors' reported average differences across the two datasets.

Config GraphCL Cora GraphCL CiteSeer Average difference SimGRACE Cora SimGRACE CiteSeer Average difference
SRP full model 68.1 ± 2.3 49.0 ± 2.3 0.0 67.2 ± 7.3 51.4 ± 5.2 0.0
SRP-NM, no null-space channel 63.4 ± 2.8 41.1 ± 3.8 −6.3 60.3 ± 5.9 43.5 ± 2.9 −7.4
SRP-NS, no spectral channel 64.4 ± 2.1 45.6 ± 3.2 −3.6 63.5 ± 6.6 47.7 ± 4.3 −3.7
SRP-NR, all masks set to 1 64.7 ± 2.7 48.0 ± 4.0 −2.2 63.4 ± 4.4 49.9 ± 4.9 −2.7
SRP-Bi, bidirectional weighting 64.7 ± 3.3 47.9 ± 3.3 −2.3 63.0 ± 4.5 50.0 ± 4.8 −2.8

Removing the null-space channel causes the largest loss, particularly on GraphCL CiteSeer, which falls from 49.0 to 41.1. Removing reverse selection also hurts, but less than removing the entire null-space channel. Selecting weak directions and recovering information invisible to the original transformation therefore provide complementary benefits.

Appendix Table 8 further matches the number of selected directions to test whether dimensionality reduction alone explains the gains. The following entries are mean accuracies for five selection strategies. They come from additional experiments and retain their own precision rather than replacing the main-table values.

Dataset Strong directions Random directions Weak directions Hard threshold STE Soft threshold, temperature 10
Cora 64.25 64.44 67.80 67.95 68.10
CiteSeer 47.10 47.46 48.78 48.90 48.99
NCI1 66.42 66.45 66.87 66.89 66.80

Key Findings

  • With matched cardinality, weak-direction selection beats strong and random selection, supporting the role of spectral choice. However, soft thresholds do not beat hard thresholds on every dataset; NCI1 is a counterexample.
  • Raising the temperature from 0.001 to 10 improves Cora from 64.55 to 68.10 and CiteSeer from 47.64 to 48.99. An excessively soft mask approaches uniform reading and loses directional discrimination.
  • In the eight GraphGPS/Graphormer settings of Appendix Table 6, SRP has the highest mean in six. On Graphormer NCI1 and ENZYMES it trails GraphLoRA by 0.82 and 0.75 percentage points, respectively, so cross-backbone wins are not universal.
  • On the heterophilic extensions, gains over GraphLoRA are 6.11, 2.06, and 2.15 percentage points for Cornell, Squirrel, and Chameleon. This is evidence from three additional datasets, not a universal heterophilic-graph guarantee.
  • Appendix Table 10 shows faster per-epoch execution than GPF on five of six datasets, but ENZYMES is slower at 0.292 versus 0.221 seconds. Offline PCA initialization is not included in this throughput comparison.

Highlights & Insights

  • Prompt directions deserve explicit design, not just parameter counts. SRP limits low-rank capacity while using frozen-weight geometry to determine which input directions the prompt reads.
  • The null space is not a set of weakly activated output dimensions; it consists of input directions eliminated by a particular weight. Reading them through a bypass differs fundamentally from adjusting a classifier on existing outputs.
  • Matched-cardinality experiments distinguish mechanisms better than simply removing a mask. A transferable research strategy is to compare strong, weak, and random subspaces under the same budget rather than only comparing full adaptation with small adapters.

Limitations & Future Work

  • Global thresholds and shared singular bases are not region- or instance-adaptive. The same feature direction may carry different graph-frequency characteristics across regions or graphs.
  • Null-space benefits depend on input dimensionality and weight rank. Low-dimensional, full-column-rank inputs may leave this channel empty, and high-variance PCA directions need not be discriminative downstream.
  • The source's “lossless fusion” claim is too strong, and some main-text products use inconsistent transposes. Its claim of PCA-basis uniqueness also needs conditions such as an eigengap; variance optimality should not be read as unconditional uniqueness or classification optimality.
  • Default hyperparameters conflict between 32/16 and 8/64, while complexity discussions mix cacheable projections with recomputation in each forward pass. Reproduction requires checking code and actual run configurations rather than copying a single complexity slogan.
  • Five-seed results can vary substantially, some gains are smaller than standard deviations, and significance tests are absent. Dataset size, feature dimensionality, and task type change together, preventing a single-factor explanation of gain differences.
  • Future evaluation could test region-adaptive thresholds, label-budget-constrained null-space selection, and end-to-end costs including initialization, caching, peak memory, and total training time.
  • vs GPF / EdgePrompt: These methods primarily prompt features or edges. SRP organizes input reading around frozen-weight reachability and injects prompts after aggregation, at the cost of offline decomposition and basis storage.
  • vs GraphLoRA / LoRA: Low-rank weight updates change the effective operator inside the message-passing path. SRP preserves the original operator and builds a representation bypass from weak singular and null-space directions. The distinction includes whether adaptation is filtered by the current layer's propagation, not merely how low-rank matrices are named.
  • vs standard PCA: Standard PCA prioritizes overall variance and may favor directions already well captured by the encoder. Null-space PCA first excludes the reachable component and then compresses its complement. A useful next comparison is unsupervised variance selection versus supervised discriminative selection under the same budget.

Rating

  • Novelty: 4/5. Combining spectral capacity misalignment, null-space reading, and post-aggregation graph prompts is more targeted than adding a generic low-rank branch.
  • Experimental Thoroughness: 4/5. Ten main datasets, two pretraining strategies, and extensive appendix analyses are included, but ablation coverage, variance, and reproduction configurations remain limitations.
  • Writing Quality: 3/5. The mechanism is clear, but transposes, default hyperparameters, and losslessness claims need clarification.
  • Value: 4/5. The paper provides an interpretable strategy for selecting adaptation directions in frozen graph encoders, especially for few-shot tasks with high-dimensional node features.