Relation-Aware Graph Foundation Model¶
Conference: NeurIPS 2026
arXiv: 2505.12027
Code: https://github.com/jianxiangyu/REEF/
Area: Graph Learning
Keywords: graph foundation model, relation tokens, semantic hypernetworks, cross-dataset transfer, few-shot learning
TL;DR¶
REEF treats transferable relation semantics as the unit of graph foundation modeling and uses semantic hypernetworks to generate aggregators, classifiers, and dataset projectors, reaching 48.66%, 48.65%, and 47.72% accuracy on three target graphs under 1-shot classification, while retaining clear limitations for unseen relations, zero-shot tasks, and computational cost.
Background & Motivation¶
Cross-graph pretraining is difficult for reasons beyond graph size: Cora and PubMed are both citation networks, but their node features represent computer science and biomedicine, respectively, with different dimensions, label spaces, and distributions. Sharing all parameters of a GCN or GAT can conflate these differences; assigning a separate expert to every dataset instead makes transfer depend on dataset identity rather than reusable mechanisms. In the main experiment, jointly trained GCN reaches only 54.97% average accuracy across five graphs, showing that pooled training does not automatically yield cross-domain capability.
The authors instead focus on relations. A citation still describes an informational connection between papers across disciplines, and webpage classification can share task semantics across universities. This sharing does not require identical node content, but it does require the model to know how to propagate information and how to determine whether a relation holds. Meanwhile, knowing that a graph contains citations cannot distinguish the attribute distributions of PubMed and Citeseer. Relation-level sharing and dataset-level adaptation must therefore coexist; a relation token is not a universal remedy for domain shift.
REEF borrows the vocabulary concept, but its tokens are neither discrete node IDs nor a conversion of graphs into language sequences. Relation names and task descriptions become semantic vectors that generate the parameters of graph predictors, while a GNN still processes graph structure. Core Idea: share the mapping from relation semantics to predictor parameters while correcting feature distributions with dataset descriptions, making knowledge about how the same relation operates across graphs transferable.
Method¶
Overall Architecture¶
Inputs comprise node attributes, relation-typed edges, relation descriptions, and a dataset description. The prediction unit is a triple \(\langle s_i,r,s_j\rangle\), whose endpoints are node-centered local subgraphs; for classification, the second endpoint is a label-prototype node. The output is a binary probability that the supplied relation holds, not freely generated text or multiclass token prediction over the relation vocabulary.
Numerical attributes are first aligned to 128 dimensions with SVD. Dataset adaptation produces an initial feature bias and a layer-specific projection of the node's own state; relation semantics generate aggregation matrices to encode local subgraphs, and another semantic hypernetwork generates the classifier that scores the two endpoints. Training mixes batches from different graphs and randomly masks edges. Inference does not follow the label-supervision or optimization paths.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
X["Node attributes and local graph<br/>SVD alignment"] --> D["Dataset Adaptation"]
T["Dataset description<br/>Sentence-BERT"] --> D
D --> A["Semantic Relation Aggregation"]
R["Relation vocabulary<br/>Sentence-BERT"] --> A
A --> C["Relation-Existence Classification"]
R --> C
C --> O["Inference: relation probability<br/>class or link decision"]
C -. "Training predictions" .-> M["Mixed-Graph Training"]
Y["Training labels and link supervision"] -. "Training only" .-> M
M -. "Edge masking and parameter updates" .-> D
M -. "Parameter updates" .-> A
M -. "Parameter updates" .-> C
The vocabulary comes from dataset schemas and task semantics rather than individual sample annotations. Main pretraining uses 237 FB15K237 relations, 11 WN18RR relations, and one aggregation/classification pair for each of the citation, webpage, and product domains, totaling 254 tokens. Appendix D states that an LLM generates natural-language relation descriptions from the background, official relation names, and task semantics before Sentence-BERT encodes them. This is a preparation step, not a requirement to invoke a generative LLM for every graph prediction.
Key Designs¶
1. Dataset Adaptation: distinguish dimensional alignment from distribution adaptation
SVD resolves incompatible input dimensions by giving numerical attributes from different graphs a common width that the predictor can accept. It does not guarantee that the first coordinate has the same meaning across graphs, nor is it supervised cross-domain semantic alignment. The text-attribute experiments instead encode node text directly with Sentence-BERT, forming a different input configuration; their 3-shot results should not be treated as results for the main SVD configuration.
Sentence-BERT then initializes a dataset-description vector, which the paper explicitly states is updated during training. A two-layer MLP maps this vector to a dataset-level feature bias added to every node's initial representation. A projector hypernetwork generates a dataset-specific matrix at each layer to transform the target node's own hidden state. Thus, this projection is not merely an input operation reducing features to 128 dimensions: it continuously adapts the self-information channel. Neighbor transformations for a citation relation can be shared, while PubMed and Citeseer retain different self-state transformations and initial biases.
These adaptations address a shared initial distribution shift and how a node retains and transforms its own attributes during propagation, respectively. They neither generate new classes directly nor replace relation aggregation. The cached version of Equation (10) ends with a formatting fragment; this note retains the identifiable additive-bias mechanism without reconstructing a damaged equation.
2. Semantic Relation Aggregation: use relation descriptions to determine neighbor transformations
Sentence-BERT initializes relation text as \(h_r\), and the aggregator hypernetwork generates a relation-specific matrix \(\Phi_r^{(l)}\) for each layer from that vector. Unlike RGCN, which directly learns a matrix for each relation ID, REEF inserts a learned mapping between relation semantics and propagation parameters. Parameters are therefore semantically conditioned rather than indexed only by unrelated IDs.
The final node update retains two channels: the dataset projector transforms the node's previous state, while neighbors are grouped by relation, transformed separately, averaged within each group, and summed. The identifiable mechanism in Equation (9) is:
Here, \(\mathcal N_i^r\) contains neighbors connected by relation \(r\), and computation only needs relation groups with actual neighbors. Within-group averaging prevents a relation from dominating solely because it has many neighbors; it is not learned attention across relations. The final center-node state represents the local subgraph, rather than a global average pool over all nodes.
The main model uses lightweight linear propagation, while the hypernetworks themselves remain nonlinear two-layer MLPs. The appendix compares ReLU and GELU and does not establish linear propagation as an indispensable innovation. The central mechanism is generating transformation matrices from semantics, not choosing one fixed activation.
3. Relation-Existence Classification: unify classification and link prediction through pair scoring
For node classification, each label is represented by the centroid of its training examples and treated as a single-node prototype subgraph. A query node's local subgraph is paired with every label prototype to ask whether a paper-classification or webpage-classification relation holds. For link prediction, both endpoints are node-centered local subgraphs, and the question concerns a supplied edge relation. Aggregation and prediction therefore need not use the same token: citation edges drive propagation, while paper classification drives class scoring.
The classifier hypernetwork generates classifier parameters \(\Psi_r\) from the semantics of the prediction relation. The two endpoint representations are multiplied elementwise and passed to a sigmoid scorer. This differs from a 254-way vocabulary softmax and does not require a unified class-index system across datasets. Equations (6)–(7) give:
Label prototypes should use only available training support examples, not test labels. With one support example per class in the 1-shot setting, a prototype reduces to that example's feature representative. The elementwise product is invariant to swapping endpoints, however, so this scoring equation alone does not establish direction-sensitive relation prediction. The cache does not explain additional encoding of direction or endpoint roles in detail.
Some hypernetwork parameters carry relation or dataset subscripts in the equations, while the narrative emphasizes semantically conditioned parameter generation. This note therefore does not assert that all relations share a single fixed-size generator, or that unseen relations obtain reliable predictions from text alone without adaptation. Appendix E.3 evaluates new relations through fine-tuning.
4. Mixed-Graph Training: share parameter-generation mechanisms without conflating all labels
Pretraining shuffles batches from seven source datasets into one queue and randomly removes edges each epoch at a mask rate of 0.2. The model repeatedly switches between graph structures and task relations, reducing the local bias of single-graph training, while edge perturbation encourages stability under topology changes. Mixing batches does not necessarily give datasets equal weight: the paper provides no dataset-uniform sampling or inverse-size weighting formula.
The source collection contains two knowledge graphs, two citation graphs, two WebKB graphs, and Photo. Relation-existence prediction connects pretraining and transfer, but supervision depends on the task: knowledge graphs provide link relations, while node classification supplies training-label prototypes. This is not purely unlabeled self-supervised pretraining on every source graph, and data augmentation alone does not make it GraphCL-style contrastive learning.
A Worked Example¶
Consider transferring from PubMed, Citeseer, and other source graphs to 1-shot classification on Cora. Cora has 7 classes with one labeled support node per class. These 7 nodes form 7 label prototypes, while the remaining nodes are divided between validation and testing at a 1:9 ratio.
For a query paper, extract its local citation subgraph, align attributes through SVD, add the bias generated from the Cora description, and encode its neighborhood using Cora's layer-specific self projections and the citation aggregation matrices. Pair the resulting center-node representation with each of the 7 prototype representations, then score them with the paper-classification classifier to obtain a class decision. The transferred knowledge concerns citation and classification relations; PubMed's three diabetes classes are not directly mapped into Cora's seven research areas.
Target adaptation fine-tunes the model and learns the target dataset representation, so this is few-shot adaptation rather than zero-shot transfer. It also differs from in-context learning that supplies support examples without parameter updates. The cache does not enumerate the exact final class-selection rule and all optimization details, so these are not supplemented as an author-specified algorithm.
Loss & Training¶
The paper specifies a binary relation-prediction objective but does not provide a complete loss equation, negative-sampling ratio, or all prototype implementation details in the main text and appendices. A conventional binary cross-entropy loss is therefore not inserted as if explicitly specified by the authors.
The main Appendix D configuration is Adam, learning rate 0.0002, a 2-layer GNN, hidden dimension 64, dropout 0.5, batch size 128, and 100 training epochs. All three hypernetworks and the feature-bias function use two-layer MLPs. Main pretraining uses only source training splits. PubMed, Citeseer, and Photo use a 6:2:2 split, while Wisconsin and Texas use official splits.
Main 1-shot transfer uses one labeled example per class on Cora, Cornell, and Computers. Appendix F instead describes the 26-graph experiment as fully-inductive and follows GraphAny splits, but does not fully specify REEF's source-graph selection or target-adaptation budget under that protocol. Its 70.75% mean should not be interpreted as the main 1-shot mean or automatically treated as a strict zero-shot result.
Key Experimental Results¶
Main Results¶
The following selection from Table 2 preserves 1-shot metrics on the 0–1 scale and retains the reported error terms; the checklist identifies them as standard deviations. Baselines follow their original protocols and do not use identical pretraining graph collections, so this is not a strictly controlled equal-pretraining-budget comparison.
| Target Dataset | Method | Accuracy | AUC | F1 |
|---|---|---|---|---|
| Cora | MDGPT | 0.4421 ± 0.08 | 0.7912 ± 0.05 | 0.4271 ± 0.08 |
| Cora | REEF | 0.4866 ± 0.04 | 0.8035 ± 0.03 | 0.4800 ± 0.05 |
| Cornell | GraphCL | 0.4175 ± 0.04 | 0.6350 ± 0.02 | 0.3500 ± 0.04 |
| Cornell | REEF | 0.4865 ± 0.08 | 0.7016 ± 0.03 | 0.3649 ± 0.07 |
| Computers | MDGPT | 0.3837 ± 0.09 | 0.7950 ± 0.04 | 0.4120 ± 0.08 |
| Computers | REEF | 0.4772 ± 0.06 | 0.8480 ± 0.03 | 0.4403 ± 0.02 |
Cora accuracy improves by 4.45 percentage points over MDGPT, and Computers by 9.35 percentage points. The paper's “+10.07%” and “+24.05%” are relative accuracy improvements, not percentage-point gains. The strongest Cornell baseline depends on the metric: GraphCL leads on accuracy and F1, while GCOPE+CL has the strongest baseline AUC at 0.6694. Thus, the single baseline shown above is not the strongest on all three metrics.
Ablation Study¶
The following table selects key means from Tables 6 and 10, all on the percentage scale. Dashes indicate that the controlled aggregator experiment does not separately report those three target-graph values here. Reported values are retained rather than replaced with recomputed results when the source table is questionable.
| Config | Cora Acc | Cornell Acc | Computers Acc | Transfer Mean | Overall Mean |
|---|---|---|---|---|---|
| REEF | 48.66 | 48.65 | 47.72 | 48.34 | 72.93 |
| Random relation initialization, REEF-LM | 46.58 | 44.72 | 43.46 | 44.92 | 69.43 |
| Without feature bias, REEF-FB | 40.07 | 44.51 | 42.09 | 42.22 | 68.21 |
| Without dataset projector, REEF-FP | 40.22 | 43.69 | 46.50 | 45.10 | 68.32 |
| Without edge augmentation, REEF-AGU | 40.81 | 43.69 | 46.68 | 43.73 | 69.98 |
| Shared aggregation matrix, Shared | — | — | — | 29.80 | 56.16 |
| Relation concatenation, Concat | — | — | — | 39.39 | 68.23 |
| Independent relation matrices, Relation-ID | — | — | — | 33.51 | 68.32 |
Removing feature bias reduces the transfer mean by 6.12 percentage points, the largest mean transfer deterioration in Table 6. Removing the projector, edge augmentation, and semantic relation initialization reduces it by 3.24, 4.61, and 3.42 points, respectively. Table 10 changes only the aggregator while keeping the classifier, projector, bias, and training protocol fixed. REEF exceeds Concat by 8.95 transfer points and relation-ID matrices by 14.83 points, more directly supporting semantic parameter generation rather than relation-specific capacity alone.
Several ablation means in Table 6 are arithmetically inconsistent with their per-dataset entries. The prose also calls REEF-LM the worst variant, although its reported transfer mean of 44.92 exceeds REEF-FB's 42.22. This table preserves the source values, and the drops use reported means; these issues should not be interpreted as evidence of statistical significance.
Key Findings¶
Different protocols address different questions. The following table marks their evidence boundaries rather than merging them into one leaderboard.
| Protocol and Source Table | REEF Result | Comparison and Interpretation |
|---|---|---|
| Node classification on five source graphs, Table 1 | Mean Acc 79.70% | RGCN joint 76.83%; REEF does not win every graph: Texas 70.27% versus RGCN 80.95% |
| Cross-domain evaluation on 26 graphs, Tables 3 / 15 | Mean Acc 70.75% | GraphAny 68.23%; REEF co-purchase-domain mean 71.00% is below GraphAny's 71.58% |
| Molecular zero-shot, Table 4 | BBBP / HIV AUROC 54.80 / 63.17 | GOFA 54.91 / 53.02; GIMLET 59.39 / 66.24, with different supervision scales |
| Text-node 3-shot, Table 5 | S1: 75.95 / 57.40 / 37.20; S2: 74.58 / 62.82 / 38.24 | Order: Cora / History / Ratings; S1 excludes E-commerce source graphs, S2 adds Photo |
| Unseen-relation fine-tuning, Table 9 | Cornell / Computers Acc 34.58 / 26.40 | Corresponding full-relation-coverage results: 48.65 / 47.72; unseen-relation adaptation is substantially harder |
| Eight heterophilous graphs, Table 13 | Mean Acc 66.06% | H2GCN 66.41%; not a comprehensive improvement over heterophily-specific methods |
- Relation sharing does not eliminate domain shift. Removing source coverage corresponding to Photo, Wisconsin, and Texas reduces Computers F1 from 44.03 to 6.10. Supporting a new relation does not imply preserving performance; this experiment includes fine-tuning and is not unseen-relation zero-shot transfer.
- The molecular experiment excludes molecular-domain pretraining and fine-tuning, but the cache does not sufficiently specify graph-level readout, task representation, or unlabeled prototype construction. The numbers can be recorded as the authors' zero-shot evidence without inventing a fully specified end-to-end implementation.
- The 26-graph evaluation contains a substantial regression on ogbn-arxiv: REEF scores 50.42%, compared with GraphAny's 57.79%. A leading overall mean does not establish suitability for every large graph or domain.
- Appendix E.10 reports 177.4M trainable REEF parameters, 8 hours on one A800 80GB, and 48.21 GiB peak memory, versus GCOPE's 93.3K parameters, 34 minutes, and 23.40 MiB. This is not an equal-budget resource comparison, but avoiding generative LLM inference does not imply inexpensive training.
Highlights & Insights¶
- Relation semantics generate predictor parameters rather than merely being appended to node features. They can change computation during both propagation and decision-making, and the controlled aggregator experiment is more informative than a broad comparison with RGCN alone.
- Separating aggregation relations from task relations avoids conflating citation with class membership. This can transfer to cross-domain graph tasks combining interaction prediction and entity classification, provided schemas and task descriptions are explicit.
- Dataset bias and layer-specific self projection clarify the appropriate sharing granularity: share relation mechanisms without forcing all node attributes into identical distributions. Generality in a foundation model can mean conditional adaptation rather than the absence of domain-specific state.
Limitations & Future Work¶
- The authors propose vocabulary expansion, higher-order compositional relations, and dynamic or temporal graphs; the current experiments do not establish these extensions as implemented capabilities. Reported scaling behavior comes from limited dataset combinations and hidden-dimension comparisons, not a fitted and validated universal power law.
- Few-shot transfer, 26-graph fully-inductive evaluation, molecular zero-shot transfer, and testing on source graphs must be read separately. Restricting labels to training examples alone does not prove a fully inductive node-classification pipeline; graph-structure visibility, SVD fitting scope, and target-adaptation budgets still require code inspection.
- Equations (4), (6), and (8) attach relation/dataset subscripts to hypernetwork parameters, while the narrative emphasizes semantic parameter generation. Parameter-sharing boundaries require implementation verification. The 177.4M parameters and high memory footprint also prevent judging the full model's size from its two GNN layers alone.
- Statistical reporting is incomplete: Table 2 includes error terms, but many extended tables omit variance and repetition counts. Inconsistent ablation means further limit rigorous significance claims.
- The cache contains a declaration conflict: Appendix D describes LLM-generated relation descriptions, while Appendix J and the checklist say LLMs were used only for writing refinement. Knowledge-graph split summaries and Citeseer/Photo statistics also differ across configurations; reproduction should verify the exact experimental versions rather than silently standardize them.
- Further experiments could test synonymous, incorrect, and semantically empty relation descriptions to separate language priors, parameter capacity, and domain identification. Lightweight generators and relation-conditioned low-rank adaptation should also be compared under matched training budgets.
Related Work & Insights¶
- vs RGCN: RGCN models multirelational propagation with relation-specific matrices. REEF generates matrices through semantic hypernetworks and adds relation classifiers and dataset adaptation. Different total models and parameter budgets mean the full gap cannot be attributed to relation text alone.
- vs GCOPE / MDGPT: GCOPE connects graphs through coordinator nodes, while MDGPT uses domain tokens and prompts for adaptation. REEF shares at the relation level, with dataset identity conditioning feature adaptation. Main transfer results are stronger, but pretraining-data protocols differ.
- vs GraphAny / TS-Mean: These methods target transferable cross-graph node classification; REEF emphasizes the semantic-to-parameter mapping. The 26-graph mean supports cross-domain effectiveness, but REEF's source graphs and adaptation details under that protocol need clarification before strict independence can be assessed.
- vs GOFA / GIMLET: These methods use graph-language modeling or molecular graph-text pretraining, whereas REEF's inference core is a GNN with generated parameters. The molecular zero-shot table compares different supervision scales and supports feasibility, not dominance over all graph-text foundation models under identical training conditions.
- Classification assessment: Retain graph_learning. The main contribution is relation-conditioned graph prediction and cross-graph transfer, not a general unlabeled self-supervised objective; self_supervised would be less precise.
Rating¶
- Novelty: 4/5 — Combines relation semantics, parameter generation, and dataset adaptation in a unified graph-transfer framework.
- Experimental Thoroughness: 4/5 — Covers multiple protocols and controlled aggregator ablations, but inconsistent means and reproduction boundaries reduce confidence.
- Writing Quality: 3/5 — The main mechanism is clear, while some declarations, statistics, and extended protocols are inconsistent.
- Value: 4/5 — Relations are an informative transfer unit, but practical benefits should account for target-relation coverage and training cost.