Understanding and Mitigating Under-Confidence in GNNs from the Final Layer¶
Conference: NeurIPS 2026
arXiv: 2505.11335
Code: https://github.com/huangJC0429/SCAR
Area: Graph Learning
Keywords: confidence calibration, under-confidence, final-layer weight decay, class prototypes, node-level calibration
TL;DR¶
SCAR explains GNN under-confidence through the final classifier, reduces only final-layer weight decay during training, and moves representations toward predicted-class prototypes at inference time using training-set adjacency groups, reducing ECE across multiple node classification settings without guaranteeing unchanged labels or accuracy.
Background & Motivation¶
A graph neural network (GNN) can classify many nodes correctly while assigning them relatively low maximum class probabilities: classification accuracy and confidence are different quantities. GCN and GAT often exhibit this under-confidence on citation networks, unlike the over-confidence seen in many deep networks. Temperature scaling (TS) globally adjusts the sharpness of predictions, while graph calibration methods such as CaGCN, GATS, and GETS exploit graph structure or additional models to generate calibration parameters. These external components, however, do not directly explain why the original model produces low confidence.
This paper focuses on the final operation that multiplies aggregated representations by classifier weights. Final-layer weight decay shrinks those weights; at an optimization stationary point, the weights can also be interpreted as class prototypes. Calibration therefore need not require learning another mapping: it can first change the classifier's training constraint and then adjust test representations relative to its prototypes. These two scales are not interchangeable. Greater overall prototype separation cannot eliminate individual representation deviations, and message passing exposes training-near and training-distant nodes to different amounts of training information.
SCAR consequently combines two complementary levels rather than making the entire method a training-free post-hoc procedure. Core idea: weaken only final-layer weight decay during training to improve overall prototype separation, then interpolate node representations toward predicted-class prototypes using structural groups, directly intervening in the final-layer geometry that generates confidence.
Method¶
Overall Architecture¶
Inputs are graph structure, node features, and training-node labels; outputs are class distributions for test nodes. SCAR retains the GCN or GAT backbone and cross-entropy objective, first performing “Class-Level Calibration” by training the classifier with less weight decay than other layers. At inference time, “Structure-Based Weighting” selects a weaker or stronger interpolation coefficient, and “Prototype Interpolation” modifies the final aggregated representation before probabilities are recomputed with the same classifier.
A “class centroid” here is not a simple average of same-class training representations. The paper uses each class column of the final weight matrix as a prototype. Its theoretical interpretation comes from a stationary condition for regularized cross-entropy: same-class representations are weighted by prediction residuals and accumulated, while weighted representations from other classes are subtracted. Since softmax probabilities depend on the weights themselves, this is an implicit condition, not a closed-form optimum obtained by inserting the data once.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Graph, features, training labels"] -->|Training: cross-entropy supervision| B["Class-Level Calibration"]
B --> C["Freeze GNN and final prototypes"]
C -->|Inference: graph and training-node locations| D["Structure-Based Weighting"]
D --> E["Prototype Interpolation"]
E --> F["Recompute logits and probabilities"]
Training labels supervise the GNN. At inference time, the initial predicted label selects a prototype, while training-node locations determine the structural group; test ground-truth labels are not read. Post-processing introduces no additional calibration network requiring gradient training. Nevertheless, selecting final-layer decay involves backbone training, and interpolation strength requires hyperparameter selection: this does not establish zero training cost for the complete procedure.
Key Designs¶
1. Class-Level Calibration: weaken final-layer decay without simultaneously perturbing every representation layer
Reducing weight decay globally changes both the classifier and preceding representation learning, potentially harming accuracy. SCAR changes only the final-layer coefficient while retaining regularization elsewhere. Theorem 3.1 separates the probability after a gradient update into weight-shrinkage and cross-entropy-gradient contributions. The former can be expressed as a temperature factor acting on the corresponding logit term:
Here \(\eta\) is the learning rate and \(\lambda^{(K)}\) is final-layer decay. A positive temperature above 1 makes a fixed logit distribution flatter, motivating smaller decay as a way to alleviate under-confidence; with zero decay, this factor equals 1. The derivation also retains a class-dependent correction from the cross-entropy gradient and incorporates updated representations. The entire training process is therefore not exactly equivalent to a single TS operation, nor does the decomposition automatically describe every step of Adam.
In the cache, Eqs. (15)–(17) in Appendix B.1 contain transcription concerns involving pre-update versus post-update weights, class indices, and repeated exponential terms. This note does not reconstruct the complete probability formula; it retains only the readable temperature factor and its scope.
Theorem 3.2 interprets each final-layer class column as a discriminative prototype: insufficiently confident same-class training nodes contribute positively, while other-class nodes contribute negatively through their erroneous probability assigned to that class. This is neither an unweighted arithmetic mean nor a relation necessarily satisfied after any finite number of training steps. Prototypes also encode exclusion of other classes, so proximity is measured in a learned discriminative space rather than against the mean input feature of a class.
Proposition 3.3 has a narrower guarantee. With final-layer input representations fixed, it compares classifiers optimized under two positive decay coefficients and uses the centered representative of softmax weights. Smaller decay yields no smaller average pairwise squared prototype distance. The proof first compares optimal Frobenius norms, then uses a centering identity to connect those norms to aggregate prototype distances. It does not guarantee that every prototype pair moves farther apart or that end-to-end retraining unconditionally lowers ECE. Class-level separation is an aggregate geometric result, not a universal monotonicity theorem for calibration error.
2. Structure-Based Weighting: correct first-order training neighbors less and remaining nodes more
Message passing gives some test nodes easier access to training-node information. A uniform interpolation strength may overcorrect training-near nodes while undercorrecting distant ones. First-order neighbors of training nodes form \(V_F\) and receive the smaller coefficient \(\alpha\); second-order and higher-order nodes form \(V_S\) and receive the larger \(\beta\), with \(0\leq\alpha<\beta\leq1\). This is a two-level structural rule, not a separately trained node-wise temperature predictor.
Proposition 3.4 motivates this choice through a linearized GNN with a nonnegative, row-normalized propagation matrix. It additionally assumes tighter prototype-deviation bounds for transformed features of same-class training nodes. Under these conditions, weaker same-class training influence yields a looser upper bound on prototype distance. When graph distance to that class's training set exceeds the propagation depth, same-class training influence is zero. A looser bound, however, does not imply that actual distance must increase or that every distant node must be more under-confident.
Theory and implementation must also be distinguished. The theoretical quantity is the total propagation weight from same-class training nodes; the implemented grouping uses adjacency to the training set, a coarser proxy. General nonlinear GNNs or symmetrically normalized GCNs cannot be assumed to satisfy all conditions of this row-normalized linear model. Applicability to heterophilic graphs is likewise different from an unconditional theoretical guarantee for arbitrary graphs.
3. Prototype Interpolation: select a prototype using the predicted class and recompute the distribution
After training, the original GNN first predicts a class \(\hat y_t\) for each node, selecting the corresponding final-layer weight column as its prototype. It then applies one interpolation:
The adjusted representation passes through the same final classifier and softmax, without adding a calibration model or back-propagating into the backbone. Moving toward the predicted prototype intuitively raises its class score. Through the stationary condition, the test logit can be interpreted as the combination of a positive same-class similarity term and a negative other-class similarity term. This connects class-level constraints and node-level corrections, but inner-product scores must not be equated with unconditional Euclidean nearest-prototype classification.
Prototype selection uses the predicted class, not the true class. An incorrect initial prediction can therefore receive stronger erroneous confidence. Unequal prototype norms or unfavorable cross-prototype inner products can also change the argmax after recomputation. Positive-temperature TS preserves class ordering, but representation interpolation does not automatically inherit that property. Appendix E.2 claims that node-level processing “does not alter” predictions; Eqs. (9) and (10) provide no general order-preservation conditions, so this note treats the statement as requiring verification rather than as a method guarantee.
A Worked Example¶
The following is a constructed arithmetic example, not reported node data. Suppose a binary node has final representation \((0.4,0.2)\) and class prototypes \((1,0)\) and \((-1,0)\). Its logits are \((0.4,-0.4)\), so the initial prediction is the first class. SCAR selects that prototype without querying the test label.
For a node adjacent to the training set, the weaker \(\alpha=0.0002\) produces \((0.40012,0.19996)\). For a training-distant node, the stronger \(\beta=0.002\) produces \((0.4012,0.1996)\). Recomputing probabilities slightly increases first-class confidence in both cases, with a larger correction for the distant node. These coefficients come from the GCN/Cora, L/C=20 configuration in Table 6.
Symmetric prototypes prevent a label change in this example, not in general. Learned prototypes can have unequal norms. A reliable evaluation should jointly inspect calibration, accuracy, and flips between original and adjusted predictions rather than checking only whether average confidence increases.
Loss & Training¶
Backbone training still uses cross-entropy with layer-wise \(L_2\) regularization. Class-level calibration only reduces the final-layer coefficient; node-level calibration runs after training and is not an additional supervised loss. Experiments use two-layer GCN/GAT and Adam. Other layers retain decay of \(5\times10^{-4}\); learning rates are selected from 0.01, 0.015, and 0.02, hidden dimensions from 8, 16, and 64, and dropout from 0.5 and 0.6.
For GCN/Cora with L/C=20, final-layer decay is \(10^{-4}\), \(\alpha=2\times10^{-4}\), and \(\beta=2\times10^{-3}\). The paper recommends binary search for final-layer decay and a partial grid search within the triangular region \(\beta>\alpha\) for interpolation coefficients. No calibration-network training does not mean no tuning. Reproduction should select these values on validation data and freeze them before testing, but the cache does not clearly specify the validation objective or complete search budget, so that selection protocol cannot be considered verified.
Hyperparameter descriptions conflict. Appendix D.3 gives the reversed interval \((0.0005,0.000005]\); Appendix E.1 describes a scan over [0,0.001], whereas Tables 6/7 contain larger \(\beta\) values such as 0.002 and 0.005, as well as smaller \(\alpha\) values. This note retains the theoretical constraint and explicit table example rather than inventing a consistent search range.
Key Experimental Results¶
Main Results¶
The main task is semi-supervised node classification on Cora, Citeseer, Pubmed, and CoraFull, using 20, 40, or 60 training labels per class, 500 validation nodes, and 1000 test nodes. Calibration uses 20-bin ECE: the absolute difference between accuracy and mean maximum class probability in each bin is weighted by that bin's sample fraction and summed. Lower is better, but this evaluates top-confidence calibration under one binning scheme, not calibration of the entire class distribution.
The following subset of Tables 1/2 uses L/C=20. Values are ECE (%), and \(\pm\) denotes the standard deviation of 10 runs, not the standard error. TS and GETS are retained as references; GETS is not assumed to be the strongest baseline in every setting.
| Backbone | Dataset | Uncalibrated | TS | GETS | SCAR |
|---|---|---|---|---|---|
| GCN | Cora | 13.47 ± 0.63 | 4.88 ± 0.55 | 3.83 ± 0.41 | 3.35 ± 0.53 |
| GCN | Citeseer | 12.48 ± 0.71 | 6.41 ± 0.87 | 6.72 ± 0.78 | 3.43 ± 0.58 |
| GCN | Pubmed | 5.86 ± 0.77 | 5.41 ± 0.38 | 3.93 ± 0.37 | 3.81 ± 0.47 |
| GCN | CoraFull | 19.86 ± 0.61 | 10.31 ± 0.61 | 7.01 ± 0.85 | 6.96 ± 0.48 |
| GAT | Cora | 15.58 ± 0.89 | 7.17 ± 0.98 | 4.53 ± 0.48 | 3.52 ± 0.74 |
| GAT | Citeseer | 15.34 ± 0.50 | 9.16 ± 0.87 | 5.85 ± 0.33 | 4.37 ± 0.83 |
| GAT | Pubmed | 8.35 ± 0.31 | 6.56 ± 0.46 | 4.16 ± 0.37 | 3.78 ± 0.84 |
| GAT | CoraFull | 21.19 ± 0.36 | 11.01 ± 0.51 | 6.79 ± 0.71 | 6.41 ± 0.63 |
The claim of always being best should not be repeated. For GAT/Pubmed at L/C=20, DCGC obtains 3.52 ± 0.21 and CaGCN obtains 3.56 ± 0.63, both below SCAR's 3.78 ± 0.84. SCAR and GETS also have very similar means on GCN/CoraFull; a small mean difference alone does not establish statistical significance.
Ablation Study¶
CLC, NLC, and full SCAR below come from Table 3; the single-coefficient variant comes from Table 10. All values are ECE (%) at L/C=20. NLC alone post-processes the original trained model, whereas the full method combines reduced final-layer decay with node-level correction.
| Config | GCN/Cora | GCN/Citeseer | GCN/CoraFull | GAT/Cora | GAT/Citeseer | GAT/CoraFull |
|---|---|---|---|---|---|---|
| CLC only | 3.72 ± 0.53 | 4.92 ± 0.79 | 7.09 ± 0.52 | 3.64 ± 0.63 | 5.98 ± 0.62 | 6.89 ± 0.83 |
| NLC only | 4.08 ± 0.62 | 3.63 ± 0.63 | 9.23 ± 0.55 | 4.18 ± 0.72 | 4.96 ± 0.66 | 7.12 ± 0.45 |
| Full SCAR | 3.35 ± 0.65 | 3.43 ± 0.58 | 6.96 ± 0.48 | 3.52 ± 0.74 | 4.37 ± 0.83 | 6.41 ± 0.63 |
| One node-correction coefficient | 3.81 ± 0.61 | 3.91 ± 0.64 | 7.21 ± 0.52 | 3.94 ± 0.71 | 4.88 ± 0.78 | 6.73 ± 0.57 |
Component contributions depend on the dataset: CLC alone clearly outperforms NLC alone on GCN/CoraFull, whereas the reverse holds on GCN/Citeseer. Two coefficients improve on one, providing empirical support for structural grouping, not evidence that all structural biases have been modeled.
Source inconsistencies must be retained. Full SCAR on GCN/Cora is 3.35 ± 0.53 in Table 1 but 3.35 ± 0.65 in Tables 3/10. GCN/Pubmed is 3.81 ± 0.47 in Table 1, 3.78 ± 0.53 in Table 3, and 3.81 ± 0.53 in Table 10. GCN/CoraFull TS is 10.31 ± 0.61 in Table 1 and 10.13 ± 0.61 in Table 3. Values are reported according to their source tables rather than silently reconciled.
Key Findings¶
Accuracy should not be summarized only as an overall improvement. The following subset of Table 8 reports accuracy (%) and illustrates both gains and regressions from the calibration procedure.
| Backbone and dataset | L/C | Uncalibrated accuracy | SCAR accuracy | Mean change (percentage points) |
|---|---|---|---|---|
| GCN/Cora | 40 | 83.04 ± 0.24 | 82.70 ± 0.16 | -0.34 |
| GAT/Pubmed | 20 | 78.88 ± 0.31 | 78.62 ± 0.36 | -0.26 |
| GCN/CoraFull | 20 | 62.03 ± 0.27 | 64.17 ± 0.52 | +2.14 |
| GAT/CoraFull | 20 | 58.87 ± 0.35 | 62.91 ± 0.32 | +4.04 |
- Table 4 extends evaluation to GraphSAGE, FAGCN, and GCNII. For example, GCNII/Cora ECE decreases from 35.51 ± 0.56 to 9.81 ± 0.48, showing that gains are not limited to the two basic backbones.
- Table 9 covers Chameleon, Squirrel, and the larger Arxiv-year graph. GCN ECE on Arxiv-year falls from 13.34 ± 0.72 to 6.03 ± 0.38, but these results do not replace evaluation under real dynamic distribution shifts.
- Table 11 constructs OOD graphs by preserving labels and rewiring edges through a stochastic block model. SCAR reaches 4.73 ± 0.58 on Cora-OOD versus 7.15 ± 0.32 without calibration. The paper also notes that a lower uncalibrated OOD ECE can result from reduced accuracy rather than greater reliability.
- Figure 8 supports low additional runtime for node-level processing, but the text provides no auditable per-method timing values. Decay search and repeated training budgets still require separate accounting; lightweight single-pass inference does not imply no additional cost for the complete procedure.
Highlights & Insights¶
- Locating calibration at the final layer avoids assuming that the entire graph or backbone must be redesigned. It provides a practical parameter-grouping perspective: representation learning may need regularization without equally compressing classification scores.
- Discriminative prototypes connect an optimization stationary condition with test-node scoring. This explains why decay and representation position can be complementary controls, while emphasizing that prototypes are not simply mean class features.
- Two-level interpolation turns structural influence into an inspectable operation without learning an extra calibration network. Its value lies in simplicity and ablatability, not in a guarantee of node-level correctness.
Limitations & Future Work¶
- The authors acknowledge that homophily, heterophily, and other structural biases are not fully exploited. First-order adjacency grouping is coarse; future work could consider training influence associated with the predicted class instead of adjacency to any training node.
- Moving toward the predicted prototype can amplify mistakes. Further work should consider node error risk, prototype norms, and class-order constraints, separately reporting calibration for correct predictions, incorrect predictions, and prediction flips.
- Theory concerns a single-step gradient decomposition, final-layer optimization with fixed representations, and a linearized propagation upper bound. It does not directly prove monotonic ECE improvement for actual Adam training or nonlinear models. Already over-confident models may also be unsuitable for further confidence increases.
- Twenty-bin ECE can obscure class-specific, structural-subgroup, and low-sample-bin failures. NLL, Brier score, class-conditional calibration, and binning sensitivity should supplement it rather than treating one ECE as a complete reliability certificate.
- Search protocols and text–table consistency affect reproducibility. Beyond the result conflicts above, Citeseer and Pubmed training counts in Table 5 conflict with the stated 20/40/60 labels per class. Actual splits require code verification; the offline cache cannot resolve their origin.
Related Work & Insights¶
- vs TS: TS applies a positive scalar temperature to fixed logits and preserves argmax. SCAR changes training at the class level and representations at the node level, without the same order-preservation guarantee. They should not simply be grouped as training-free calibration methods.
- vs CaGCN / GATS / GETS: These methods use additional calibration components to generate temperatures or combine calibration outputs. SCAR intervenes directly in backbone constraints and node representations, reducing inference components while still requiring tuning and evaluation budgets.
- vs AU-LS / GCL: These methods alter training objectives through confidence-related constraints. SCAR adds no separate calibration loss, instead changing layer-wise regularization and adding post-processing. Accuracy and ECE require joint comparison.
- vs graph uncertainty quantification and conformal prediction: Such methods may target uncertainty decomposition or prediction-set coverage. SCAR targets empirical calibration of maximum class probabilities; its theoretical interpretation must not be upgraded into a coverage guarantee.
- Research direction: With representations fixed, compare prototype-norm constraints, prototype centering, and order-preserving node corrections to separate confidence increases from calibration improvements. This is a follow-up direction proposed by the note, not a result established by the paper.
Rating¶
- Novelty: 4/5. Connects final-layer regularization, discriminative prototypes, and node corrections through a clear mechanism, although the operations are simple.
- Experimental Thoroughness: 4/5. Includes multiple backbones, structural ablations, heterophilic graphs, and limited OOD settings, but lacks broader calibration metrics and a transparent search budget.
- Writing Quality: 3/5. The main argument is accessible, but some formula transcriptions, table values, and universal claims conflict.
- Value: 4/5. Offers a lightweight, reusable GNN calibration baseline whose applicability should be checked through validation and joint accuracy evaluation.