Skip to content

HoTS: Homophily-Aware Temperature Scaling for Graph Neural Network Calibration

Conference: NeurIPS2026
arXiv: 2609.32426
Code: https://github.com/inu0104/HoTS
Area: Graph Learning
Keywords: graph neural networks, probability calibration, local homophily, temperature scaling, selective classification

TL;DR

HoTS generates positive node-wise temperatures from predictive entropy and local homophily estimated by an auxiliary GCN, using an idealized CSBM inverse-homophily temperature law as a structural prior; it preserves predicted classes and achieves the best mean ECE of 4.79% across 18 datasets, but does not win on every dataset or coverage level.

Background & Motivation

Confidence in graph node classification is not merely a displayed probability: it determines which nodes are deferred to human review and which predictions enter risk-sensitive decisions. Global temperature scaling (TS) applies one temperature to every node, correcting overall overconfidence without changing classes, but cannot distinguish neighborhood-specific reliability. In a message-passing GNN, identical logits may reflect support from same-class neighbors or an incidental strong response in a mixed neighborhood; confidence magnitude alone need not reveal their true correctness probabilities.

CaGCN, GATS, GETS, and WATS already incorporate graph structure into calibration, typically through auxiliary networks, attention, mixtures of experts, or graph wavelets. Rather than adding another graph feature, this paper asks what relationship should hold between local structure and optimal temperature. It first shows that if homophily explains optimal-temperature variation after conditioning on logits, any logit-only temperature rule leaves a cross-entropy risk gap. It then calculates a posterior in an analytically tractable contextual stochastic block model (CSBM), connecting neighborhood homophily to aggregated signal strength.

The theoretical law is not directly deployable: true classes are unknown, and practical GNNs are not the theoretical template classifier. The paper therefore trains a homophily predictor from observed labels and combines predictive entropy with the estimated structural signal in a constrained temperature function. Core Idea: determine node-wise temperature from how concentrated the current prediction is relative to how much support its neighborhood supplies, and learn the sensitivity of this structural correction instead of assigning every node one global temperature.

Method

Overall Architecture

The inputs are a trained node classifier, its frozen logits, and the full graph structure and node features. HoTS first obtains a structural signal through “Homophily Estimation,” summarizes uncalibrated prediction concentration through “Entropy Concentration Proxy,” and combines them in “Structural Temperature Map” to produce a positive scalar temperature and a temperature-scaled softmax.

The homophily predictor is trained separately and its outputs are cached before fitting the temperature function's three scalars. No true node labels are required at inference time; test-node features and edges may participate in transductive GCN computation, but test labels cannot participate in fitting. “Three parameters” refers only to the temperature map, not the auxiliary GCN weights.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    G["Graph and node features"] --> H["Homophily Estimation"]
    L["Training and validation labels"] -.->|Training supervision only| H
    Z["Frozen classifier logits"] --> E["Entropy Concentration Proxy"]
    H --> T["Structural Temperature Map"]
    E --> T
    V["Validation labels"] -.->|Temperature CE fitting| T
    R["Training labels"] -.->|Checkpoint selection| T
    T --> P["Positive-temperature softmax<br/>Predicted class unchanged"]
    Z --> P

Solid edges represent computational data flow; dashed edges represent label access during fitting or checkpoint selection. All label-supervision edges are removed at inference time. The two input branches merge at Structural Temperature Map, and the key designs follow the order Homophily Estimation, Entropy Concentration Proxy, and Structural Temperature Map.

Key Designs

1. Homophily Estimation: learn an inference-time structural signal from observed neighbor labels

Local homophily is the fraction of neighbors sharing a node's class, but computing it at test time cannot rely on the true classes of that node and its neighbors. HoTS uses a two-layer auxiliary GCN taking raw features and the full transductive graph, with hidden width 32, ReLU, and a scalar sigmoid output as estimated homophily. This network predicts an interpretable structural intermediate rather than calibration temperature directly.

Supervision targets are constructed only for known-label nodes in the training and validation sets: collect neighbors whose labels are also known, exclude the center node itself, and calculate the same-class fraction. Known nodes without such neighbors are excluded from the predictor loss. Target construction excludes self-loops, whereas GCN message passing retains them; these are different choices. Features and graph structure let the predictor estimate homophily for nodes without labeled neighbors, without assigning pseudo-label supervision to those nodes.

This avoids test-label leakage, but does not constitute a fully independent held-out calibration protocol: validation labels contribute to homophily targets and subsequent temperature fitting, and also select the base classifier. The auxiliary predictor is frozen after training, and its estimates are cached for all nodes; temperature optimization does not update it. Results should therefore be interpreted in terms of both the structural prior and auxiliary supervision, rather than describing the complete system as learning only three numbers.

2. Entropy Concentration Proxy: extract a node-wise scale from multiclass probabilities

Temperature cannot depend only on homophily because raw prediction concentration still varies within a structural group. In binary classification, a positive logit gap \(\delta_i\) and a desired homophily-group correctness rate \(q(h_i)\) give \(T_i^{\mathrm{match}}=\delta_i/\operatorname{logit}(q(h_i))\). This is a group-reliability matching surrogate, not the pointwise cross-entropy optimum under the full CSBM posterior; positive-temperature matching also requires \(q(h_i)>1/2\).

Multiclass classification has no unique binary gap, so the paper constructs a concentration proxy from normalized predictive entropy:

\[ e_i=\frac{H(\operatorname{softmax}(z_i))}{\log K},\qquad c_i=\sqrt{2K\log K(1-e_i)}. \]

Here \(H(p)=-\sum_k p_k\log p_k\), \(K\) is the number of classes, and entropy is computed before temperature scaling. Near uniform probabilities, the second-order entropy deficit is proportional to the squared norm of centered logits, making this square root a proxy for logit concentration. Appendix A.5 explicitly restricts that expansion to a neighborhood of the uniform distribution; using \(c_i\) over the full entropy range is a monotone proxy, not a globally exact identity.

The numerator increases with prediction concentration, allowing sharply peaked raw predictions to receive a stronger temperature correction. This does not imply that every more confident prediction is less reliable: the structural denominator and learned parameters determine the final correction. For uniform predictions, the proxy approaches zero and temperature returns to the base offset.

3. Structural Temperature Map: turn a restricted theoretical law into stable positive temperatures

The theory assumes uniform class priors, independent Gaussian features, equidistant class means, and a one-layer linear GCN with self-loops. It additionally uses a population-concentration approximation: conditional on oracle local homophily, degree is fixed at its population value and off-class neighbors are balanced across the other classes, excluding finite-degree and composition fluctuations. These assumptions make the class-conditional aggregate distribution Gaussian with a common covariance.

Let \(m_c=\mu_c-\bar\mu\) and \(\rho_0=\|m_c\|\); the theoretical class-template score is \(s_c=\rho_0^{-1}m_c^\top(\bar x_i-\bar\mu)\). Normalized homophily moves the class-random neighborhood baseline to zero, and the signal and exact posterior are:

\[ \tilde h=\frac{Kh-1}{K-1},\quad v=\frac{\sigma^2}{\bar d+1},\quad \Gamma(h)=\frac{\rho_0(1+\bar d\tilde h)}{\bar d+1},\qquad P(Y_i=c\mid s,H_i=h)=\operatorname{softmax}\!\left(\frac{\Gamma(h)}{v}s\right)_c. \]

The proof does not assume that any trained GCN is intrinsically calibrated; it cancels class-independent terms in the Gaussian likelihood. Equidistant centered means form a regular simplex and have equal norms, leaving likelihood ratios proportional to template inner products; feature components orthogonal to the template span carry no class information. The resulting signed Bayes inverse-temperature \(\Gamma(h)/v\) applies only to these template scores, not to an arbitrary learned linear head.

When \(\Gamma(h)>0\), the positive Bayes temperature is \(\sigma^2/[\rho_0(1+\bar d\tilde h)]\). Only in the homophilic regime \(\bar d\tilde h\gg1\) is it approximately proportional to \(\tilde h^{-1}\). A class-random neighborhood with \(\tilde h=0\) still retains positive signal from the node's own feature, so the asymptotic inverse law cannot justify claiming that exact temperature diverges there. If \(\Gamma(h)<0\), the exact Bayes rule reverses the template-class ordering, which no positive temperature scaling can implement.

The strict improvement in Proposition 1 is also conditional: after fixing logits, the optimal positive inverse-temperature must retain nonzero conditional variance with homophily, with a unique interior optimum, nonconstant logits, and local derivative regularity. Its proof uses strict convexity of cross-entropy in inverse-temperature; a local quadratic expansion expresses the risk gap as a curvature-weighted conditional variance. Appendix A.3 establishes strictness within positive-signal template-score groups through overlapping Gaussian supports for different homophily values, not a universal gap on every real graph.

The practical map substitutes estimated for oracle homophily and adds an absolute value, stabilizer, and positive offset:

\[ \hat{\tilde h}_i=\frac{K\hat h_i-1}{K-1},\qquad T_i=T_{\mathrm{base}}+\frac{\beta\sqrt{2K\log K(1-e_i)}}{(|\hat{\tilde h}_i|+\varepsilon)^\alpha},\qquad p_i^{\mathrm{cal}}=\operatorname{softmax}(z_i/T_i). \]

The absolute value allows stable positive temperatures with weak or negative estimated signals; positive and negative signals of equal magnitude receive the same denominator, so this is not an exact signed Bayes correction in heterophilic regions. The entropy proxy, base offset, learned exponent, and absolute value form a heuristic stable extension: the theory supplies a structural prior, not an exact optimal formula for real networks.

The learnable exponent \(\alpha\) is motivated by measurement error. Under the positive-temperature inverse approximation, let log-homophily be \(u\) and its observation be \(\hat u=u+\eta\), with independent, zero-mean noise of variance \(\tau^2\). Population least-squares regression of \(\log T^*=a-u\) gives:

\[ \alpha^*=\frac{\operatorname{Var}(u)}{\operatorname{Var}(u)+\tau^2}\leq1. \]

The proof retains negative true-signal variance in the covariance while adding noise variance to the explanatory variable's variance. This is an attenuation law for log-space OLS, not a theorem for the cross-entropy objective used by HoTS; implementation requires only \(\alpha>0\) and imposes no upper bound of 1. The synthetic fitted values 1.009 without noise and 0.182 at maximum noise are empirical results, not evidence of an enforced \(\alpha\leq1\) constraint.

Finally, every class logit for a node is divided by the same positive number, preserving within-node ordering and the unique predicted class; confidence ordering across nodes can nevertheless change. This is the source of selective-classification gains and also means that HoTS cannot improve accuracy when every node is retained.

Loss & Training

The base GCN/GAT models have two layers and first optimize cross-entropy on training labels, usually with validation-loss early stopping; PubMed, ogbn-arxiv, and Reddit run the full 200 epochs. All datasets receive newly generated random 20/10/70 training/validation/test splits for each seed, including ogbn-arxiv rather than its official temporal split.

The auxiliary homophily predictor optimizes mean squared error on eligible known-label nodes, using Adam with learning rate \(10^{-2}\), weight decay \(10^{-4}\), at most 200 epochs, and supervised-training-loss early-stopping patience 30. Its targets combine training and validation labels; no additional independent predictor-validation set is described.

With the base classifier and homophily estimates frozen, temperature parameters optimize cross-entropy on the validation set; training-set cross-entropy selects checkpoints and controls early stopping. Adam uses learning rate \(10^{-2}\), no weight decay, at most 1,000 epochs, and patience 50. The temperature-fitting and early-stopping sets must not be reversed, and the protocol should not be described as fully independent held-out calibration.

Softplus keeps all three parameters positive: the base temperature receives an additional 0.1, and the scale and exponent each receive 0.01; the denominator stabilizer is \(\varepsilon=0.02\). Effective initial values are approximately 1.07 for base temperature, 0.98 for scale, and 0.70 for exponent. These parameters let data weaken or strengthen the homophily correction without forcing real graphs to follow the theoretical asymptotic exponent.

Key Experimental Results

Main Results

The main results pool 18 datasets, GCN and GAT backbones, and 10 seeds per backbone. ECE uses 15 equal-width confidence bins, summing each bin's sample proportion times the absolute difference between mean confidence and empirical accuracy; table values are percentages, with lower being better. The following selection from source Table 1 includes both successes and failures. Best baseline denotes the lowest baseline mean on that dataset, not one fixed method.

Dataset / summary HoTS ECE Best baseline ECE Baseline method
Cora 2.50 ± 0.72 2.90 ± 0.78 VS
PubMed 1.26 ± 0.30 0.83 ± 0.30 WATS
CoraFull 2.73 ± 0.81 2.20 ± 0.63 ETS
Texas 15.37 ± 5.90 17.69 ± 6.71 HTS
Wisconsin 14.67 ± 5.38 16.81 ± 6.81 ETS
Roman-Empire 4.36 ± 0.78 2.10 ± 0.83 VS
tolokers 3.25 ± 0.78 2.04 ± 0.63 GETS
Mean across all datasets 4.79 5.37 HTS
Average rank 3.33 3.94 HTS

The results support best average performance, not wins on all homophilic or heterophilic graphs. Beyond the counterexamples above, CiteSeer, Computers, Photo, CS, Physics, Chameleon, Squirrel, Actor, ogbn-arxiv, and Reddit also have baselines with lower means. Large standard deviations on small WebKB graphs caution against inferring statistical significance from mean ordering alone.

Appendix Table 3 reports mean NLL of 0.92 for HoTS versus 0.94 for HTS and ETS; degree-stratified ECE instead favors CaGCN at 5.81 over HoTS at 5.90, so HoTS is not best on every metric. GATS is missing on Reddit because of the 24 GB memory limit, making its average comparison incomplete over the same dataset set. The additional-backbone experiment in Appendix B.3 covers 17 datasets, excludes Reddit, and ranks HoTS second on GCNII rather than first on every backbone.

Selective classification retains the highest-ranked nodes by calibrated confidence and macro-averages retained-node accuracy across datasets, backbones, and seeds. Source Table 2 shows only selected class-preserving baselines; values are percentages:

Coverage Uncalibrated TS HTS HoTS
100% 67.39 67.39 67.39 67.39
95% 68.87 68.90 68.92 68.94
90% 70.01 70.08 70.13 70.15
85% 71.04 71.12 71.17 71.18
80% 71.94 72.05 72.13 72.11
75% 72.81 72.92 72.94 72.96
70% 73.60 73.73 73.75 73.81

HoTS is highest at most listed coverage levels, but trails HTS at 80%, and all methods tie at 100%. Some advantages are very small percentage-point differences, and the table supplies no significance test for these macro-average gaps.

Ablation Study

Source Table 10 again reports mean ECE percentages. The full and fixed-exponent methods use matched checkpoints, estimated homophily, splits, and optimization settings.

Configuration Mean ECE Evidence boundary
TS 6.04 One global temperature
Entropy-only 4.83 Removes homophily denominator; independently fits positive parameters
Homophily-only 5.56 Historical result, not rerun with the current observed-label-only predictor
HoTS, exponent fixed at 1 5.90 Matched comparison with the full method
HoTS, learned exponent 4.79 Full method

The fixed exponent is 1.11 percentage points worse than the full method when calculated from rounded table values; the text reports +1.10 percentage points, and both precision levels are retained. The mean gap between full HoTS and entropy-only is just 0.04 percentage points, so the average improvement should not be attributed primarily to the homophily denominator; the older homophily-only row is not a strictly matched ablation either.

Key Findings

  • Controlled CSBM experiments obtain a Pearson correlation of −0.846 between temperature and homophily from 8 sweep-point means with 20 seeds each, over \(0.25\leq h\leq0.95\); this does not summarize the entire negative-signal region.
  • As estimation noise increases from 0 to 0.5, the fitted exponent decreases from 1.009 to 0.182, supporting attenuation, not proving that real-graph cross-entropy optima obey the OLS theorem.
  • Appendix B.8 reports homophily residual correlations of pooled +0.384 and macro +0.283 after controlling for raw confidence and log-degree, but uses true-label oracle homophily: this is a retrospective diagnostic, not a measurement of the deployed predictor's effectiveness.
  • The edge-perturbation experiment recomputes logits and refits calibrators on each perturbed graph, and switches to predicted-label neighbor agreement for homophily estimation; it does not demonstrate direct distribution-shift robustness with a fixed model and fixed calibrator.

Highlights & Insights

  • Temperature scaling can preserve within-node class ordering while changing reliability rankings across nodes. Separating accuracy improvements from gains after rejecting unreliable predictions is essential to evaluating calibration methods correctly.
  • The theory has three distinct levels: necessity of structural information, an idealized template posterior, and a practical temperature function. Explicit conditions and surrogate steps are more reusable than presenting the heuristic function as universally Bayes-optimal.
  • A learned exponent converts structural estimation noise into adjustable correction strength. This idea may transfer to other reliability side information, but the objective, noise dependence, and signal direction require fresh validation rather than importing the OLS bound directly.

Limitations & Future Work

  • The exact posterior requires uniform priors, equidistant means, Gaussian features, population-concentrated degree and composition, one-layer linear aggregation, and oracle homophily; real multilayer nonlinear GNNs, class imbalance, and low-degree nodes do not satisfy the full premise.
  • Positive temperatures cannot repair class-order reversal under negative signal, and the absolute-value denominator discards direction. Future work could compare signed structural features or class-dependent calibration, while separately reporting any resulting prediction changes.
  • The auxiliary GCN makes the complete pipeline larger than a three-parameter model, and validation labels participate in multiple training and selection stages. Independent calibration splits or cross-fitting, predictor-error reporting, and total runtime costs are needed to separate structural form, supervision, and capacity effects.
  • Broad random-split coverage does not establish temporal-shift or inductive new-node generalization. Synthetic validation, refitted edge-perturbation experiments, and oracle retrospective diagnostics do not replace deployment tests.
  • The average gap between entropy-only and HoTS is small, and homophily-only was not rerun under the current protocol. Matched ablations, paired uncertainty analysis, and sensitivity to ECE binning remain important.
  • Appendix A.8's product approximation for neighbor logits controls only Gaussian residual dependence under sparse overlap, subject to spectral-norm and integrability conditions; it is not a conditional-independence theorem for the full finite CSBM and does not establish that other structural statistics contain no additional information.
  • vs TS / HTS: TS learns one global temperature and HTS builds node temperatures from predictive entropy; HoTS adds an estimated-homophily denominator and a learnable exponent. The structural correction is interpretable, but requires an auxiliary predictor and delivers a limited mean gain over entropy-only scaling.
  • vs RBS / SimCalib: RBS already groups nodes by same-class-neighbor ratios, while SimCalib uses node similarity, so structurally informed calibration is not itself new. The distinction is an explicit CSBM posterior law and continuous temperature form; neither method is included in the experiments, preventing a claim of comprehensive superiority over these close predecessors.
  • vs CaGCN / GATS / WATS: These methods learn node temperatures through graph networks, attention, or wavelets, whereas HoTS constrains the final temperature function. Complete cost comparisons must still include its homophily GCN, and CaGCN performs better on degree-stratified ECE.
  • vs VS / GETS: Class-dependent transformations can change predicted labels, mixing calibration with classification refinement. HoTS deliberately preserves labels, making its selective-classification improvements interpretable as ranking changes rather than reclassification.

Rating

  • Novelty: 4/5. Connects an idealized structural posterior to node-wise temperature design, although homophily-informed calibration has predecessors.
  • Experimental Thoroughness: 4/5. Broad dataset, backbone, and diagnostic coverage, with remaining gaps in independent calibration, matched ablations, and significance evidence.
  • Writing Quality: 4/5. Clearly states theoretical scope, although some textual conclusions and table precision require appendix-aware reading.
  • Value: 4/5. Useful for graph confidence and rejection mechanisms that must preserve classifier labels, not for improving full-coverage accuracy.