Skip to content

Keep Your Friends Close, and the Right Neighbours Closer: Disaster-Conditioned Kernel-Regularized Graph Attention for Building Damage Classification

Conference: ECCV 2026
Paper: ECCV 2026 Oral/Poster
Area: Remote Sensing / Graph Learning
Keywords: Building Damage Assessment, xBD Dataset, Kernel-Regularized Graph Attention, Moran's I, Zero-Shot Cross-Event Transfer

TL;DR

Addressing the vulnerability of isolated building damage classifiers to appearance domain shifts and the propensity of naive graph attention to oversmooth damage boundaries, this paper proposes a graph attention framework incorporating a disaster-type-conditioned multi-scale spatial kernel prior and a residual Moran's I de-correlation regularizer, preserving local evidence while adaptively aggregating neighbourhood context across diverse disaster types.

Background & Motivation

Rapid post-disaster building damage assessment (BDA) using high-resolution satellite imagery is a vital operational prerequisite for emergency rescue dispatch, casualty triage, and post-disaster recovery planning. While the xBD dataset and the xView2 challenge established scalable bitemporal satellite benchmarks with polygon-level annotations, operational deployments continue to struggle with two interconnected bottlenecks: poor cross-event generalization and spatially structured prediction errors. Predominant paradigms isolate localized building instances as independent image patches, classifying them via bitemporal convolutional or Transformer encoders. However, satellite captures across disparate disaster zones suffer from significant domain shifts in rooftop materials, solar azimuth, shadow occlusion, and atmospheric conditions, causing isolated patch models to incur spatially clustered systematic errors.

The fundamental tension stems from Tobler's First Law of Geography: disaster damage is inherently spatial, meaning damaged structures naturally cluster; yet the physical scale of this spatial autocorrelation fluctuates drastically across different hazard regimes. In widespread hurricanes or expansive wildfires, destruction exhibits extensive spatial autocorrelation over hundreds of meters, making distant neighbours highly informative. Conversely, in flash floods or complex earthquakes, localized hydrodynamic pressures, micro-topography, and structural resilience produce high spatial heterogeneity, where spatial autocorrelation decays sharply over short distances. Applying an unconstrained, globally uniform vanilla graph attention network (GAT) often triggers indiscriminate neighbourhood smoothing. This oversmooths damage class boundaries and propagates structured appearance errors across mixed-severity neighbourhoods, generating deceptively cohesive damage maps riddled with systematic misclassifications.

The paper's entry point is that spatial context should neither be ignored nor enforced as an unconstrained smoothing heuristic; rather, the model must anchor to per-building evidence while adaptively drawing closer only the truly informative neighbours tailored to the hazard regime. Core idea: regularize graph attention logits using an explicit multi-scale spatial distance kernel prior conditioned on disaster-type embeddings, coupled with a residual Moran's I de-correlation loss that penalizes positive spatial autocorrelation in prediction errors, compelling the network to leverage spatial context as an adaptive predictive signal rather than a naive smoothing shortcut.

Method

Overall Architecture

The framework establishes a controlled, post-localization building damage classification pipeline. The inputs comprise paired pre/post-disaster satellite imagery, annotated building footprints, and the event disaster-type token. Building polygons are cropped into pre/post combined (PPC) bitemporal patches, while polygon centroids are projected onto a local metric coordinate system (UTM) to construct static \(k\)-nearest-neighbour (\(k\text{NN}\)) spatial graphs. A frozen patch encoder extracts initial node representations. These features are iteratively updated across graph attention layers regularized by disaster-conditioned multi-scale spatial kernel priors. Finally, the model is optimized end-to-end using a joint objective encompassing class-balanced cross-entropy, Earth Mover's Distance (EMD) ordinal consistency, and residual Moran's I spatial de-correlation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Bitemporal Satellite Imagery & Building Polygons<br/>Pre/Post RGB + Footprints"] --> B["Preprocessing & Graph Construction<br/>PPC Patch Extraction & Metric UTM kNN Graph"]
    B --> C["Patch Feature Encoder<br/>Frozen ResNet-50 Extracts Node Embeddings"]
    C --> D["Disaster-Conditioned Multi-Scale Kernel Prior<br/>Disaster Token Predicts Mixture Weights & Length Scales"]
    D --> E["Kernel-Regularized Graph Attention<br/>Logit Fusion of Feature Compatibility & Spatial Prior"]
    E --> F["Residual Moran De-correlation & Multi-Task Heads<br/>CE + EMD Ordinal Loss + Residual Moran's I Penalty"]

Key Designs

1. Disaster-Conditioned Multi-Scale Kernel Prior: Adapting Spatial Correlation Scales across Hazard Regimes Disaster mechanisms manifest fundamentally disparate spatial decay rates; hence, static Euclidean distance priors fail across global hazard distributions. The model maps a categorical disaster token (e.g., hurricane, wildfire, flood, earthquake) into a continuous embedding \(z_d\). This vector modulates two sets of parameters through lightweight neural projections: nonnegative mixture weights \(\pi(e)\) and \(M\) disaster-specific correlation length scales \(\rho_m(e)\): $\(\pi(e) = \text{softmax}(W_\pi z_d + b_\pi), \quad \rho_m(e) = \rho_{\min} + \text{softplus}(w_m^\top z_d + b_m), \quad m \in \{1, \dots, M\}\)$ where \(\pi_m(e) \ge 0\) and \(\sum_{m=1}^M \pi_m(e) = 1\), scaling emphasis across short-, medium-, and long-range neighbourhood contexts. Given the projected centroid distance \(d_{ij} = \|p_i - p_j\|_2\) between buildings \(i\) and \(j\), individual exponential kernel bases are computed as \(\kappa_m(d_{ij} \mid e) = \exp\left(-\frac{d_{ij}}{\rho_m(e) + \varepsilon}\right)\). The Disaster-Conditioned Multi-Scale (DCMS) kernel prior becomes: $\(K_{ij}(e) = \sum_{m=1}^M \pi_m(e) \kappa_m(d_{ij} \mid e)\)$ For hazards with broad spatial reach (e.g., hurricanes), the network adaptively elevates mixture weights for large length scales \(\rho_m\), encouraging broader contextual aggregation; for localized hazards (e.g., urban flooding), the distribution contracts toward minimal length scales, suppressing long-range interference.

2. Kernel-Regularized Graph Attention: Balancing Feature-Driven Compatibility with Physical Distance Priors Unconstrained attention mechanisms in remote sensing graphs are highly vulnerable to visual artifacts and localized domain shifts, which destabilize attention scores across damage boundaries. This design injects the geostatistical kernel prior directly into the attention logit as a log-prior regularizer. At layer \(\ell\), the raw compatibility score \(\psi_{ij}^{(\ell)}\) is computed via an MLP \(a_\phi\) over node features \(v_i^{(\ell)}\), neighbour features \(v_j^{(\ell)}\), static geometric edge features \(u_{ij}\) (comprising relative directional angles and normalized distances), and the disaster embedding \(z_d\): $\(\psi_{ij}^{(\ell)} = a_\phi\left([W_q v_i^{(\ell)} \parallel W_k v_j^{(\ell)} \parallel u_{ij} \parallel z_d]\right)\)$ The regularized attention logit incorporates the spatial kernel prior scaled by hyperparameter \(\lambda_K\): $\(\ell_{ij}^{(\ell)} = \psi_{ij}^{(\ell)} + \lambda_K \log\left(K_{ij}(e) + \varepsilon\right)\)$ Normalized attention coefficients \(\alpha_{ij}^{(\ell)}\) are evaluated over the fixed spatial neighbourhood \(\mathcal{N}(i)\) using Softmax: $\(\alpha_{ij}^{(\ell)} = \frac{\exp(\ell_{ij}^{(\ell)})}{\sum_{j' \in \mathcal{N}(i)} \exp(\ell_{ij'}^{(\ell)})}\)$ This formulation anchors dynamic message passing to smooth physical distance decay and hazard-specific correlation footprints, preventing unconstrained attention weights from drifting toward spurious distant correlations or destabilizing sharp damage boundaries.

3. Residual Moran Spatial De-correlation: Suppressing Spatially Clustered Misclassifications and Oversmoothing Standard graph classification objectives optimize cross-entropy independently per node. When attention oversmooths heterogeneously damaged regions, entire neighbourhoods inherit identical misclassifications, creating pronounced positive spatial autocorrelation in the prediction residuals. To eradicate this failure mode, the paper introduces a differentiable spatial autocorrelation penalty using Moran's I. Discrete damage severities are mapped to ordinals \(s_i \in \{0, 1, 2, 3\}\), while the model's continuous expected severity prediction is \(\hat{s}_i = \sum_{c=0}^3 c \cdot \hat{y}_{i,c}\). The prediction residual is defined as \(r_i = s_i - \hat{s}_i\), with mean-centered residual \(\tilde{r}_i = r_i - \bar{r}\). Using a row-normalized spatial weight matrix \(W^{(e)}\) defined over the graph edges, the differentiable residual Moran's I statistic is formulated as: $\(I_r^{(e)} = \frac{n_e}{\sum_{ij} W_{ij}^{(e)}} \cdot \frac{\sum_{i,j} W_{ij}^{(e)} \tilde{r}_i \tilde{r}_j}{\sum_i \tilde{r}_i^2 + \varepsilon}\)$ Because positive Moran's I values (\(I_r^{(e)} > 0\)) signify spatially clustered error patterns, the residual loss penalizes only positive spatial autocorrelation: $\(\mathcal{L}_{\text{res}} = \frac{1}{|\mathcal{E}_{\text{batch}}|} \sum_{e \in \mathcal{E}_{\text{batch}}} \max(0, I_r^{(e)})\)$ This regularizer discourages the network from achieving local visual coherence through indiscriminate smoothing, penalizing errors that clump spatially and compelling message passing to act as a genuine denoiser.

Loss & Training

The cumulative optimization objective unifies classification accuracy, ordinal consistency, and residual spatial de-correlation: $\(\mathcal{L} = \mathcal{L}_{\text{CE}} + \lambda_{\text{emd}} \mathcal{L}_{\text{EMD}} + \lambda_{\text{res}} \mathcal{L}_{\text{res}}\)$ 1. Class-Balanced Cross-Entropy Loss \(\mathcal{L}_{\text{CE}}\): Employs effective-number weights \(w_c = \frac{1 - \beta}{1 - \beta^{n_c}}\) (normalized such that \(\sum_c w_c = 4\)) to combat the extreme class imbalance between dominant classes (No Damage, Destroyed) and minority categories (Minor, Major Damage). 2. Earth Mover's Distance Loss \(\mathcal{L}_{\text{EMD}}\): Enforces ordinal consistency across cumulative class probability vectors: $\(\mathcal{L}_{\text{EMD}} = \frac{1}{B} \sum_{i=1}^B \sum_{t=0}^3 \left( \hat{F}_i(t) - F_i^\star(t) \right)^2, \quad \hat{F}_i(t) = \sum_{c \le t} \hat{y}_{i,c}, \quad F_i^\star(t) = \sum_{c \le t} \mathbf{1}[y_i = c]\)$ 3. Hyperparameters and Training Strategy: Tuned loss weights are \(\lambda_{\text{emd}} = 0.25\) and \(\lambda_{\text{res}} = 0.1\). The bitemporal ResNet-50 patch encoder is frozen throughout training to strictly isolate the contextual contribution of the spatial graph module from backbone representation learning.

Key Experimental Results

Main Results

To establish external reference calibration, the model is evaluated on the official xView2 holdout set under fixed building instance crops (Table 1). It is subsequently subjected to a rigorous Leave-One-Event-Out (LOEO) zero-shot cross-event transfer benchmark on xBD (Table 2).

Table 1: xView2 Holdout External Reference Comparison (Fixed building instance patch classification protocol)

Method Macro-F1 No Damage F1 Minor F1 Major F1 Destroyed F1 Pipeline / Protocol
xView2 Official Winner #1 † 0.80446 0.92344 0.64445 0.78591 0.86403 End-to-end localization + classification
Patch-Only Encoder (ResNet-50) 0.82205 0.92880 0.67320 0.81520 0.87100 Fixed polygon crops without graph
Vanilla GAT on Patch Encoder 0.84102 0.93160 0.70050 0.84030 0.89168 Standard unconstrained \(k\text{NN}\) graph attention
Ours (Kernel + Disaster + Residual) 0.87288 0.93400 0.76200 0.89200 0.90352 Full kernel-regularized graph model

Note: † Official Solution #1 is listed as an external reference since its score incorporates localization errors. Under identical ground-truth instances, the proposed framework substantially outclasses the patch-only baseline and vanilla GAT, yielding notable absolute gains of +8.88% on Minor and +7.68% on Major damage classes.

Table 2: xBD Zero-Shot Cross-Event Transfer & Component Ablation (LOEO Protocol)

Variant Kernel Prior Disaster Condition Residual Loss LOEO Macro-F1 ↑ Per-Class F1 (No, Mi, Ma, D) Residual Moran's I ↓ Severity MAE ↓
Patch-Only Baseline \(\times\) \(\times\) \(\times\) 0.433 (0.61, 0.37, 0.41, 0.34) 0.248 0.562
Vanilla GAT (\(k\text{NN}\) baseline) \(\times\) \(\times\) \(\times\) 0.453 (0.57, 0.40, 0.44, 0.40) 0.256 0.541
+ Kernel Prior only \(\checkmark\) \(\times\) \(\times\) 0.473 (0.60, 0.42, 0.46, 0.41) 0.182 0.507
+ Disaster Conditioning only \(\times\) \(\checkmark\) \(\times\) 0.460 (0.59, 0.41, 0.45, 0.39) 0.205 0.521
+ Residual De-correlation only \(\times\) \(\times\) \(\checkmark\) 0.450 (0.58, 0.39, 0.44, 0.39) 0.139 0.529
+ Kernel Prior + Disaster Conditioning \(\checkmark\) \(\checkmark\) \(\times\) 0.488 (0.61, 0.44, 0.48, 0.42) 0.116 0.484
+ Kernel Prior + Residual Loss \(\checkmark\) \(\times\) \(\checkmark\) 0.490 (0.60, 0.44, 0.49, 0.43) 0.096 0.474
Full Model (Ours) \(\checkmark\) \(\checkmark\) \(\checkmark\) 0.503 (0.62, 0.45, 0.50, 0.44) 0.079 0.458

Table 3: Zero-Shot Cross-Dataset Transfer from xBD to Ida-BD (Hurricane Ida)

Method Pipeline / Protocol Disaster Token Macro-F1 ↑ Per-Class F1 (No, Mi, Ma, D)
Siam-UNet Dense prediction (Destroyed merged into Major) — 0.458 (0.916, 0.208, 0.251, —)
xView2 Winner Per-class weighted F1 pipeline — 0.268 (0.667, 0.211, 0.154, 0.041)
Patch-Only Baseline (Ours) Polygon patch classification \(\checkmark\) 0.275 (0.670, 0.190, 0.160, 0.080)
Vanilla GAT (Ours) Patch classification + \(k\text{NN}\) graph \(\checkmark\) 0.295 (0.690, 0.220, 0.180, 0.090)
Ours (Full Model) Patch classification + regularized graph \(\checkmark\) 0.335 (0.710, 0.270, 0.230, 0.130)

Key Findings

  1. Spatial Kernel Prior Yields the Strongest Standalone Gain: In the LOEO ablation, introducing the multi-scale distance kernel prior increases Macro-F1 from 0.453 to 0.473 while reducing residual Moran's I from 0.256 to 0.182. This confirms that grounding attention weights in distance-decay geometry eliminates chaotic attention dispersion across distant nodes.
  2. Residual De-correlation Drastically Cuts Clustered Errors: Adding \(\mathcal{L}_{\text{res}}\) alone reduces residual Moran's I by nearly half (from 0.256 down to 0.139). Combining all components pushes residual Moran's I to 0.079, proving that directly penalizing spatial error clustering effectively prevents oversmoothing shortcuts.
  3. Vanilla GAT Suffers from Severe Boundary Oversmoothing: LOEO confusion analysis reveals that unconstrained GAT shifts error mass toward adjacent damage classes (misclassifying 29% of No Damage as Minor and 38% of Destroyed as Major in mixed regions). The proposed framework tightens the diagonal accuracy to 70% for No Damage and 68% for Destroyed, maintaining crisp boundaries in mixed-severity neighbourhoods.
  4. Superior Zero-Shot Cross-Dataset Transferability: On the unobserved Ida-BD hurricane benchmark, the full model achieves 0.335 Macro-F1, outperforming patch-only (0.275) and Vanilla GAT (0.295). Crucially, it lifts Destroyed class F1 to 0.130 (over 60% relative gain against patch-only 0.080), confirming that conditioned geostatistical priors transfer robustly across dataset domains.

Highlights & Insights

  • Transforming Geostatistical Diagnostics into Differentiable Regularizers: Rather than relegating Moran's I to a post-hoc analytical statistic, the authors construct a differentiable surrogate loss directly on prediction residuals, elegantly solving the chronic dilemma of "visually coherent but systematically wrong" predictions in spatial GNNs.
  • Disaster-Conditioned Effective Receptive Fields: Modulating multi-scale exponential kernel weights using a global disaster token allows the network to dynamically adapt its spatial reach—expanding aggregation across wide-spread wildfires and hurricanes while contracting radius in heterogeneous flood environments.
  • Generalizability to Spatial Structured Vision: The combination of domain-conditioned multi-scale distance priors and residual de-correlation provides an extensible template for diverse geospatial vision domains where interaction scales vary with geography, terrain, or environmental conditions (e.g., crop yield forecasting, urban heat island mapping).

Limitations & Future Work

  • Reliance on Ground-Truth Footprints and Known Metadata: The controlled evaluation eliminates localization error by relying on pre-existing building polygons and assumes the event disaster type is provided. In real-world operational deployments, detector inaccuracies, boundary jitter, and absent metadata require robust uncertainty estimation and self-supervised hazard inference.
  • Euclidean \(k\text{NN}\) Graph Ignores Topographical and Built Barriers: Constructing edges solely on Euclidean centroid distances overlooks physical urban features like rivers, major highways, and firebreaks that partition damage propagation in reality.
  • Exclusion of Non-Optical Modalities: The framework operates strictly on optical bitemporal patches, remaining vulnerable to cloud cover and thick smoke plumes. Incorporating Synthetic Aperture Radar (SAR) modalities within the spatial graph architecture represents a vital future trajectory.
  • vs. xView2 Challenge Winners and Bitemporal CNN/Transformers (Weber & Kané 2020, BDANet 2022): Prior baselines heavily prioritize sophisticated convolutional fusion backbones on isolated building crops. The proposed approach proves that once local features are extracted, explicit geostatistical graph modeling delivers substantial orthogonal gains under extreme event shift.
  • vs. Vanilla Graph Attention Networks (GAT / GATv2): Standard GAT relies purely on feature matching, frequently causing boundary collapse and error contagion in remote sensing scenes. The proposed method establishes structured geostatistical bounds via conditional kernel priors and residual Moran penalties, achieving adaptive context without boundary degradation.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Incorporates geostatistical Moran's I as a differentiable residual loss and introduces hazard-conditioned multi-scale kernel regularizers for GNNs.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Includes xView2 holdout calibration, comprehensive leave-one-event-out (LOEO) cross-event testing, detailed 8-row ablations, and zero-shot cross-dataset evaluation on Ida-BD.
  • Writing Quality: ⭐⭐⭐⭐⭐ The prose rigorously decouples local representation from spatial contextual reasoning, clearly elucidating the theoretical tension between spatial context and boundary oversmoothing.
  • Value: ⭐⭐⭐⭐☆ Offers valuable methodological insights for post-disaster remote sensing and graph neural networks applied to spatial-temporal earth observation.