Skip to content

Learning Semantic-Robust Change Detection via Semantic-Invariant Self-Distillation

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/elecreak/SCDistill
Area: Remote Sensing / Change Detection
Keywords: Change Detection, Self-Distillation, Vision Foundation Model, Diffusion Perturbation Simulation, Representation Learning

TL;DR

To alleviate false alarms and missed detections caused by non-semantic environmental disturbances (e.g., illumination, shadows, and weather) in bitemporal remote sensing change detection, this paper introduces SCDistill, which simulates realistic appearance variations via diffusion models and guides a vision foundation model encoder via semantic-invariant self-distillation, achieving state-of-the-art results across semantic change detection, binary change detection, and change captioning tasks.

Background & Motivation

Bitemporal change detection (CD) plays a pivotal role in earth observation, supporting vital applications including urban expansion monitoring, land-cover mapping, and disaster assessment. Powered by deep learning, architectures combining dual-branch feature extraction and temporal differencing have achieved substantial accuracy improvements. However, multi-temporal optical imagery captured under open-world conditions inevitably exhibits drastic non-semantic disturbances—such as changing solar azimuth and elevation angles, ground shadows, haze, rain, snow, and seasonal vegetation phenology. Mainstream CD frameworks typically rely on general-purpose encoders pre-trained on single-view classification or segmentation tasks (e.g., ResNets or standard ViTs). These encoders remain inherently sensitive to pixel-level appearance shifts, often interpreting radiometrically altered but semantically unchanged land cover as genuine category transitions, thereby triggering severe false alarms or masking subtle ground changes under heavy shadow.

This vulnerability stems from a dual bottleneck across both data distribution and feature representations. On the data side, existing synthetic CD datasets (such as FSC-180k or Changen2) predominantly focus on inserting or modifying semantic foreground objects while leaving unchanged background regions under pristine, artificially aligned conditions. As a consequence, models trained on synthetic benchmarks lack exposure to physically plausible negative disturbance cues. On the representation side, although vision foundation models possess rich high-level semantic priors, their feature spaces lack explicit constraints enforcing invariance across environmental perturbations; propagating such disturbance-sensitive features into downstream difference decoders directly compounds prediction errors.

To address both challenges simultaneously, this paper attacks the problem from synergistic data and representation perspectives. By combining generative physical simulation with self-supervised feature alignment, the model learns to isolate invariant semantics from transient visual artifacts. Core idea: leverage relighting and instruction-guided diffusion models to synthesize diverse non-semantic environmental perturbations on bitemporal samples, and train a student encoder via semantic-invariant self-distillation against a frozen clean-input vision foundation model teacher, instilling disturbance-resistant representations without requiring manual annotations.

Method

Overall Architecture

SCDistill consists of two cooperative stages: diffusion-based non-semantic perturbation simulation and semantic-invariant self-distillation, followed by seamlessly integrating the resulting distilled vision foundation model encoder (Distilled VFME) into a multi-task change detection framework. First, realistic physical perturbations are synthesized on unedited or synthetic remote sensing imagery to generate paired samples sharing identical semantics under drastic appearance shifts. Next, a frozen foundation model teacher processes the clean reference image and provides high-level guidance, supervising the student network to align its representations extracted from perturbed inputs. Finally, the robust distilled encoder serves as the backbone feature extractor, coupled with shallow temporal difference features to feed a unified decoder for semantic change detection, binary change detection, or change captioning.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Bitemporal RS Images / Synthetic Dataset"] --> B["Non-semantic Perturbation Simulation<br/>DIR Relighting + InstructPix2Pix Weather Editing"]
    B --> C["Noise-Augmented Semantic Pairs<br/>Semantically invariant under appearance shifts"]
    C --> D["Semantic-Invariant Self-Distillation<br/>Frozen teacher aligns student representations"]
    D --> E["Distilled VFME Backbone<br/>Extracts disturbance-resistant temporal features"]
    E --> F["Downstream Change Detection & Decoding<br/>Task-specific heads output robust predictions"]

Key Designs

1. Non-semantic Perturbation Simulation: Synthesizing Physically Plausible Environmental Variations Synthetic change detection datasets mitigate the scarcity of pixel-level bitemporal annotations, but prior generative pipelines alter only changed areas while assuming unchanged areas remain photometrically static. To bridge this gap, this paper constructs an environmental perturbation simulation pipeline that models realistic physical variations. Given a sample triplet \(\{X_0, X_1, L\}\), non-semantic inconsistencies across temporal images can be induced by perturbing one temporal image alone. Therefore, continuous perturbations are applied to the pre-change image \(X_0\). Specifically, for solar angle and illumination changes, a single-image relighting model (DIR) is applied with azimuth angles sampled uniformly in \(0^\circ \sim 360^\circ\) and elevation angles in \(20^\circ \sim 70^\circ\). For weather and atmospheric shifts such as haze, rain, or snow, InstructPix2Pix is employed conditioned on predefined descriptive textual prompts \(p_i\). The cascaded simulation yields multi-perturbed variants: $\(\tilde{X}_0^{(i)} = \text{Relighting}(\text{Editing}(X_0, p_i))\)$ Augmented further with Gaussian blur, grayscale mapping, and brightness/saturation jitter, this pipeline produces realistic bitemporal pairs with rich environmental diversity, providing the necessary cross-condition supervision.

2. Semantic-Invariant Self-Distillation: Enforcing Feature Robustness Without Manual Supervision Even with diverse simulated perturbations, training an encoder purely with task-specific supervised losses often allows the network to exploit superficial low-level shortcuts. To instill genuine invariance into the representations, the authors introduce a semantic-invariant self-distillation objective. A pre-trained vision foundation model (e.g., DINOv2 or DINOv3) serves as a frozen teacher \(E_T\), and the student \(E_S\) is initialized from the same weights. The clean input image \(X\) is fed into the teacher to obtain the pristine semantic target \(F_T = E_T(X)\), while the student receives the perturbed counterpart \(\tilde{X}\), outputting \(F_S = E_S(\tilde{X})\). The student is optimized by minimizing the mean squared error: $\(\mathcal{L}_{distill} = \|F_S - F_T\|_2^2\)$ This formulation requires no ground-truth change labels or dense land-cover masks, allowing the distillation to scale across massive unlabeled remote sensing imagery. Crucially, the teacher is evaluated only on clean images where its semantic guidance is stable; hence, the student's resilience is not passively inherited from the teacher, but actively cultivated through the cross-condition invariance constraint.

3. Robust Feature Fusion and Multi-Task Change Decoding: Guiding Temporal Differencing To deploy the distilled disturbance-resistant representations in downstream tasks, the distilled VFME is incorporated into a video modeling-based change detection framework (Change3D). For a bitemporal input pair \((I_0, I_1)\), the distilled encoder produces stable high-level semantic embeddings \(F_0\) and \(F_1\). These features are merged with shallow spatiotemporal representations from an X3D backbone, acting as a semantic gate that suppresses spurious difference activations caused by surface reflectance and transient shadows. The refined features are then forwarded to task-specific decoders to support semantic change detection (SCD), binary change detection (BCD), and change captioning under a unified architecture.

Loss & Training

The training pipeline consists of self-distillation pre-training, synthetic change detection pre-training, and real-world benchmark fine-tuning. In self-distillation, the student encoder is trained using AdamW with an initial learning rate of \(2 \times 10^{-5}\), weight decay of 0.05, and a batch size of 8 per GPU (requiring only 5 hours on an RTX 3090). Pre-training on the augmented FSC-180k dataset takes 20 hours as a one-time cost. For downstream fine-tuning, the network is trained with task-specific multi-task loss functions: - Semantic Change Detection (SCD): optimized jointly with cross-entropy loss \(\mathcal{L}_{ce}\), Dice loss \(\mathcal{L}_{dice}\), and a cosine similarity loss \(\mathcal{L}_{cos}\) enforcing temporal feature consistency in unchanged regions: $\(\mathcal{L}_{SCD} = \mathcal{L}_{ce} + \mathcal{L}_{dice} + \mathcal{L}_{cos}\)$ - Binary Change Detection (BCD): supervised with \(\mathcal{L}_{ce} + \mathcal{L}_{dice}\). - Change Captioning: optimized with autoregressive token cross-entropy loss \(\mathcal{L}_{ce}\).

Key Experimental Results

Main Results

On two standard, highly challenging high-resolution real-world semantic change detection benchmarks—SECOND and HRSCD—SCDistill is comprehensively evaluated against leading methods. Metrics include change-region F1 score (\(F_{scd}\)), mean Intersection over Union (mIoU), Overall Accuracy (OA), and Separate Kappa (SeK).

Method Backbone SECOND \(F_{scd}\) (%) SECOND mIoU (%) SECOND SeK (%) HRSCD \(F_{scd}\) (%) HRSCD mIoU (%) HRSCD SeK (%)
SSCD-L (TGRS'22) ResNet 61.22 72.60 21.86 69.69 66.21 21.88
Bi-SRNet (TGRS'22) ResNet 61.85 72.08 21.36 71.72 67.83 24.59
ChangeMamba (TGRS'24) Mamba-based 61.59 72.32 20.30 70.04 67.93 24.50
Changen2 (TPAMI'24) ViT 62.79 72.39 22.09 73.03 68.73 26.25
Change3D (CVPR'25) VideoEncoder 62.83 72.95 22.98 73.29 68.67 26.85
SCDistill (Ours) Distilled VFME 65.64 74.01 25.12 74.29 70.28 28.38

In generalization benchmarks, SCDistill also establishes new state-of-the-art results across binary change detection (CLCD, LEVIR-CD, WHU-CD) and change captioning (DUBAI-CC). On DUBAI-CC, SCDistill improves the CIDEr score from Change3D's 86.19 to 95.59 (+9.40 points) and reaches a BLEU-2 of 59.77%, highlighting the high precision of its textual descriptions under environmental shifts.

Ablation Study

The synergistic effect of perturbation simulation and self-distillation is systematically evaluated on the SECOND benchmark (Table 4), alongside representation alignment metrics computed on unchanged real-world HRSCD pairs (Table 6).

Setting Simulate Distill \(F_{scd}\) (%) mIoU (%) OA (%) SeK (%)
Vanilla baseline ✗ ✗ 62.86 72.54 87.53 22.30
+ Perturbation Simulation ✔ ✗ 63.55 73.10 87.87 23.39
+ Self-Distillation ✗ ✔ 63.91 73.13 88.24 23.37
Full Model (Ours) ✔ ✔ 65.64 74.01 88.59 25.12

Feature invariance analysis on unchanged real-world HRSCD image pairs: | Distillation Status | MSE (↓) | Cosine Sim. (↑, %) | CKA Alignment (↑, %) | |---|---|---|---| | Before Distillation | 6.41 | 84.05 | 87.08 | | After Distillation | 3.32 | 89.78 | 90.89 |

Key Findings

  • Synergistic Closed-Loop Effect: Incorporating perturbation simulation alone gains \(+0.69\) \(F_{scd}\), and self-distillation alone yields \(+1.05\); however, combining both achieves \(+2.78\) \(F_{scd}\) and pushes SeK from 22.30 to 25.12. This verifies that realistic physical negative data and self-supervised distillation form a mutually reinforcing closed loop.
  • Substantial Feature Stabilization on Real Scenes: On unchanged real-world bitemporal pairs, self-distillation halves feature MSE from 6.41 to 3.32 and improves CKA alignment to 90.89%. Qualitative PCA heatmaps confirm that spurious activations triggered by clouds and sun angle shifts are strongly attenuated.
  • Robust Few-Shot Transfer: When fine-tuning with only 1% of labeled real-world data on SECOND, the full model achieves an \(F_{scd}\) of 55.6%, outperforming the non-distilled ablation (48.9%) by +6.7%, indicating that invariant pre-training instills fundamental structural priors resilient to extreme data scarcity.

Highlights & Insights

  • Physics-Grounded Generative Augmentation: Rather than relying on simple photometric color jittering or unconstrained GAN inversion, using 3D single-image relighting (DIR) and promptable diffusion editing (InstructPix2Pix) models physical sun angles and atmospheric scattering with high fidelity, creating reliable disturbance supervision.
  • "Self as Teacher" Invariance Learning: By establishing the clean-image foundation model as a trusted teacher, the framework avoids costly dense manual annotations while compelling the student to strip away appearance-level camouflage, offering a general recipe for all-weather vision models.
  • Single Distillation, Universal Downstream Reuse: The self-distillation phase requires only 5 hours on an RTX 3090 and can be performed once; the resulting encoder serves as a drop-in backbone across binary detection, multi-class semantic detection, and multimodal change captioning.

Limitations & Future Work

  • Lack of Geometric / Perspective Distortion Modeling: The current simulation pipeline focuses predominantly on radiometric, lighting, and atmospheric fluctuations. Non-semantic variations caused by satellite off-nadir tilt angles and building perspective layover are not explicitly addressed.
  • Multi-Stage Offline Processing Pipeline: Generating realistic weather and relighting conditions relies on external diffusion models and offline dataset preprocessing; exploring an end-to-end differentiable latent perturbation generator represents a promising future avenue.
  • vs. Changen2 / ChangeDiff: Prior diffusion-based CD generators concentrate on synthesizing new object categories and foreground changes. SCDistill identifies that real-world failures mostly occur on unchanged backgrounds, purposefully introducing realistic non-semantic noise to preserve true semantic invariance.
  • vs. Change3D: While Change3D reformulates change detection as video modeling, its feature backbone remains vulnerable to optical shifts. SCDistill complements Change3D by equipping its spatiotemporal decoder with an invariant, disturbance-resistant feature extractor.

Rating

  • Novelty: ⭐⭐⭐⭐☆ (Combines physical perturbation simulation with cross-condition self-distillation to tackle open-world environmental shifts)
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ (Validates performance across semantic CD, binary CD, change captioning, few-shot regimes, and internal representation metrics)
  • Writing Quality: ⭐⭐⭐⭐⭐ (Well-structured methodology, crisp motivation, and clear experimental validation)
  • Value: ⭐⭐⭐⭐⭐ (Provides a highly practical, plug-and-play foundation feature learning paradigm for robust earth observation)