Skip to content

Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment

Conference: ECCV 2026
Paper: ECCV 2026 Official
Code: https://github.com/wzy6055/GACR
Area: Segmentation
Keywords: cloud removal, observation-anchored residual flow, geo-contextual alignment, vision foundation model, remote sensing interpretation

TL;DR

The paper proposes GACR, an interpretation-oriented cloud removal framework that formulates recovery as an observation-anchored residual flow matching process coupled with a vision foundation model semantic manifold constraint, mitigating hallucinated artifacts and semantic drift while drastically improving downstream interpretation accuracy.

Background & Motivation

Optical satellite imagery serves as a foundational data source for Earth observation missions, ranging from urban expansion monitoring and environmental management to land-cover classification. However, widespread cloud coverage in the atmosphere severely disrupts surface reflectance signals, obscuring land structures and injecting critical uncertainty into downstream processing pipelines. Consequently, cloud removal (CR) has shifted from a superficial image enhancement preprocessing step into a vital structural and semantic reconstruction task within remote sensing pipelines. Contemporary deep learning CR frameworks primarily fall into two categories: residual denoising models and generative diffusion models. Residual approaches treat cloud coverage as additive residual noise and regress clear reflectance directly, which inevitably yields over-smoothed textures and structural ambiguities under heavy occlusion where ground truth surface signals are entirely missing. Conversely, diffusion models synthesize visually plausible patterns from pure Gaussian noise, but their stochastic sampling trajectories lack physical observation anchoring, frequently hallucinating geographically implausible textures that contradict true spatial land-cover distributions.

The fundamental culprit behind these failures is the pervasive "visual-fidelity-oriented" optimization paradigm. Prevailing approaches primarily minimize pixel-level reconstruction errors (e.g., L1/L2 losses tailored for conventional PSNR and SSIM benchmarks), under the implicit yet flawed assumption that visual clarity equals semantic correctness. Because real-world geographic landscapes follow strict spatial-semantic rules (e.g., contiguous forest distributions, regular urban geometry, hydrological topography), unconstrained pixel-level optimization under thick cloud occlusion allows synthetic details to drift away from the authentic geographical context. When reconstructed images are fed into downstream networks for semantic segmentation, building extraction, or height estimation, subtle structural deviations and boundary distortions accumulate, severely compromising downstream interpretation reliability.

The angle of attack in this paper is to eliminate unanchored stochastic search while moving beyond narrow pixel-level supervision. The core idea is to reformulate cloud removal as an observation-anchored residual flow matching process grounded in the cloudy observation, and enforce a geo-contextual manifold constraint induced by a vision foundation model, ensuring that the generative trajectory simultaneously delivers high-fidelity texture recovery and strictly preserved downstream interpretation reliability.

Method

Overall Architecture

The GACR framework consists of three synergistic core components: the physics-inspired Observation-Anchored Residual Flow (OAR-Flow), the Vision Foundation Model (VFM)-based Geo-Contextual Prior Alignment (GCPA), and a downstream interpretation verification protocol.

The full inference pipeline originates from the degraded cloudy observation \(x_c\). OAR-Flow anchors the starting state of a deterministic probability flow ordinary differential equation (ODE) to the cloudy image perturbed by structured residuals. A lightweight and scalable Hourglass Diffusion Transformer (HDiT) serves as the velocity field network, progressively integrating along a deterministic flow path toward the cloud-free clear state \(x_*\). During the training phase, the model is supervised in pixel space via a velocity matching loss \(\mathcal{L}_{\mathrm{vel}}\) to capture transport dynamics from cloudy perturbations to clear terrain. Simultaneously, intermediate bottleneck features are extracted from HDiT and mapped via an Adaptive Projector (ADP) to match multi-scale geo-contextual representations from a frozen DINOv3 vision foundation model under a patch-wise cosine similarity constraint (\(\mathcal{L}_{\mathrm{GCI}}\)), strictly confining reconstruction within a geographically consistent semantic manifold.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Cloudy Observation xc and Ground Truth x*"] --> B["Observation-Anchored Residual Flow<br/>Construct anchored interpolant trajectory xt and decompose velocity"]
    B --> C["Efficient Decoupled Architecture<br/>HDiT velocity field prediction and bottleneck feature extraction"]
    C --> D["Geo-Contextual Prior Alignment<br/>Frozen VFM feature extraction and adaptive projection alignment"]
    D --> E["Unified Objective Optimization and ODE Sampling<br/>Lvel coupled with LGCI for deterministic surface reconstruction"]
    E --> F["Downstream Interpretation Evaluation<br/>Land-cover classification / building extraction / segmentation / height estimation"]

Key Designs

1. Observation-Anchored Residual Flow: Physically Grounded Deterministic Flow Matching Standard diffusion and generative flow methods typically initiate sampling from unconditioned isotropic Gaussian noise, leading to prolonged trajectories and stochastic drift. In contrast, OAR-Flow constructs a continuous forward interpolant trajectory explicitly anchored in the cloudy observation \(x_c\): $\(x_t = \alpha_t x_* + \beta_t x_c + \sigma_t \epsilon, \quad \epsilon \sim \mathcal{N}(0, \mathbf{I})\)$ with linear time scheduling \(\alpha_t = 1 - t\), \(\beta_t = \rho t\), and \(\sigma_t = t\) for \(t \in [0, 1]\). Under thin cloud coverage, where ground signals remain partially observable, the physical anchor \(x_c\) dominates the dynamics, preserving low-frequency cues and structural boundaries without unguided random perturbation. Under thick cloud coverage, the stochastic perturbation \(\sigma_t \epsilon\) provides the generative flexibility required for semantic inpainting. The corresponding ideal deterministic probability velocity field is: $\(v_t(x) = \mathbb{E}[\dot{\alpha}_t x_* + \dot{\beta}_t x_c + \dot{\sigma}_t \epsilon \mid x_t = x]\)$ The parameterized velocity network \(\mathbf{u}_t(x) = \mathrm{Net}_\theta(x_t, t, x_c)\) directly predicts this transport direction conditioned on \(x_c\). During inference, numerical ODE integration via \(\hat{x}_{t-1} = \hat{x}_t - \Delta t \cdot \mathbf{u}_t(\hat{x}_t)\) deterministically recovers \(\hat{x}_0\) from \(t=1\), accelerating convergence while keeping the generative path strictly aligned with the observation.

2. Geo-Contextual Prior Alignment: VFM Manifold Regularization and GCI Loss To eliminate semantic hallucination and category drift under severe cloud coverage, GCPA leverages a pretrained DINOv3 remote sensing foundation model (e.g., ViT-L/16-SAT-300M or ViT-L/16-LVD-1689M) to construct an authentic geographic semantic manifold. Given the clear image \(x_*\), the frozen VFM encoder extracts dense representations \(z_* = f_{\mathrm{vfm}}(x_*) \in \mathbb{R}^{B \times C' \times H' \times W'}\). Intermediate hidden states \(h_t\) from the HDiT bottleneck are projected into the identical manifold space via an Adaptive Projector (ADP), yielding \(z_t\). Geo-contextual coherence is enforced by maximizing patch-wise cosine similarity via the Geo-Contextual Integrity (GCI) loss: $\(\mathcal{L}_{\mathrm{GCI}} = - \mathbb{E} \left[ \frac{1}{N} \sum_{n=1}^{N} \frac{\langle z_*^{[n]}, z_t^{[n]} \rangle}{\| z_*^{[n]} \|_2 \| z_t^{[n]} \|_2} \right]\)$ This constraint forces the intermediate representations of the generative model to respect large-scale spatial structures (e.g., contiguous forests, road connectivity, water-land boundaries), effectively preventing the model from hallucinating geographically incompatible artifacts.

3. Efficient Decoupled Architecture: HDiT Backbone with Adaptive Projection Conventional Diffusion Transformers (DiTs) incur prohibitive computational complexity (hundreds of GFLOPs) when applied directly in pixel space on high-resolution satellite imagery. The framework adopts the Hourglass Diffusion Transformer (HDiT), which leverages hierarchical multiscale downsampling and upsampling to reduce computational overhead to approximately one-quarter of standard DiT. Crucially, the deepest bottleneck of HDiT inherently aggregates global semantic abstractions that naturally correspond to token representations in foundation models. The lightweight ADP module bridges these representations without requiring fine-tuning of the external foundation model, successfully decoupling pixel-level flow regression from high-level semantic manifold alignment.

Loss & Training

The overall training objective is formulated as a unified multi-task loss: $\(\mathcal{L} = \mathcal{L}_{\mathrm{vel}} + \lambda \mathcal{L}_{\mathrm{GCI}}\)$ where the velocity matching loss \(\mathcal{L}_{\mathrm{vel}}\) supervises deterministic pixel-level trajectory fitting: $\(\mathcal{L}_{\mathrm{vel}} = \mathbb{E}_{x_*, \epsilon, x_c, t} \left[ \left\| \mathbf{u}_t(x_t) - (\dot{\alpha}_t x_* + \dot{\beta}_t x_c + \dot{\sigma}_t \epsilon) \right\|^2 \right]\)$ The hyperparameter \(\lambda\) balances pixel-level restoration fidelity and high-level geo-contextual alignment. Empirical analysis identifies \(\lambda = 0.5\) as the optimal trade-off point between pixel sharpness and semantic consistency.

Key Experimental Results

Main Results

Quantitative evaluations across six benchmark datasets spanning real and synthetic cloud conditions confirm that GACR achieves state-of-the-art reconstruction fidelity, especially in thick cloud environments:

Method CUHKCR-GZ PSNRโ†‘ / SSIMโ†‘ CUHKCR-CS PSNRโ†‘ / SSIMโ†‘ Potsdam-Thick PSNRโ†‘ / SSIMโ†‘ Vaihingen-Thick PSNRโ†‘ / SSIMโ†‘
MPRNet (WACV 2021) 23.454 / 0.712 23.365 / 0.695 26.418 / 0.902 27.805 / 0.935
Restormer (CVPR 2022) 25.839 / 0.743 23.632 / 0.710 28.831 / 0.922 28.867 / 0.922
AST (CVPR 2024) 25.482 / 0.735 23.365 / 0.695 25.886 / 0.890 26.471 / 0.916
MambaIR (ECCV 2024) 25.626 / 0.733 23.445 / 0.704 27.027 / 0.903 29.852 / 0.945
DFCFormer (JSTARS 2025) 25.816 / 0.746 23.876 / 0.711 28.196 / 0.917 30.396 / 0.951
EMRDM (CVPR 2025) 25.862 / 0.747 23.736 / 0.712 27.199 / 0.923 28.979 / 0.951
GACR-SAT/2 (Ours) 25.964 / 0.736 24.230 / 0.709 30.578 / 0.934 33.018 / 0.964
GACR-SAT/1 (Ours) 26.100 / 0.744 24.354 / 0.713 31.049 / 0.938 34.048 / 0.970

On twelve downstream tasks evaluated with independent frozen DINOv3 encoders, GACR delivers consistent and substantial accuracy gains:

Model CLS-1 (Acc.โ†‘) BLD-2 (IoUโ†‘) SEG-2 (mIoUโ†‘) SEG-4 (mIoUโ†‘) HE-2 (RMSEโ†“) HE-4 (RMSEโ†“)
Without CR (Lower Bound) 0.746 0.596 0.490 0.550 2.703 2.095
EMRDM (CVPR 2025) 0.776 0.696 0.668 0.692 2.133 1.629
DFCFormer (JSTARS 2025) 0.763 0.657 0.630 0.671 2.317 1.687
GACR-SAT/2 (Ours) 0.833 0.710 0.699 0.737 2.014 1.554
Clear Image (Upper Bound) 0.882 0.762 0.733 0.755 1.868 1.477

Ablation Study

Ablation analysis on CUHKCR-EXT-CS isolates the contribution of generative modeling and semantic manifold alignment:

Configuration PSNR โ†‘ SSIM โ†‘ CLS Acc. โ†‘ BLD IoU โ†‘ GFLOPs Note
DiT + MRDM 24.037 0.707 0.670 0.693 697.20 Standard diffusion transformer with heavy compute overhead
HDiT + MRDM 23.736 0.712 0.726 0.696 166.72 Hourglass backbone reduces compute but keeps lengthy trajectory
HDiT + OAR-Flow 23.986 0.698 0.690 0.700 56.05 Observation-anchored flow sharply cuts complexity to 56G
HDiT + OAR-Flow + GCPA (Full GACR) 24.230 0.709 0.781 0.710 56.05 Geo-contextual alignment improves classification by +9.1%

Hyperparameter sensitivity on Vaihingen-CR-Thick across \(\lambda\) and patch size \(p\):

Config Parameter PSNR โ†‘ SSIM โ†‘ SEG (mIoU) โ†‘ HE (RMSE) โ†“ GFLOPs Note
\(\lambda = 0.2\) (\(p=2\)) 33.073 0.965 0.728 1.578 56.05 Weak semantic alignment leads to suboptimal downstream scores
\(\lambda = 0.5\) (\(p=2\)) 33.018 0.964 0.737 1.554 56.05 Optimal balance, maximizing downstream metrics
\(\lambda = 1.0\) (\(p=2\)) 32.838 0.964 0.730 1.563 56.05 Slightly over-constrained high-frequency texture recovery
\(\lambda = 5.0\) (\(p=2\)) 32.615 0.962 0.727 1.596 56.05 Over-regularization degrades reconstruction fidelity
\(p = 4\) (\(\lambda=0.5\)) 30.303 0.949 0.693 1.616 14.15 Coarse partitioning limits fine detail recovery
\(p = 1\) (\(\lambda=0.5\)) 33.799 0.969 0.720 1.548 223.60 Finer granularity improves PSNR at \(4\times\) compute cost

Key Findings

  • Observation Anchoring Accelerates Convergence: OAR-Flow reaches high-PSNR regimes with roughly one-third of the iterations required by the mean-reverting diffusion baseline (EMRDM). When augmented with GCPA, the full GACR model achieves an overall \(5\times\) training acceleration.
  • Dense Downstream Tasks Are Highly Sensitive to Cloud Degradation: While global classification can be learned reasonably well end-to-end directly from cloudy inputs, dense prediction tasks such as semantic segmentation and height estimation experience catastrophic degradation under heavy cloud coverage (mIoU plunges to 0.490 on Vaihingen-CR-thick). GACR recovers mIoU to 0.699 (approaching the clean upper bound of 0.733), proving that geo-contextual alignment is indispensable for preserving dense geographic category boundaries.
  • Feature Distance Distribution Alignment: Feature distance distribution analysis indicates that conventional CR methods produce high-level representations with noticeable distribution shifts compared to clean references. GACR tightly matches the feature distribution of clear imagery, validating that the reconstructed textures carry authentic geographic meaning.

Highlights & Insights

  • From Pure Noise to Observation-Anchored Flow: By anchoring the deterministic ODE trajectory to the cloudy observation \(x_c\), the model adaptively handles thin-cloud transmission and thick-cloud inpainting, completely avoiding the unanchored hallucinations typical of standard diffusion models.
  • Bridging Low-Level Restoration and High-Level Interpretation: Leveraging the latent representation of a frozen vision foundation model via patch-wise cosine alignment injects geographic common sense into low-level restoration at virtually zero extra training overhead.
  • Optimal Compute-Accuracy Trade-off: The combination of the HDiT architecture and \(p=2\) patch partitioning maintains an efficient footprint of only 56 GFLOPs while outperforming models with hundreds of GFLOPs across both low-level restoration and high-level interpretation tasks.

Limitations & Future Work

  • Author-Acknowledged Limitations: The framework depends on a pretrained VFM (such as DINOv3) for high-level semantic priors. In unconventional geographical terrains or novel hyperspectral bands underrepresented in foundation model pretraining, guidance efficacy may diminish.
  • Scope Limitations: The study primarily focuses on single-temporal optical satellite imagery. Operational satellite monitoring often exploits multitemporal sequences and cloud-penetrating synthetic aperture radar (SAR) modalities, which are not currently integrated into the single-image flow model.
  • Improvement Directions: Extending observation-anchored flow matching to multi-modal condition settings by combining SAR backscatter and multitemporal constraints represents a promising path toward all-weather remote sensing interpretation foundation models.
  • vs EMRDM (CVPR 2025): EMRDM uses a mean-reverting stochastic differential equation for diffusion CR, which involves extensive sampling steps, slow convergence, and purely pixel-level objectives; GACR introduces deterministic OAR-Flow with \(5\times\) faster convergence and enforces explicit downstream semantic alignment via GCPA.
  • vs DFCFormer (JSTARS 2025): DFCFormer relies on dual-domain frequency Transformers for direct residual regression, resulting in over-smoothed blur under thick clouds where signals are missing; GACR preserves generative flexibility via flow residuals while anchoring structures within a semantic manifold.

Rating

  • Novelty: โญโญโญโญโญ [Pioneers an interpretation-oriented cloud removal paradigm unifying observation-anchored flow matching and foundation model manifold alignment]
  • Experimental Thoroughness: โญโญโญโญโญ [Comprehensive benchmarking across 6 CR datasets and 12 downstream interpretation tasks across both pixel and semantic metrics]
  • Writing Quality: โญโญโญโญโญ [Clear motivation, rigorous mathematical formulation, and well-structured empirical validation]
  • Value: โญโญโญโญโญ [Directly resolves the persistent industrial bottleneck where cloud-removed imagery fails in downstream mapping and surveying tasks]