Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment¶
Conference: ECCV 2026
Paper: ECCV 2026 Official
Code: https://github.com/wzy6055/GACR
Area: Segmentation
Keywords: cloud removal, observation-anchored residual flow, geo-contextual alignment, vision foundation model, remote sensing interpretation
TL;DR¶
The paper proposes GACR, an interpretation-oriented cloud removal framework that formulates recovery as an observation-anchored residual flow matching process coupled with a vision foundation model semantic manifold constraint, mitigating hallucinated artifacts and semantic drift while drastically improving downstream interpretation accuracy.
Background & Motivation¶
Optical satellite imagery serves as a foundational data source for Earth observation missions, ranging from urban expansion monitoring and environmental management to land-cover classification. However, widespread cloud coverage in the atmosphere severely disrupts surface reflectance signals, obscuring land structures and injecting critical uncertainty into downstream processing pipelines. Consequently, cloud removal (CR) has shifted from a superficial image enhancement preprocessing step into a vital structural and semantic reconstruction task within remote sensing pipelines. Contemporary deep learning CR frameworks primarily fall into two categories: residual denoising models and generative diffusion models. Residual approaches treat cloud coverage as additive residual noise and regress clear reflectance directly, which inevitably yields over-smoothed textures and structural ambiguities under heavy occlusion where ground truth surface signals are entirely missing. Conversely, diffusion models synthesize visually plausible patterns from pure Gaussian noise, but their stochastic sampling trajectories lack physical observation anchoring, frequently hallucinating geographically implausible textures that contradict true spatial land-cover distributions.
The fundamental culprit behind these failures is the pervasive "visual-fidelity-oriented" optimization paradigm. Prevailing approaches primarily minimize pixel-level reconstruction errors (e.g., L1/L2 losses tailored for conventional PSNR and SSIM benchmarks), under the implicit yet flawed assumption that visual clarity equals semantic correctness. Because real-world geographic landscapes follow strict spatial-semantic rules (e.g., contiguous forest distributions, regular urban geometry, hydrological topography), unconstrained pixel-level optimization under thick cloud occlusion allows synthetic details to drift away from the authentic geographical context. When reconstructed images are fed into downstream networks for semantic segmentation, building extraction, or height estimation, subtle structural deviations and boundary distortions accumulate, severely compromising downstream interpretation reliability.
The angle of attack in this paper is to eliminate unanchored stochastic search while moving beyond narrow pixel-level supervision. The core idea is to reformulate cloud removal as an observation-anchored residual flow matching process grounded in the cloudy observation, and enforce a geo-contextual manifold constraint induced by a vision foundation model, ensuring that the generative trajectory simultaneously delivers high-fidelity texture recovery and strictly preserved downstream interpretation reliability.
Method¶
Overall Architecture¶
The GACR framework consists of three synergistic core components: the physics-inspired Observation-Anchored Residual Flow (OAR-Flow), the Vision Foundation Model (VFM)-based Geo-Contextual Prior Alignment (GCPA), and a downstream interpretation verification protocol.
The full inference pipeline originates from the degraded cloudy observation \(x_c\). OAR-Flow anchors the starting state of a deterministic probability flow ordinary differential equation (ODE) to the cloudy image perturbed by structured residuals. A lightweight and scalable Hourglass Diffusion Transformer (HDiT) serves as the velocity field network, progressively integrating along a deterministic flow path toward the cloud-free clear state \(x_*\). During the training phase, the model is supervised in pixel space via a velocity matching loss \(\mathcal{L}_{\mathrm{vel}}\) to capture transport dynamics from cloudy perturbations to clear terrain. Simultaneously, intermediate bottleneck features are extracted from HDiT and mapped via an Adaptive Projector (ADP) to match multi-scale geo-contextual representations from a frozen DINOv3 vision foundation model under a patch-wise cosine similarity constraint (\(\mathcal{L}_{\mathrm{GCI}}\)), strictly confining reconstruction within a geographically consistent semantic manifold.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Cloudy Observation xc and Ground Truth x*"] --> B["Observation-Anchored Residual Flow<br/>Construct anchored interpolant trajectory xt and decompose velocity"]
B --> C["Efficient Decoupled Architecture<br/>HDiT velocity field prediction and bottleneck feature extraction"]
C --> D["Geo-Contextual Prior Alignment<br/>Frozen VFM feature extraction and adaptive projection alignment"]
D --> E["Unified Objective Optimization and ODE Sampling<br/>Lvel coupled with LGCI for deterministic surface reconstruction"]
E --> F["Downstream Interpretation Evaluation<br/>Land-cover classification / building extraction / segmentation / height estimation"]
Key Designs¶
1. Observation-Anchored Residual Flow: Physically Grounded Deterministic Flow Matching Standard diffusion and generative flow methods typically initiate sampling from unconditioned isotropic Gaussian noise, leading to prolonged trajectories and stochastic drift. In contrast, OAR-Flow constructs a continuous forward interpolant trajectory explicitly anchored in the cloudy observation \(x_c\): $\(x_t = \alpha_t x_* + \beta_t x_c + \sigma_t \epsilon, \quad \epsilon \sim \mathcal{N}(0, \mathbf{I})\)$ with linear time scheduling \(\alpha_t = 1 - t\), \(\beta_t = \rho t\), and \(\sigma_t = t\) for \(t \in [0, 1]\). Under thin cloud coverage, where ground signals remain partially observable, the physical anchor \(x_c\) dominates the dynamics, preserving low-frequency cues and structural boundaries without unguided random perturbation. Under thick cloud coverage, the stochastic perturbation \(\sigma_t \epsilon\) provides the generative flexibility required for semantic inpainting. The corresponding ideal deterministic probability velocity field is: $\(v_t(x) = \mathbb{E}[\dot{\alpha}_t x_* + \dot{\beta}_t x_c + \dot{\sigma}_t \epsilon \mid x_t = x]\)$ The parameterized velocity network \(\mathbf{u}_t(x) = \mathrm{Net}_\theta(x_t, t, x_c)\) directly predicts this transport direction conditioned on \(x_c\). During inference, numerical ODE integration via \(\hat{x}_{t-1} = \hat{x}_t - \Delta t \cdot \mathbf{u}_t(\hat{x}_t)\) deterministically recovers \(\hat{x}_0\) from \(t=1\), accelerating convergence while keeping the generative path strictly aligned with the observation.
2. Geo-Contextual Prior Alignment: VFM Manifold Regularization and GCI Loss To eliminate semantic hallucination and category drift under severe cloud coverage, GCPA leverages a pretrained DINOv3 remote sensing foundation model (e.g., ViT-L/16-SAT-300M or ViT-L/16-LVD-1689M) to construct an authentic geographic semantic manifold. Given the clear image \(x_*\), the frozen VFM encoder extracts dense representations \(z_* = f_{\mathrm{vfm}}(x_*) \in \mathbb{R}^{B \times C' \times H' \times W'}\). Intermediate hidden states \(h_t\) from the HDiT bottleneck are projected into the identical manifold space via an Adaptive Projector (ADP), yielding \(z_t\). Geo-contextual coherence is enforced by maximizing patch-wise cosine similarity via the Geo-Contextual Integrity (GCI) loss: $\(\mathcal{L}_{\mathrm{GCI}} = - \mathbb{E} \left[ \frac{1}{N} \sum_{n=1}^{N} \frac{\langle z_*^{[n]}, z_t^{[n]} \rangle}{\| z_*^{[n]} \|_2 \| z_t^{[n]} \|_2} \right]\)$ This constraint forces the intermediate representations of the generative model to respect large-scale spatial structures (e.g., contiguous forests, road connectivity, water-land boundaries), effectively preventing the model from hallucinating geographically incompatible artifacts.
3. Efficient Decoupled Architecture: HDiT Backbone with Adaptive Projection Conventional Diffusion Transformers (DiTs) incur prohibitive computational complexity (hundreds of GFLOPs) when applied directly in pixel space on high-resolution satellite imagery. The framework adopts the Hourglass Diffusion Transformer (HDiT), which leverages hierarchical multiscale downsampling and upsampling to reduce computational overhead to approximately one-quarter of standard DiT. Crucially, the deepest bottleneck of HDiT inherently aggregates global semantic abstractions that naturally correspond to token representations in foundation models. The lightweight ADP module bridges these representations without requiring fine-tuning of the external foundation model, successfully decoupling pixel-level flow regression from high-level semantic manifold alignment.
Loss & Training¶
The overall training objective is formulated as a unified multi-task loss: $\(\mathcal{L} = \mathcal{L}_{\mathrm{vel}} + \lambda \mathcal{L}_{\mathrm{GCI}}\)$ where the velocity matching loss \(\mathcal{L}_{\mathrm{vel}}\) supervises deterministic pixel-level trajectory fitting: $\(\mathcal{L}_{\mathrm{vel}} = \mathbb{E}_{x_*, \epsilon, x_c, t} \left[ \left\| \mathbf{u}_t(x_t) - (\dot{\alpha}_t x_* + \dot{\beta}_t x_c + \dot{\sigma}_t \epsilon) \right\|^2 \right]\)$ The hyperparameter \(\lambda\) balances pixel-level restoration fidelity and high-level geo-contextual alignment. Empirical analysis identifies \(\lambda = 0.5\) as the optimal trade-off point between pixel sharpness and semantic consistency.
Key Experimental Results¶
Main Results¶
Quantitative evaluations across six benchmark datasets spanning real and synthetic cloud conditions confirm that GACR achieves state-of-the-art reconstruction fidelity, especially in thick cloud environments:
| Method | CUHKCR-GZ PSNRโ / SSIMโ | CUHKCR-CS PSNRโ / SSIMโ | Potsdam-Thick PSNRโ / SSIMโ | Vaihingen-Thick PSNRโ / SSIMโ |
|---|---|---|---|---|
| MPRNet (WACV 2021) | 23.454 / 0.712 | 23.365 / 0.695 | 26.418 / 0.902 | 27.805 / 0.935 |
| Restormer (CVPR 2022) | 25.839 / 0.743 | 23.632 / 0.710 | 28.831 / 0.922 | 28.867 / 0.922 |
| AST (CVPR 2024) | 25.482 / 0.735 | 23.365 / 0.695 | 25.886 / 0.890 | 26.471 / 0.916 |
| MambaIR (ECCV 2024) | 25.626 / 0.733 | 23.445 / 0.704 | 27.027 / 0.903 | 29.852 / 0.945 |
| DFCFormer (JSTARS 2025) | 25.816 / 0.746 | 23.876 / 0.711 | 28.196 / 0.917 | 30.396 / 0.951 |
| EMRDM (CVPR 2025) | 25.862 / 0.747 | 23.736 / 0.712 | 27.199 / 0.923 | 28.979 / 0.951 |
| GACR-SAT/2 (Ours) | 25.964 / 0.736 | 24.230 / 0.709 | 30.578 / 0.934 | 33.018 / 0.964 |
| GACR-SAT/1 (Ours) | 26.100 / 0.744 | 24.354 / 0.713 | 31.049 / 0.938 | 34.048 / 0.970 |
On twelve downstream tasks evaluated with independent frozen DINOv3 encoders, GACR delivers consistent and substantial accuracy gains:
| Model | CLS-1 (Acc.โ) | BLD-2 (IoUโ) | SEG-2 (mIoUโ) | SEG-4 (mIoUโ) | HE-2 (RMSEโ) | HE-4 (RMSEโ) |
|---|---|---|---|---|---|---|
| Without CR (Lower Bound) | 0.746 | 0.596 | 0.490 | 0.550 | 2.703 | 2.095 |
| EMRDM (CVPR 2025) | 0.776 | 0.696 | 0.668 | 0.692 | 2.133 | 1.629 |
| DFCFormer (JSTARS 2025) | 0.763 | 0.657 | 0.630 | 0.671 | 2.317 | 1.687 |
| GACR-SAT/2 (Ours) | 0.833 | 0.710 | 0.699 | 0.737 | 2.014 | 1.554 |
| Clear Image (Upper Bound) | 0.882 | 0.762 | 0.733 | 0.755 | 1.868 | 1.477 |
Ablation Study¶
Ablation analysis on CUHKCR-EXT-CS isolates the contribution of generative modeling and semantic manifold alignment:
| Configuration | PSNR โ | SSIM โ | CLS Acc. โ | BLD IoU โ | GFLOPs | Note |
|---|---|---|---|---|---|---|
| DiT + MRDM | 24.037 | 0.707 | 0.670 | 0.693 | 697.20 | Standard diffusion transformer with heavy compute overhead |
| HDiT + MRDM | 23.736 | 0.712 | 0.726 | 0.696 | 166.72 | Hourglass backbone reduces compute but keeps lengthy trajectory |
| HDiT + OAR-Flow | 23.986 | 0.698 | 0.690 | 0.700 | 56.05 | Observation-anchored flow sharply cuts complexity to 56G |
| HDiT + OAR-Flow + GCPA (Full GACR) | 24.230 | 0.709 | 0.781 | 0.710 | 56.05 | Geo-contextual alignment improves classification by +9.1% |
Hyperparameter sensitivity on Vaihingen-CR-Thick across \(\lambda\) and patch size \(p\):
| Config Parameter | PSNR โ | SSIM โ | SEG (mIoU) โ | HE (RMSE) โ | GFLOPs | Note |
|---|---|---|---|---|---|---|
| \(\lambda = 0.2\) (\(p=2\)) | 33.073 | 0.965 | 0.728 | 1.578 | 56.05 | Weak semantic alignment leads to suboptimal downstream scores |
| \(\lambda = 0.5\) (\(p=2\)) | 33.018 | 0.964 | 0.737 | 1.554 | 56.05 | Optimal balance, maximizing downstream metrics |
| \(\lambda = 1.0\) (\(p=2\)) | 32.838 | 0.964 | 0.730 | 1.563 | 56.05 | Slightly over-constrained high-frequency texture recovery |
| \(\lambda = 5.0\) (\(p=2\)) | 32.615 | 0.962 | 0.727 | 1.596 | 56.05 | Over-regularization degrades reconstruction fidelity |
| \(p = 4\) (\(\lambda=0.5\)) | 30.303 | 0.949 | 0.693 | 1.616 | 14.15 | Coarse partitioning limits fine detail recovery |
| \(p = 1\) (\(\lambda=0.5\)) | 33.799 | 0.969 | 0.720 | 1.548 | 223.60 | Finer granularity improves PSNR at \(4\times\) compute cost |
Key Findings¶
- Observation Anchoring Accelerates Convergence: OAR-Flow reaches high-PSNR regimes with roughly one-third of the iterations required by the mean-reverting diffusion baseline (EMRDM). When augmented with GCPA, the full GACR model achieves an overall \(5\times\) training acceleration.
- Dense Downstream Tasks Are Highly Sensitive to Cloud Degradation: While global classification can be learned reasonably well end-to-end directly from cloudy inputs, dense prediction tasks such as semantic segmentation and height estimation experience catastrophic degradation under heavy cloud coverage (mIoU plunges to 0.490 on Vaihingen-CR-thick). GACR recovers mIoU to 0.699 (approaching the clean upper bound of 0.733), proving that geo-contextual alignment is indispensable for preserving dense geographic category boundaries.
- Feature Distance Distribution Alignment: Feature distance distribution analysis indicates that conventional CR methods produce high-level representations with noticeable distribution shifts compared to clean references. GACR tightly matches the feature distribution of clear imagery, validating that the reconstructed textures carry authentic geographic meaning.
Highlights & Insights¶
- From Pure Noise to Observation-Anchored Flow: By anchoring the deterministic ODE trajectory to the cloudy observation \(x_c\), the model adaptively handles thin-cloud transmission and thick-cloud inpainting, completely avoiding the unanchored hallucinations typical of standard diffusion models.
- Bridging Low-Level Restoration and High-Level Interpretation: Leveraging the latent representation of a frozen vision foundation model via patch-wise cosine alignment injects geographic common sense into low-level restoration at virtually zero extra training overhead.
- Optimal Compute-Accuracy Trade-off: The combination of the HDiT architecture and \(p=2\) patch partitioning maintains an efficient footprint of only 56 GFLOPs while outperforming models with hundreds of GFLOPs across both low-level restoration and high-level interpretation tasks.
Limitations & Future Work¶
- Author-Acknowledged Limitations: The framework depends on a pretrained VFM (such as DINOv3) for high-level semantic priors. In unconventional geographical terrains or novel hyperspectral bands underrepresented in foundation model pretraining, guidance efficacy may diminish.
- Scope Limitations: The study primarily focuses on single-temporal optical satellite imagery. Operational satellite monitoring often exploits multitemporal sequences and cloud-penetrating synthetic aperture radar (SAR) modalities, which are not currently integrated into the single-image flow model.
- Improvement Directions: Extending observation-anchored flow matching to multi-modal condition settings by combining SAR backscatter and multitemporal constraints represents a promising path toward all-weather remote sensing interpretation foundation models.
Related Work & Insights¶
- vs EMRDM (CVPR 2025): EMRDM uses a mean-reverting stochastic differential equation for diffusion CR, which involves extensive sampling steps, slow convergence, and purely pixel-level objectives; GACR introduces deterministic OAR-Flow with \(5\times\) faster convergence and enforces explicit downstream semantic alignment via GCPA.
- vs DFCFormer (JSTARS 2025): DFCFormer relies on dual-domain frequency Transformers for direct residual regression, resulting in over-smoothed blur under thick clouds where signals are missing; GACR preserves generative flexibility via flow residuals while anchoring structures within a semantic manifold.
Rating¶
- Novelty: โญโญโญโญโญ [Pioneers an interpretation-oriented cloud removal paradigm unifying observation-anchored flow matching and foundation model manifold alignment]
- Experimental Thoroughness: โญโญโญโญโญ [Comprehensive benchmarking across 6 CR datasets and 12 downstream interpretation tasks across both pixel and semantic metrics]
- Writing Quality: โญโญโญโญโญ [Clear motivation, rigorous mathematical formulation, and well-structured empirical validation]
- Value: โญโญโญโญโญ [Directly resolves the persistent industrial bottleneck where cloud-removed imagery fails in downstream mapping and surveying tasks]