Skip to content

Ice Cloud Geometry Retrieval with Calibrated Uncertainty from Passive Satellite Imagery

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/ayushprd/sparse2cloud
Area: Remote Sensing / Earth Science (recommended reclassification from autonomous_driving)
Keywords: Ice cloud geometry retrieval, Sparse-to-dense prediction, Conformalized quantile regression, Uncertainty quantification, Cross-track generalization

TL;DR

Addressing the coverage gap between narrow active radar/lidar nadir tracks (~1.6% supervision) and wide passive satellite swaths lacking direct vertical sensitivity, this paper introduces an ERA5-conditioned ConvNeXt-UNet that produces dense predictions of eight ice cloud geometry and microphysics targets with guaranteed 90% coverage via conformalized quantile regression (CQR), demonstrating that the retrieval operates as an independent per-pixel spectral mapping rather than spatial interpolation.

Background & Motivation

Ice clouds play a decisive role in regulating the Earth's radiation budget by reflecting incoming solar shortwave radiation and absorbing outgoing terrestrial longwave radiation. In successive Intergovernmental Panel on Climate Change (IPCC) assessment reports, cloud vertical geometry and its associated radiative feedbacks have consistently remained the largest source of uncertainty in global climate projections. Accurately characterizing cloud top height, cloud base height, geometric thickness, and vertical ice water content is imperative for resolving atmospheric radiative transfer, precipitation initiation, and large-scale atmospheric dynamics.

The gold standard for observing cloud vertical structure relies on active spaceborne instruments, such as the radar and lidar aboard the recently launched EarthCARE satellite. Active sensors deliver unmatched vertical resolution and physical accuracy; however, operating strictly along a narrow nadir track of roughly 5 km width, they cover less than 1% of the globe daily, resulting in extreme spatial sparsity. Conversely, passive optical and infrared radiometers, such as VIIRS on Suomi-NPP, provide wide swaths (~3000 km) and achieve twice-daily global coverage at 750 m resolution. Yet, due to the vertically integrated nature of passive radiances, operational algorithms based on lookup tables only estimate cloud-top properties and cannot infer cloud base, geometric thickness, or vertical ice water distribution.

This fundamental dichotomy between vertical physical fidelity and spatial swath coverage motivates learning dense 3D cloud geometry across wide passive swaths from extremely sparse active supervision (<1.6% of pixels). Nevertheless, existing models face unanswered questions regarding their cross-track generalization across thousands of kilometers, and downstream climate models cannot assimilate uncalibrated predictions. The core idea is to frame sparse-to-dense ice cloud geometry retrieval as a per-pixel spectral mapping conditioned on thermodynamic background profiles, leveraging a ConvNeXt-UNet quantile ensemble with post-hoc conformalized quantile regression (CQR) to provide distribution-free 90% coverage guarantees and eliminating vertical ordering violations via zero-cost post-hoc projection.

Method

Overall Architecture

The retrieval pipeline takes as input ten VIIRS channels at 750 m spatial resolution (five thermal infrared brightness temperature bands M12โ€“M16, four shortwave reflectance bands M07/M08/M10/M11, and solar zenith angle) alongside a 104-dimensional ERA5 reanalysis profile (temperature, specific humidity, relative humidity, and geopotential height across 26 pressure levels). Multi-scale spatial-spectral features are extracted through a three-stage ConvNeXt encoder, conditioned at the bottleneck by projected atmospheric profiles and linear self-attention, and decoded through transposed convolutions with skip connections. The network outputs predictions for 8 cloud geometry and microphysics targets across three quantile levels (\(\tau \in \{0.1, 0.5, 0.9\}\)). Finally, predictions are calibrated on an independent calibration set via CQR and projected to enforce physical ordering constraints, yielding dense physical variables with calibrated 90% prediction intervals.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: VIIRS 10-ch Multi-spectral + ERA5 104-dim Profiles"] --> B["Atmosphere-Aware ConvNeXt-UNet"]
    B --> C["Multi-Quantile Ensemble & Post-Hoc CQR Calibration"]
    C --> D["Zero-Loss Post-Hoc Physical Ordering Projection"]
    E["Output: 8 Ice Cloud Geometry Targets + 90% Prediction Intervals"]
    D --> E

Key Designs

1. Atmosphere-Aware ConvNeXt-UNet: Fusing Broad-Swath Spectra with Thermodynamic Background Profiles

Passive vertical cloud retrieval is inherently ill-posed, as local pixel radiances suffer from severe spectral degeneracies under varying synoptic regimes. To resolve this, the architecture employs a three-stage ConvNeXt encoder (channel dimensions [96, 192, 384]), where each stage comprises two ConvNeXt blocks with 7ร—7 depthwise separable convolutions to capture large-scale cloud texture at modest parameter cost, regularized by stochastic depth (drop rate 0.25) and dropout (0.15). At the bottleneck, a 2-layer MLP projects the 104-dimensional ERA5 profile into the 384-dimensional feature space, injecting it additively into the feature maps alongside a 4-head linear self-attention module to capture global atmospheric context. Conditioning on ERA5 provides an 8-point boost in average \(R^2\), and with only 14M parameters, the network matches the performance of a 2.6B-parameter SatVision foundation model.

2. Multi-Quantile Ensemble & Post-Hoc CQR Calibration: Balancing Distributional Sharpness and Coverage Guarantees

Because labels exist solely along the ~1.6% nadir track of EarthCARE, standard models or deep ensembles exhibit severe miscalibration when predicting under sparse supervision (deep ensembles produce an empirical raw coverage of only 37.1% against an 80% nominal interval). The proposed G-QE method averages quantile predictions across five independently trained ConvNeXt-UNet models, further enhanced at inference by D4 test-time augmentation (averaging across 4 rotations and 2 flips). To eliminate residual calibration drift without distributional assumptions, conformalized quantile regression (CQR) is applied on a held-out calibration set using nonconformity scores:

\[s_i = \max \left( \hat{q}_{0.1}(x_i) - y_i,\; y_i - \hat{q}_{0.9}(x_i) \right)\]

Evaluating the score quantile at level \(\lceil (n+1)(1-\alpha) \rceil / n\) yields calibrated prediction intervals \([\hat{q}_{0.1}(x) - \hat{q}, \; \hat{q}_{0.9}(x) + \hat{q}]\). This mathematically guarantees marginal coverage \(P(y \in C(x)) \ge 1-\alpha\) (\(\alpha=0.1\)). Because G-QE produces superior raw intervals (raw coverage 79.6% vs. nominal 80%), its calibrated intervals are significantly sharper (MPIW=25.1) than competing approaches.

3. Zero-Loss Post-Hoc Physical Ordering Projection: Enforcing Vertical Monotonicity

Atmospheric physics imposes strict monotonic relations on ice cloud geometry: cloud base \(\le\) centroid \(\le\) cloud top, and cloud base \(\le\) peak ice level \(\le\) cloud top. When unconstrained, independent quantile predictions occasionally violate these physical boundaries (0.293% violation rate in G-Q). Adding a soft ordering penalty during training (\(L_{\text{order}} = \lambda \max(0, \hat{y}_{\text{base}} - \hat{y}_{\text{top}})\)) causes gradient conflicts that degrade overall retrieval accuracy (\(R^2\) drops from 0.731 to 0.712). Instead, a parameter-free post-hoc projection is applied at inference time by clipping: \(\hat{y}_{\text{base}} \leftarrow \min(\hat{y}_{\text{base}}, \hat{y}_{\text{top}})\) with corresponding cascading adjustments for centroid and peak levels. This completely eliminates physical violations (0% violations) while preserving full regression accuracy across all targets.

Loss & Training

The network predicts 8 target variables across three quantile levels (\(\tau \in \{0.1, 0.5, 0.9\}\)) and is optimized using the pinball loss computed exclusively over the ~1.6% supervised pixel columns \(\Omega_{\text{sup}}\) along the EarthCARE nadir track:

\[\mathcal{L}_{\text{pinball}}(\hat{q}_\tau, y) = \begin{cases} \tau \cdot (y - \hat{q}_\tau), & \text{if } y \ge \hat{q}_\tau \\ (1 - \tau) \cdot (\hat{q}_\tau - y), & \text{otherwise} \end{cases}\]

Models are trained for 200 epochs using AdamW with an initial learning rate of \(10^{-4}\) cosine-annealed to \(10^{-6}\), weight decay of \(5 \times 10^{-4}\), batch size 32, gradient clipping norm 5.0, and bfloat16 mixed precision on an NVIDIA H200 GPU, applying early stopping with a patience of 25 epochs based on validation loss.

Key Experimental Results

Main Results

The models were evaluated on the co-located EarthCARE-VIIRS benchmark (4.5M training, 990K validation, and 924K test pixels) across point estimation baselines and eight uncertainty quantification techniques under identical 90% CQR nominal coverage.

Method Params \(R^2\) โ†‘ MAE โ†“ PICP (Post-CQR) MPIW โ†“ Raw PICP Correlation \(\rho\) โ†‘ Note
G-QE (Ours) 70M (5ร—14M) 0.742 6.08 0.908 25.1 0.796 0.415 5-member quantile ensemble + TTA; top overall performance
G-Q 14M 0.731 6.27 0.912 26.8 0.791 0.388 Single quantile regression; competitive and fast
G-DE 70M (5ร—14M) 0.732 6.20 0.911 27.9 0.371 0.242 Deep ensemble; poor raw calibration forces wide CQR intervals
G-HET 14M 0.718 6.55 0.901 27.1 0.811 0.373 Heteroscedastic regression with learned input variance
G-EVI 14M 0.720 6.39 0.912 28.9 0.511 0.328 Evidential deep learning estimating epistemic uncertainty
G-MCD 14M 0.716 6.41 0.918 31.1 โ€” โ€” MC-Dropout with 20 passes; widest calibrated intervals
G-QP 14M 0.712 6.84 0.907 25.9 0.777 0.383 Soft ordering penalty penalizes gradient descent optimization
G-QPC 14M 0.705 6.95 0.911 30.2 0.683 0.318 Differentiable ordering head constrains network capacity
G-B (point baseline) 14M 0.716 6.41 โ€” โ€” โ€” โ€” Deterministic ConvNeXt-UNet trained with \(L_1\) loss
MLP (5ร—5 ctx + ERA5) 0.9M 0.641 โ€” โ€” โ€” โ€” โ€” Spatial multi-layer perceptron baseline
MLP (pixel only) 0.8M 0.548 โ€” โ€” โ€” โ€” โ€” Point-wise baseline lacking spatial context
BT quadratic regression โ€” 0.590 โ€” โ€” โ€” โ€” โ€” Physics-informed 5 thermal band baseline (MODIS operational counterpart)

Per-Target Accuracy and Physics Constraint Enforcement

The breakdown across all eight vertical geometry and microphysics targets, alongside physical constraint violation rates, is detailed below.

Configuration Centroid Cloud Top Cloud Base Peak Level Thickness Core IWC Mean IWC Log IWP Average \(R^2\) Violations (%)
G-QE + Projection 0.847 0.826 0.859 0.795 0.407 0.812 0.692 0.697 0.742 0.005 โ†’ 0
G-Q + Projection 0.836 0.812 0.847 0.781 0.370 0.800 0.671 0.680 0.725 0.293 โ†’ 0
G-B (baseline) 0.829 0.805 0.848 0.779 0.339 0.788 0.660 0.679 0.716 โ€”
MLP (contextual) 0.782 0.750 0.830 0.726 0.244 0.706 0.530 0.563 0.641 โ€”
G-QPC (structural head) 0.819 0.791 0.831 0.763 0.344 0.774 0.640 0.650 0.702 0 (built-in)

Key Findings

  • Cross-Track Distance Invariance Establishes Spectral Mapping: Binning test pixels by distance from the EarthCARE supervision track (from 0โ€“1 pixels up to 5โ€“8 pixels, ~6 km) reveals that \(R^2\) (0.730 vs. 0.731), PICP (0.90), and interval widths vary by less than 3%. This proves the network performs independent spectral retrieval from per-pixel radiative signatures rather than spatial interpolation from nearby labels.
  • Information-Theoretic Thickness Ceiling: While vertical altitude targets (centroid, top, base) achieve \(R^2 > 0.82\), geometric thickness reaches an empirical ceiling of \(R^2 = 0.407\) (up from 0.244 for MLP). This conforms to atmospheric physics limits where passive thermal infrared radiometers possess only 2โ€“3 degrees of freedom for vertical profiling.
  • Deep Convective Clouds Represent the Primary Failure Mode: Conditional coverage remains remarkably stable across latitudes (polar 0.898 to mid-level 0.920) and diurnal cycles (day 0.905 vs. night 0.911), verifying the dominance of thermal channels. However, deep convective systems suffer an under-coverage of 3.8% (PICP = 0.862) due to intense vertical spatial heterogeneity.

Highlights & Insights

  • Selective Prediction Enhances Climate Utility: Leveraging calibrated interval widths to filter out the 50% most uncertain retrievals boosts cloud top \(R^2\) from 0.742 to 0.935 and thickness \(R^2\) from 0.407 to 0.501, allowing climate researchers to selectively query high-confidence subsets.
  • Uncertainty Features Boost Downstream Cloud Classification: Incorporating predicted interval widths into an ISCCP 6-class cloud classifier increases accuracy from 82.1% to 84.9% and macro F1 from 72.9% to 77.6%, proving that physical uncertainty carries meaningful structural semantics.
  • High-Throughput Planetary Deployment: On a single GPU, the entire inference pipeline processes an entire day of VIIRS global imagery (241 granules, 609 million dense pixel predictions) in 2.4 hours, demonstrating operational readiness.

Limitations & Future Work

  • Author-Acknowledged Limitations: The framework is trained exclusively on EarthCARE ATL_ICE_2A ice cloud products and cannot retrieve warm water clouds or precipitation. In addition, operational execution requires ERA5 reanalysis profiles, which have latency, necessitating operational numerical weather prediction (NWP) forecast substitutes.
  • Negative Empirical Results: Pretrained geospatial foundation models (2.6B SatVision Swin-v2) failed to outperform the 14M ConvNeXt-UNet, indicating that retrieval quality is constrained by input spectral information rather than model capacity. Adversarial PatchGAN training was numerically unstable, and cross-attention ERA5 fusion degraded \(R^2\) by 0.9%.
  • Future Directions: Addressing the 3.8% under-coverage in deep convective clouds requires cloud-type-stratified CQR calibration. Furthermore, surpassing the thickness retrieval ceiling may necessitate multi-angle stereo compositing from overlapping satellite overpasses.
  • vs IceCloudNet (Jeggle et al., 2024): While IceCloudNet pioneered 3D ice water content retrieval from SEVIRI using DARDAR data (\(R^2 \approx 0.69\) on cloudy voxels), this work leverages the higher vertical fidelity of EarthCARE, expands retrieval to eight geometric and microphysical parameters, and establishes a rigorous conformal uncertainty framework.
  • vs Dense Canopy Height from GEDI LiDAR (Pauls et al., ICML 2024): Both tackle sparse active track supervision over dense imagery. Whereas Pauls et al. focused on loss functions resilient to spatial shifts, this paper rigorously proves cross-track distance invariance, demonstrating that sparse active tracks covering diverse weather regimes provide sufficient supervision for global retrieval.

Rating

  • Novelty: โญโญโญโญโ˜† [First dense ice cloud retrieval combining sparse EarthCARE supervision with CQR uncertainty quantification]
  • Experimental Thoroughness: โญโญโญโญโญ [Exhaustive comparison across 8 uncertainty schemes, multi-dimensional conditional calibration, and global deployment]
  • Writing Quality: โญโญโญโญโญ [Clear, rigorous, and transparently documents negative results and physical information boundaries]
  • Value: โญโญโญโญโญ [Provides a scalable, calibrated 3D ice cloud geometry pipeline for operational climate modeling]