Video Can Teach PAN-Sharpening: PSF-Aware Cross-Domain Supervision¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/brachiohyup/ViPS
Area: Remote Sensing
Keywords: PAN-sharpening, cross-domain supervision, point spread function, video prior, remote sensing image fusion
TL;DR¶
ViPS introduces a cross-domain supervision framework that disentangles supervision sources into video-derived spatial supervision for high-frequency detail and misregistration recovery, and real satellite observations for spectral fidelity, bridged by sensor-specific PSF banks to achieve SOTA PAN-sharpening without real HRMS labels.
Background & Motivation¶
PAN-sharpening is a foundational task in Earth observation and remote sensing, aiming to fuse a high-resolution panchromatic (HRPAN) image with a low-resolution multispectral (LRMS) image to synthesize a high-resolution multispectral (HRMS) product exhibiting both sharp spatial boundaries and faithful spectral radiance. Constrained by physical hardware limits, onboard satellite sensors face an inevitable trade-off between spatial sampling and spectral bandwidth. Consequently, modern imaging satellites rely on separate sensor arrays. However, because real HRMS ground-truth imagery is physically unobtainable during orbital acquisitions, the overwhelming majority of deep learning PAN-sharpening models are trained on synthetically downsampled PANโMS pairs under the classic Wald protocol. This synthetic training regime creates an acute domain gap when deploying models on real-world in-orbit data: physical satellite acquisitions inevitably exhibit inter-sensor mounting angles, orbital jitter, and line-delay misregistration from pushbroom scanning mechanisms, alongside complex sensor-specific optical responses. Under full-resolution testing, synthetic-trained models suffer from noticeable double-edge artifacts, structural blurring, and severe spectral distortion.
Prior works attempt to mitigate these artifacts primarily through architectural remedies, including explicit edge and gradient penalties, deformation-aware convolution modules, and modality-consistent attention alignments. Nevertheless, these strategies remain trapped within the closed ecosystem of remote sensing data: panchromatic gradient edges do not always align with true multispectral luminance transitions, meaning that over-constraining spatial alignment often triggers spectral distortion. Crucially, without access to real high-resolution ground truth, satellite-only supervision lacks the diverse, high-frequency spatial gradients and realistic displacement variations required to teach the network robust deblurring and alignment.
This paper breaks free from the closed remote sensing data paradigm by recognizing that high-frame-rate natural video sequences are natural, scalable spatial teachers for satellite PAN-sharpening. Consecutive frames undergoing smooth translational camera motion provide abundant fine-grained geometry and natural textures, while their inter-frame sub-pixel shifts statistically emulate the scanline misregistration characteristic of satellite pushbroom imaging. Core idea: propose the ViPS cross-domain training framework, which disentangles supervision into video-derived pseudo spatial supervision and real satellite-derived spectral self-consistency supervision, coupled with sensor-specific Point Spread Function (PSF) banks to reconcile the optical gap without requiring any high-resolution multispectral labels.
Method¶
Overall Architecture¶
ViPS decouples the training supervision into two cooperative branches: a video branch providing rich spatial structure and displacement robustness, and a remote sensing branch enforcing physical spectral consistency. To bridge the optical domain discrepancy between terrestrial consumer cameras and satellite telescope optics, ViPS pre-constructs sensor-specific Point Spread Function (PSF) banks. In each training iteration, a physically plausible PSF kernel is sampled and applied synchronously to degrade both video frames and real satellite imagery. This matched degradation guarantees that the high-frequency restoration learned from natural scenes aligns with the spectral attenuation of the target satellite sensor.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
subgraph S1["Disentangled Cross-Domain Input & Shift Emulation"]
direction TB
A["Natural Video Frame Pair (It, It+a)<br/>4K 1000 FPS Translation Sampling"]
B["Real Satellite Observation (HRPAN, LRMS)<br/>Target Sensor s"]
end
subgraph S2["PSF Bank Optical Degradation Alignment"]
direction TB
C["Sensor-Specific PSF Bank<br/>Uniform Kernel Sampling Ks ~ Bs"]
end
subgraph S3["Pseudo-Pair Synthesis & Backbone Forward Pass"]
direction TB
D["Video Branch: Pseudo PAN & Pseudo LRMS<br/>Grayscale It + Band Projection & Ks Blur"]
E["RS Branch: LRMS Self-Supervised Scale Degradation<br/>Ks Blur Produces Reduced-Scale Pairs"]
F["Hybrid Convolution-Transformer Backbone<br/>S2D Transform + Swin-T Interaction + D2S Upsampling"]
end
subgraph S4["Disentangled Loss Supervision"]
direction TB
G["Luminance Gradient Spatial Loss Lspa<br/>Enforces Edge Consistency on Pseudo HRMS"]
H["LRMS-Domain Spectral Fidelity Loss Lspe<br/>L1 Self-Consistency on RS Branch"]
end
A --> D
B --> E
C -->|Identical Optical Blur| D
C -->|Identical Optical Blur| E
D --> F
E --> F
F --> G
F --> H
Key Designs¶
1. Disentangled Cross-Domain Supervision: Video-Driven Spatial Restoration Meets Satellite Spectral Fidelity
Conventional PAN-sharpening relies on a single data source to learn both spatial detail and spectral response, leaving networks brittle under real-world misregistration. ViPS resolves this by assigning complementary roles to two domains: spatial edge sharpness is taught by natural video, while spectral radiometric consistency is anchored by real satellite observations. In the video branch, sequences recorded with a 4K, 1000 FPS camera under translational motion are sampled with a temporal offset of \(a=10\) frames to form \((I_t, I_{t+a})\). The reference frame \(I_t\) is converted to single-channel luminance via channel averaging \(\bar{I}_{pan}^{h,p} = \text{Gray}(I_t)\) to serve as the pseudo HRPAN input, while the shifted frame \(I_{t+a}\) undergoes channel projection and PSF degradation to form the pseudo LRMS input \(I_{ms}^{l,p}\). This setup exposes the network to abundant high-frequency structures while naturally simulating inter-band pushbroom timing offsets. Concurrently, the remote sensing branch takes real satellite pairs \((I_{pan}^h, I_{ms}^l)\) and operates at a further degraded scale using observed \(I_{ms}^l\) as the supervisory target. Alternating between these branches prevents the network from inheriting terrestrial video color distributions while conferring superior deblurring and misregistration resilience.
2. Physically Grounded PSF Bank Optical Alignment with Random Kernel Sampling
Directly mixing natural video frames and satellite images causes severe domain collapse due to their vastly divergent point spread functions. Satellite optics exhibit anisotropic blur and Modulation Transfer Function (MTF) cutoffs that simple bicubic or isotropic Gaussian kernels fail to capture. ViPS constructs offline PSF banks \(\mathcal{B}_s\) containing \(M_s = 10,000\) unique \(16\times 16\) kernels for each target satellite (e.g., WorldView-3, QuickBird, GaoFen-2). During training, a kernel is uniformly sampled per mini-batch:
The degradation operator jointly models convolution and scale downsampling by factor \(r=4\):
In the video branch, the kernel degrades projected video features into pseudo LRMS: \(I_{ms}^{l,p} = \mathcal{D}(\mathcal{P}_{C_s}(I_{t+a}); \mathcal{K}_s, r)\). In the satellite branch, the exact same kernel creates reduced-resolution training pairs: \(I_{pan}^l = \mathcal{D}(I_{pan}^h; \mathcal{K}_s, r)\) and \(I_{ms}^{vl} = \mathcal{D}(I_{ms}^l; \mathcal{K}_s, r)\). Sharing the identical optical kernel across domains ensures that high-frequency super-resolution priors learned from video seamlessly transfer to the satellite frequency response, while random sampling prevents overfitting to any single static blur kernel.
3. Misregistration-Tolerant Luminance Gradient Spatial Loss
Because frame pairs \((I_t, I_{t+a})\) contain deliberate physical displacement, and because channel projection \(\mathcal{P}_{C_s}(\cdot)\) is a dimension-matching tool rather than an exact radiometric emulator, computing standard pixel-level \(\mathcal{L}_1\) loss against video frames would cause chromatic errors and penalize structural alignment. ViPS introduces a luminance gradient magnitude loss. The predicted pseudo HRMS \(\hat{I}_{ms}^{h,p}\) is first converted to grayscale \(\bar{I}_{ms}^{h,p} = \text{Gray}(\hat{I}_{ms}^{h,p})\), and its gradient magnitude is matched against the pseudo panchromatic input:
By supervising only edge magnitudes rather than signed pixel intensities, the loss enforces sharp structural boundaries while granting local shift-invariance, training the network to resolve sharp edges even when inputs exhibit sub-pixel misregistration.
4. Lightweight Hybrid Convolution-Transformer Backbone with Direct Texture Injection
To meet the high-throughput requirements of remote sensing imagery while retaining global context modeling, ViPS employs an efficient hybrid architecture. Rather than processing large HR feature maps directly, the panchromatic image \(I_{pan}^h\) is downsampled via a Space-to-Depth (S2D) transform by a factor of 4, expanding its channels 16-fold to match the spatial resolution of \(I_{ms}^l\). The concatenated PAN-MS representation is processed by a wide-receptive-field \(9\times 9\) convolution followed by stacked Swin Transformer Blocks (STB) to model spatial-spectral cross-attention. A second \(9\times 9\) convolution and another set of STB blocks further refine the latent features. Features are then upsampled back to PAN resolution via a Depth-to-Space (D2S) pixel shuffle. Finally, the raw HRPAN image is concatenated with the upsampled feature map to inject pristine high-frequency textures directly before three consecutive \(1\times 1\) convolutions yield the final HRMS prediction. This design keeps the majority of computation in the compact low-resolution space, requiring only 19.5 GFLOPs and 0.003 seconds per \(256\times 256\) frame on an RTX 4090.
Loss & Training¶
The overall training objective combines the video spatial loss and satellite spectral loss:
where the spectral loss is supervised by the observed LRMS image on downsampled inputs:
The loss weights are set to \(\lambda_{spa} = 10\) and \(\lambda_{spe} = 1\). Optimization uses AdamW with an initial learning rate of \(1\times 10^{-4}\) and weight decay of 0.01 for \(10^5\) iterations per sensor, alternating satellite and video mini-batches at a strict 1:1 ratio.
Key Experimental Results¶
Main Results¶
Quantitative evaluations are conducted on the PAN-Collection benchmark across WorldView-3 (8-band), GaoFen-2 (4-band), and QuickBird (4-band) sensors under both Full-Resolution (FR) and Reduced-Resolution (RR) protocols. Below are the quantitative results on WorldView-3, reporting HQNR along with spatial distortion \(D_s\) and spectral distortion \(D_\lambda\) for FR, and standard full-reference metrics (ERGAS, SAM, PSNR, SSIM) alongside runtime and FLOPs for RR:
| Method | Venue | FR: HQNRโ | FR: \(D_s\)โ | FR: \(D_\lambda\)โ | RR: ERGASโ | RR: SAMโ | RR: PSNRโ | RR: SSIMโ | Infer. Time(s)โ | FLOPs(G)โ |
|---|---|---|---|---|---|---|---|---|---|---|
| PanNet | ICCV 2017 | 0.918 | 0.049 | 0.035 | 2.538 | 3.402 | 36.148 | 0.966 | 0.001 | 5.17 |
| MSDCNN | JSTARS 2018 | 0.924 | 0.050 | 0.028 | 2.489 | 3.300 | 36.329 | 0.967 | 0.002 | 14.96 |
| FusionNet | ICCV 2021 | 0.920 | 0.053 | 0.029 | 2.428 | 3.188 | 36.569 | 0.968 | 0.002 | 5.13 |
| LAGConv | AAAI 2022 | 0.915 | 0.055 | 0.033 | 2.380 | 3.153 | 36.732 | 0.970 | 0.004 | 8.43 |
| S2DBPN | TGRS 2023 | 0.946 | 0.030 | 0.025 | 2.245 | 3.019 | 37.216 | 0.972 | 0.005 | 158.94 |
| PanDiff | TGRS 2023 | 0.952 | 0.034 | 0.014 | 2.276 | 3.058 | 37.029 | 0.971 | 2.955 | 62.07 |
| DCPNet | TGRS 2024 | 0.923 | 0.036 | 0.043 | 2.301 | 3.083 | 37.009 | 0.972 | 0.109 | 105.40 |
| TMDiff | TGRS 2024 | 0.924 | 0.059 | 0.018 | 2.151 | 2.885 | 37.477 | 0.973 | 9.997 | 1284.42 |
| CANConv | CVPR 2024 | 0.951 | 0.030 | 0.020 | 2.163 | 2.927 | 37.441 | 0.973 | 0.451 | 52.21 |
| U-Know | CVPR 2025 | 0.955 | 0.029 | 0.016 | 2.046 | 2.797 | 37.934 | 0.976 | - | - |
| PAN-Crafter | ICCV 2025 | 0.958 | 0.027 | 0.016 | 2.040 | 2.787 | 37.956 | 0.976 | 0.009 | 79.03 |
| ViPS (Ours) | ECCV 2026 | 0.959 | 0.026 | 0.016 | 2.021 | 2.770 | 38.021 | 0.978 | 0.003 | 19.50 |
On GaoFen-2 and QuickBird, ViPS establishes state-of-the-art results: on GaoFen-2, ViPS attains 0.976 HQNR and 45.122 dB PSNR; on QuickBird, ViPS achieves 0.932 HQNR and 38.451 dB PSNR.
Ablation Study¶
The table below isolates the contributions of video spatial supervision, PSF-bank degradation alignment, random kernel sampling, and evaluates backbone transferability on WorldView-3:
| Variant | Video \(\mathcal{L}_{spa}\) | PSF-Bank Degrad. | Random Sampling | FR: HQNRโ | FR: \(D_s\)โ | FR: \(D_\lambda\)โ | RR: ERGASโ | RR: SAMโ | RR: PSNRโ |
|---|---|---|---|---|---|---|---|---|---|
| (A) RS-only | โ | โ | โ | 0.949 | 0.036 | 0.015 | 2.420 | 2.992 | 37.118 |
| (B) w/o PSF-bank | โ | โ | โ | 0.956 | 0.028 | 0.016 | 2.235 | 2.892 | 37.277 |
| (C) w/o Random Sel. | โ | โ | โ | 0.948 | 0.029 | 0.023 | 2.213 | 3.190 | 37.613 |
| (D) ViPS (full) | โ | โ | โ | 0.959 | 0.026 | 0.016 | 2.020 | 2.770 | 38.021 |
| PanNet (baseline) | - | - | - | 0.918 | 0.049 | 0.035 | 2.538 | 3.402 | 36.148 |
| PanNet + ViPS Super. | โ | โ | โ | 0.954 | 0.028 | 0.018 | 2.217 | 2.814 | 37.229 |
Sensitivity analysis on the frame interval offset \(a\) shows: \(a=1\) yields 37.359 dB PSNR; \(a=5\) yields 37.808 dB; \(a=10\) provides the optimal trade-off at 38.021 dB; and \(a=20\) drops performance to 37.013 dB due to severe correspondence breakdown. Thus, \(a=10\) serves as the default setting.
Key Findings¶
- Cross-domain supervision provides massive spatial restoration gains: Compared with the RS-only baseline (A), adding video spatial supervision drops spatial distortion \(D_s\) from 0.036 to 0.026 and boosts PSNR by over 0.9 dB, proving that natural video sequences effectively resolve the lack of real HRMS labels.
- Optics alignment is mandatory for stable domain transfer: Disabling the PSF bank (B) or freezing the kernel without random sampling (C) severely degrades spectral fidelity, increasing SAM from 2.770 to 3.190. Physically grounded optical alignment is the linchpin that allows the network to ingest natural video without corrupting satellite spectral dynamics.
- Model-agnostic supervision superiority: Equipping the basic 2017 CNN architecture (PanNet) with the ViPS supervision framework elevates its HQNR from 0.918 to 0.954 and PSNR from 36.148 dB to 37.229 dB. This demonstrates that overcoming data supervision bottlenecks produces larger gains than simply scaling architectural depth.
Highlights & Insights¶
- Shifting the paradigm from network design to supervision: Rather than contriving complex modules to mitigate synthetic data artifacts, ViPS demonstrates that high-frame-rate terrestrial video translation naturally mimics orbital pushbroom line-rates and inter-band offsets.
- Minimalist physical degradation bridge: By deploying sensor-specific unit-sum normalized 16ร16 PSF kernels across both training branches, ViPS bridges the camera-satellite domain gap without adversarial objectives or unstable multi-stage pipelines.
- Exceptional efficiency for operational deployment: In contrast to diffusion-based competitors requiring tens of sampling steps (e.g., TMDiff taking ~10 seconds per image), ViPS runs in just 0.003 seconds with 19.5 GFLOPs, rendering it viable for satellite edge computing and large-scale ground-station production.
Limitations & Future Work¶
- Static scene and planar translation assumptions: The curated ViPS dataset assumes static scenes with pure translation. Real satellite observations over mountainous terrain or dynamic scenes (e.g., maritime waves, high-speed vehicles) present parallax and non-rigid displacements that video translation cannot fully emulate.
- Naive spectral band projection: Video RGB frames are mapped to multispectral channel counts using a linear projection matrix \(\mathcal{P}_{C_s}(\cdot)\), without modeling specific Spectral Response Functions (SRF) for near-infrared or red-edge bands.
- Future directions: Integrating optical flow warping to simulate topographic relief parallax, and incorporating spectral reflectance synthesis models into the video pseudo-generation pipeline.
Related Work & Insights¶
- vs Wald Protocol (Wald et al., 2000): Conventional supervised methods overfit to artificially downsampled pairs; ViPS bypasses Wald's limitations by exploiting video translation for edge deblurring and real observations for spectral fidelity.
- vs PAN-Crafter (ICCV 2025) & U-Know (CVPR 2025): Recent methods rely on attention-guided feature warping or uncertainty diffusion distillation; ViPS shows that addressing the supervision source directly achieves superior quality with vastly lower inference cost.
- vs Spatial Data Augmentation (Chen et al., TGRS 2023): While Chen et al. introduced PSF banks as a spatial data augmentation for satellite images, ViPS elevates PSF banks into a cross-domain bridge connecting natural video and satellite sensors.
Rating¶
- Novelty: โญโญโญโญโญ [Pioneering use of natural video as a spatial and misregistration teacher for satellite PAN-sharpening]
- Experimental Thoroughness: โญโญโญโญโญ [Rigorous evaluation across 3 satellites under both FR and RR protocols, accompanied by thorough ablation and backbone transfer tests]
- Writing Quality: โญโญโญโญโญ [Clear motivation, mathematically rigorous degradation formulation, and transparent experimental reporting]
- Value: โญโญโญโญโญ [Resolves the longstanding real-world HRMS ground truth shortage with an ultra-fast inference footprint]