Skip to content

Degradation-Robust and Temporally Consistent Infrared–Visible Video Fusion via One-step Diffusion Framework

Conference: ECCV 2026
Paper: ECCV 2026
Area: Image Restoration
Keywords: Infrared-visible video fusion, One-step diffusion model, Rectified flow, Temporal consistency, Local motion compensation

TL;DR

Addressing the vulnerability of conventional video fusion pipelines where modality-specific degradations impair temporal motion estimation, DRT-VF advocates a "fusion-first" paradigm using a one-step rectified-flow diffusion framework paired with hierarchical temporal modeling, achieving high-fidelity and flicker-free video fusion.

Background & Motivation

Infrared and visible sensors offer complementary strengths for all-weather environment perception: visible imaging captures abundant high-frequency textures and chromatic details under favorable illumination, but degrades sharply in nighttime, extreme shadows, or severe haze; conversely, infrared sensors reliably capture thermal signatures immune to illumination changes, but lack fine visual textures and are prone to sensor noise. Although deep-learning-based image fusion methods have achieved considerable success in static scenes, directly deploying image models frame by frame onto continuous video streams introduces severe temporal flicker, ghosting artifacts, and frame-to-frame visual incoherence in practical applications like autonomous driving and 24/7 surveillance.

To alleviate temporal inconsistency, existing video fusion frameworks typically incorporate recurrent feature aggregation or optical-flow-based spatial warping. However, the overwhelming majority follow a "temporal-first" processing paradigm, conducting intra-modality temporal alignment prior to cross-modal feature fusion. This design suffers from an intrinsic flaw: it relies heavily on motion estimation derived from single-modality observations. In real-world degraded scenarios, visible frames suffer from underexposure, shadows, and occlusion, while infrared sequences frequently exhibit low signal-to-noise ratios and sensor flicker. Calculating motion fields or optical flows on degraded single-modality streams inevitably leads to substantial estimation drift and geometric errors, which catastrophically propagate into the subsequent cross-modal fusion stage, causing severe temporal collapse and localized distortion.

This work rethinks the temporal-cross-modal pipeline and advocates a "fusion-first" strategy: performing cross-modal complementary integration before temporal restoration to build a degradation-robust joint semantic representation, providing reliable anchors for subsequent inter-frame stabilization. Core idea: build a progressive video fusion framework on a one-step rectified-flow diffusion backbone, wherein Stage I leverages Modality-Collaborative Attention to synthesize a degradation-robust intra-frame representation in the latent manifold, and Stage II hierarchically decomposes temporal consistency into global channel statistics modulation and local gradient-domain motion compensation.

Method

Overall Architecture

DRT-VF adopts a two-stage progressive pipeline (Stage I for intra-frame cross-modal representation learning and single-step generative refinement, Stage II for hierarchical temporal consistency regularization and motion compensation), illustrated in the flowchart below:

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Input: Consecutive Visible & Infrared Frames<br/>{I_t^v, I_t^i}"] --> MCA["Modality-Collaborative Attention (MCA)<br/>Cross-modal interaction builds joint semantic anchor"]
    MCA --> DiT["One-step Rectified Flow DiT<br/>Generative prior denoising on latent manifold"]
    DiT --> GTC["Global Temporal Consistency (GTC)<br/>Channel statistics residual suppresses global flicker"]
    GTC --> LMC["Local Motion Compensation (LMC)<br/>DCN alignment + context aggregation & boundary refinement"]
    LMC --> Out["Output: High-fidelity Temporally Consistent Fused Video"]

In Stage I, visible frames are encoded into compact structural latents via a pre-trained variational autoencoder (VAE), while infrared frames are processed by a lightweight trainable Infrared Adapter to capture thermal saliency. Both feature streams are fed into the Modality-Collaborative Attention (MCA) module to dynamically synthesize a unified semantic anchor. This joint representation conditions a one-step Diffusion Transformer (DiT) governed by rectified flow, mapping the degraded representation to a clean latent manifold in a single inference step before the VAE decoder outputs the refined intra-frame result.

In Stage II, the spatial parameters are stabilized, and hierarchical temporal modeling is applied across neighboring frames. First, in the low-resolution latent space, the Global Temporal Consistency (GTC) module regularizes channel-wise feature statistics across adjacent frames, eliminating large-scale photometric flicker without requiring explicit spatial alignment. Second, embedded right before the final full-resolution decoder layer, the Local Motion Compensation (LMC) module performs deformable alignment (DCN) and decouples context aggregation from gradient-domain boundary refinement, restoring sharp dynamic contours and eliminating ghosting artifacts.

Key Designs

1. Modality-Collaborative Attention (MCA): Adaptive cross-modal synergy via a shared semantic anchor Direct concatenation or naive summation often injects modality-specific noise (e.g., infrared sensor noise or visible underexposure artifacts) into fused features. MCA circumvents rigid early alignment by adopting a query-driven cross-attention mechanism. Latent visible features \(F^v\) and infrared features \(F^i\) are concatenated along the channel dimension and passed through a convolutional projection to form a unified query anchor \(Q\). Modality-specific linear projections subsequently yield independent key-value pairs \((K^v, V^v)\) and \((K^i, V^i)\). Using \(Q\) as a shared semantic reference, cross-attention independently aggregates complementary cues from both modalities:

\[A^v = \text{Softmax}\left(\frac{Q (K^v)^\top}{\sqrt{d}}\right)V^v, \quad A^i = \text{Softmax}\left(\frac{Q (K^i)^\top}{\sqrt{d}}\right)V^i\]

The attended features are passed through feed-forward networks (FFNs) to generate the unified intra-frame representation \(F^{\text{fuse}} = \text{FFN}(A^v) + \text{FFN}(A^i)\), filtering single-modality degradations while preserving high-contrast thermal targets and background structures.

2. One-step Rectified Flow DiT Backbone: Generative diffusion priors with deterministic single-step inference While multi-step diffusion models excel at synthesizing realistic microstructures, their iterative sampling overhead (tens to hundreds of steps) is impractical for continuous video fusion. DRT-VF adopts Rectified Flow (RF) to enforce straight-line Ordinary Differential Equation (ODE) trajectories between Gaussian noise \(z_1\) and the clean fused latent \(z_0\), training the network \(v_\theta\) to predict the velocity vector \((z_1 - z_0)\):

\[\mathcal{L}_{RF} = \mathbb{E}_{z_0, z_1, t}\left[ \| v_\theta(z_t, t, c) - (z_1 - z_0) \|^2 \right]\]

where \(c\) is the cross-modal condition supplied by the MCA module. Because the learned flow trajectory is linear, inference requires only a single Euler step: \(\hat{z}_0 = z_1 - v_\theta(z_1, t=1, c)\). This retains the generative restoration capability of diffusion priors for severely degraded scenes while enabling ultra-fast, real-time-compatible inference.

3. Global Temporal Consistency (GTC): Channel-wise response statistics regularization for flicker suppression Photometric flicker in multimodal video fusion predominantly stems from rapid inter-frame fluctuations in global feature activations. GTC suppresses this instability directly in the latent space without relying on error-prone optical flow. Given features from three consecutive frames \(\{F_{t-1}, F_t, F_{t+1}\}\), global average pooling (GAP) compresses spatial dimensions into channel vectors \(u_i = \text{GAP}(F_i)\). An explicit deviation residual \(\delta_t = u_t - \frac{u_{t-1} + u_{t+1}}{2}\) acts as an anomaly indicator for frame-to-frame intensity jumps. A lightweight MLP with Sigmoid activation maps the pooled statistics and deviation vector to channel-wise weights \(w_t\), recalibrating the center frame features residually via \(F_t^{\text{out}} = F_t + F_t \odot w_t\). This effectively levels global brightness oscillations without degrading spatial high-frequency details.

4. Local Motion Compensation (LMC): Decoupled dual-branch dynamic boundary and context refinement Because latent global statistics cannot resolve local pixel-level displacements of dynamic objects, LMC operates at full spatial resolution immediately before the final decoder layer. Neighboring features are coarsely aligned to the center reference frame using Deformable Convolution Networks (DCN) with learned offsets: \(\hat{D}_k = \text{DCN}(D_k, \Delta p_k)\). To prevent the interpolation blur and ghosting inherent in direct aggregation, LMC splits refinement into two decoupled branches: the Context Aggregation Branch (CAB) utilizes Mutual Self-Attention (MutualSA) using center features as Query and aligned neighbors as Keys/Values to filter misaligned artifacts and enrich local textures; in parallel, the Boundary Refinement Branch (BRB) computes the temporal gradient residual \(\mathcal{G}_t = | \nabla \hat{D}_{t-1} - \nabla D_t | + | \nabla \hat{D}_{t+1} - \nabla D_t |\), isolating high-frequency dynamic contours from the static background and calibrating edges through a residual block. The sum of both branches yields crisp boundaries and temporally smooth dynamic regions.

Loss & Training

The framework is trained in a progressive two-stage manner on a single NVIDIA GeForce RTX 5090 GPU: - Stage I (Spatial Integration): Trained for 100 epochs with batch size 8 using AdamW (learning rate \(1 \times 10^{-4}\)). The objective incorporates intensity fidelity, gradient preservation, color retention, and the velocity matching loss \(\mathcal{L}_{RF}\), ensuring high structural and thermal fidelity. - Stage II (Temporal Consistency): Trained for 100 epochs with batch size 1 across consecutive frame sequences. Retaining the spatial fusion losses, it incorporates an inter-frame temporal consistency regularization term \(\mathcal{L}_{\text{temporal}}\) following motion compensation to jointly optimize GTC and LMC parameters.

Key Experimental Results

Main Results

Evaluation is conducted across four challenging infrared-visible video fusion benchmarks: HDO (infrared flicker), M3SVD (multiple dynamic targets), NOT-156 (nighttime low-light), and VTMOT (fast motion blur). Metrics include Mutual Information (MI↑), Visual Information Fidelity (VIF↑), and Structural Similarity (SSIM↑) for spatial fidelity, alongside Bidirectional Short-term Weighted Error (BiSWE↓) and Mean Squared Sum of Reciprocal Flow (MS2R↓) for temporal stability.

Dataset Metric DRT-VF (Ours) UniVF (Prev. SOTA) TemCOCO (Video Baseline) DCEVO (Image SOTA)
HDO MI ↑ / VIF ↑ / SSIM ↑ 3.521 / 0.792 / 0.680 3.483 / 0.806 / 0.671 3.250 / 0.647 / 0.568 3.409 / 0.781 / 0.671
BiSWE ↓ / MS2R ↓ 6.232 / 0.210 6.378 / 0.227 6.414 / 0.213 6.335 / 0.223
M3SVD MI ↑ / VIF ↑ / SSIM ↑ 3.373 / 0.934 / 0.712 3.366 / 0.931 / 0.688 2.218 / 0.466 / 0.441 3.421 / 0.931 / 0.691
BiSWE ↓ / MS2R ↓ 6.479 / 0.222 7.316 / 0.222 6.488 / 0.258 7.395 / 0.227
NOT-156 MI ↑ / VIF ↑ / SSIM ↑ 4.449 / 0.934 / 0.642 4.410 / 0.926 / 0.628 3.713 / 0.703 / 0.605 4.219 / 0.884 / 0.641
BiSWE ↓ / MS2R ↓ 7.542 / 0.574 8.551 / 0.591 8.122 / 0.826 8.579 / 0.577
VTMOT MI ↑ / VIF ↑ / SSIM ↑ 4.232 / 0.981 / 0.662 4.212 / 0.977 / 0.656 2.814 / 0.712 / 0.615 3.822 / 0.953 / 0.654
BiSWE ↓ / MS2R ↓ 4.822 / 0.332 4.690 / 0.338 4.593 / 0.382 4.848 / 0.345

Across all four benchmarks, DRT-VF consistently secures top-two rankings. In challenging nighttime scenarios (NOT-156), DRT-VF drops BiSWE from 8.551 (UniVF) down to 7.542, verifying its robustness against illumination collapse. In downstream zero-shot object tracking using ByteTrack with a frozen YOLOv11n on NOT-156, DRT-VF reaches 0.244 AUC and 0.236 DP@20, substantially outperforming all competing models.

Ablation Study

Ablations on M3SVD and NOT-156 examine the architectural paradigm, Stage I/II modules, and the decoupled LMC components:

Config M3SVD: MI ↑ M3SVD: SSIM ↑ M3SVD: BiSWE ↓ NOT-156: MI ↑ NOT-156: SSIM ↑ NOT-156: BiSWE ↓ Note
Full model (Ours) 3.373 0.712 6.479 4.449 0.642 7.542 Complete two-stage fusion-first model
Temporal-First 3.152 0.651 7.854 4.125 0.594 8.956 Moving temporal modeling prior to fusion causes sharp metric drop
w/o Stage II 3.392 0.724 8.542 4.466 0.653 9.684 High spatial metrics but complete loss of temporal stability
w/o MCA 3.214 0.686 6.551 4.253 0.616 7.625 Simple concatenation degrades cross-modal complementarity
w/o DiT 3.187 0.672 6.613 4.204 0.608 7.702 Bypassing diffusion backbone impairs structural detail recovery
w/o GTC 3.361 0.706 8.216 4.432 0.636 9.458 Removing GTC drastically increases BiSWE (severe flicker)
w/o LMC 3.315 0.682 6.582 4.385 0.612 7.683 Removing LMC degrades MS2R (0.574 to 0.694 on NOT-156)
w/o DCN (in LMC) 3.332 0.693 6.514 4.406 0.622 7.605 Omitting alignment causes severe ghosting artifacts
w/o CA (CAB in LMC) 3.295 0.688 6.492 4.367 0.627 7.564 Discarding MutualSA weakens context aggregation
w/o BR (BRB in LMC) 3.344 0.698 6.486 4.415 0.631 7.552 Omitting gradient refinement yields blurry motion boundaries

Key Findings

  • Superiority of the "Fusion-First" paradigm: The Temporal-First baseline suffers severe performance degradation (NOT-156 BiSWE worsening to 8.956 and MS2R to 0.652) because modality-specific noise and low illumination derail motion estimation before cross-modal fusion can compensate.
  • Clear decomposition in hierarchical temporal consistency: Removing GTC causes severe large-scale photometric flicker (BiSWE spikes by nearly 2 points on NOT-156), whereas removing LMC causes fine-scale motion blur and ghosting (MS2R jumps from 0.574 to 0.694), proving that macroscopic and microscopic temporal modeling are orthogonal and mutually necessary.
  • Value of single-step diffusion priors: Removing the DiT backbone degrades MI and SSIM across all benchmarks, proving that rectified-flow diffusion priors are essential for reconstructing fine structural details from degraded inputs.

Highlights & Insights

  • "Fusion-First" paradigm shift: By performing cross-modal complementary fusion prior to inter-frame temporal modeling, the method provides a clean, degradation-robust joint semantic representation for temporal stabilization, cutting off the error propagation chain inherent in single-modality motion tracking.
  • One-step rectified flow for video-rate diffusion: Integrating single-step ODE rectified flow into video fusion harnesses the rich generative priors of diffusion models to reconstruct subtle textures from degraded inputs without the prohibitive computational delay of iterative reverse sampling.
  • Decoupled global and local temporal stabilization: Decomposing temporal consistency into low-resolution latent channel statistics regularization (GTC for flicker suppression) and high-resolution gradient-domain boundary refinement (LMC for edge preservation) delivers a highly balanced and modular architecture.

Limitations & Future Work

  • Temporal sliding window restricted to three frames: The current GTC and LMC designs operate over a symmetric 3-frame local window, lacking long-term memory to resolve persistent occlusions or extended slow-moving dynamics.
  • Sensitivity to severe hardware parallax: While DCN provides non-rigid geometric adaptation, scenes with large physical binocular parallax or uncalibrated camera rigs may still exhibit residual boundary artifacts.
  • Future directions: Integrating state-space models (Mamba) or recurrent memory queues could extend temporal coherence over long horizons, while integrating self-adaptive stereo parallax rectification into Stage I could handle uncalibrated dual-camera streams.
  • vs UniVF [CVPR 2024]: UniVF enhances single-modality features with temporal context prior to fusion, which tends to amplify noise under severe modality degradation (e.g., extreme underexposure); DRT-VF establishes a fusion-first paradigm with diffusion priors, achieving much stronger degradation robustness.
  • vs TemCOCO [ICCV 2023]: While TemCOCO utilizes visual-semantic collaboration and recurrent constraints, it lacks alignment-free channel statistics regularization to handle global exposure flicker; DRT-VF's hierarchical GTC and LMC modules provide superior flicker elimination and edge sharpness.
  • vs DDFM [ICCV 2023] / Mask-Dif [TPAMI 2024]: Traditional diffusion-based fusion relies on dozens of stochastic reverse steps, incurring multi-second latency per frame and severe inter-frame jitter; DRT-VF employs deterministic one-step rectified flow combined with explicit two-stage temporal training, making diffusion-based video fusion practical and coherent.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Advocates the principled "fusion-first" paradigm, effectively integrating one-step rectified flow diffusion with decoupled hierarchical temporal consistency.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluations across four degraded video benchmarks, detailed quantitative spatial and temporal metrics, extensive ablations, and zero-shot downstream tracking validation.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear exposition, insightful motivation, rigorous mathematical formulation, and well-structured experimental analysis.
  • Value: ⭐⭐⭐⭐☆ Highly practical and generalizable for real-world all-weather vision tasks, including autonomous driving, robotics, and low-light surveillance.