P²Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/YiShi99/P2Fusion
Area: 3D Vision
Keywords: Infrared-visible image fusion, Prompt-based learning, Knowledge distillation, Multimodal perception, Mixture of experts
TL;DR¶
To tackle optimization conflicts from static prior penalties and granularity mismatch from foundation models in infrared-visible image fusion, P²Fusion introduces the "Teach-to-Fuse" paradigm that distills intrinsic thermal saliency and spatial quality into dynamic learnable prompts, combining with a Gated Dynamic Expert Recalibration module for decoupled feature refinement.
Background & Motivation¶
Infrared-visible image fusion (IVIF) aims to integrate complementary modalities by merging thermal radiation targets captured by infrared sensors with intricate background textures captured by visible sensors, providing all-weather visual representations for downstream applications like autonomous driving, target re-identification, and security surveillance. Because paired ground-truth images do not exist in real-world scenes, IVIF is fundamentally an unsupervised generative problem. Conventional deep architectures rely heavily on hand-crafted loss functions and heuristic designs, often struggling to strike an optimal balance between preserving fine textural fidelity and highlighting high-contrast thermal signatures. Consequently, prior-guided fusion has emerged as a promising direction.
However, existing prior-guided paradigms suffer from severe limitations across three mainstream routes. First, explicit hard-constraint approaches (e.g., STDFusionNet, PIAFusion) employ static saliency masks or segmentation labels as rigid penalty losses. This forces networks to fit prior distributions over pixel-level interactions, sacrificing subtle textural boundaries in regions where modalities compete. Second, joint multi-task optimization pipelines (e.g., TarDal, SuperFusion) integrate downstream detection or segmentation networks into the training loop, which introduces severe multi-objective gradient conflicts—optimizing for bounding-box saliency frequently compromises natural background details, inducing a persistent trade-off between aesthetic fidelity and task accuracy. Third, recent extrinsic semantic guidance frameworks (e.g., CLIP- or DINO-based methods) adopt large-scale vision foundation models. Although they alleviate rigid penalties, high-level abstract semantic representations suffer from a profound "granularity mismatch" with the low-level pixel reconstruction required by image fusion, failing to provide precise guidance for fine-grained local textures.
This paper's angle of attack is to return to the intrinsic properties of multimodal images rather than chasing static priors or external foundation models, creating a lightweight and adaptive prompt learning mechanism. Core idea: propose the "Teach-to-Fuse" paradigm that distills intrinsic infrared thermal saliency and visible spatial quality into dynamic, learnable prompts, coupled with a Gated Dynamic Expert Recalibration module that decouples feature refinement and rectifies biased signals through specialized experts.
Method¶
Overall Architecture¶
The P²Fusion pipeline consists of dual-teacher prompt distillation, bidirectional cross-attention modulation, and Gated Dynamic Expert Recalibration (GDER). First, input image pairs are processed by feature encoders while domain teachers extract intrinsic thermal saliency and spatial quality priors, which are distilled into learnable dynamic prompt embeddings. Next, the dynamic prompts are injected into the modal features to condition bidirectional cross-attention for cross-modal interaction. Finally, the interactive features are passed into the GDER module, which synergizes a modality specialist and an attention specialist under dynamic gating to perform multi-round progressive refinement before reconstruction via the image decoder.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multimodal Inputs<br/>Infrared Image I_ir and Visible Image I_vis"] --> B["Prior Extraction Teachers<br/>Infrared Saliency IST and Visible Quality VQT"]
B --> C["Dual-Teacher Prompt Distillation<br/>Distill intrinsic priors into dynamic soft prompts"]
C --> D["Bidirectional Cross-Attention<br/>Feature modulation conditioned by prompts F'"]
D --> E["Gated Dynamic Expert Recalibration (GDER)<br/>Dynamic arbitration between Modality & Attention Experts"]
E --> F["Progressive Reconstruction<br/>High-fidelity Fused Image I_fused and Downstream Tasks"]
Key Designs¶
1. Dual-Teacher Prompt Distillation: Soft dynamic prompts resolving hard constraints and granularity mismatch
Conventional methods impose hard spatial penalties using binary masks or category boundaries, which limits optimization flexibility and degrades fine-grained textures. Meanwhile, borrowing priors from foundation models introduces an abstract granularity mismatch with pixel-level reconstruction. To resolve both issues, this design proposes the Teach-to-Fuse mechanism, extracting endogenous properties from within the modalities and distilling them into dynamic prompt vectors that softly guide the feature representations. In the infrared branch, an Infrared-Saliency Teacher (IST) built on SegFormer-B2 pre-trained on MSRS and FMB extracts foreground target saliency distributions \(T_{ir} \in \mathbb{R}^{H \times W}\) covering pedestrians and vehicles. In the visible branch, a no-reference Visible-Quality Teacher (VQT) using BRISQUE evaluates spatial degradation scores \(s \in [0, 100]\) and normalizes them into continuous quality prior scalars \(T_{vis} = (100 - s) / 100\). The learnable prompt features \(P_{ir}\) and \(P_{vis}\) are aligned with the teachers via BCE and MSE distillation losses, and added to the shallow modal features: $\(F_m = f_m + \text{Embed}(P_m), \quad m \in \{vis, ir\}\)$ This converts rigid constraints into continuous soft conditioning, eliminating gradient conflicts while steering the network toward salient foregrounds and reliable textures.
2. Gated Dynamic Expert Recalibration (GDER): Decoupled specialist cooperation for self-correcting arbitration
Infrared and visible feature spaces are inherently heterogeneous: infrared captures thermal thermal contours while visible captures rich textural details. Standard cross-attention features often suffer from modal confusion and interference. Unlike standard Mixture-of-Experts (MoE) architectures that focus purely on scaling parameter capacity, the GDER module serves as a decoupled feature recalibrator built on functional specialization, pairing a prior-responsive modality expert \(E_{mod}\) with a prior-agnostic attention expert \(E_{att}\). The prior-responsive modality expert \(E_{mod}\) explicitly ingests dynamic prompts to focus on salient foreground targets and high-contrast edges. Conversely, the prior-agnostic attention expert \(E_{att}\) operates independently of the prompts using convolutional block attention (CBAM) to capture long-range contextual dependencies and subtle background textures that might be missed by the priors. A gating network \(G(\cdot)\) assesses local feature characteristics and prompt alignment to allocate adaptive weights \(w = \text{softmax}(G(F', P))\): $\(F'' = F' + w_1 \cdot E_{mod}(F') + w_2 \cdot E_{att}(F')\)$ Empirical correlation between both specialists is only \(r = 0.46\), proving functional decoupling. When severe overexposure or smoke degrades prior accuracy, the gating network dynamically suppresses the modality expert and elevates the attention expert, providing robust error correction.
Loss & Training¶
The overall objective integrates perceptual reconstruction constraints and prompt distillation terms: $\(\mathcal{L}_{total} = \lambda_1 \mathcal{L}_{int} + \lambda_2 \mathcal{L}_{ssim} + \lambda_3 \mathcal{L}_{grad} + \lambda_4 \mathcal{L}_{ir}^{distill} + \lambda_5 \mathcal{L}_{vis}^{distill}\)$ The perceptual terms include: 1. Intensity loss \(\mathcal{L}_{int} = \|I_{fused} - \max(I_{ir}, I_{vis})\|_1\) to balance thermal luminance and visible lighting; 2. Gradient loss \(\mathcal{L}_{grad} = \|\nabla I_{fused} - \max(|\nabla I_{ir}|, |\nabla I_{vis}|)\|_1\) to enforce dominant edge preservation; 3. Structural similarity loss \(\mathcal{L}_{ssim} = 1 - (\omega_{ir} \text{SSIM}(I_{ir}, I_{fused}) + \omega_{vis} \text{SSIM}(I_{vis}, I_{fused}))\), where weights \(\omega\) adapt according to mean gradient magnitudes. The distillation losses comprise MSE for the visible quality prompt \(\mathcal{L}_{vis}^{distill} = \|P'_{vis} - T_{vis}\|_2^2\) and binary cross-entropy for the infrared saliency prompt \(\mathcal{L}_{ir}^{distill} = - (T_{ir} \odot \log P_{ir}^{map} + (1 - T_{ir}) \odot \log(1 - P_{ir}^{map}))\). The hyperparameters are set to \(\lambda_1=8, \lambda_2=25, \lambda_3=20, \lambda_4=0.5, \lambda_5=5\). The model is trained on 4 NVIDIA RTX 3090 GPUs using Adam for 100 epochs with an initial learning rate of \(1 \times 10^{-4}\) decayed at specified intervals.
Key Experimental Results¶
Main Results¶
P²Fusion is evaluated against 11 SOTA methods on MSRS, M3FD, FMB, and RoadScene. Performance is measured using Fusion Quality Index Evaluation (FQIE), Visual Information Fidelity (VIF), gradient-based feature transfer (\(Q^{AB/F}\)), and Mutual Information (MI).
Table 1: Quantitative image fusion results on MSRS, M3FD, FMB, and RoadScene benchmarks
| Dataset | Method | FQIE ↑ | VIF ↑ | Qabf ↑ | MI ↑ |
|---|---|---|---|---|---|
| MSRS | DiTFuse (TPAMI'25) | 0.642 | 0.281 | 0.363 | 2.146 |
| FreqGAN (TCSVT'25) | 0.679 | 0.285 | 0.492 | 2.871 | |
| Freefusion (TPAMI'25) | 0.590 | 0.281 | 0.430 | 1.984 | |
| LutFuse (ICCV'25) | 0.783 | 0.424 | 0.610 | 3.625 | |
| SAGE (CVPR'25) | 0.763 | 0.335 | 0.535 | 3.217 | |
| Ours | 0.850 | 0.445 | 0.682 | 3.905 | |
| M3FD | DiTFuse (TPAMI'25) | 0.389 | 0.188 | 0.274 | 2.514 |
| Freefusion (TPAMI'25) | 0.621 | 0.473 | 0.539 | 2.712 | |
| LutFuse (ICCV'25) | 0.466 | 0.319 | 0.480 | 3.987 | |
| SAGE (CVPR'25) | 0.659 | 0.363 | 0.575 | 3.115 | |
| Ours | 0.743 | 0.429 | 0.635 | 3.819 | |
| FMB | DiTFuse (TPAMI'25) | 0.536 | 0.215 | 0.363 | 2.620 |
| Freefusion (TPAMI'25) | 0.696 | 0.501 | 0.587 | 3.004 | |
| LutFuse (ICCV'25) | 0.575 | 0.319 | 0.527 | 3.897 | |
| SAGE (CVPR'25) | 0.763 | 0.384 | 0.634 | 3.429 | |
| Ours | 0.826 | 0.448 | 0.677 | 4.030 | |
| RoadScene | DiTFuse (TPAMI'25) | 0.445 | 0.320 | 0.426 | 2.938 |
| DDFM (ICCV'23) | 0.485 | 0.372 | 0.471 | 2.940 | |
| LutFuse (ICCV'25) | 0.328 | 0.265 | 0.381 | 3.644 | |
| SAGE (CVPR'25) | 0.379 | 0.237 | 0.334 | 2.919 | |
| Ours | 0.617 | 0.388 | 0.557 | 3.021 |
To verify downstream task enhancement, YOLOv7s is evaluated for object detection and SegFormer-B5 for semantic segmentation, as summarized in Table 2.
Table 2: Downstream evaluation: Object detection on M3FD and semantic segmentation on FMB
| Task / Dataset | Method | Key Class Metrics (Person / Car / Truck) | Headline Performance |
|---|---|---|---|
| M3FD Detection | DiTFuse (TPAMI'25) | Person: 0.514, Car: 0.684, Truck: 0.663 | mAP: 60.46% |
| (YOLOv7s) | FreqGAN (TCSVT'25) | Person: 0.544, Car: 0.700, Truck: 0.673 | mAP: 61.85% |
| Freefusion (TPAMI'25) | Person: 0.527, Car: 0.702, Truck: 0.674 | mAP: 61.87% | |
| SAGE (CVPR'25) | Person: 0.531, Car: 0.707, Truck: 0.658 | mAP: 60.86% | |
| Ours | Person: 0.534, Car: 0.704, Truck: 0.690 | mAP: 62.37% (+0.5%) | |
| FMB Segmentation | DiTFuse (TPAMI'25) | T.Sign: 73.15, Car: 81.76, Pole: 43.21 | mPA: 63.340%, mIoU: 56.808% |
| (SegFormer-B5) | FreqGAN (TCSVT'25) | T.Sign: 73.56, Car: 82.23, Pole: 44.73 | mPA: 64.471%, mIoU: 57.284% |
| Freefusion (TPAMI'25) | T.Sign: 71.75, Car: 80.16, Pole: 41.97 | mPA: 63.469%, mIoU: 56.421% | |
| SAGE (CVPR'25) | T.Sign: 72.39, Car: 81.39, Pole: 41.86 | mPA: 63.798%, mIoU: 56.479% | |
| Ours | T.Sign: 74.15, Car: 82.27, Pole: 43.91 | mPA: 64.672%, mIoU: 57.456% |
Ablation Study & Out-of-Distribution (OOD) Verification¶
To analyze component efficacy and cross-domain resilience, rigorous out-of-distribution verification was conducted on the DroneVehicle aerial dataset without domain fine-tuning.
Table 3: Out-of-distribution (OOD) verification and ablation analysis on DroneVehicle
| Config / Method | Fusion Metrics (FQIE / VIF / Qabf) | Detection mAP50→95 (Car / Truck / Bus / All) | Diagnostic Findings |
|---|---|---|---|
| Full model (Ours) | FQIE: 0.6282, VIF: 0.2655, Qabf: 0.5269 | Car: 0.632, Truck: 0.576, All: 0.604 | Optimal visual fidelity and top downstream detection |
| w/o Visible Prompt (w/o Pvis) | Severe drop in FQIE, largest drop in VIF | Gating collapses to near-uniform spatial distribution | Spatial quality prior is critical for textural sharpness |
| w/o Infrared Prompt (w/o Pir) | Sharp drop in Qabf, weak thermal contrast | Infrared response blurred, target recall drops | Thermal saliency prior anchors foreground targets |
| w/o Dual Distillation (w/o both) | Largest overall degradation (>15% across metrics) | Modality competition unmediated, causing confusion | Validates the indispensability of dual intrinsic priors |
| TIM (TPAMI'24) Baseline | FQIE: 0.2728, VIF: 0.2211, Qabf: 0.3182 | Car: 0.622, Truck: 0.567, All: 0.595 | Significant performance drop under aerial shifts |
| FreqGAN (TCSVT'25) Baseline | FQIE: 0.5340, VIF: 0.2775, Qabf: 0.4643 | Car: 0.601, Truck: 0.490, All: 0.540 | Frequency modeling produces artifacts in complex lighting |
Key Findings¶
- Synergy of Dual Intrinsic Priors: Removing \(P_{vis}\) collapses the spatial gating allocation into an undifferentiated map and sharply degrades VIF; removing \(P_{ir}\) attenuates target boundaries and increases downstream false negatives. Only bimodal co-conditioning balances foreground targets and background textures.
- Specialist Orthogonality in GDER: The modality expert and attention expert exhibit low feature correlation (\(r = 0.46\)). In overexposed or smoke-degraded conditions, the attention expert captures fine contours while the modality expert isolates salient thermal shapes, demonstrating effective self-correcting arbitration.
- Zero-Shot Aerial Generalization: On the unseen DroneVehicle benchmark with severe bird's-eye perspective shifts and tiny objects, P²Fusion achieves an mAP50-95 of 60.4% in downstream detection, surpassing the runner-up by +0.9%.
Highlights & Insights¶
- From Hard Penalties to Soft Prompt Distillation: Transforming static saliency masks and quality scores into dynamic prompt vectors preserves adaptive scene conditioning while liberating pixel-level generative optimization from rigid gradient interference.
- Decoupled Expert Recalibration over Capacity Scaling: Repurposing MoE from simple parameter scaling into functional feature decoupling (prior-responsive vs. prior-agnostic) enables the network to arbitrate between prior cues and global textures dynamically.
- Prior-Agnostic Extensibility: The Teach-to-Fuse architecture treats guidance cues as dynamic prompt regulators, providing an extensible interface that can seamlessly accommodate alternative cues such as depth, surface normals, or optical flow.
Limitations & Future Work¶
- Author-Acknowledged Limitations: The infrared saliency teacher (IST) relies on pre-trained segmentation weights for foreground binarization; severe out-of-domain failure of the teacher may inject initial noise into the dynamic prompt.
- Experimental & Scope Limitations: The visible quality prior uses global image-level BRISQUE scores, which lack dense local spatial granularity to capture localized motion blurs or localized lens glare.
- Future Directions: Developing patch-level self-supervised quality assessors and enabling bidirectional feedback where fusion features iteratively refine teacher representations during joint training.
Related Work & Insights¶
- vs STDFusionNet / PIAFusion: Prior works rely on hard-coded saliency masks or illumination-aware loss penalties that suppress fine details; P²Fusion modulates features via soft continuous prompts, avoiding optimization conflicts.
- vs SAGE (CVPR 2025) / CMFS (IJCAI 2025): Foundation-model-based approaches suffer from high-level semantic granularity mismatch with pixel reconstruction; P²Fusion utilizes modality-intrinsic physical priors for precise low-level texture recovery.
- vs TarDal (CVPR 2022) / SuperFusion (IJAS 2022): Downstream co-training introduces gradient competition between detection accuracy and visual aesthetic quality; P²Fusion maintains an unsupervised generative formulation that naturally enhances downstream perception without multi-objective conflict.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Elegant shift from static prior constraints to dynamic prompt distillation with decoupled expert self-correction.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarking across 5 datasets, 12 SOTA baselines, downstream detection/segmentation, and zero-shot aerial OOD testing.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear motivation, well-articulated pipeline, and cohesive experimental validation.
- Value: ⭐⭐⭐⭐⭐ Establishes an effective paradigm for physical-prior-guided multimodal fusion and robust perception.