TurboMPLE: Joint Infrared Turbulence Mitigation and Physical Fields Estimation via Mutual Progressive Layered Extraction¶
Conference: ECCV 2026
Paper: ECCV 2026 Official
Code: https://github.com/Ayt777/TurboMPLE
Area: Optimization & Theory
Keywords: atmospheric turbulence, turbulence mitigation, physical fields estimation, multi-task learning, progressive layered extraction
TL;DR¶
To tackle optical distortions, blur, and grayscale drift induced by atmospheric turbulence in thermal infrared imaging, TurboMPLE establishes an end-to-end multi-task learning framework that achieves bidirectional feature-level mutual guidance and physical law constraints, jointly delivering state-of-the-art video turbulence mitigation and accurate 2D dynamic physical fields estimation.
Background & Motivation¶
In long-wave infrared (8–13 \(\mu\text{m}\)) imaging systems, atmospheric temperature gradients near the ground induce violent fluctuations in refractive index across space and time. These fluctuations inflict severe spatio-temporal blur and geometric distortions on the focal plane, while atmospheric absorption, scattering, and thermal dissipation produce pronounced grayscale drift. Conventional turbulence mitigation algorithms primarily rely on frame registration and "lucky-region" temporal fusion, which easily break down under dynamic foregrounds or strong long-path turbulence. Recently, deep learning models have achieved remarkable restoration gains, yet almost all of them treat the problem as a purely data-driven image restoration task, completely decoupling image degradation from the underlying atmospheric fluid physics and leading to fragile generalization in unconstrained real-world environments.
Although some attempts have sought to inject physical concepts—such as learning latent phase distortion vectors via variational autoencoders—the resulting latent codes lack physical units and interpretability. Another pioneering framework, PBCL (Nature Computational Science 2023), proposed a physically boosted cooperative learning paradigm; however, it adopts a multi-stage pipeline: estimating static turbulence strength first, restoring images conditioned on those estimates, and finally refining the fields using the restored images. This sequential design suffers from severe error accumulation, prevents end-to-end joint optimization, and assumes turbulence fields are static, fundamentally ignoring the essential temporal dynamics in video imaging.
This paper's angle of attack is that atmospheric turbulence degradation is inherently the physical interaction between optical wave propagation and turbulent fluid fields; degraded frames retain physical projection clues, while accurate physical fields provide strong structural priors for restoration. The core idea is to unify infrared turbulence mitigation and physical fields estimation within an end-to-end multi-task learning architecture, introducing mutual progressive layered extraction blocks and turbulent physics constraints to establish deep feature-level synergy between physical priors and visual restoration.
Method¶
Overall Architecture¶
TurboMPLE takes a sequence of \(T\) degraded infrared frames \(I \in \mathbb{R}^{T \times 1 \times H \times W}\) as input, and simultaneously predicts the restored clean video sequence \(I_{\text{out}} \in \mathbb{R}^{T \times 1 \times H \times W}\) and four dynamic 2D atmospheric physical parameter fields \(P_{\text{out}} \in \mathbb{R}^{T \times 4 \times \frac{H}{16} \times \frac{W}{16}}\). These physical fields comprise the refractive index structure constant \(C_n^2\), temperature structure constant \(C_T^2\), turbulent kinetic energy dissipation rate \(\varepsilon\), and Reynolds stress \(\text{RS}\).
The overall network operates through three interconnected stages: first, the input video passes through an encoder embedded with the Bouguer–Lambert–Beer law (BLB), compensating for thermal radiation attenuation and extracting spatio-temporal features via temporal attention; next, these spatio-temporal representations are fed into \(N\) stacked Mutual Progressive Layered Extraction (MutualPLE) blocks, where shared experts and task-specific experts collaborate under the guidance of cross-task mutual attention and dynamic gating; finally, task-specific features are decoded by dedicated mitigation and estimation heads, supervised under joint physical-consistency and visual restoration loss functions.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input Sequence<br/>T×1×H×W"] --> Enc["BLB-Embedded Encoder<br/>Radiation Compensation + Spatio-temporal Features"]
Enc --> PLE["MutualPLE Extraction Blocks<br/>Dual-branch Shared + Task-specific Experts"]
PLE --> CTMB["Cross-Task Mutual Guidance<br/>Bidirectional Cross-Attention + Dynamic Gating"]
CTMB --> Heads["Dual Decoding Prediction Heads<br/>Restoration Head + Field Estimation Head"]
Heads --> OutMit["Restored Image Sequence<br/>T×1×H×W"]
Heads --> OutPhys["Physical Fields Output<br/>Cn2, CT2, ε, RS"]
Key Designs¶
1. BLB-Embedded Encoder: Alleviating Thermal Grayscale Drift and Modeling Spatio-temporal Dynamics
To counter the grayscale drift and contrast degradation caused by atmospheric absorption and scattering in infrared propagation, the encoder explicitly embeds the Bouguer–Lambert–Beer (BLB) law in its initial layer. This mechanism applies radiation attenuation compensation directly to the degraded image \(I\): $\(I_1 = I \times \exp\left((\mu_a + \mu_s) \times d\right)\)$ where \(\mu_a\) and \(\mu_s\) denote the atmospheric absorption and scattering coefficients, and \(d\) represents the effective transmission path length. All three parameters are learnable, constrained via Sigmoid functions to physically meaningful ranges of \([0, 1]\) and \([0, 10]\). Subsequently, the compensated sequence \(I_1\) is processed by stacked convolutions and a Temporal Attention Unit (TAU), downsampling fourfold to yield rich spatio-temporal representations \(F_{\text{st}} \in \mathbb{R}^{T \times C \times \frac{H}{4} \times \frac{W}{4}}\).
2. Shared and Task-Specific Expert Decoupling: Combining State Space Models with Turbulence Modulation
To avoid gradient conflict and negative transfer typical in hard parameter sharing, MutualPLE decouples representation extraction into shared and task-specific experts at each layer. Given that turbulence exhibits both localized spatial distortions and long-range temporal motion, the shared expert incorporates a dual-branch design: a convolutional branch employs a lightweight estimator to predict a spatial turbulence map and generate an adaptive weight mask \(W\), modulating features via element-wise multiplication \(F_{\text{conv}} = F_{\text{st}} \odot W\) to emphasize turbulence-sensitive regions while suppressing background redundancy; concurrently, a state-space model (SSM) branch utilizes 2D-Selective-Scan (SS2D) to model long-range temporal dynamics with linear complexity. The two branches are merged via channel concatenation and channel-shuffle operations into \(F_s^i\), while mitigation and estimation experts employ dual-branch convolutions to capture global-local task-specific representations \(F_m^i\) and \(F_e^i\).
3. Cross-Task Mutual Guidance: Bidirectional Cross-Attention and Progressive Gating Fusion
To eliminate information bottlenecks between the two branches, the Cross-Task Mutual Block (CTMB) establishes a bidirectional interaction highway. In block \(i\), two complementary cross-attention mechanisms exchange representations: the mitigation expert queries the estimation representations to inject physics-grounded degradation priors, while the estimation expert queries high-resolution visual representations to enrich turbulence intensity sensing: $\(\tilde{F}_m^i = \text{CrossAttn}_{m \leftarrow e}(F_m^i, F_e^i), \quad \tilde{F}_e^i = \text{CrossAttn}_{e \leftarrow m}(F_e^i, F_m^i)\)$ To avoid premature noise injection during early representation stages, a depth-dependent progressive mutual weight \(\omega_i = 0.2 + i \times 0.4\) dynamically balances cross-task interaction against self-representations: $\(F_m^{i\prime} = F_m^i \odot (1 - W_m^i) + \tilde{F}_m^i \odot W_m^i, \quad W_m^i = \omega_i \cdot \Phi_m([F_m^i, \tilde{F}_m^i])\)$ Task-specific and shared representations are then adaptively aggregated through global pooling, linear projection, and Softmax gating functions \(\mathcal{G}\), achieving seamless collaboration while preserving task-specific purity.
Loss & Training¶
TurboMPLE utilizes a multi-objective end-to-end training strategy optimizing visual fidelity alongside turbulent fluid consistency.
The turbulence mitigation objective is formulated as: $\(\mathcal{L}_m = \mathcal{L}_{\text{mit}} + \beta_1 \mathcal{L}_{\text{re-deg}}\)$ where \(\mathcal{L}_{\text{mit}} = \|I_{\text{out}} - I_{\text{GT}}\|_1\), and \(\mathcal{L}_{\text{re-deg}}\) is a re-degradation loss applying a differentiable forward simulator (sequentially inducing grayscale drift, Gaussian blur, and tilt) on \(I_{\text{out}}\) to match the degraded input \(I\). The coefficient \(\beta_1\) is set to 0 initially and switched to 1 in final epochs.
The physical fields estimation objective is defined as: $\(\mathcal{L}_e = \mathcal{L}_{\text{est}} + \beta_2 \mathcal{L}_{\text{phys}}\)$ where \(\mathcal{L}_{\text{est}} = \|P_{\text{out}} - P_{\text{GT}}\|_1\) with \(\beta_2 = 0.5\). The physics regularization term \(\mathcal{L}_{\text{phys}}\) directly incorporates Kolmogorov atmospheric optical physics governing \(C_n^2\), \(C_T^2\), \(\varepsilon\), and \(\text{RS}\): $\(\mathcal{L}_{\text{phys}} = \mathbb{E}\left[ \left\| \hat{C}_n^2 - \left(79 \times \frac{P}{\hat{T}^2}\right)^2 \hat{C}_T^2 \right\|_1 + \left\| \hat{\varepsilon} - \gamma (\hat{C}_n^2)^{3/2} \left(\frac{T}{P}\right) \right\|_1 + \left\| \widehat{\text{RS}} - \frac{(\hat{\varepsilon} L)^{2/3}}{L} \right\|_1 \right]\)$ This physics-informed regularization enforces strict inter-variable proportionality, guaranteeing physically faithful field estimation even on unseen real-world sequences.
Key Experimental Results¶
Main Results¶
Evaluations are conducted on the newly constructed large-scale infrared turbulence benchmark (3,357 sequences, 100,710 frames with paired \(16 \times 16\) 4-channel physical fields). All algorithms are evaluated under an identical protocol using 30-frame sequences at \(256 \times 256\) resolution.
Table 1: Quantitative evaluation of turbulence mitigation (PSNR (dB) / SSIM)
| Category | Method | Weak Turbulence | Medium Turbulence | Strong Turbulence | Overall | GMACs |
|---|---|---|---|---|---|---|
| Video Baseline | BasicVSR (CVPR'21) | 29.2718 / 0.8695 | 25.8040 / 0.7474 | 24.1317 / 0.6714 | 26.4142 / 0.7632 | 8355.0 |
| Video Deblurring | ESTRNN (IJCV'23) | 25.1907 / 0.7898 | 24.1390 / 0.7178 | 23.1995 / 0.6609 | 24.1810 / 0.7231 | 1450.2 |
| Single-frame Restoration | ConvIR (TPAMI'24) | 26.9256 / 0.8007 | 23.9879 / 0.6641 | 22.8554 / 0.6016 | 24.5989 / 0.6893 | 120.4 |
| Single-frame Restoration | MaIR (CVPR'25) | 29.3921 / 0.8705 | 25.5369 / 0.7443 | 24.0329 / 0.6675 | 26.3329 / 0.7612 | 134.8 |
| Turbulence Restoration | TMT (TCI'24) | 29.3165 / 0.8697 | 26.0301 / 0.7498 | 24.7026 / 0.6825 | 26.6936 / 0.7678 | 496.0 |
| Turbulence Restoration | DATUM (CVPR'24) | 29.1793 / 0.8683 | 26.2379 / 0.7613 | 24.9530 / 0.6964 | 26.7997 / 0.7757 | 1582.0 |
| Physics Cooperative | PBCL (Nat. Comput. Sci.'23) | 28.3768 / 0.8502 | 26.2520 / 0.7693 | 25.0672 / 0.7265 | 26.5729 / 0.7823 | 642.0 |
| State Space Model | MambaTM (CVPR'25) | 26.8179 / 0.7848 | 24.9059 / 0.6923 | 23.9590 / 0.6445 | 25.2341 / 0.7075 | 312.0 |
| Joint Multi-Task | TurboMPLE (Ours) | 29.8887 / 0.8790 | 27.1542 / 0.7918 | 25.6436 / 0.7274 | 27.5718 / 0.7998 | 145.1 |
Table 2: Physical fields estimation performance (spatial and temporal metrics)
| Method | Spatial \(C_n^2\) (MAE / \(R^2\)) | Spatial \(C_T^2\) (MAE / \(R^2\)) | Spatial \(\varepsilon\) (MAE / \(R^2\)) | Spatial \(\text{RS}\) (MAE / \(R^2\)) | Temporal \(C_n^2\) (\(R^2\)) | Temporal \(C_T^2\) (\(R^2\)) | Temporal \(\varepsilon\) (\(R^2\)) | Temporal \(\text{RS}\) (\(R^2\)) |
|---|---|---|---|---|---|---|---|---|
| UNet | 0.076 / 0.604 | 0.071 / 0.564 | 0.053 / 0.523 | 0.074 / 0.586 | 0.893 | 0.890 | 0.870 | 0.891 |
| GRCNN | 0.060 / 0.771 | 0.054 / 0.763 | 0.039 / 0.712 | 0.057 / 0.767 | 0.888 | 0.883 | 0.853 | 0.886 |
| UwU-Net | 0.118 / 0.124 | 0.105 / 0.109 | 0.072 / <0.1 | 0.112 / 0.117 | 0.386 | 0.387 | 0.368 | 0.387 |
| FoVFormer | 0.036 / 0.891 | 0.034 / 0.882 | 0.024 / 0.860 | 0.035 / 0.888 | 0.948 | 0.944 | 0.929 | 0.947 |
| PBCL | 0.048 / 0.754 | 0.044 / 0.734 | 0.037 / 0.621 | 0.046 / 0.745 | 0.932 | 0.927 | 0.909 | 0.930 |
| TurboMPLE | 0.024 / 0.943 | 0.023 / 0.935 | 0.017 / 0.919 | 0.023 / 0.940 | 0.973 / 0.997 | 0.972 / 0.997 | 0.963 / 0.995 | 0.973 / 0.997 |
Ablation Study¶
Table 3: Ablation study of key components in TurboMPLE
| Configuration | Components | PSNR (dB) | SSIM | Spatial \(R^2_s\) | Temporal \(R^2_t\) |
|---|---|---|---|---|---|
| (1) Baseline | Residual Dense Blocks replacing Dual-branch Experts; no CTMB / BLB / \(\mathcal{L}_{\text{phys}}\) | 26.470 | 0.781 | 0.890 | 0.967 |
| (2) + Dual-branch Experts | (1) + SS2D temporal dynamics + turbulence-aware modulation branch | 26.726 | 0.786 | 0.900 | 0.974 |
| (3) + CTMB | (2) + Cross-Task Mutual Block (bidirectional cross-attention) | 27.113 | 0.787 | 0.909 | 0.983 |
| (4) + BLB embedding | (3) + Bouguer–Lambert–Beer law embedding in encoder | 27.452 | 0.797 | 0.916 | 0.972 |
| (5) Full model | (4) + Turbulent physics consistency loss \(\mathcal{L}_{\text{phys}}\) | 27.572 | 0.800 | 0.934 | 0.985 |
Table 4: Influence of MutualPLE block number \(N\)
| Number of Blocks \(N\) | PSNR (dB) | SSIM | Spatial \(R^2_s\) | Temporal \(R^2_t\) | GMACs | Params (M) |
|---|---|---|---|---|---|---|
| \(N = 1\) | 26.482 | 0.783 | 0.896 | 0.966 | 96.3 | 0.645 |
| \(N = 2\) (Default) | 27.572 | 0.800 | 0.934 | 0.985 | 145.1 | 0.942 |
| \(N = 3\) | 27.477 | 0.795 | 0.936 | 0.981 | 193.8 | 1.239 |
Key Findings¶
- Bidirectional Mutual Guidance Unlocks Mutual Gains: Incorporating the CTMB module produces a notable jump in restoration PSNR from 26.726 dB to 27.113 dB, while simultaneously boosting field spatial fit \(R^2_s\) by 0.009. This confirms that physical priors effectively guide geometric de-warping, while fine visual details ground turbulence strength estimation.
- Physics-Informed Embedding Grounds Generalization: Coupling BLB radiation compensation and \(\mathcal{L}_{\text{phys}}\) brings an additional +0.459 dB in PSNR. In real-world thermal infrared experiments, predicted physical fields align remarkably well with ground measurements from four-probe thermocouple thermometers (TFM), attaining Pearson correlation coefficients of \(r = 0.944\) (\(C_n^2\)) and \(r = 0.968\) (\(C_T^2\)).
- Lightweight Architecture with Exceptional Efficiency: Requiring merely 0.942M parameters and 145.1 GMACs for a 30-frame sequence, TurboMPLE is over \(50\times\) more computationally efficient than BasicVSR (8355 GMACs) and over \(10\times\) faster than DATUM (1582 GMACs), making it highly suitable for resource-constrained edge thermal platforms.
Highlights & Insights¶
- Progressive Mutual Multi-Task Paradigm: Overcomes the traditional dichotomy of treating turbulence mitigation and physical estimation as isolated or cascaded problems, leveraging progressive weights \(\omega_i\) to coordinate cross-attention representations seamlessly.
- Unified Hard and Soft Physics Integration: Directly integrates the Bouguer–Lambert–Beer optical transmission law into the front-end encoder and enforces Kolmogorov turbulence relations in the loss function, anchoring deep learning representations within classical atmospheric physics.
- Comprehensive Benchmark Construction: Introduces a benchmark containing 80,580 real-world thermal video frames across diverse topographies (Beijing, Hubei, Chongqing) and 100,710 synthetic frames with complete spatio-temporal physical field ground truth.
Limitations & Future Work¶
- Spatial Resolution of Physical Fields: Due to fluid simulation constraints and computational budgets, the estimated physical fields are resolved at \(\frac{1}{16} \times \frac{1}{16}\) of the image resolution (\(16 \times 16\) grid), rather than dense per-pixel predictions.
- Abrupt Non-Stationary Turbulence: Under violent local turbulent events (e.g., rocket plumes or localized high-heat combustion), the assumption of slowly varying ambient pressure \(P\) and characteristic scale \(L\) may necessitate dynamic spatial adaptation.
- Future Directions: Extending the framework to multi-spectral (visible, SWIR, LWIR) collaborative turbulence sensing, and exploring diffusion-based models for dense per-pixel physical field super-resolution.
Related Work & Insights¶
- vs PBCL (Nature Computational Science 2023): PBCL uses a multi-stage disconnected architecture, assumes static turbulence fields, and suffers from error accumulation. TurboMPLE achieves true end-to-end multi-task learning, captures temporal dynamics via SS2D, and outperforms PBCL by 1.00 dB in PSNR with far lower complexity.
- vs MambaTM (CVPR 2025): MambaTM models latent phase distortion via VAE, resulting in uninterpretable latent codes. TurboMPLE predicts four verifiable, dimensionally grounded physical fields with clear fluid dynamic validity.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First joint, end-to-end mutual progressive multi-task learning framework for infrared turbulence mitigation and physical fields estimation.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous benchmark encompassing over 180,000 frames across synthetic and real-world thermal datasets, validated against hardware thermocouple probes.
- Writing Quality: ⭐⭐⭐⭐⭐ Cohesive theoretical derivation, clean mathematical formulation, and well-structured architectural descriptions.
- Value: ⭐⭐⭐⭐⭐ Bridges the gap between atmospheric optics, fluid physics, and computer vision, offering substantial utility for remote sensing and infrared surveillance.