Skip to content

SupIR-GS: Thermal Infrared Super-Resolution Novel View Synthesis with Imaging-Calibrated 3D Gaussian Splatting

Conference: ECCV 2026
Paper: CVF Open Access
Area: 3D Vision
Keywords: thermal infrared imaging, 3D Gaussian Splatting, super-resolution, novel view synthesis, imaging calibration

TL;DR

Addressing the core challenge of "signal-to-noise ratio inversion" in thermal infrared imaging—where thermal diffusion smooths true boundaries while sensor non-uniformity introduces structured noise—SupIR-GS presents the first imaging-calibrated super-resolution 3D Gaussian Splatting framework, decoupling hardware artifacts via a differentiable physical calibration layer, boosting flat thermal gradients via a local tone adapter, and progressively recovering fine geometric boundaries via frequency-aware curriculum learning.

Background & Motivation

Thermal infrared (TIR) imaging measures long-wave infrared radiation to characterize scene thermal signatures, offering robust all-weather perception under low light, smoke, and degraded visibility across autonomous driving, disaster response, and industrial inspection. However, constrained by detector manufacturing costs and physical diffraction limits, uncooled microbolometer thermal cameras typically yield low-resolution (LR) observations. When modern continuous 3D representations like neural radiance fields (NeRF) or 3D Gaussian Splatting (3DGS) are applied directly to multi-view LR thermal images, they inevitably encounter a severe geometry-radiance ambiguity.

The physical origin of this failure lies in a pronounced "signal-to-noise ratio inversion" intrinsic to thermal infrared sensors. On one hand, physical heat conduction, thermal diffusion, and atmospheric scattering attenuate true geometric boundaries into low-contrast, low-frequency transitions. On the other hand, sensor imperfections such as non-uniform response (NUR) and fixed pattern noise (FPN) manifest as sharp, structured high-frequency textures on the sensor plane. Standard 3DGS optimized under photometric supervision is easily misled by these spurious signals: it overfits hardware noise by spawning dense high-frequency floaters in 3D space, while collapsing in texture-scarce regions due to gradient starvation. Off-the-shelf 3D visible super-resolution pipelines (e.g., SRGS) or 2D image restoration models fail because they mistake sensor noise for real geometry and lack cross-view consistency.

The core idea is to natively render continuous 3D Gaussians at the target high resolution, explicitly decouple sensor-induced degradation from latent thermal radiance through a differentiable forward-imaging calibration chain and bounded local tone modulation, and steer geometric optimization from coarse low-frequency scaffolds to sharp high-frequency thermal boundaries using frequency-aware curriculum learning.

Method

Overall Architecture

SupIR-GS establishes an end-to-end framework integrating forward physical degradation modeling with inverse optimization of continuous 3D radiance fields. Given multi-view low-resolution infrared sequences \(\{I_v^{lr}\}\), the model directly rasterizes continuous 3D Gaussian primitives on a high-density target sampling grid \(\Omega^{hr}\) (\(r^2\) times more query points, where \(r\) is the super-resolution scale), synthesizing an ideal, degradation-free latent high-resolution radiance image \(I_{base}\). This clean rendering then passes sequentially through the Physically-constrained Infrared Calibration module (PhysIR-Cal) and the lightweight Intensity-conditioned Tone Adapter (IRTone-Adapter) to simulate sensor-specific NUR, additive bias, lens radial vignetting, and local dynamic range compression, producing a calibrated prediction \(I_{pred}\). Supervised loss backpropagates through this forward chain to update 3D Gaussian positions, scales, rotations, and opacities in the clean radiance domain. Meanwhile, a Frequency-Aware Curriculum Learning strategy incorporates structural priors from a pretrained super-resolution teacher, leveraging confidence masks and a cosine ramp-up schedule to smoothly transition optimization from low-frequency scaffolds to fine thermal details.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multi-view Low-Resolution TIR Inputs<br/>Multi-view sequences and high-density HR sampling grid"] --> B["Resolution-Agnostic 3DGS Radiance Rendering<br/>Continuous 3D Gaussians rasterized into latent clean HR radiance"]
    B --> C["Physical-Constrained Infrared Calibration<br/>Parametric gain-bias grid and radial vignetting decouple NUR and FPN"]
    C --> D["Intensity-conditioned Tone Adapter<br/>Lightweight MLP predicts bounded local affine residuals to enhance contrast"]
    D --> E["Frequency-Aware Curriculum Learning Strategy<br/>Teacher-guided confidence-masked high-frequency loss with cosine ramp-up"]
    E --> F["High-Fidelity Thermal Super-Resolution NVS<br/>Sharp geometric boundaries and accurate thermodynamic fidelity"]

Key Designs

1. Physical-Constrained Infrared Calibration: Decoupling Sensor Non-Uniformity and Fixed Pattern Noise

Uncooled microbolometer arrays suffer from non-uniform response (NUR) and fixed pattern striping noise (FPN). Directly fitting pixel-wise noise severely overfits low-SNR observations, pulling 3D Gaussians away from actual physical surfaces. PhysIR-Cal introduces a parameter-efficient, differentiable forward calibration layer to explicitly absorb these systematic hardware aberrations. The module defines two low-resolution learnable parameter grids \(G_{grid}, B_{grid} \in \mathbb{R}^{16 \times 16}\), which are bilinearly upsampled to target dimensions and bounded by \(\tanh(\cdot)\) activations to parameterize multiplicative spatial gain \(G\) and additive spatial bias \(B\). To model optical attenuation from infrared lenses, a learnable 4th-order radial vignetting function is formulated as \(V(r) = \exp(a_2 r^2 + a_4 r^4)\), where \(r\) denotes the normalized radial distance. For a base 3DGS rendering \(I_{base}\), the calibrated intensity is computed as:

\[I_{calib} = \big[I_{base} \cdot (1 + G) + B\big] \cdot V\]

To accommodate automatic gain control (AGC) or per-frame histogram stretching common in thermal sensors, an optional monotonic tone mapping in the logit domain absorbs inter-frame gain fluctuations: \(I_{phy} = \sigma\big(s_k \cdot \text{logit}(\tilde{I}_{calib}) + t_k\big)\), where \(s_k, t_k\) are per-view parameters and \(\tilde{I}_{calib}\) is numerically clamped. By isolating hardware noise inside this low-frequency parametric layer, PhysIR-Cal shields 3D Gaussian geometry and densification gradients from spurious high-frequency artifacts.

2. Intensity-conditioned Tone Adapter: Surfacing Latent Temperature Gradients via Local Affine Residuals

Thermal infrared images frequently suffer from narrow dynamic range and flattened intensity histograms, leaving cold background and low-contrast regions texture-scarce. In these regions, standard photometric reprojection errors generate negligible and noisy gradients that fail to guide sub-pixel Gaussian refinement. IRTone-Adapter employs a lightweight two-layer MLP (64 hidden channels) conditioned on local rendered intensity \(I_{in}(p)\) to predict adaptive multiplicative gain \(\delta_m\) and additive bias \(\delta_o\) residuals per pixel:

\[[\delta_m, \delta_o] = \Phi_{tone}\big(\text{mean}(I_{in}(p)); \theta_{tone}\big)\]
\[I_{pred}(p) = I_{in}(p) \cdot \big(1 + \lambda_m \tanh(\delta_m)\big) + \lambda_o \tanh(\delta_o)\]

To prevent the network from masking true 3D geometric inaccuracies through excessive intensity remapping, the scaling bounds are strictly restricted to \(\lambda_m = 0.08\) and \(\lambda_o = 0.04\), bounding perturbations within \(\pm 8\%\) and \(\pm 4\%\). This conservative, bounded affine modulation gently amplifies subtle temperature variations in flat regions, generating informative and consistent backpropagated gradients for sharpening fine geometry.

3. Frequency-Aware Curriculum Learning Strategy: Confidence-Masked Annealing from Coarse to Fine

Because thermal diffusion intrinsically blurs physical boundaries, high-frequency cues in low-resolution inputs are inherently ambiguous. Naively injecting high-frequency details from 2D super-resolution networks risks propagating single-view hallucinations. SupIR-GS first constructs a spatial confidence mask \(W_{conf}\) by evaluating consistency between the pretrained SwinIR teacher output \(I_{teacher}\) and the bilinearly upsampled input \(I_{up}\) across both intensity and gradient domains:

\[W_{conf}(\mathbf{p}) = \exp\Big(-\gamma_p |I_{teacher}(\mathbf{p}) - I_{up}(\mathbf{p})| - \gamma_g \|\nabla I_{teacher}(\mathbf{p}) - \nabla I_{up}(\mathbf{p})\|_1\Big)\]

This exponential formulation guarantees that \(W_{conf}(\mathbf{p})\) approaches 1 only when the teacher prediction strictly agrees with low-frequency observation structures. The confidence-masked high-frequency residual loss is defined via an average-pooling high-pass filter \(H(\cdot)\):

\[\mathcal{L}_{hf} = \frac{1}{N} \sum_{\mathbf{p}} W_{conf}(\mathbf{p}) \, \big|H(I_{pred}(\mathbf{p})) - H(I_{teacher}(\mathbf{p}))\big|\]

During training iteration \(t\), the high-frequency loss weight \(\lambda_{hf}(t)\) is governed by a cosine ramp-up schedule: kept at 0 during early stages so that photometric and SSIM losses establish a stable low-frequency geometric scaffold, and then smoothly increased across iterations (65% to 95%) to inject sharp edge details safely without destabilizing early convergence.

Loss & Training

The entire framework is trained end-to-end for 30,000 iterations. The unified loss function integrates photometric error, structural similarity, confidence-weighted high-frequency residuals, and parameter regularization:

\[\mathcal{L}_{total} = \mathcal{L}_{pix} + \lambda_{ssim}\mathcal{L}_{ssim} + \lambda_{hf}(t)\mathcal{L}_{hf} + \mathcal{R}_{cal} + \mathcal{R}_{tone}\]

Here, \(\mathcal{L}_{pix}\) enforces \(\ell_1\) photometric consistency; \(\mathcal{L}_{ssim}\) preserves structural integrity; \(\mathcal{R}_{cal}\) and \(\mathcal{R}_{tone}\) are regularizers with weights \(\lambda_{cal} = 5 \times 10^{-4}\) and \(\lambda_{tone} = 10^{-3}\) to prevent degradation parameters from drifting. The SwinIR-S teacher operates purely offline prior to training to generate high-frequency targets and masks, incurring zero extra memory or compute overhead during training and novel-view rendering.

Key Experimental Results

Main Results

Quantitative evaluations are conducted on TI-NSD (20 scenes across Indoor, Outdoor, and UAV settings) and ThermoScenes (10 scenes featuring high-contrast indoor objects and building facades). All models take \(4\times\) bicubic-downsampled images as input.

Table 1: Quantitative comparisons on the TI-NSD dataset under 4× super-resolution

Method Indoor SSIM↑ Indoor PSNR↑ (dB) Indoor LPIPS↓ Outdoor SSIM↑ Outdoor PSNR↑ (dB) Outdoor LPIPS↓ UAV SSIM↑ UAV PSNR↑ (dB) UAV LPIPS↓
Bicubic + 3DGS 0.9222 31.16 0.3095 0.8466 27.47 0.3311 0.8399 29.26 0.3127
Lanczos + 3DGS 0.9237 31.04 0.3085 0.8375 26.94 0.3360 0.8444 29.39 0.3069
DifIISR + 3DGS 0.9278 30.45 0.3097 0.8453 26.77 0.3235 0.8449 28.96 0.2899
3DGS (4×) 0.9060 28.23 0.3147 0.7825 21.86 0.3491 0.7555 23.37 0.3562
Thermal3DGS (4×) 0.9154 30.15 0.3117 0.7852 23.36 0.3418 0.7287 23.40 0.3733
Mip-Splatting (4×) 0.9232 30.63 0.3020 0.8280 25.37 0.3304 0.8582 29.92 0.2786
Scaffold-GS (4×) 0.9145 28.70 0.3077 0.7998 22.05 0.3331 0.7405 21.64 0.3612
SRGS 0.9312 32.09 0.2765 0.8701 27.38 0.2790 0.9078 31.78 0.2014
Ours (SupIR-GS) 0.9455 34.81 0.2597 0.9066 29.94 0.2336 0.9428 34.24 0.1453

Table 2: Quantitative comparisons and temperature fidelity on ThermoScenes dataset (illustrating smoothness bias)

Method Indoor SSIM↑ Indoor PSNR↑ (dB) Indoor LPIPS↓ Indoor MAE↓ Indoor \(\text{MAE}_{roi}\)↓ Outdoor SSIM↑ Outdoor PSNR↑ (dB) Outdoor LPIPS↓ Outdoor MAE↓ Outdoor \(\text{MAE}_{roi}\)↓
Bicubic + 3DGS 0.9650 31.19 0.0415 0.3536 1.6318 0.9663 30.24 0.1577 0.6166 0.5411
Lanczos + 3DGS 0.9665 30.73 0.0414 0.3574 1.7223 0.9686 30.54 0.1555 0.7988 0.7441
DifIISR + 3DGS 0.9486 28.95 0.0530 0.7947 1.9993 0.9617 29.21 0.1636 1.1811 1.1606
3DGS (4×) 0.9612 29.23 0.0556 0.3581 2.0971 0.9238 25.71 0.2164 0.8346 0.7380
Thermal3DGS (4×) 0.9122 27.17 0.0504 0.8558 2.1585 0.9201 23.92 0.2205 1.8974 2.0263
Mip-Splatting (4×) 0.9662 30.59 0.0476 0.3685 1.8094 0.9585 29.83 0.1659 0.7149 0.6316
Scaffold-GS (4×) 0.9656 28.93 0.0545 0.3219 2.1495 0.9315 26.45 0.1998 0.6888 0.6311
SRGS 0.9537 29.84 0.0454 0.4048 2.1793 0.9679 30.76 0.1472 0.6619 0.5969
Ours (SupIR-GS) 0.9672 31.32 0.0438 0.3368 1.7903 0.9662 30.40 0.1546 0.7957 0.7329

Ablation Study

Table 3: Progressive ablation study on core modules (TI-NSD dataset)

Config PhysIR-Cal IRTone-Adapter Frequency-Aware Indoor SSIM↑ Indoor PSNR↑ (dB) Indoor LPIPS↓ Outdoor SSIM↑ Outdoor PSNR↑ (dB) Outdoor LPIPS↓ UAV SSIM↑ UAV PSNR↑ (dB) UAV LPIPS↓
(1) Baseline × × × 0.9312 32.086 0.2765 0.8701 27.381 0.2790 0.8888 30.291 0.2299
(2) Tone only × ✓ × 0.9400 31.700 0.2920 0.8655 27.202 0.2836 0.8988 30.860 0.2087
(3) Freq only × × ✓ 0.9354 32.281 0.2662 0.8727 27.263 0.2680 0.9121 31.550 0.1868
(4) Tone+Freq × ✓ ✓ 0.9401 32.936 0.2638 0.8925 28.514 0.2491 0.9231 32.203 0.1745
(5) Calib only ✓ × × 0.9420 33.692 0.2837 0.8795 28.449 0.2681 0.9118 32.218 0.1928
(6) Calib+Freq ✓ × ✓ 0.9453 34.276 0.2788 0.9031 29.800 0.2377 0.9305 33.142 0.1668
(7) Full Ours ✓ ✓ ✓ 0.9455 34.806 0.2597 0.9066 29.935 0.2336 0.9320 33.326 0.1653

Key Findings

  1. Physical Module Dependency and Sequence: Enabling IRTone-Adapter in isolation drops Indoor PSNR by 0.39 dB because uncalibrated contrast stretching amplifies sensor FPN and NUR into pseudo-geometry. Similarly, applying Frequency-Aware constraints without calibration degrades Outdoor PSNR by 0.12 dB due to gradient conflict between noise and structural edges. Only when PhysIR-Cal first strips sensor noise to form a reliable scaffold (+1.61 dB Indoor, +1.93 dB UAV) can tone adaptation and high-frequency priors unleash their full synergy, boosting PSNR by over 2.5 to 3.0 dB across subsets.
  2. Smoothness Bias in Thermal Evaluation: In large uniform thermal backgrounds (e.g., cold skies and building walls), naive 2D interpolations (Bicubic+3DGS) mathematically minimize global pixel variance, showing deceptively high PSNR while obliterating window frames and structural edges. SupIR-GS resists this smoothness bias, outperforming all baselines on perceptual metric LPIPS and region-of-interest temperature error (\(\text{MAE}_{roi}\)).
  3. High Efficiency and Low Memory: Optimized in 16.42 minutes on an RTX 3080, SupIR-GS achieves 381.10 FPS rendering speed with peak memory of only 2792 MB during novel-view synthesis, notably surpassing 3DGS (276.36 FPS / 3580 MB) and SRGS (313.75 FPS / 3484 MB) because all calibration and teacher networks are stripped at inference time.

Highlights & Insights

  • Physics-Informed Sensor Decoupling: Introduces a differentiable parameter-efficient grid to absorb multiplicative NUR and additive FPN, directly solving the "SNR inversion" bottleneck and preventing high-frequency floaters from corrupting 3D geometry.
  • Bounded Intensity-Conditioned Adaptation: Restricts local affine tone modifications to narrow bounds (\(\pm 8\%\) gain, \(\pm 4\%\) bias), surfacing subtle gradients in flat thermal scenes without permitting color remapping to mask geometric errors.
  • Joint Metric-Masked Frequency Annealing: Combines intensity and gradient deviation into a confidence mask along with cosine ramp-up scheduling, safely distilling teacher high-frequency edge priors while filtering out cross-view hallucinations.

Limitations & Future Work

  • Static Thermal Equilibrium Assumption: Assumes objects remain in steady thermodynamic equilibrium. Dynamic friction heating, transient convection, or rapid cooling are not currently modeled.
  • Teacher Domain Gap in Extreme Thermal Industrial Scenarios: The teacher network originates from natural RGB/infrared super-resolution; when applied to extreme non-standard industrial thermal defects, its prior confidence weighting may become less effective.
  • Future Work: Extending the framework to 4D dynamic thermal radiance fields by coupling transient heat conduction differential equations into Gaussian temporal updates.
  • vs SRGS: Designed for the visible spectrum, SRGS assumes all high-frequency cues represent true textures. On thermal images, it overfits sensor noise, whereas SupIR-GS decouples NUR/FPN via PhysIR-Cal, outperforming SRGS by 2.6~2.7 dB in PSNR with drastically fewer floaters.
  • vs Thermal3D-GS: Thermal3D-GS incorporates infrared radiation modeling but directly optimizes on noisy LR observations, suffering from blurred boundaries and striping noise. SupIR-GS uses native continuous high-density rasterization aligned with forward physical degradation to maintain crisp boundaries.
  • vs DifIISR + 3DGS: Cascaded 2D diffusion super-resolution introduces view-inconsistent geometric hallucinations that degrade multi-view 3D convergence. SupIR-GS embeds super-resolution constraints directly into 3D continuous Gaussian space, ensuring multi-view consistency and real-time inference.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering framework addressing SNR inversion in thermal super-resolution 3DGS via differentiable physical calibration.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across indoor, outdoor, and UAV benchmarks with temperature fidelity and ablation checks.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical formulations, rigorous physical rationale, and well-structured experimental analysis.
  • Value: ⭐⭐⭐⭐⭐ Significant practical value for all-weather autonomous navigation, drone surveillance, and 3D thermal metrology.