SupIR-GS: Thermal Infrared Super-Resolution Novel View Synthesis with Imaging-Calibrated 3D Gaussian Splatting¶
Conference: ECCV 2026
Paper: CVF Open Access
Area: 3D Vision
Keywords: thermal infrared imaging, 3D Gaussian Splatting, super-resolution, novel view synthesis, imaging calibration
TL;DR¶
Addressing the core challenge of "signal-to-noise ratio inversion" in thermal infrared imaging—where thermal diffusion smooths true boundaries while sensor non-uniformity introduces structured noise—SupIR-GS presents the first imaging-calibrated super-resolution 3D Gaussian Splatting framework, decoupling hardware artifacts via a differentiable physical calibration layer, boosting flat thermal gradients via a local tone adapter, and progressively recovering fine geometric boundaries via frequency-aware curriculum learning.
Background & Motivation¶
Thermal infrared (TIR) imaging measures long-wave infrared radiation to characterize scene thermal signatures, offering robust all-weather perception under low light, smoke, and degraded visibility across autonomous driving, disaster response, and industrial inspection. However, constrained by detector manufacturing costs and physical diffraction limits, uncooled microbolometer thermal cameras typically yield low-resolution (LR) observations. When modern continuous 3D representations like neural radiance fields (NeRF) or 3D Gaussian Splatting (3DGS) are applied directly to multi-view LR thermal images, they inevitably encounter a severe geometry-radiance ambiguity.
The physical origin of this failure lies in a pronounced "signal-to-noise ratio inversion" intrinsic to thermal infrared sensors. On one hand, physical heat conduction, thermal diffusion, and atmospheric scattering attenuate true geometric boundaries into low-contrast, low-frequency transitions. On the other hand, sensor imperfections such as non-uniform response (NUR) and fixed pattern noise (FPN) manifest as sharp, structured high-frequency textures on the sensor plane. Standard 3DGS optimized under photometric supervision is easily misled by these spurious signals: it overfits hardware noise by spawning dense high-frequency floaters in 3D space, while collapsing in texture-scarce regions due to gradient starvation. Off-the-shelf 3D visible super-resolution pipelines (e.g., SRGS) or 2D image restoration models fail because they mistake sensor noise for real geometry and lack cross-view consistency.
The core idea is to natively render continuous 3D Gaussians at the target high resolution, explicitly decouple sensor-induced degradation from latent thermal radiance through a differentiable forward-imaging calibration chain and bounded local tone modulation, and steer geometric optimization from coarse low-frequency scaffolds to sharp high-frequency thermal boundaries using frequency-aware curriculum learning.
Method¶
Overall Architecture¶
SupIR-GS establishes an end-to-end framework integrating forward physical degradation modeling with inverse optimization of continuous 3D radiance fields. Given multi-view low-resolution infrared sequences \(\{I_v^{lr}\}\), the model directly rasterizes continuous 3D Gaussian primitives on a high-density target sampling grid \(\Omega^{hr}\) (\(r^2\) times more query points, where \(r\) is the super-resolution scale), synthesizing an ideal, degradation-free latent high-resolution radiance image \(I_{base}\). This clean rendering then passes sequentially through the Physically-constrained Infrared Calibration module (PhysIR-Cal) and the lightweight Intensity-conditioned Tone Adapter (IRTone-Adapter) to simulate sensor-specific NUR, additive bias, lens radial vignetting, and local dynamic range compression, producing a calibrated prediction \(I_{pred}\). Supervised loss backpropagates through this forward chain to update 3D Gaussian positions, scales, rotations, and opacities in the clean radiance domain. Meanwhile, a Frequency-Aware Curriculum Learning strategy incorporates structural priors from a pretrained super-resolution teacher, leveraging confidence masks and a cosine ramp-up schedule to smoothly transition optimization from low-frequency scaffolds to fine thermal details.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multi-view Low-Resolution TIR Inputs<br/>Multi-view sequences and high-density HR sampling grid"] --> B["Resolution-Agnostic 3DGS Radiance Rendering<br/>Continuous 3D Gaussians rasterized into latent clean HR radiance"]
B --> C["Physical-Constrained Infrared Calibration<br/>Parametric gain-bias grid and radial vignetting decouple NUR and FPN"]
C --> D["Intensity-conditioned Tone Adapter<br/>Lightweight MLP predicts bounded local affine residuals to enhance contrast"]
D --> E["Frequency-Aware Curriculum Learning Strategy<br/>Teacher-guided confidence-masked high-frequency loss with cosine ramp-up"]
E --> F["High-Fidelity Thermal Super-Resolution NVS<br/>Sharp geometric boundaries and accurate thermodynamic fidelity"]
Key Designs¶
1. Physical-Constrained Infrared Calibration: Decoupling Sensor Non-Uniformity and Fixed Pattern Noise
Uncooled microbolometer arrays suffer from non-uniform response (NUR) and fixed pattern striping noise (FPN). Directly fitting pixel-wise noise severely overfits low-SNR observations, pulling 3D Gaussians away from actual physical surfaces. PhysIR-Cal introduces a parameter-efficient, differentiable forward calibration layer to explicitly absorb these systematic hardware aberrations. The module defines two low-resolution learnable parameter grids \(G_{grid}, B_{grid} \in \mathbb{R}^{16 \times 16}\), which are bilinearly upsampled to target dimensions and bounded by \(\tanh(\cdot)\) activations to parameterize multiplicative spatial gain \(G\) and additive spatial bias \(B\). To model optical attenuation from infrared lenses, a learnable 4th-order radial vignetting function is formulated as \(V(r) = \exp(a_2 r^2 + a_4 r^4)\), where \(r\) denotes the normalized radial distance. For a base 3DGS rendering \(I_{base}\), the calibrated intensity is computed as:
To accommodate automatic gain control (AGC) or per-frame histogram stretching common in thermal sensors, an optional monotonic tone mapping in the logit domain absorbs inter-frame gain fluctuations: \(I_{phy} = \sigma\big(s_k \cdot \text{logit}(\tilde{I}_{calib}) + t_k\big)\), where \(s_k, t_k\) are per-view parameters and \(\tilde{I}_{calib}\) is numerically clamped. By isolating hardware noise inside this low-frequency parametric layer, PhysIR-Cal shields 3D Gaussian geometry and densification gradients from spurious high-frequency artifacts.
2. Intensity-conditioned Tone Adapter: Surfacing Latent Temperature Gradients via Local Affine Residuals
Thermal infrared images frequently suffer from narrow dynamic range and flattened intensity histograms, leaving cold background and low-contrast regions texture-scarce. In these regions, standard photometric reprojection errors generate negligible and noisy gradients that fail to guide sub-pixel Gaussian refinement. IRTone-Adapter employs a lightweight two-layer MLP (64 hidden channels) conditioned on local rendered intensity \(I_{in}(p)\) to predict adaptive multiplicative gain \(\delta_m\) and additive bias \(\delta_o\) residuals per pixel:
To prevent the network from masking true 3D geometric inaccuracies through excessive intensity remapping, the scaling bounds are strictly restricted to \(\lambda_m = 0.08\) and \(\lambda_o = 0.04\), bounding perturbations within \(\pm 8\%\) and \(\pm 4\%\). This conservative, bounded affine modulation gently amplifies subtle temperature variations in flat regions, generating informative and consistent backpropagated gradients for sharpening fine geometry.
3. Frequency-Aware Curriculum Learning Strategy: Confidence-Masked Annealing from Coarse to Fine
Because thermal diffusion intrinsically blurs physical boundaries, high-frequency cues in low-resolution inputs are inherently ambiguous. Naively injecting high-frequency details from 2D super-resolution networks risks propagating single-view hallucinations. SupIR-GS first constructs a spatial confidence mask \(W_{conf}\) by evaluating consistency between the pretrained SwinIR teacher output \(I_{teacher}\) and the bilinearly upsampled input \(I_{up}\) across both intensity and gradient domains:
This exponential formulation guarantees that \(W_{conf}(\mathbf{p})\) approaches 1 only when the teacher prediction strictly agrees with low-frequency observation structures. The confidence-masked high-frequency residual loss is defined via an average-pooling high-pass filter \(H(\cdot)\):
During training iteration \(t\), the high-frequency loss weight \(\lambda_{hf}(t)\) is governed by a cosine ramp-up schedule: kept at 0 during early stages so that photometric and SSIM losses establish a stable low-frequency geometric scaffold, and then smoothly increased across iterations (65% to 95%) to inject sharp edge details safely without destabilizing early convergence.
Loss & Training¶
The entire framework is trained end-to-end for 30,000 iterations. The unified loss function integrates photometric error, structural similarity, confidence-weighted high-frequency residuals, and parameter regularization:
Here, \(\mathcal{L}_{pix}\) enforces \(\ell_1\) photometric consistency; \(\mathcal{L}_{ssim}\) preserves structural integrity; \(\mathcal{R}_{cal}\) and \(\mathcal{R}_{tone}\) are regularizers with weights \(\lambda_{cal} = 5 \times 10^{-4}\) and \(\lambda_{tone} = 10^{-3}\) to prevent degradation parameters from drifting. The SwinIR-S teacher operates purely offline prior to training to generate high-frequency targets and masks, incurring zero extra memory or compute overhead during training and novel-view rendering.
Key Experimental Results¶
Main Results¶
Quantitative evaluations are conducted on TI-NSD (20 scenes across Indoor, Outdoor, and UAV settings) and ThermoScenes (10 scenes featuring high-contrast indoor objects and building facades). All models take \(4\times\) bicubic-downsampled images as input.
Table 1: Quantitative comparisons on the TI-NSD dataset under 4× super-resolution
| Method | Indoor SSIM↑ | Indoor PSNR↑ (dB) | Indoor LPIPS↓ | Outdoor SSIM↑ | Outdoor PSNR↑ (dB) | Outdoor LPIPS↓ | UAV SSIM↑ | UAV PSNR↑ (dB) | UAV LPIPS↓ |
|---|---|---|---|---|---|---|---|---|---|
| Bicubic + 3DGS | 0.9222 | 31.16 | 0.3095 | 0.8466 | 27.47 | 0.3311 | 0.8399 | 29.26 | 0.3127 |
| Lanczos + 3DGS | 0.9237 | 31.04 | 0.3085 | 0.8375 | 26.94 | 0.3360 | 0.8444 | 29.39 | 0.3069 |
| DifIISR + 3DGS | 0.9278 | 30.45 | 0.3097 | 0.8453 | 26.77 | 0.3235 | 0.8449 | 28.96 | 0.2899 |
| 3DGS (4×) | 0.9060 | 28.23 | 0.3147 | 0.7825 | 21.86 | 0.3491 | 0.7555 | 23.37 | 0.3562 |
| Thermal3DGS (4×) | 0.9154 | 30.15 | 0.3117 | 0.7852 | 23.36 | 0.3418 | 0.7287 | 23.40 | 0.3733 |
| Mip-Splatting (4×) | 0.9232 | 30.63 | 0.3020 | 0.8280 | 25.37 | 0.3304 | 0.8582 | 29.92 | 0.2786 |
| Scaffold-GS (4×) | 0.9145 | 28.70 | 0.3077 | 0.7998 | 22.05 | 0.3331 | 0.7405 | 21.64 | 0.3612 |
| SRGS | 0.9312 | 32.09 | 0.2765 | 0.8701 | 27.38 | 0.2790 | 0.9078 | 31.78 | 0.2014 |
| Ours (SupIR-GS) | 0.9455 | 34.81 | 0.2597 | 0.9066 | 29.94 | 0.2336 | 0.9428 | 34.24 | 0.1453 |
Table 2: Quantitative comparisons and temperature fidelity on ThermoScenes dataset (illustrating smoothness bias)
| Method | Indoor SSIM↑ | Indoor PSNR↑ (dB) | Indoor LPIPS↓ | Indoor MAE↓ | Indoor \(\text{MAE}_{roi}\)↓ | Outdoor SSIM↑ | Outdoor PSNR↑ (dB) | Outdoor LPIPS↓ | Outdoor MAE↓ | Outdoor \(\text{MAE}_{roi}\)↓ |
|---|---|---|---|---|---|---|---|---|---|---|
| Bicubic + 3DGS | 0.9650 | 31.19 | 0.0415 | 0.3536 | 1.6318 | 0.9663 | 30.24 | 0.1577 | 0.6166 | 0.5411 |
| Lanczos + 3DGS | 0.9665 | 30.73 | 0.0414 | 0.3574 | 1.7223 | 0.9686 | 30.54 | 0.1555 | 0.7988 | 0.7441 |
| DifIISR + 3DGS | 0.9486 | 28.95 | 0.0530 | 0.7947 | 1.9993 | 0.9617 | 29.21 | 0.1636 | 1.1811 | 1.1606 |
| 3DGS (4×) | 0.9612 | 29.23 | 0.0556 | 0.3581 | 2.0971 | 0.9238 | 25.71 | 0.2164 | 0.8346 | 0.7380 |
| Thermal3DGS (4×) | 0.9122 | 27.17 | 0.0504 | 0.8558 | 2.1585 | 0.9201 | 23.92 | 0.2205 | 1.8974 | 2.0263 |
| Mip-Splatting (4×) | 0.9662 | 30.59 | 0.0476 | 0.3685 | 1.8094 | 0.9585 | 29.83 | 0.1659 | 0.7149 | 0.6316 |
| Scaffold-GS (4×) | 0.9656 | 28.93 | 0.0545 | 0.3219 | 2.1495 | 0.9315 | 26.45 | 0.1998 | 0.6888 | 0.6311 |
| SRGS | 0.9537 | 29.84 | 0.0454 | 0.4048 | 2.1793 | 0.9679 | 30.76 | 0.1472 | 0.6619 | 0.5969 |
| Ours (SupIR-GS) | 0.9672 | 31.32 | 0.0438 | 0.3368 | 1.7903 | 0.9662 | 30.40 | 0.1546 | 0.7957 | 0.7329 |
Ablation Study¶
Table 3: Progressive ablation study on core modules (TI-NSD dataset)
| Config | PhysIR-Cal | IRTone-Adapter | Frequency-Aware | Indoor SSIM↑ | Indoor PSNR↑ (dB) | Indoor LPIPS↓ | Outdoor SSIM↑ | Outdoor PSNR↑ (dB) | Outdoor LPIPS↓ | UAV SSIM↑ | UAV PSNR↑ (dB) | UAV LPIPS↓ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| (1) Baseline | × | × | × | 0.9312 | 32.086 | 0.2765 | 0.8701 | 27.381 | 0.2790 | 0.8888 | 30.291 | 0.2299 |
| (2) Tone only | × | ✓ | × | 0.9400 | 31.700 | 0.2920 | 0.8655 | 27.202 | 0.2836 | 0.8988 | 30.860 | 0.2087 |
| (3) Freq only | × | × | ✓ | 0.9354 | 32.281 | 0.2662 | 0.8727 | 27.263 | 0.2680 | 0.9121 | 31.550 | 0.1868 |
| (4) Tone+Freq | × | ✓ | ✓ | 0.9401 | 32.936 | 0.2638 | 0.8925 | 28.514 | 0.2491 | 0.9231 | 32.203 | 0.1745 |
| (5) Calib only | ✓ | × | × | 0.9420 | 33.692 | 0.2837 | 0.8795 | 28.449 | 0.2681 | 0.9118 | 32.218 | 0.1928 |
| (6) Calib+Freq | ✓ | × | ✓ | 0.9453 | 34.276 | 0.2788 | 0.9031 | 29.800 | 0.2377 | 0.9305 | 33.142 | 0.1668 |
| (7) Full Ours | ✓ | ✓ | ✓ | 0.9455 | 34.806 | 0.2597 | 0.9066 | 29.935 | 0.2336 | 0.9320 | 33.326 | 0.1653 |
Key Findings¶
- Physical Module Dependency and Sequence: Enabling IRTone-Adapter in isolation drops Indoor PSNR by 0.39 dB because uncalibrated contrast stretching amplifies sensor FPN and NUR into pseudo-geometry. Similarly, applying Frequency-Aware constraints without calibration degrades Outdoor PSNR by 0.12 dB due to gradient conflict between noise and structural edges. Only when PhysIR-Cal first strips sensor noise to form a reliable scaffold (+1.61 dB Indoor, +1.93 dB UAV) can tone adaptation and high-frequency priors unleash their full synergy, boosting PSNR by over 2.5 to 3.0 dB across subsets.
- Smoothness Bias in Thermal Evaluation: In large uniform thermal backgrounds (e.g., cold skies and building walls), naive 2D interpolations (Bicubic+3DGS) mathematically minimize global pixel variance, showing deceptively high PSNR while obliterating window frames and structural edges. SupIR-GS resists this smoothness bias, outperforming all baselines on perceptual metric LPIPS and region-of-interest temperature error (\(\text{MAE}_{roi}\)).
- High Efficiency and Low Memory: Optimized in 16.42 minutes on an RTX 3080, SupIR-GS achieves 381.10 FPS rendering speed with peak memory of only 2792 MB during novel-view synthesis, notably surpassing 3DGS (276.36 FPS / 3580 MB) and SRGS (313.75 FPS / 3484 MB) because all calibration and teacher networks are stripped at inference time.
Highlights & Insights¶
- Physics-Informed Sensor Decoupling: Introduces a differentiable parameter-efficient grid to absorb multiplicative NUR and additive FPN, directly solving the "SNR inversion" bottleneck and preventing high-frequency floaters from corrupting 3D geometry.
- Bounded Intensity-Conditioned Adaptation: Restricts local affine tone modifications to narrow bounds (\(\pm 8\%\) gain, \(\pm 4\%\) bias), surfacing subtle gradients in flat thermal scenes without permitting color remapping to mask geometric errors.
- Joint Metric-Masked Frequency Annealing: Combines intensity and gradient deviation into a confidence mask along with cosine ramp-up scheduling, safely distilling teacher high-frequency edge priors while filtering out cross-view hallucinations.
Limitations & Future Work¶
- Static Thermal Equilibrium Assumption: Assumes objects remain in steady thermodynamic equilibrium. Dynamic friction heating, transient convection, or rapid cooling are not currently modeled.
- Teacher Domain Gap in Extreme Thermal Industrial Scenarios: The teacher network originates from natural RGB/infrared super-resolution; when applied to extreme non-standard industrial thermal defects, its prior confidence weighting may become less effective.
- Future Work: Extending the framework to 4D dynamic thermal radiance fields by coupling transient heat conduction differential equations into Gaussian temporal updates.
Related Work & Insights¶
- vs SRGS: Designed for the visible spectrum, SRGS assumes all high-frequency cues represent true textures. On thermal images, it overfits sensor noise, whereas SupIR-GS decouples NUR/FPN via PhysIR-Cal, outperforming SRGS by 2.6~2.7 dB in PSNR with drastically fewer floaters.
- vs Thermal3D-GS: Thermal3D-GS incorporates infrared radiation modeling but directly optimizes on noisy LR observations, suffering from blurred boundaries and striping noise. SupIR-GS uses native continuous high-density rasterization aligned with forward physical degradation to maintain crisp boundaries.
- vs DifIISR + 3DGS: Cascaded 2D diffusion super-resolution introduces view-inconsistent geometric hallucinations that degrade multi-view 3D convergence. SupIR-GS embeds super-resolution constraints directly into 3D continuous Gaussian space, ensuring multi-view consistency and real-time inference.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Pioneering framework addressing SNR inversion in thermal super-resolution 3DGS via differentiable physical calibration.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across indoor, outdoor, and UAV benchmarks with temperature fidelity and ablation checks.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical formulations, rigorous physical rationale, and well-structured experimental analysis.
- Value: ⭐⭐⭐⭐⭐ Significant practical value for all-weather autonomous navigation, drone surveillance, and 3D thermal metrology.