Learning Spectral and Polarimetric Clues for One-to-Multimodal Novel View Synthesis¶
Conference: ECCV 2026
Paper: ECCV Official Link
Project Page: https://medialab.dei.unipd.it/paper_data/SPoILeR/
Area: 3D Vision
Keywords: Multimodal Neural Rendering, Novel View Synthesis, Dictionary Basis Decomposition, Spectral & Polarimetric Priors, Implicit Radiance Fields
TL;DR¶
Addressing the prohibitive acquisition costs of specialized sensors, SPoILeR decouples implicit radiance fields into shared bases and scene-specific coefficients with multi-scene pre-training, enabling photorealistic and strictly multi-view consistent synthesis of multispectral, infrared, and polarimetric novel views using only RGB supervision.
Background & Motivation¶
Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have demonstrated exceptional fidelity in novel view synthesis and high-precision 3D geometry reconstruction, a revolution propelled by the vast availability of RGB camera datasets. However, in scientific and industrial domains such as non-destructive material inspection, reflection separation, and remote sensing, imaging modalities beyond visible light—such as Multispectral (MS), Near-Infrared (NIR), and Polarization (Pol)—capture fundamental physical and electromagnetic properties that RGB cameras inherently miss. Emerging multimodal neural rendering frameworks such as X-NeRF, NeSpoF, and MultimodalStudio have extended radiance fields to handle multiple sensors. Nonetheless, they universally require deploying expensive, bulky, and meticulously calibrated multimodal camera rigs to capture extensive multi-view frames for every single target scene, drastically hindering real-world deployment.
A deeper technical barrier emerges when trying to infer non-conventional modalities without full sensor coverage. Existing cross-modality synthesis strategies fall into a critical dilemma. On the one hand, feed-forward 2D conversion models (such as MST++ for spectral recovery or PolarAnything for polarization estimation) operate on single RGB frames independently. Because they lack 3D awareness, their predictions across different camera poses suffer from severe multi-view geometric and physical inconsistencies; feeding these noisy 2D predictions into volumetric rendering pipelines results in severe blurriness and spatial artifacts. On the other hand, attempting to reconstruct missing modalities directly within a standard 3D neural field from single-modality RGB supervision leads to catastrophic representation drift and collapse in unobserved latent channels.
This paper tackles the challenge by decoupling cross-scene shared physical radiance correlations from scene-specific geometry and texture: because different optical sensors observe the electromagnetic scattering of identical physical materials, the cross-modality correlations remain invariant across scenes. Core idea: decompose the multimodal radiance field into globally shared basis grids and scene-specific coefficient grids, pre-training a robust multimodal latent space and sensor decoders across diverse scenes, and fine-tuning only the lightweight scene coefficients and geometry on a target scene with RGB frames alone—regularized by latent geometry, inverse mapping, and modality-to-luma losses to achieve zero-shot, multi-view consistent novel view synthesis of unseen modalities.
Method¶
Overall Architecture¶
SPoILeR builds upon an SDF-based neural implicit surface framework and a decoupled dictionary field representation. During the multi-scene pre-training (PT) phase, the model is trained across a dataset of scenes captured with a full suite of sensors (RGB, NIR, Monochrome, Polarization, and Multispectral). It alternates between optimizing scene-specific coefficients/geometry and globally shared bases/radiance decoders. During the target-scene fine-tuning (FT) phase, all shared basis grids, encoders, and multimodal decoders are frozen; the model optimizes only the scene-specific coefficient field and geometry using dense RGB images of the target scene, thereby enabling multi-view consistent rendering of any missing modality.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input: 3D point x, view direction v, geometry feature g"] --> DecoupledField["Decoupled Basis & Coefficient Fields<br/>Inner product of scene c(x) and shared b(γ(x))"]
DecoupledField --> LatentEnc["Shared Latent Encoder & Multimodal Decoders<br/>Projection J + Latent Encoder Z + Sensor Decoders D_l"]
LatentEnc --> Regularization["Triple Latent Regularization Mechanisms<br/>Geometry Distance L_lsg + Inverse Anchor L_inv + Modality Luma L_m2l"]
Regularization --> Sched["Two-Stage Optimization & Gradient Gating<br/>Cross-scene PT + Single-modality FT with Dropout"]
Sched --> Out["Output: Multi-view consistent RGB / NIR / Mono / Pol / MS renderings"]
Key Designs¶
1. Decoupled Basis and Coefficient Fields: Separating Cross-Scene Physical Priors from Scene Content Conventional single-scene NeRFs lock model capacity within per-scene hash grids, preventing knowledge transfer across different scenes. Drawing inspiration from Dictionary Fields, SPoILeR parameterizes the radiance feature at spatial position \(\mathbf{x}\) as the dot product between a scene-specific coefficient field \(\mathbf{c}(\mathbf{x}) \in \mathbb{R}^D\) and a globally shared basis field \(\mathbf{b}(\gamma(\mathbf{x})) \in \mathbb{R}^D\), where \(\gamma(\cdot)\) is a multi-frequency periodic coordinate mapping enabling spatial basis reusability: $\(f = \mathbf{c}(\mathbf{x})^\top \mathbf{b}(\gamma(\mathbf{x}))\)$ This structure provides a strict inductive bias: high-frequency surface textures and scene-specific albedo variations are absorbed by the coefficient field \(\mathbf{c}\), while cross-spectral and polarimetric physical coupling laws are preserved in the shared basis field \(\mathbf{b}\), establishing a transferable foundation for unobserved modality extrapolation.
2. Shared Latent Encoder and Multimodal Decoders: Building a Compact Radiance Pathway Following the feature combination \(f\), the radiance module employs a projection MLP \(\mathcal{J}\), a latent encoder \(\mathcal{Z}\), and sensor-specific decoders \(\mathcal{D}_l\). The encoder \(\mathcal{Z}\) concatenates the projected feature with the viewing direction \(\mathbf{v}\) and a geometric feature \(\mathbf{g}\) extracted from the SDF geometry field, producing a latent vector \(\mathbf{z}\). Modality-specific decoders \(\mathcal{D}_l\) then decode \(\mathbf{z}\) into individual spectral or polarimetric channels: $\(\mathcal{P}(f, \mathbf{v}, \mathbf{g}, l) = \mathcal{D}_l(\mathcal{Z}(\mathbf{g}, \mathcal{J}(f), \mathbf{v}))\)$ All mapping functions are implemented as shallow MLPs to concentrate capacity inside the basis and coefficient grids. This keeps the latent space moderately compressed and easily steerable during single-modality fine-tuning, preventing high-dimensional overfitting and ensuring robust decoder activation.
3. Triple Latent Regularization Mechanisms: Preventing Representation Drift in Missing Channels When fine-tuning under RGB supervision alone, latent representations corresponding to unobserved modalities risk severe divergence. SPoILeR introduces three complementary regularization objectives: - Latent Space Geometry Loss \(\mathcal{L}_{lsg}\): To ensure decoder stability against latent perturbations, this loss matches the pairwise Euclidean distance distributions of ray-integrated latents \(\hat{\mathbf{z}}\) to those of decoded multimodal radiances \(\hat{\mathbf{m}}\) via KL divergence: $\(\mathcal{L}_{lsg} = \mathrm{KL}\left(\mathrm{softmax}_{\mathrm{row}}(-D_z), \, \mathrm{softmax}_{\mathrm{row}}(-D_m)\right)\)$ - Inverse Function Loss \(\mathcal{L}_{inv}\): An inverse MLP \(\mathcal{I}_\Theta(\hat{\mathbf{m}}) \to \hat{\mathbf{z}}_e\) is trained during PT to reconstruct latent codes from decoded radiances. During FT, \(\mathcal{I}_\Theta\) is frozen and provides an anchoring constraint \(\mathcal{L}_{inv} = \mathrm{MSE}(\hat{\mathbf{z}}_e, \hat{\mathbf{z}})\), guiding the coefficient optimization and preventing latent drift. - Modality-to-Luma Loss \(\mathcal{L}_{m2l}\): A shallow MLP \(\mathcal{M}_\Theta(\hat{\mathbf{m}}_l, l) \to g_e\) maps decoded channels of modality \(l\) to a grayscale luminance matching the RGB ground-truth luma \(g\): \(\mathcal{L}_{m2l} = \mathrm{MSE}(g_e, g)\). This aligns infrared-sensitive channels (Mono, NIR, Pol) with visible luminance, preventing them from collapsing into decoupled orthogonal sub-spaces.
4. Two-Stage Optimization and Gradient Gating: Safeguarding Shared Priors from Geometry Leakage During pre-training, the model alternates scene-specific and shared module updates in cycles of \(B=15\) iterations (80% steps updating scene coefficients/geometry and 20% updating shared bases/decoders). To prevent scene-specific SDF geometry features \(\mathbf{g}\) from memorizing radiance clues (which would fail to generalize during FT when novel geometry is learned), SPoILeR applies gradient dropout to \(\mathbf{g}\): only a single random modality is allowed to backpropagate gradients into the geometry feature per iteration. In the FT stage, all shared modules are frozen, optimizing only 300k parameters (~7% of total) in 30k iterations—reducing training time to one-fourth of training from scratch.
Loss & Training¶
The overall training objective combines photometric loss \(\mathcal{L}_{mod}\), SDF Eikonal regularizer \(\mathcal{L}_{eik}\) (\(\lambda_{eik} = 0.1\)), curvature smoothness \(\mathcal{L}_{curv}\) (\(\lambda_{curv} = 0.0005\)), and the three latent regularizations: $\(\mathcal{L} = \mathcal{L}_{mod} + \lambda_{eik}\mathcal{L}_{eik} + \lambda_{curv}\mathcal{L}_{curv} + \lambda_{lsg}\mathcal{L}_{lsg} + \lambda_{inv}\mathcal{L}_{inv} + \lambda_{m2l}\mathcal{L}_{m2l}\)$ Pre-training runs for 1M iterations on a single NVIDIA RTX 6000 Ada with \(\lambda_{lsg} = 0.1, \lambda_{inv} = 1.0, \lambda_{m2l} = 1.0\). Scene-specific fine-tuning runs for 30k iterations with \(\mathcal{L}_{lsg}\) disabled (\(\lambda_{lsg} = 0\)), using \(\lambda_{inv} = 0.01\) and \(\lambda_{m2l} = 0.02\).
Key Experimental Results¶
Main Results¶
Experiments are evaluated on MMS-DATA (32 object-centric scenes; 27 for PT, 5 for FT; covering RGB, NIR, Mono, Pol, and MS modalities). Metrics include PSNR (dB) and SSIM (for single-channel/demosaicked images). The baseline MMS-FW requires 100k iterations and ~12M parameters trained from scratch per scene. In contrast, SPoILeR FT optimizes only ~300k parameters (~7% of its 4.1M total capacity) for 30k iterations.
The table below summarizes novel view synthesis quality under standard RGB-only supervision and fully supervised upper bounds:
| Evaluation Setup | Train Mod. | Test Mod. | SPoILeR (Ours) PSNR ↑ | SPoILeR SSIM ↑ | MMS-FW [42] PSNR ↑ | MMS-FW SSIM ↑ | Gain / Remark |
|---|---|---|---|---|---|---|---|
| Standard (RGB Only) | RGB | RGB | 30.07 | — | 29.53 | — | +0.54 dB (Pre-trained priors benefit RGB) |
| Standard (RGB Only) | RGB | Mono | 25.78 | 0.88 | N/A | — | Zero-shot non-conventional rendering |
| Standard (RGB Only) | RGB | NIR | 26.55 | 0.87 | N/A | — | Zero-shot near-infrared extrapolation |
| Standard (RGB Only) | RGB | Pol | 24.25 | — | N/A | — | Zero-shot polarimetric synthesis |
| Standard (RGB Only) | RGB | MS | 25.45 | — | N/A | — | Zero-shot multispectral rendering |
| Upper Bound (All Mod.) | ALL | RGB | 30.54 | — | 32.77 | — | -2.23 dB (Capacity gap vs scratch model) |
| Upper Bound (All Mod.) | ALL | Mono | 29.21 | 0.93 | 32.98 | 0.94 | Comparable structural fidelity |
| Upper Bound (All Mod.) | ALL | NIR | 32.12 | 0.92 | 34.25 | 0.93 | 2.13 dB margin with 1/3 parameters |
| Upper Bound (All Mod.) | ALL | Pol | 29.32 | — | 30.66 | — | 1.34 dB margin |
| Upper Bound (All Mod.) | ALL | MS | 28.46 | — | 31.42 | — | High-fidelity upper bound retained |
In comparison with 2D conversion pipelines combined with NeRF: - On Multispectral synthesis, MMS-FW trained on MST++ predictions achieves 21.99 dB PSNR, whereas SPoILeR RGB-FT achieves 25.45 dB (+3.46 dB gain). - On Polarization synthesis, MMS-FW trained on PolarAnything estimates achieves an AoP Mean Angular Error (MAngE) of 44.13° and DoP Mean Absolute Error (MAbsE) of 0.089. SPoILeR improves these to 30.05° (reducing error by 14.08°) and 0.045 (reducing error by ~50%), demonstrating the essential value of 3D-consistent prior learning.
Ablation Study¶
Ablation analysis on the three proposed latent regularizers during RGB-only fine-tuning:
| Configuration | RGB (dB) | Mono (dB) | NIR (dB) | Pol (dB) | MS (dB) | Observation & Physical Interpretation |
|---|---|---|---|---|---|---|
| Full Model (Ours) | 30.07 | 25.78 | 26.55 | 24.25 | 25.45 | Best comprehensive multi-band synthesis |
| w/o Latent Geometry \(\mathcal{L}_{lsg}\) | 29.90 (-0.17) | 25.44 (-0.34) | 26.30 (-0.24) | 23.08 (-1.18) | 23.14 (-2.31) | Major drops in MS and Pol (-2.31 dB / -1.18 dB) |
| w/o Inverse Function \(\mathcal{L}_{inv}\) | 29.96 (-0.11) | 25.54 (-0.25) | 25.96 (-0.58) | 24.01 (-0.24) | 25.28 (-0.17) | Moderate global degradation across modalities |
| w/o Modality-to-Luma \(\mathcal{L}_{m2l}\) | 29.95 (-0.12) | 21.04 (-4.38) | 22.91 (-3.64) | 23.73 (-0.52) | 25.36 (-0.09) | Catastrophic collapse in Mono and NIR (-4.38 dB / -3.64 dB) |
Key Findings¶
- Modality-to-Luma Loss \(\mathcal{L}_{m2l}\) is crucial for bridging infrared modalities: Because Mono and NIR record light in the infrared band, unconstrained latent spaces tend to decouple them from visible RGB representations. Removing \(\mathcal{L}_{m2l}\) causes Mono and NIR performance to drop by 4.38 dB and 3.64 dB, respectively.
- Latent Space Geometry Loss \(\mathcal{L}_{lsg}\) governs high-dimensional channels (MS & Pol): Multispectral and polarimetric modalities contain numerous correlated channels. Aligning the latent distance metric with output radiance space ensures that imperfect latent estimates smoothly map to physically consistent spectral and polarimetric curves.
- Robustness in Unbalanced Few-Shot Regimes: Under extreme setups with 45 RGB views but only 1 view of a secondary modality, SPoILeR surpasses MMS-FW by ~8 dB on MS, ~7 dB on Pol, and ~6 dB on NIR, demonstrating that pre-trained physical bases eliminate the need for dense multi-view capture across all sensors.
Highlights & Insights¶
- Decoupled Basis-Coefficient Inductive Bias: Adapting Dictionary Fields to multimodal rendering successfully decouples general electromagnetic correlations from per-scene geometric textures, breaking the dependency on dense specialized captures.
- 3D Implicit Priors Eliminate 2D Error Accumulation: Operating directly in 3D implicit space inherently guarantees multi-view consistency, avoiding the perspective flicker and spatial averaging artifacts inherent to 2D image conversion methods like MST++ and PolarAnything.
- Geometry Feature Gradient Gating: Applying gradient dropout to geometric features during pre-training prevents scene-specific SDF representations from leaking radiative priors, an elegant and practical design for cross-scene representations.
Limitations & Future Work¶
- Dependency on Pre-trained Material Distributions: Extrapolations are bounded by the material and reflectance distributions observed in the pre-training set; novel meta-materials or atypical polarizers may exhibit degraded accuracy.
- Controlled Illumination Constraints: The method assumes static white-balance, fixed camera exposure, and controlled indoor lighting across training views.
- Future Directions: Extending the decoupled framework to uncontrolled in-the-wild illumination, and adapting the dictionary basis design to real-time 3D Gaussian Splatting architectures.
Related Work & Insights¶
- vs MultimodalStudio (MMS-FW): MMS-FW requires capturing all 5 sensor modalities for 100k iterations per scene; SPoILeR pre-trains transferable priors, allowing 30k-iteration RGB fine-tuning with only ~300k trainable parameters, drastically lowering acquisition costs.
- vs 2D Conversion Baselines (MST++ / PolarAnything): 2D image-to-modality generators produce view-inconsistent predictions that collapse when optimized in 3D volumetric fields; SPoILeR ensures strict multi-view consistency by decoding directly from unified 3D neural fields.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Pioneering framework for one-to-multimodal multi-view consistent novel view synthesis.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive 5-modality evaluation, including unbalanced regimes, 2D SOTA comparisons, and exhaustive ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical modeling, crystal-clear motivation, and meticulous experimental reporting.
- Value: ⭐⭐⭐⭐⭐ Significantly lowers the barrier for multimodal 3D digital twins, inspection, and robotic perception.