Skip to content

Physics-Guided Deep Learning for Linear Mueller Matrix Acquisition

Conference: ECCV 2026
Paper: ECCV Official
Area: Others
Keywords: Polarization, Mueller Matrix, Polarimetric BRDF, Physics-Guided Deep Learning, Inverse Rendering

TL;DR

Addressing the low efficiency of traditional Mueller matrix polarimetry that requires multiple incident polarization switches, this paper proposes a physics-guided two-stage framework that estimates pBRDF parameters from single-shot four-channel polarization images to construct an analytical initialization and refines it via residual learning into an accurate, physically plausible \(3 \times 3\) linear Mueller matrix.

Background & Motivation

Polarization, an intrinsic wave property of light orthogonal to intensity and phase, encodes rich structural and physical information regarding surface microgeometry, material composition, and refractive properties. In polarimetric optics and computer vision, the Mueller matrix provides a comprehensive linear transformation operator that maps an incident Stokes vector to an emergent Stokes vector upon reflection or scattering. Because its elements are governed strictly by the target object's intrinsic physical characteristics and remain independent of the incident polarization state, acquiring high-fidelity Mueller matrices is invaluable across biomedical tissue diagnosis, microstructural defect inspection, and 3D normal measurement.

However, practical acquisition of high-precision Mueller matrices has long been constrained by bottlenecks in acquisition efficiency and physical interpretability. Mathematically, resolving a local Mueller matrix requires at least four non-collinear, independent pairs of incident and emergent Stokes vectors. Even with modern division-of-focal-plane (DoFP) polarimetric cameras capable of capturing four linear polarization orientations (\(0^\circ, 45^\circ, 90^\circ, 135^\circ\)) in a single exposure, conventional time-division setups must mechanically or electro-optically modulate the illumination polarization at least four successive times, which leads to slow acquisition and motion vulnerability. On the other hand, forward analytical models using polarimetric Bidirectional Reflectance Distribution Functions (pBRDF) offer physically structured matrix representations, yet they suffer from heavy reliance on inaccessible ground-truth normals, roughness, and illumination geometry, while failing under real-world model mismatch and complex scattering. Conversely, purely data-driven black-box neural networks often generate unconstrained matrices that violate fundamental optical realizability conditions.

This paper's angle of attack is that analytical physical models and neural networks are mutually complementary: an analytical pBRDF model provides a structured, physically consistent initialization from sparse observations, while neural networks excel at modeling high-frequency non-ideal residual errors and complex material variations. Core idea: formulate single-shot linear Mueller matrix recovery as a two-stage physics-guided process combining pBRDF initialization with neural residual refinement, where PIPNet predicts physical surface parameters to synthesize an initial pBRDF Mueller matrix, and PMMNet refines it via polarization-color decoupling and redundant-channel separation.

Method

Overall Architecture

The framework accepts as input a single set of four linear polarization images \(P = (I_0, I_{45}, I_{90}, I_{135})\) captured under a fixed, known incident polarization state using a DoFP camera, along with an object binary mask \(\text{Mask}\). The end-to-end pipeline consists of two cascaded stages: Stage 1 utilizes the Polarization Initialization Parameters Network (PIPNet) to predict the dense surface normal \(\mathbf{n}\), diffuse albedo \(\mathbf{k}_d\), surface roughness \(\sigma\), and incident illumination direction \(\boldsymbol{\omega}_i\). These parameters are fed into an analytical pBRDF model to construct structured initial specular and diffuse Mueller matrices (\(\mathbf{M}^s_{\text{init}}\) and \(\mathbf{M}^d_{\text{init}}\)). In Stage 2, the Polarization Mueller Matrix Network (PMMNet) ingests the measured Stokes vectors and the physical initializations, employing specialized specular and diffuse branches under polarization-color decoupling and channel separation to predict residual corrections, ultimately assembling the 27-dimensional full-color linear Mueller matrix \(\mathbf{M}_{\text{opt}}\).

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Single-Shot 4 Polarization Images P + Mask<br/>(Fixed known incident polarization state)"] --> B["Stage 1: PIPNet Physical Parameter Estimation<br/>Predict normal n, diffuse albedo kd, roughness σ, light direction ωi"]
    B --> C["Analytical pBRDF Mueller Initialization<br/>Construct structured specular Ms_init and diffuse Md_init"]
    C --> D["Stage 2: PMMNet Dimensionality-Reduced Residual Optimization<br/>Polarization-color decoupling and redundant-channel separation"]
    E["Color & Channel Reassembly<br/>Output 3×3 full-color linear Mueller matrix Mopt (27-dim)"]
    D --> E

Key Designs

1. PIPNet Physical Parameter Estimation: Multi-branch shared backbone for geometry and reflectance decoupling

Inverting an under-constrained polarimetric transfer operator directly from single-shot observations without prior regularization often collapses into degenerate local minima. PIPNet adopts a U-Net-like encoder-decoder backbone that shares multi-scale feature representations while branching into three parallel pixel-wise heads for normalized surface normal \(\mathbf{n} \in \mathbb{R}^{H \times W \times 3}\), RGB diffuse albedo \(\mathbf{k}_d \in \mathbb{R}^{H \times W \times 3}\), and light direction \(\boldsymbol{\omega}_i \in \mathbb{R}^{H \times W \times 3}\). For the image-wide scalar microfacet roughness \(\sigma\), an average pooling layer and fully connected network branch off the encoder bottleneck. This structured parameter regression maps raw polarization measurements into well-defined pBRDF state variables, sidestepping the coaxial illumination or multi-view photometric requirements common in prior fitting pipelines.

2. Analytical pBRDF Mueller Initialization: Anchoring structural optics via physics-based reflection

To prevent the neural network from wandering through unconstrained high-dimensional matrix spaces, the framework computes a coarse yet physically sound initial Mueller matrix for each valid pixel using the classic microfacet pBRDF model. The refractive index is fixed to \(\eta = 1.5\), specular albedo \(k_s = 1\), and viewing direction \(\boldsymbol{\omega}_o\) is derived directly from the calibrated camera model. Evaluating the analytical formulation on PIPNet's output \((\mathbf{n}_{\text{init}}, \mathbf{k}_{d,\text{init}}, \sigma_{\text{init}}, \boldsymbol{\omega}_{i,\text{init}})\) generates the specular component \(\mathbf{M}^s_{\text{init}}\) and diffuse component \(\mathbf{M}^d_{\text{init}}\): $\(\mathbf{M} = k_s \mathbf{M}^s + \mathbf{k}^d \mathbf{M}^d\)$ Although this analytical initialization inevitably contains model mismatch errors due to idealized Fresnel and Lambertian assumptions, it rigorously satisfies fundamental optical reciprocity, energy conservation, and trace invariants, providing a robust physical anchor for subsequent refinement.

3. PMMNet Dimensionality-Reduced Residual Optimization: Polarization-color decoupling and redundant-channel separation

Directly optimizing a 27-dimensional RGB Mueller matrix (\(3 \text{ colors} \times 3 \times 3 \text{ matrix}\)) introduces extreme optimization complexity and risks cross-color polarization artifacts. Authors resolve this through two complementary dimensionality-reduction strategies. First, polarization-color decoupling leverages the empirical observation that polarization transformations exhibit minimal chromatic dispersion across visible wavelengths for typical materials; the 27-channel space is collapsed into a single wavelength-independent \(3 \times 3\) polarization block, with color handled exclusively by diffuse albedo weights. Second, redundant-channel separation targets only physically relevant non-zero residual elements: the specular branch refines 5 key elements \(\{m_{00}, m_{11}, m_{12}, m_{21}, m_{22}\}\) using high-frequency attention to focus on highlight zones, while the diffuse branch updates 5 elements \(\{m_{00}, m_{01}, m_{02}, m_{10}, m_{20}\}\) with multi-head channel attention to capture global depolarization and attenuation. The refined 5-channel vectors are expanded back to 9-channel matrices, multiplied by color albedo, and recombined into the final 27-dimensional matrix \(\mathbf{M}_{\text{opt}}\).

Loss & Training

PIPNet is trained with a composite loss balancing cosine similarity on normal vectors with \(L_1\) distances on radiometric parameters: $\(\mathcal{L}_{\text{PIPNet}} = \mathcal{L}_{\cos}(\mathbf{n}_{\text{init}}, \mathbf{n}_{\text{GT}}) + 0.5\mathcal{L}_1(\mathbf{k}_{d,\text{init}}, \mathbf{k}_{d,\text{GT}}) + 0.5\mathcal{L}_1(\boldsymbol{\omega}_{i,\text{init}}, \boldsymbol{\omega}_{i,\text{GT}}) + 0.3\mathcal{L}_1(\sigma_{\text{init}}, \sigma_{\text{GT}})\)$

PMMNet optimizes both matrix parameter accuracy and emergent Stokes consistency across synthetic incident test probes \(\mathcal{K} = \{[1,0,0]^T, [1,1,0]^T, [1,0,1]^T\}\) representing unpolarized, horizontal linear, and \(45^\circ\) linear polarizations: $\(\mathcal{L}_{\text{PMMNet}} = \mathcal{L}_1(\mathbf{M}, \mathbf{M}_{\text{GT}}) + \mathcal{L}_{\text{Stokes}}(\mathbf{S}^{\text{out}}_{\text{pred}}, \mathbf{S}^{\text{out}}_{\text{GT}})\)$ $\(\mathcal{L}_{\text{Stokes}} = \frac{1}{|\mathcal{K}| \cdot |\mathcal{V}|} \sum_{c \in \{R,G,B\}} \sum_{\mathbf{S}^{\text{in}} \in \mathcal{K}} \sum_{p \in \mathcal{V}} \left\| \mathbf{S}^{\text{out}}_{c,\text{pred}}(p) - \mathbf{S}^{\text{out}}_{c,\text{GT}}(p) \right\|_1\)$ The models are trained sequentially in PyTorch on an NVIDIA RTX 4080 Super with Adam optimizer and \(256 \times 256\) resolution. PIPNet is trained for 40 epochs with initial learning rate \(2 \times 10^{-4}\) (linearly decaying after 20 epochs); PMMNet is trained for 20 epochs with a learning rate reduction in the final 10 epochs.

Key Experimental Results

Main Results

Evaluation was conducted on both synthetic datasets (rendered via Mitsuba 3 with KAIST real polarimetric measurements) and a real-world dataset comprising 12 diverse objects captured across multiple poses and illumination directions on a turntable setup. Baselines represent adapted analytical pBRDF fitting methods (Baek et al., Hwang et al., Ichikawa et al.) provided with ground-truth geometric parameters. Metrics include matrix PSNR, observed-state rendering PSNR and SSIM, Degree of Linear Polarization (DoLP) MAE, and Angle of Linear Polarization (AoLP) wrapped angular error in degrees.

Dataset Method Matrix PSNR (dB) ↑ Rendering PSNR (dB) ↑ Rendering SSIM ↑ DoLP MAE ↓ AoLP Angular Err. (deg) ↓
Synthetic Baek et al. [5] 46.15 38.27 0.8892 0.0830 57.98
Synthetic Hwang et al. [19] 48.56 38.12 0.9393 0.0801 57.08
Synthetic Ichikawa et al. [21] / 38.95 0.9058 0.0789 48.21
Synthetic Ours 51.70 43.87 0.9950 0.0782 34.51
Real Baek et al. [5] 26.63 23.11 0.8998 0.1610 71.62
Real Hwang et al. [19] 28.27 23.80 0.8984 0.0865 74.03
Real Ichikawa et al. [21] / 26.67 0.9051 0.1057 65.49
Real Ours 40.42 32.33 0.9895 0.0749 24.34

Ablation Study: Held-Out Forward Stokes Prediction

To verify whether the recovered matrix functions as an authentic physical transfer operator rather than simply memorizing the single observed state, forward Stokes prediction was evaluated across 11 held-out, unseen incident polarization states. Baseline comparisons include pure analytical prediction (PIPNet-pBRDF), analytical fitting upper bound (GT/Best-fit pBRDF), parallel multi-task supervision (Joint pBRDF/MM multi-task), and unconstrained data-driven mapping (Direct Stokes→MM).

Dataset Strategy / Config Forward Stokes \(L_2\) ↓ DoLP MAE ↓ AoLP Angular Err. (deg) ↓ Note
Synthetic PIPNet-pBRDF 0.4065 0.1346 45.75 Pure analytical reconstruction from predicted parameters
Synthetic GT/Best-fit pBRDF 0.2715 0.0963 44.60 Fitting upper bound still restricted by model mismatch
Synthetic Joint pBRDF/MM multi-task 0.0457 0.0614 44.92 Parallel multi-task lacks explicit sequential physical anchoring
Synthetic Direct Stokes→MM 0.0433 0.0471 42.03 Unconstrained data-driven direct regression
Synthetic Ours (Full Model) 0.0214 0.0391 37.88 Physical prior + residual refinement yields lowest error
Real PIPNet-pBRDF 0.6450 0.1318 50.44 Compounded parameter and model errors on real captures
Real Best-fit pBRDF 0.4953 0.0994 44.72 Inadequate modeling of non-ideal real depolarization
Real Joint pBRDF/MM multi-task 0.1965 0.0786 28.36 Struggles to resolve severe under-constrained ambiguity
Real Direct Stokes→MM 0.1615 0.0721 29.84 Pure neural mapping degrades under unseen incident states
Real Ours (Full Model) 0.0926 0.0642 26.57 Achieves 42.7% forward error reduction over Direct Stokes→MM

Key Findings

  • In matrix-level fidelity and real-world rendering, the proposed framework achieves 40.42 dB PSNR on real captures—a substantial 12.15 dB leap over the best analytical baseline Hwang et al. (28.27 dB)—while slashing AoLP error from \(74.03^\circ\) to \(24.34^\circ\).
  • Under held-out forward Stokes evaluation, pure black-box regression (Direct Stokes→MM) exhibits severe operator extrapolation errors (0.1615 on real data), whereas our physics-guided two-stage framework drops the error to 0.0926, confirming that analytical pBRDF initialization supplies essential geometric inductive bias.
  • Physical plausibility checks reveal that under the rigorous analytical Givens-Kostinski (GK) diagnostic, traditional fitting methods achieve only a 33.91% pass rate on real data, whereas our method reaches 86.94% (and 98.25% on synthetic data), with a 100% pass rate under positive semidefiniteness covariance testing.

Highlights & Insights

  • Physics-Anchored Residual Formulation: Deconstructs an ill-posed four-state acquisition bottleneck into a sequential pipeline of "analytical pBRDF initialization + neural residual correction", marrying physical bounds with data-driven non-ideal modeling.
  • Orthogonal Dimensionality Reduction: Mitigates 27-dimensional optimization instability via polarization-color decoupling and redundant-channel selection, concentrating capacity on physically active specular and diffuse elements.
  • Operator-Level Stokes Consistency Loss: Probing the emergent response against virtual multi-state incident vectors provides an elegant, transferable regularization mechanism for polarimetric inverse problems.

Limitations & Future Work

  • Active Point-Light Illumination Assumption: Relies on a calibrated single point light and known incident polarization, which limits direct applicability under uncontrolled, ambient, or spatially extended natural lighting.
  • Absence of Circular Polarization Components: Focuses exclusively on the \(3 \times 3\) linear Mueller block measurable by DoFP polarimeters, omitting circular polarization states (e.g., \(m_{33}\) and chiral coupling terms) necessary for highly optically active media.
  • Weak Dispersion Assumption: The polarization-color decoupling assumes negligible chromatic dispersion in polarization transfer; materials exhibiting strong wavelength-selective dichroism or deep subsurface chromatic scattering may introduce residual color-edge errors.
  • vs Baek et al. [ACM TOG 2018]: Baek et al. require coaxial polarized illumination and structured light with expensive iterative non-linear optimization. This method recovers linear Mueller blocks from single-shot DoFP images within milliseconds while achieving substantially higher physical fidelity.
  • vs Hwang et al. [ACM TOG 2022] / Ichikawa et al. [CVPR 2023]: Prior polarimetric BRDF methods either fit isolated low-order metrics (e.g., DoLP) or require dense multi-view flash captures. This work achieves full object-level linear Mueller matrix recovery, showing clear downstream gains in Shape from Polarization (SfP) and material classification.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ First framework to combine analytical pBRDF models with dual-branch residual learning for single-shot linear Mueller matrix recovery.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Evaluated across synthetic and calibrated multi-source real datasets, including unseen state forward prediction, strict GK physical diagnostics, and downstream SfP/material classification.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulation, clear operator definitions, and comprehensive ablation design.
  • Value: ⭐⭐⭐⭐⭐ Bridges specialized optical polarimetry and mainstream computer vision, turning conventional DoFP cameras into fast, physically grounded Mueller imaging sensors.