Skip to content

Differentiable Polarized Path Tracing

Conference: ECCV 2026
Paper: ECCV Official
Project: https://vcai.mpi-inf.mpg.de/projects/DPPT/
Area: 3D Vision
Keywords: Differentiable Rendering, Polarized Light Transport, Path Replay Backpropagation, Mueller-Stokes Calculus, Inverse Rendering

TL;DR

To resolve numerical breakdowns in path replay backpropagation caused by the rank-deficiency of polarimetric Mueller matrices, this paper presents an unbiased differentiable polarized path tracing framework using cached suffix radiance and hybrid recomputation, achieving constant memory usage and higher efficiency for material, normal, and optical element optimization.

Background & Motivation

Physically based differentiable rendering (PBDR) leverages reverse-mode automatic differentiation to jointly optimize scene geometry, surface reflectance, and illumination parameters from image observations. However, the vast majority of existing PBDR formulations operate solely on radiometric intensity, completely discarding rich polarimetric cues embedded in the electromagnetic wave's oscillation state. In real-world physics, polarized light transport exhibits acute sensitivity to Fresnel surface reflection, diffuse depolarization, and microfacet surface normal orientations, offering crucial physical constraints to disentangle specular and diffuse reflection, resolve surface normal azimuth ambiguities, and disambiguate photometrically identical materials.

Incorporating polarization into differentiable path tracing introduces formidable computational and mathematical hurdles. In forward light transport simulation, polarized light is governed by Mueller-Stokes calculus, where scalar radiance values become 4D Stokes vectors and surface scattering events (BSDFs) or optical filters are characterized by \(4 \times 4\) Mueller matrices. Naively applying conventional reverse-mode automatic differentiation (Conv. AD) requires retaining enormous computational graphs across Monte Carlo path iterations, which quickly exhausts system memory (OOM). Path replay backpropagation (PRB) was originally designed to achieve constant-memory gradient evaluation by exploiting the local invertibility of scalar path throughput during the adjoint pass. However, common polarimetric operators are fundamentally singular and non-invertible: an ideal diffuse reflection depolarizes light and retains only its scalar albedo component in entry \((0,0)\), while a linear polarizer completely eliminates orthogonal polarization states. Applying naive regularized matrix inversion within PRB leads to severe gradient bias, ill-conditioned numeric oscillations, or zero-division failures.

The central tension lies in computing mathematically unbiased gradients for polarized light transport without retaining the massive computation graphs of reverse-mode AD. Core idea: abandon the reliance on inverting singular Mueller matrices during adjoint replay, and instead introduce "Cached Suffix Replay" by using a compact backward fold at the end of the primal pass to store suffix Stokes radiances, which are directly reloaded during the adjoint pass to evaluate unbiased gradients with constant-like memory and accelerated by hybrid depolarization-aware recomputation.

Method

Overall Architecture

The computational pipeline consists of two coordinated passes: a primal forward polarized path tracing pass and an adjoint backpropagation replay pass. During the primal pass, the integrator traces path vertices \(i\), logging local emissions, Stokes radiances, and detached Mueller transport matrices; once the forward path finishes, a backward fold reconstructs and caches the suffix polarization radiances. During the adjoint replay pass, the identical geometric path is reconstructed using the same random seed, directly retrieving cached suffix radiances to evaluate local BSDF parameter gradients via automatic differentiation—completely sidestepping explicit Mueller matrix inversions and delivering unbiased parameter gradients.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Scene Parameters and Incident Polarized Rays"] --> B["Primal Sampling & Local State Recording<br/>Emissions Le and Mueller Matrix M"]
    B --> C["Cached Suffix Radiance via Backward Fold<br/>Construct Suffix Polarized Radiance L[i+1]"]
    C --> D["Adjoint Path Replay & Unbiased Gradient Computation<br/>Direct Suffix Lookup, Inversion-Free"]
    D --> E["Depolarization Index Check & Hybrid Recomputation<br/>Block-wise Checkpoints for Deep Paths"]
    E --> F["Output: Unbiased Scene Parameter Gradients δπ"]

Key Designs

1. Rank-Deficient Mueller Operator Breakdown: Formulating the Failure Modes of Replay Inversion

Standard scalar PRB achieves constant memory during the adjoint pass by dividing out the local BSDF contribution to recover the incident radiance from downstream bounces. In polarized transport, however, light is represented by a 4D Stokes vector \(\mathbf{s} = [s_0, s_1, s_2, s_3]^\top\), and surface interactions correspond to \(4 \times 4\) Mueller matrices \(\mathbf{M}\). For instance, an ideal Lambertian diffuse BSDF completely depolarizes light: its Mueller matrix contains only a single non-zero entry \(M_{00} = \alpha\) (the surface albedo) while the remaining 15 entries are strictly zero, yielding a matrix rank of 1. Similarly, a linear polarizer filters out circular and orthogonal linear polarization components, yielding a rank-2 matrix. Applying standard PRB with ad-hoc noise regularization (e.g., adding diagonal jitter \(\mathbf{M} + u\mathbf{I}_4\)) produces severely ill-conditioned matrices, resulting in relative gradient errors exceeding 140% and catastrophic divergence when optimizing roughness textures or polarizer orientations.

2. Cached Suffix Replay: Eliminating Matrix Inversion via Local Backward Folding

To guarantee unbiased gradient estimation without ever computing matrix inverses, this method introduces a lightweight suffix caching strategy. In the primal tracing routine (sample_polarized_primal), the path tracer records detached emitted radiances \(\mathbf{L}_e\) and normalized Mueller factors \(\boldsymbol{\beta}_i = \mathbf{M}_i / p(\omega_i)\) in compact local arrays of size \(N+1\). Upon completing forward path generation, a backward fold is executed from the terminal vertex back to the origin:

\[\mathbf{L}[j] \leftarrow \mathbf{L}[j] + \boldsymbol{\beta}[j] \mathbf{L}[j+1], \quad \text{for } j = N-2, \dots, 0\]

This operation summarizes the downstream polarized light contribution at each vertex into a compact 4D Stokes vector \(\mathbf{L}[j]\). During the adjoint pass (sample_polarized_adjoint), when the replay reaches vertex \(i\), it directly loads the cached suffix radiance \(\mathbf{L}_i = \mathbf{L}[i+1]\) from the array and evaluates the local parameter derivative:

\[\delta \pi \mathrel{+}= \text{backward}\left( \delta \mathbf{L}^\top \boldsymbol{\beta}_{\text{replay}} \mathbf{M}_i \mathbf{L}[i+1] \right)\]

This formulation bypasses Mueller matrix inversion entirely. Because it only stores 4D Stokes vectors rather than full autograd computational nodes, memory usage remains modest and linear in path depth rather than scene complexity, while preserving strict mathematical unbiasedness.

3. Hybrid Cached Replay: Balancing Storage and Recomputation via Depolarization Heuristics

While per-vertex suffix caching is optimal for moderate path depths (\(N \le 32\)), storing suffix radiances at every bounce could become memory-intensive for extremely deep paths with multiple internal bounces. The hybrid replay variant establishes checkpoints at intervals of \(k\) bounces. For intermediate vertices between checkpoints, the adjoint pass computes the depolarization index \(\text{DI}(\bar{\mathbf{M}})\) on the detached accumulated block throughput \(\bar{\mathbf{M}}\):

\[\text{suffix recovery} = \begin{cases} \text{scalar PRB update}, & \text{if } \text{DI}(\bar{\mathbf{M}}) < \gamma \\ \bar{\mathbf{M}}^{-1} \mathbf{L}, & \text{if } \text{DI}(\bar{\mathbf{M}}) \ge \gamma \text{ and } \sigma_{\min}(\bar{\mathbf{M}}) > \varepsilon \\ \text{recursive forward recomputation}, & \text{otherwise} \end{cases}\]

When multiple scattering events render the light completely depolarized (\(\text{DI}(\bar{\mathbf{M}}) < \gamma\)), the algorithm seamlessly transitions to scalar PRB updates. If the transport retains polarization and the accumulated Mueller matrix is well-conditioned (\(\sigma_{\min} > \varepsilon\)), direct inversion is permitted. Otherwise, a local forward recomputation is triggered from the previous checkpoint, bounding memory consumption in scenes with over 80 bounces.

Key Experimental Results

Main Results

All benchmarks were performed on a workstation equipped with an NVIDIA GeForce RTX 4090 GPU (24 GiB). In gradient correctness tests against the conventional automatic differentiation (Conv. AD) reference, relative error (RE) was measured across multiple inverse rendering targets:

Scene Target Parameter P-PRB (RE) P-RB (RE) Ours (RE) Conv. AD (RE)
Living Room Diffuse Texture 1.4939 0.3907 0.3815 Reference (0.0000)
Aluminium Vase Normal Map 0.6905 0.1398 0.1398 Reference (0.0000)
Clock Roughness Texture 2.7589 0.1235 0.1235 Reference (0.0000)
Cornell Box Polarizer Angle \(\theta\) 0.3269 0.0209 0.0210 Reference (0.0000)

Across end-to-end inverse rendering optimization tasks, parameter reconstruction accuracy was evaluated using root mean square error (RMSE):

Inverse Task Objective & Parameter Scalar PRB (RMSE) Ours Polarized (RMSE) Conv. AD
Kitchen Glare Removal Polarizer orientation \(\theta\) suppressing cooktop specular glare 12.1848 3.9550 OOM (Out of Memory)
Veach Complex Illumination Joint diffuse reflectance and roughness disentanglement 0.2491 0.1274 0.1272
Living Room Floor Diffuse texture recovery under indirect bounces 0.2679 0.0069 OOM (Out of Memory)
Marble Bust Normal Map 12-view 1K-resolution high-frequency surface normal recovery 0.04761 0.04196 OOM (Out of Memory)

Ablation Study & Performance Analysis

Integration with Projective Sampling for single-view 3D geometry reconstruction across seven mesh datasets:

Configuration Chamfer Distance ↓ Normal Error ↓ Reconstruction Fidelity
Unpolarized Baseline 0.1874 65.61° Oversmoothed contours, blurred micro-geometry
Polarized (Ours) 0.0904 (-51.8%) 46.72° (-18.89°) Sharp geometric silhouettes, faithful surface normals

Performance evaluation in the Modern Hall scene equipped with 85 linear polarizers (resolution \(1280 \times 720\), 1 spp): - Memory Consumption: Conv. AD rapidly exhausted GPU memory on an NVIDIA RTX 5090 as bounce count increased; the proposed method maintained nearly flat, constant memory up to depth 32, exhibiting only modest sublinear growth thereafter. - Computational Runtime: P-PRB was fast but produced corrupted gradients (RE > 2.0); P-RB yielded unbiased gradients but incurred high computational latency; the proposed method matched Conv. AD in accuracy while executing approximately \(2\times\) faster than P-RB.

Key Findings

  • Mueller singularity causes complete failure in regularized PRB: Adding diagonal noise to singular Mueller operators (P-PRB) fails to resolve ill-conditioning, yielding a severe relative error of 2.7589 when differentiating roughness and causing optimization divergence.
  • Polarization cues resolve material ambiguity: In the Veach benchmark, scalar PRB conflates diffuse albedo with specular roughness (RMSE 0.2491), whereas polarized observations decouple orthogonal and parallel reflectance components, reducing RMSE by nearly half to 0.1274.
  • Inversion-free replay overcomes AD memory limits: Conv. AD runs out of memory on 1K-resolution multi-view normal optimization and complex diffuse transport scenes; our suffix-cached replay comfortably optimizes full-resolution parameters within a single consumer GPU.

Highlights & Insights

  • Inversion-free adjoint via backward folding: By exploiting the known topology at the conclusion of the primal pass, performing a single backward fold converts what would have been an intractable matrix inversion into a deterministic array lookup during replay.
  • Physics-grounded hybrid checkpointing: Employing the depolarization index (\(\text{DI}\)) as a branching heuristic aligns algorithmic recomputation decisions with the physical dissipation of polarized states in scattering media.
  • Orthogonal compatibility with geometric differentiability: The framework seamlessly embeds into projective sampling integrators, proving that physical polarization modeling naturally augments visibility boundary tracking for sharper 3D reconstruction.

Limitations & Future Work

  • Lack of polarized subsurface scattering (BSSRDF): The current implementation focuses on surface scattering (BSDF) and does not yet solve the vector radiative transfer equation for participating or translucent media. Extending to BSSRDF is vital for separating superficial gloss from internal scattering in marble, jade, and human skin.
  • Expanded local register footprint: Although global computational graphs are eliminated, each path vertex stores a \(4 \times 4 \times 3\) tensor in local registers, incurring higher register pressure than scalar path tracers under dense sampling budgets.
  • Application to computational optical design: While verified on inverse appearance problems, future avenues include applying the framework to the inverse design of metasurfaces, liquid-crystal spatial light modulators, and AR near-eye display waveguides.
  • vs Path Replay Backpropagation (PRB) [Vicini et al. 2021]: Scalar PRB relies on the algebraic invertibility of scalar throughput factors. This paper demonstrates that polarimetric Mueller matrices are fundamentally rank-deficient, replacing matrix inversion with primal suffix caching.
  • vs Radiative Backpropagation (RB) [Nimier-David et al. 2020]: While a polarized variant of RB (P-RB) computes unbiased gradients, its bidirectional evaluation overhead is substantial; our cached replay delivers comparable accuracy with an approximate \(2\times\) speedup and significantly lower memory than Conv. AD.
  • vs Shape from Polarization (SfP) & NeRSP [Han et al. 2024]: Previous SfP and neural field formulations incorporate AoP or DoLP as auxiliary geometric loss priors. In contrast, this work establishes the foundational transport gradients, integrating full Mueller-Stokes wave optics directly into path tracing backpropagation.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates the first mathematically sound, inversion-free differentiable polarized path tracing framework, overcoming the singular Mueller matrix challenge in PRB.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous validation encompassing gradient relative error analysis, multiple inverse rendering applications, single-view 3D reconstruction, and stress-test scaling benchmarks.
  • Writing Quality: ⭐⭐⭐⭐⭐ Problem motivation is exceptionally clear, the mathematical derivations and pseudocode are rigorous, and the physical insights are thoroughly articulated.
  • Value: ⭐⭐⭐⭐⭐ Bridges the fundamental gap between wave polarization physics and physically based differentiable rendering, providing a general-purpose tool for high-fidelity inverse graphics and optical inverse design.