Skip to content

Poppy: Polarization-Based Plug-and-Play Guidance for Enhancing Surface Normal Estimation

Conference: ECCV 2026
Paper: ECCV Official
Code: https://irnkim.github.io/poppy/
Area: 3D Vision
Keywords: Surface Normal Estimation, Polarization Imaging, Test-Time Guidance, Plug-and-Play, Reflectance Decomposition

TL;DR

To tackle the systemic breakdown of monocular RGB surface normal estimators on reflective, textureless, and dark surfaces, Poppy presents a training-free test-time guidance framework that keeps arbitrary pretrained backbones frozen and refines geometry via input image offsets, output normal offsets, and differentiable polarization rendering.

Background & Motivation

Monocular surface normal estimation serves as a foundational pillar for robotics manipulation, augmented reality, scene understanding, and 3D geometric reconstruction. Despite significant strides made by modern deep learning architectures trained on large-scale RGB-normal datasets—spanning diffusion models like Marigold, flow-matching architectures like Lotus-v2, and feed-forward vision transformers like MoGe-2—these estimators frequently fail on edge cases characterized by extreme optical properties. Such failures cluster into three representative scenarios: highly reflective (specular) surfaces, where view-dependent highlights are mistaken for geometric relief; expansive textureless regions, where the lack of spatial intensity gradients yields oversmoothed, non-committal predictions; and low-illumination or low-albedo dark objects, where low signal-to-noise ratios severely degrade visual priors. Fundamentally, inferring 3D surface orientations purely from 2D photometric shading is an inherently ill-posed inverse problem.

Polarization imaging provides a physically grounded complement to resolve these ambiguities. When natural unpolarized light reflects off a physical surface, the polarization state of the emergent light wave is directly governed by the local surface normal and material refractive properties, remaining completely independent of surface macroscopic textures and spatial albedos. However, traditional Shape from Polarization (SfP) is notoriously bottlenecked by ambiguities embedded in the Fresnel reflection equations: an azimuthal flip ambiguity of \(\pi\) and an orthogonal diffuse-specular ambiguity of \(\pi/2\), which collectively generate four candidate normal orientations per pixel measurement. Classical methods attempt to break these ambiguities by capturing multiple lighting conditions, viewpoints, or auxiliary modalities, forfeiting single-snapshot portability. Conversely, learning-based SfP networks rely on curated paired polarization-normal training data; because collecting such datasets is prohibitive and limited in material diversity, these models suffer from fragile generalization across novel domains.

A stark dichotomy exists between the two paradigms: RGB foundation models possess broad semantic and topological priors across open-domain scenes but lack physical fidelity on extreme surfaces, whereas SfP methods offer rigorous physical orientation constraints but remain trapped by multi-view capture burdens and narrow training distributions. The key insight of this paper is to bypass expensive polarimetric data collection and retraining altogether by treating single-shot polarization physics as a plug-and-play test-time guidance objective. The frozen RGB backbone supplies global orientation priors that resolve the \(\pi\) azimuthal ambiguity, while dense polarimetric measurements backpropagate corrective gradients to eliminate geometric errors. Core idea: formulate single-shot polarization guidance as a training-free test-time optimization problem where frozen backbones are guided by input-space image offsets for global orientation steering, output-space normal offsets for high-frequency detail recovery, and differentiable specular radiance decomposition under a staged activation schedule.

Method

Overall Architecture

The input to Poppy consists of four multi-angle polarization intensity images \(I_{0^\circ}, I_{45^\circ}, I_{90^\circ}, I_{135^\circ}\) acquired by a linear polarization camera in a single snapshot. From these, the unpolarized intensity \(S_0\) is extracted as the RGB input image \(x\), alongside the observed linear Stokes vector \(S = [S_0, S_1, S_2]^\top\). The system outputs a refined, physically consistent surface normal map \(\hat{n} \in \mathbb{R}^{H \times W \times 3}\). The entire optimization is performed at inference time via iterative backpropagation while keeping all weights of the pretrained monocular normal estimation backbone \(f\) strictly frozen.

At the input side, a learnable per-pixel image offset \(O_x\) is added to form a perturbed input \(x + O_x\), which steers the backbone's global geometric predictions. The forward pass yields base surface normals \(\hat{n}_{\text{base}} = f(x + O_x)\). At the output side, a learnable per-pixel normal offset \(O_n\) is added to capture high-frequency surface relief, producing the refined normal \(\hat{n}_t = \hat{n}_{\text{base}} + O_n\). To project \(\hat{n}_t\) into polarization space for comparison against ground-truth measurements, a learnable per-pixel specular radiance map \(L_s\) is maintained. A differentiable polarimetric rendering layer combines \(\hat{n}_t\), viewing directions, and \(L_s\) via the Fresnel equations to analytically synthesize predicted Stokes parameters \(\hat{S}\). An L1 polarization consistency loss evaluated under a physical validity mask computes error residuals, and gradients are backpropagated to update all learnable guidance parameters simultaneously.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multi-angle Polarization Input<br/>I_0, I_45, I_90, I_135"] --> B["Stokes Extraction<br/>Intensity x and Observed Stokes S"]
    B --> C["Image Offset Guidance<br/>Global perturbation x + Ox"]
    C --> D["Frozen Normal Estimator Backbone<br/>Marigold / Lotus-v2 / MoGe-2"]
    D --> E["Normal Offset Guidance<br/>High-frequency refinement n_base + On"]
    E --> F["Differentiable Rendering & Decomposition<br/>Synthesize S_hat with Ls via Fresnel equations"]
    F --> G["Polarization Consistency Loss<br/>Physical validity mask M filtering invalid pixels"]
    G -->|backpropagate gradient| C
    G -->|backpropagate gradient| E
    G -->|backpropagate gradient| F

Key Designs

1. Image Offset Guidance: Input-Sensitivity-Driven Global Orientation Steering

When monocular geometry estimators encounter large reflective patches or dark shadows, they tend to produce severe systematic orientation biases or warped surface curvatures spanning extensive contiguous regions. Attempting to correct these errors solely by tweaking output normal vectors pixel-by-pixel easily gets trapped in poor local minima and fails to restore global topology. Poppy exploits the broad sensitivity of deep neural networks to input-space perturbations (characterized by input-to-normal Jacobian maps): a single pixel change at the input layer propagates through convolutional receptive fields and self-attention channels to affect normal predictions across wide spatial domains.

Accordingly, Poppy introduces a learnable, continuous per-pixel offset \(O_x \in \mathbb{R}^{H \times W \times 3}\) applied directly to the input image, feeding \(x + O_x\) into the backbone. During test-time backpropagation, the polarization consistency gradient flows backward through the frozen network weights into \(O_x\). This acts as a global guidance lever that converts polarimetric discrepancies into subtle adversarial perturbations in image space, effectively steering the foundation model's learned geometric priors to rotate and level out large-scale misaligned surfaces into alignment with physical polarization measurements.

2. Normal Offset Guidance: Residual-Decoupled High-Frequency Detail Recovery

Pretrained monocular estimators exhibit an inherent tendency toward oversmoothing and detail hallucination due to training losses (such as cosine similarity) and inductive biases in generative diffusion or transformer backbones. Furthermore, because image-space offsets \(O_x\) undergo aggressive feature downsampling and non-linear compression inside the network, they struggle to resolve fine-grained, high-frequency geometric structures such as dragon scales, shell ridges, and sharp facet creases.

To bridge this gap, Poppy introduces a learnable per-pixel normal residual offset \(O_n \in \mathbb{R}^{H \times W \times 3}\) coupled directly to the output of the backbone: $\(\hat{n}_t = f(x + O_x) + O_n\)$ Because \(O_n\) is linearly added to the output normals, it exhibits extremely high gradient sensitivity to the polarimetric loss. Activating \(O_n\) from the very beginning causes it to overfit prematurely to uncalibrated specular radiance noise, producing severe surface speckling. Poppy resolves this with a global-then-local guidance schedule: during the initial phase (\(t < 50\) iterations), \(O_n\) is kept inactive (fixed at its initial value of \(10^{-2}\) without receiving gradients), allowing \(O_x\) and \(L_s\) to stabilize global surface orientation and specular radiance. At step \(t = 50\), \(O_n\) is enabled for joint optimization until step \(T\). The final surface normal is projected onto the unit sphere via L2 normalization: $\(\hat{n} = \frac{f(x + O_x) + O_n}{\|f(x + O_x) + O_n\|_2}\)$

3. Differentiable Polarization Rendering: Closed-Loop Ambiguity Resolution via Radiance Decomposition

Synthesizing Stokes vectors from surface normals is a non-linear forward physical transport process. Given a surface normal \(n = [n_x, n_y, n_z]^\top\) and viewing direction \(v\) (approximated as \([0, 0, 1]^\top\) under orthographic projection), the elevation angle is \(\theta = \arccos(n \cdot v)\) and the azimuth angle is \(\psi = \arctan(n_y, n_x)\). However, emergent surface radiance comprises both diffuse reflection (from subsurface scattering) and specular reflection (from direct surface reflections), which exhibit mutually orthogonal polarization angles separated by \(\pi/2\) (\(\phi_d = \psi\) for diffuse vs. \(\phi_s = \psi + \pi/2\) for specular). Without decoupling these reflectance components, the forward rendering cannot match polarimetric measurements.

Poppy decomposes total intensity \(S_0\) into a learnable per-pixel specular radiance map \(L_s\) and diffuse radiance \(L_d = S_0 - L_s\). Assuming a typical refractive index of \(\eta = 1.5\), diffuse and specular degrees of linear polarization, \(\rho_d(\theta)\) and \(\rho_s(\theta)\), are analytically evaluated using the Fresnel reflection equations: $\(\rho_d(\theta) = \frac{(\eta - 1/\eta)^2 \sin^2\theta}{2 + 2\eta^2 - (\eta + 1/\eta)^2 \sin^2\theta + 4\cos\theta\sqrt{\eta^2 - \sin^2\theta}}\)$ $\(\rho_s(\theta) = \frac{2\sin^2\theta\cos\theta\sqrt{\eta^2 - \sin^2\theta}}{\eta^2 - \sin^2\theta - \eta^2\sin^2\theta + 2\sin^4\theta}\)$ The combined Stokes parameters are then synthesized in closed form: $\(\hat{S}_0 = L_d + L_s = S_0\)$ $\(\hat{S}_1 = L_d \rho_d \cos(2\phi_d) + L_s \rho_s \cos(2\phi_s)\)$ $\(\hat{S}_2 = L_d \rho_d \sin(2\phi_d) + L_s \rho_s \sin(2\phi_s)\)$ Through this differentiable operator \(\hat{S} = \mathcal{F}(\hat{n}_t, L_s)\), the framework establishes an end-to-end analytical bridge between normal space and Stokes measurement space, simultaneously yielding physically disentangled diffuse and specular radiance components.

Loss & Training

During the test-time guidance phase, the parameter set \(\Theta = \{O_x, O_n, L_s\}\) is iteratively optimized to minimize an L1 consistency loss against observed Stokes vectors: $\(\mathcal{L}(\Theta) = \sum_{p} M(p) \sum_{i=0}^2 \left| S_i(p) - \hat{S}_i(p) \right|\)$ where \(p\) denotes a spatial pixel location and \(M(p)\) represents a physical validity mask. The mask filters out unphysical and uninformative measurements across three categories: ① low-SNR underexposed regions (\(S_0 \le 0.01\)); ② overexposed saturated pixels (\(S_0 \ge 1.0\)); and ③ unphysical sensor measurements violating the polarization energy conservation constraint (\(S_1^2 + S_2^2 > S_0^2\)).

Optimization runs for \(T = 100\) steps using the Adam optimizer. All learnable parameters are initialized to \(10^{-2}\). Distinct learning rates are assigned to balance optimization dynamics: \(\lambda_{L_s} = 10^{-2}\) for specular radiance, \(\lambda_{O_n} = 10^{-3}\) for normal offsets (activated at step 50), and backbone-dependent learning rates for the image offset: \(\lambda_{O_x} = 10^{-3}\) for Marigold, \(5 \times 10^{-4}\) for Lotus-v2, and \(10^{-5}\) for MoGe-2. For the diffusion model Marigold, guidance operates across 4 denoising steps with 25 optimization iterations per step; for the flow-matching model Lotus-v2, guidance precedes the detail-sharpening stage; for MoGe-2, gradients backpropagate directly through the feed-forward transformer layers.

Key Experimental Results

Main Results

Poppy is comprehensively evaluated on 7 benchmark datasets, spanning 3 real-world datasets (14 objects total: SfPUEL, NeRSP, PISR, featuring specular highlights, low illumination, and textureless surfaces) and 4 synthetic datasets (18 objects total: SfPUEL, NeRSP, NeISF, DeepPol, featuring low-SNR, diffuse, and mixed materials). Metrics include Mean Angular Error (MAE in degrees, lower is better), Median Angular Error, RMSE, and thresholded accuracies measuring percentage of pixels with error below \(11.25^\circ\), \(22.5^\circ\), and \(30^\circ\) (higher is better).

Benchmark Split Method / Backbone Mean (MAE) ↓ Median ↓ RMSE ↓ Acc11.25 ↑ Acc22.5 ↑ Acc30 ↑
Real Benchmarks SfPUEL [38] (SfP Baseline) 20.68° 17.08° 26.15° 0.37 0.70 0.80
DSINE [6] 23.06° 19.07° 28.83° 0.27 0.62 0.76
Lotus [26] 16.69° 13.96° 21.70° 0.42 0.78 0.88
StableNormal [53] 17.80° 14.77° 23.29° 0.41 0.75 0.86
Marigold [32] 18.18° 15.25° 23.30° 0.36 0.74 0.86
Marigold + Poppy (Ours) 15.28° 12.57° 20.26° 0.47 0.83 0.92
Lotus-v2 [25] 14.68° 12.05° 19.51° 0.49 0.85 0.93
Lotus-v2 + Poppy (Ours) 12.65° 10.05° 17.62° 0.61 0.89 0.94
MoGe-2 [50] 13.10° 10.51° 18.11° 0.57 0.87 0.94
MoGe-2 + Poppy (Ours) 12.26° 9.88° 16.99° 0.61 0.90 0.95
Synthetic Benchmarks SfPUEL [38] (SfP Baseline) 17.03° 12.91° 23.33° 0.46 0.79 0.87
DSINE [6] 24.13° 19.94° 30.00° 0.26 0.60 0.75
Lotus [26] 19.99° 17.04° 25.06° 0.30 0.69 0.84
StableNormal [53] 18.45° 14.65° 24.55° 0.39 0.75 0.86
Marigold [32] 20.99° 17.40° 26.43° 0.30 0.68 0.82
Marigold + Poppy (Ours) 15.60° 11.89° 22.08° 0.56 0.82 0.89
Lotus-v2 [25] 16.52° 13.44° 22.08° 0.42 0.81 0.90
Lotus-v2 + Poppy (Ours) 12.26° 8.81° 18.70° 0.66 0.89 0.93
MoGe-2 [50] 14.13° 11.33° 20.01° 0.51 0.87 0.94
MoGe-2 + Poppy (Ours) 10.89° 8.05° 16.41° 0.70 0.92 0.96

Ablation Study

1. Guidance Mechanism Ablation (None vs. Image Offset vs. Joint Guidance) The table below illustrates MAE values across representative synthetic and real test scenes:

Scene / Evaluation Condition Backbone Configuration Baseline (None) Image Offset Only (\(O_x\)) Joint Guidance (\(O_x + O_n\)) Relative Reduction
Deschaintre (Synthetic) Marigold 44.63° 18.87° 17.73° 60.3%
NeISF (Synthetic) Lotus-v2 16.88° 10.99° 10.29° 39.0%
NeRSP (Synthetic) MoGe-2 13.04° 11.65° 10.36° 20.6%
PISR Textureless, Low Albedo (Real) Marigold 27.53° - 16.19° 41.2%
PISR Textureless, Low Albedo (Real) Lotus-v2 15.47° - 12.53° 19.0%
PISR Textureless, Low Albedo (Real) MoGe-2 17.66° - 13.69° 22.5%
NeRSP Low Illumination (Real) Marigold 20.72° - 12.12° 41.5%
NeRSP Low Illumination (Real) Lotus-v2 19.32° - 12.87° 33.4%
NeRSP Low Illumination (Real) MoGe-2 19.82° - 11.66° 41.2%
SfPUEL High Measurement Noise (Real) MoGe-2 12.81° - 7.56° 41.0%

2. Runtime Complexity and Peak Memory Overhead (768×768 resolution, NVIDIA RTX PRO 6000 95.6GB VRAM)

Backbone Architecture Single Backbone Pass Poppy 100-Step Time (Per-step Overhead) Backbone Peak VRAM Poppy Peak VRAM (Memory Overhead)
MoGe-2 (Feed-forward ViT) 3.00 s 57 s (+0.57 s/step) 3.07 GB 15.33 GB (+12.26 GB)
Marigold (Latent Diffusion) 4.02 s 95 s (+0.95 s/step) 7.47 GB 35.12 GB (+27.65 GB)
Lotus-v2 (Flow-matching) 3.33 s 173 s (+1.73 s/step) 38.14 GB 66.64 GB (+28.50 GB)

Key Findings

  • Distinct Division of Labor Between Offset Parameters: The image offset \(O_x\) accounts for the vast majority of MAE reduction across both synthetic and real scenes (accounting for 70%–80% of error reduction as shown in Fig. 5(a)), steering macroscopic surface orientations into correct physical alignment. Activating the normal offset \(O_n\) at step 50 further resolves high-frequency structures, boosting Acc11.25 by 37%–87% on synthetic data.
  • Graceful Degradation Against Sensor Noise: Injecting zero-mean Gaussian noise with increasing standard deviation \(\sigma\) into polarization inputs shows that Poppy yields substantial gains under low-to-moderate noise. As noise variance increases, the physical validity mask and foundation priors prevent catastrophic failure, allowing accuracy to gracefully converge back to the unguided baseline rather than blowing up.
  • Material-Dependent Guidance Gains: Specular surfaces exhibit the largest improvements due to strong, well-defined polarization signals (high DoLP). Diffuse surfaces achieve secondary gains from subtler polarization cues. Mixed diffuse-specular surfaces require accurate radiance decomposition \(L_s\), yielding slightly smaller yet consistently positive refinements.
  • Downstream 3D Mesh Reconstruction Gains: Feeding polarization-refined normals into multi-view surface reconstruction (VCR-GauS) achieves an approximate 6% improvement in Chamfer Distance over unguided RGB baselines, effectively rescuing glossy, textureless geometry where COLMAP feature tracking fails.

Highlights & Insights

  • Adversarial Perturbation Inverted as Physical Guidance: While adversarial attacks exploit input perturbations to mislead network predictions, Poppy turns this sensitivity on its head, using input image offsets as a global steering wheel to realign macroscopic surface normals via the network's multi-scale receptive field.
  • Staged Annealing Optimization Strategy: Under unknown ambient lighting, polarization ambiguities and reflectance mixing are severely coupled. Poppy freezes high-frequency normal offsets during the first 50 iterations to allow global geometry and specular radiance to converge first, cleanly sidestepping chaotic multi-parameter co-adaptation.
  • A Zero-Retraining Paradigm for Physical Modalities: Rather than collecting expensive paired polarimetric datasets to retrain massive vision foundation models, Poppy establishes that frozen foundation models can be dynamically steered at test time using lightweight physical constraints, providing a blueprint adaptable to depth completion, photometric stereo, and multispectral vision.

Limitations & Future Work

  • Computational Overhead from Iterative Backpropagation: Executing 100 optimization steps requires 50 seconds to nearly 3 minutes per single image and demands 12–28 GB of additional VRAM to store intermediate feature activations, hindering real-time robotic deployment. Future work should investigate lightweight neural adapters or amortized feed-forward distillation.
  • Orthographic Projection Approximation Under Extreme Perspective: Assuming parallel orthographic viewing rays introduces geometric bias in close-up, wide-angle captures (such as the PISR dataset). Incorporating explicit camera intrinsics into the ray tracing formulation is necessary to eliminate perspective distortion.
  • Simplified Dielectric Material Assumption: The framework assumes a constant isotropic dielectric refractive index (\(\eta = 1.5\)). Accurately modeling complex refractive media, such as metallic surfaces with complex refractive indices or translucent materials with volumetric subsurface scattering, requires more advanced polarimetric BRDF extensions.
  • vs. SfPUEL [38] / DeepSfP [5]: Learning-based SfP methods require paired polarization-normal supervision and fail to generalize beyond narrow training domains; Poppy requires zero polarimetric training pairs, directly leveraging broad open-world RGB priors to achieve superior cross-dataset generalization.
  • vs. PPFT [29]: PPFT applies prompt fusion tuning to adapt depth models to polarization inputs, modifying model weights and requiring polarization training data; Poppy focuses on surface normals and operates completely at test time while keeping all backbone weights frozen.
  • vs. Marigold-DC [48] / Test-Time Guidance: Marigold-DC steers reverse diffusion with sparse depth points; Poppy extends test-time guidance to dense, non-linear polarization physics, introducing a dual-offset mechanism (\(O_x\) and \(O_n\)) and differentiable reflectance decomposition to resolve azimuthal and diffuse-specular ambiguities.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneering training-free, test-time polarization guidance for surface normals using input/output dual offsets]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluated across 7 benchmark datasets, 3 diverse architectures, synthetic and real objects, with downstream 3D reconstruction]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear mathematical derivation of Fresnel polarization transport, insightful Jacobian analysis, and solid empirical validation]
  • Value: ⭐⭐⭐⭐⭐ [Provides a practical, zero-retraining blueprint for integrating physical sensing modalities into frozen foundation vision models]