Skip to content

Don’t Mask Out the Background! Natural-Light Photometric Stereo via Illumination Reconstruction

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Full-text Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-5632.txt
Area: 3D Vision
Keywords: Photometric Stereo / Natural Illumination / Illumination Reconstruction / Inverse Rendering / 3D Gaussian Splatting

TL;DR

Addressing the severe ill-posedness of uncalibrated natural-light photometric stereo (NaPS) with unknown lighting, this paper discards the standard practice of masking out image backgrounds, leveraging background cues captured under flexible rigid camera-object motions to explicitly reconstruct a spatially aware, near-field 3D Gaussian Splatting (3DGS) illumination emitter, and subsequently optimizes surface normals, Disney BRDF reflectance, and depth via physics-based inverse rendering.

Background & Motivation

Photometric stereo (PS) aims to recover fine-grained surface normal variations and intrinsic reflectance by observing shading variations across multiple images under shifting lighting conditions. Classical PS techniques typically presuppose strictly calibrated directional light sources in specialized laboratory darkrooms, severely restricting their utility in daily environments and practical scanning scenarios. To overcome these operational hurdles, natural-light photometric stereo (NaPS) has emerged as an attractive alternative: it operates under fixed, uncontrolled ambient illumination (such as common indoor lighting) by rigidly coupling the camera and the target object while moving or rotating the pair as a unified rig, thereby inducing rich shading variations without requiring any manual light control.

However, uncalibrated NaPS—where the environmental lighting is entirely unknown—remains an inherently ill-posed inverse problem. Indoor environments are filled with near-field illumination sources such as desk lamps, ceiling light panels, and monitors, which completely violate traditional far-field parallel lighting or low-frequency distant environment map assumptions (e.g., spherical harmonics or spherical Gaussians). These near-field effects induce complex spatial radiance attenuation, cast shadows, and specular highlights that cause severe ambiguities among illumination, geometry, and reflectance. Existing universal learning-based PS methods (such as the UniPS family) attempt to regress surface normals, albedo, and roughness independently through global context features or data-driven priors. In doing so, they often suffer from severe physical inconsistencies, baking shadows into albedo maps. Crucially, almost all traditional and modern PS pipelines habitually segment out the target object and mask out the background as distracting noise, discarding the very observations that directly record the surrounding light distribution, source positions, and radiance.

This paper identifies the key geometric property of the NaPS acquisition rig: since the camera-object relative pose is locked while moving through space, the background regions observed across viewpoints represent a dense, multi-view capture of the static surrounding illumination field. Core idea: instead of masking out the background as a nuisance, exploit it as a direct multi-view observation of the spatial lighting field, explicitly reconstruct a continuous, distance-aware, and radiation-faithful 3D Gaussian Splatting (3DGS) emitter, and convert uncalibrated NaPS into a physically consistent, tractable inverse-rendering joint optimization.

Method

Overall Architecture

The input to the method is a sequence of \(N\) multi-view images \(\{I_i\}_{i=1}^N\) (approximately 100 images) acquired while rotating or translating the camera and object together as a rigid pair under static indoor lighting. The outputs are the target object's surface normal field \(\mathbf{n}\), per-pixel two-parameter Disney BRDF parameters (albedo \(\mathbf{a}\) and roughness \(r\)), and a depth map \(z\). The pipeline operates in two sequential stages: first, background-driven 3D Gaussian illumination reconstruction, where background pixels supervised by SfM camera poses are used within the 3DGRT framework to fit 3D Gaussian primitives, refined through a two-stage color optimization to faithfully capture high-dynamic-range (HDR) radiance; second, emitter-frozen inverse rendering of object geometry and materials, where the reconstructed 3DGS light emitter remains fixed while differentiable Monte Carlo path tracing—powered by Gaussian mixture model (GMM)/vMF light importance sampling and ray-traced mesh self-shadow visibility—jointly solves for normals, reflectance, and depth in an end-to-end self-supervised loop.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Multi-view RGB Image Sequence<br/>Rigid camera-object pair + background views"] --> B["Two-stage 3D Gaussian Emitter Reconstruction<br/>Full scene structure + Otsu luminance HDR tuning"]
    B --> C["Gaussian-guided vMF Light Importance Sampling<br/>GMM on Gaussian centers + MIS combination"]
    C --> D["Occlusion-aware Path Tracing & Disney BRDF Decoupling<br/>Depth mesh cast shadows + diffuse & specular split"]
    D --> E["Multi-task Geometry & Reflectance Optimization<br/>Log-intensity photometric loss + normal-depth cosine alignment + TV regularizers"]
    E --> F["Output High-fidelity Geometry & Materials<br/>Surface normals / Albedo / Roughness / Depth map"]

Key Designs

1. Two-stage 3D Gaussian Emitter Reconstruction: Recovering Faithful HDR Spatial Lighting from Backgrounds Traditional inverse rendering relies on distant environment maps or low-order spherical harmonics, which cannot model inverse-square distance falloff or sharp localized emitter boundaries. Conversely, directly fitting 3D Gaussian Ray Tracing (3DGRT) to high-dynamic-range (HDR) inputs leads to numerical instability and floating artifacts around extreme highlights. To address this, the environment is modeled with \(M\) anisotropic 3D Gaussian primitives (characterized by center \(\boldsymbol{\mu}_k\), covariance \(\boldsymbol{\Sigma}_k\), opacity \(\sigma_k\), and spherical harmonics color \(\mathbf{c}_k\)), optimized via a two-stage training strategy. In Stage 1, all Gaussian parameters are trained strictly over background mask pixels using a robust log-intensity L1 loss: $\(\mathcal{L}_{\text{stage1}} = \frac{1}{N_b}\sum_{j=1}^{N_b}\bigl|\log(\hat{\mathbf{m}}_j + 1) - \log(\mathbf{m}_j + 1)\bigr|\)$ where \(N_b\) is the number of background pixels, and \(\hat{\mathbf{m}}_j\) and \(\mathbf{m}_j\) denote observed and rendered colors. This reliably stabilizes 3D spatial geometry without blowing up on saturated pixels. Because natural indoor radiance is overwhelmingly dominated by sparse, high-luminance light sources (e.g., luminaires, windows), Stage 2 applies Otsu's bi-level thresholding on the luminance \(Y\) derived from each primitive's RGB color to isolate the subset of high-luminance Gaussians. Freezing all spatial shapes and low-luminance primitives, Stage 2 refines solely the color parameters \(\mathbf{c}\) of these intense emitters via a stop-gradient normalized loss: $\(\mathcal{L}_{\text{stage2}} = \frac{1}{N_b}\sum_{j=1}^{N_b}\frac{\bigl|\hat{\mathbf{m}}_j - \mathbf{m}_j\bigr|}{\text{sg}(\mathbf{m}_j) + \epsilon}\)$ This explicit radiometric fitting ensures the reconstructed emitter outputs accurate physical flux for downstream shading.

2. Gaussian-guided vMF Light Importance Sampling: Slashing Monte Carlo Path Tracing Variance Evaluating the rendering equation under complex natural illumination requires hemisphere integration of incoming radiance at each surface point. However, indoor environmental illumination is highly directional and concentrated in compact emitter regions. Standard cosine-weighted hemisphere sampling or pure BRDF sampling wastes the majority of Monte Carlo (MC) ray budgets on dark background directions, creating extreme gradient noise and high variance. To steer samples toward bright, nearby emitters, the method fits a Gaussian Mixture Model (GMM) with \(K=64\) components directly onto the 3D spatial center coordinates of the reconstructed Gaussians. For a given surface shading point \(\mathbf{x}\), each GMM component's center \(\boldsymbol{\mu}_i\) and concentration proxy \(\kappa_i\) are projected onto the local unit direction sphere, establishing a von Mises–Fisher (vMF) directional mixture: $\(\hat{\boldsymbol{\mu}}_i = \frac{\boldsymbol{\mu}_i - \mathbf{x}}{\|\boldsymbol{\mu}_i - \mathbf{x}\|}, \quad \hat{\kappa}_i = s \cdot \frac{\kappa_i}{\|\boldsymbol{\mu}_i - \mathbf{x}\|}, \quad \hat{\lambda}_i = \lambda_i\)$ where the sharpness factor \(s\) is set to 10, naturally allocating sharper concentration to emitters closer to the surface. Combining this vMF emitter distribution with surface cosine sampling via Multiple Importance Sampling (MIS) concentrates rays along dominant illumination directions while maintaining smooth diffuse integration, achieving low-variance rendering with only 256 samples per pixel.

3. Occlusion-aware Path Tracing & Disney BRDF Decoupling: Disentangling Intrinsic Material from Cast Shadows Under near-field lighting, non-convex objects exhibit severe cast shadows and view-dependent specular glints. Neglecting occlusions forces inverse rendering to bake dark cast shadows into albedo textures or mistake specular glints for surface normal perturbations. The framework models surface interaction via a two-parameter Disney BRDF, split into a diffuse albedo component \(\mathbf{a}(\mathbf{x}) \in [0, 1]^3\) and an isotropic GGX microfacet specular term \(f_s(\mathbf{x}, \boldsymbol{\omega}_o, \boldsymbol{\omega}_i, r(\mathbf{x}))\) governed by per-pixel roughness \(r(\mathbf{x}) \in [0, 1]\). Crucially, at every optimization epoch, the current depth map \(z\) is triangulated into a continuous 3D surface mesh. For each sample direction \(\boldsymbol{\omega}_i\), ray-mesh intersection determines a binary visibility scalar \(V(\mathbf{x}, \boldsymbol{\omega}_i) \in \{0, 1\}\). Modulating incident radiance by visibility yields \(V(\mathbf{x}, \boldsymbol{\omega}_i) L_i(\mathbf{x}, \boldsymbol{\omega}_i)\), incorporating sharp cast shadows directly into the differentiable forward path and preventing illumination artifacts from polluting estimated albedo.

4. Multi-task Geometry & Reflectance Optimization: Building a Self-Supervised Physical Loop To jointly constrain the high-dimensional parameter space of surface normals, albedo, roughness, and depth without collapsing into local noisy minima, the inverse rendering objective integrates photometric fidelity, normal-depth differential consistency, and spatial smoothness: $\(\mathcal{L} = \mathcal{L}_{\text{color}} + \mathcal{L}_{\text{normal}} + \lambda_{\text{smooth}}\mathcal{L}_{\text{smooth}}\)$ The photometric term \(\mathcal{L}_{\text{color}}\) applies a log-intensity L1 loss over the target object mask \(N_f\) to balance gradients across dynamic highlights and shadowed areas. The geometric term \(\mathcal{L}_{\text{normal}}\) enforces cosine alignment between the directly optimized surface normal \(\mathbf{n}\) and a pseudo-normal \(\hat{\mathbf{n}}\) computed via analytical spatial differentiation of the depth map \(z\): $\(\mathcal{L}_{\text{normal}} = \frac{1}{N_f}\sum_{j=1}^{N_f}\left(1 - \frac{\hat{\mathbf{n}}_j^\top \mathbf{n}_j}{\|\hat{\mathbf{n}}_j\|\|\mathbf{n}_j\|}\right)\)$ Total variation regularizers \(\mathcal{L}_{\text{smooth}} = \mathrm{TV}(\mathbf{a}) + \mathrm{TV}(r) + \mathrm{TV}(z)\) suppress high-frequency checkerboard noise, with \(\lambda_{\text{smooth}}=0.01\) active during the first 50 epochs to establish stable global geometry before being set to 0 for the final 50 epochs to capture delicate surface textures.

Loss & Training

All optimization routines run on a single NVIDIA RTX 6000 Ada GPU. Background 3DGS emitter training takes 30K iterations per stage, using default 3DGRT learning rates for Stage 1 and a learning rate of \(1 \times 10^{-1}\) for high-luminance Gaussian colors in Stage 2. Object geometry and reflectance inverse rendering runs for 100 epochs (batch size 4, totaling ~2 hours) with 256 ray samples per pixel. Using the Adam optimizer, initial learning rates are \(5 \times 10^{-4}\) for normals \(\mathbf{n}\), \(1 \times 10^{-2}\) for albedo \(\mathbf{a}\), \(1 \times 10^{-2}\) for roughness \(r\), and \(1 \times 10^{-4}\) for depth \(z\), each halved every 25 epochs. Albedo and roughness heads employ sigmoid activations to guarantee valid \([0, 1]\) ranges. Normal and roughness maps are initialized from LINO-UniPS predictions, while depth is initialized via bilateral normal integration and continuously updates the ray-tracing visibility mesh throughout training.

Key Experimental Results

Main Results

Quantitative evaluations are conducted on a synthetic benchmark (6 diverse geometries: SPHERE, BEAR, BUDDHA, COW, POT2, READING; across 3 photorealistic indoor lighting environments: LIVING ROOM, STAIRCASE, WHITE ROOM, using 150 HDR views at \(512 \times 512\)) and real-world captures (CAR, SNOWMAN, LION across darkroom and natural indoor environments with bracketed exposure HDR capture, ~130 views). Metrics include Mean Angular Error for surface normals (MAngE in degrees, lower is better), Peak Signal-to-Noise Ratio for albedo (PSNR in dB, higher is better), and Mean Absolute Error for roughness (MAbsE, lower is better). Baselines include state-of-the-art universal photometric stereo networks SDM-UniPS and LINO-UniPS.

Method SPHERE
Norm↓ / Alb↑ / Rgh↓
BEAR
Norm↓ / Alb↑ / Rgh↓
BUDDHA
Norm↓ / Alb↑ / Rgh↓
COW
Norm↓ / Alb↑ / Rgh↓
POT2
Norm↓ / Alb↑ / Rgh↓
READING
Norm↓ / Alb↑ / Rgh↓
Average
Norm↓ / Alb↑ / Rgh↓
SDM-UniPS [15] 2.94 / 27.9 / 0.1070 4.05 / 34.4 / 0.0305 6.52 / 26.2 / 0.1730 4.48 / 29.3 / 0.0700 4.91 / 23.5 / 0.0603 6.60 / 27.8 / 0.0573 4.92 / 28.2 / 0.0830
LINO-UniPS [23] 3.47 / 28.9 / 0.0728 5.81 / 33.7 / 0.0189 7.52 / 24.8 / 0.1090 5.64 / 31.3 / 0.0896 6.36 / 22.6 / 0.0556 7.90 / 29.2 / 0.0502 6.12 / 28.4 / 0.0659
Ours 1.38 / 34.1 / 0.0341 2.97 / 30.2 / 0.0536 7.31 / 29.3 / 0.1440 1.53 / 30.4 / 0.0276 2.57 / 34.1 / 0.0422 4.32 / 36.8 / 0.0643 3.35 / 32.5 / 0.0610

On real-world objects, qualitative results mirror synthetic findings: baseline feedforward models suffer from severe shading bake-in in albedo maps due to unconstrained near-field lighting, whereas this method extracts clean, shadow-free albedo and recovers roughness maps that correspond closely with visual specular gloss.

Effect of Viewpoint Count

Evaluating performance scaling across varying input view counts (40, 80, 150 views) on the synthetic dataset:

#views SDM-UniPS [15]
Normal↓ / Albedo↑ / Roughness↓
LINO-UniPS [23]
Normal↓ / Albedo↑ / Roughness↓
Ours
Normal↓ / Albedo↑ / Roughness↓
40 views 4.84 / 27.7 / 0.0778 5.96 / 27.9 / 0.0601 5.10 / 23.6 / 0.1250
80 views 4.79 / 28.2 / 0.0822 6.02 / 28.2 / 0.0659 3.59 / 32.1 / 0.0676
150 views 4.92 / 28.2 / 0.0830 6.12 / 28.4 / 0.0659 3.35 / 32.5 / 0.0610

Ablation Study

Ablation analysis on the synthetic benchmark evaluating the impact of emitter-guided vMF light sampling and the two-stage 3DGS emitter optimization:

Config Normal MAngE ↓ Albedo PSNR ↑ Roughness MAbsE ↓ Note
Full model (Ours) 3.35° 32.5 dB 0.0610 Complete pipeline with two-stage 3DGS & vMF/MIS sampling
w/o light sampling 3.78° 32.2 dB 0.0731 Replaced with standard MIS baseline (cosine + GGX sampling)
w/o two-stage 3.59° 25.9 dB 0.0684 Single-stage 3DGS without high-luminance color refinement

Key Findings

  • Surpassing Feedforward Learning Baselines: The method achieves an average surface normal angular error of 3.35°, outperforming SDM-UniPS (4.92°) by 31.9% and LINO-UniPS (6.12°) by 45.3%. Albedo PSNR jumps from 28.2 dB to 32.5 dB (+4.3 dB gain). Feedforward models lack physical light propagation constraints and predict parameters independently, whereas the inverse rendering formulation guarantees rigorous physical consistency.
  • Criticality of Two-Stage HDR Fitting: Removing two-stage emitter optimization causes a massive 6.6 dB drop in albedo PSNR (plunging from 32.5 dB to 25.9 dB). Single-stage training fails to capture high-dynamic-range radiance of intense luminaires, causing the emitter intensity to be severely underestimated; inverse rendering compensates by artificially boosting surface albedo, introducing severe color degradation.
  • Positive Scaling with Background Coverage: While learning-based baselines plateau regardless of view count (remaining flat from 40 to 150 views), the proposed method benefits directly from view density: more background views yield a more complete, gap-free 3D spatial illumination field, driving albedo from 23.6 dB to 32.5 dB and normal error down to 3.35°.

Highlights & Insights

  • Turning "Background Nuisance" into an Illumination Asset: Overturning four decades of conventional PS wisdom that automatically masks out background pixels, this work demonstrates that in rigidly coupled NaPS acquisitions, the background serves as the ideal multi-view sensor observing the surrounding near-field lighting environment.
  • Principled Bridge Between 3D Gaussians and vMF Importance Sampling: Instead of relying on brute-force ray queries, clustering 3D Gaussian centers into a spatial GMM and projecting them into directional vMF distributions enables fast, low-variance path tracing tailored to complex natural lighting.
  • Generalizable Dual-Stage Paradigm: The pipeline of "reconstructing spatial environmental radiance \(\to\) freezing emitters \(\to\) differentiably decomposing foreground reflectance and geometry" is readily applicable beyond photometric stereo to multi-view neural relighting and handheld object digitization.

Limitations & Future Work

  • Omission of Inter-Reflections: The forward rendering model accounts for direct illumination and cast shadows, but omits indirect light bouncing between object surfaces, causing residual inaccuracies on concave geometries.
  • Computational Overhead of Differentiable Ray Tracing: Iterative inverse rendering takes approximately 2 hours on an RTX 6000 Ada GPU, which is substantially slower than near-instantaneous forward inference in neural networks.
  • Future Directions: Future research could incorporate neural radiance caching to accelerate Monte Carlo convergence and explore learned priors to compensate for high-order indirect bounces.
  • vs SDM-UniPS [15] / LINO-UniPS [23]: These universal learning-based methods estimate normals and materials from segmented foregrounds using global lighting tokens. Without explicit spatial ray tracing, they frequently bake near-field lighting gradients into albedo maps; this work provides ground-truth physical light transport, achieving superior disentanglement.
  • vs Spin-UP [26]: Spin-UP operates under natural lighting but requires a turntable setup rotating about a single axis and relies on object silhouette heuristics to guess light directions; the proposed method supports arbitrary 6-DoF rigid motion of the camera-object rig and reconstructs full spatial illumination directly from background pixels.
  • vs NeRF-based Emitter Inverse Rendering (e.g., Ling et al. [28] / EnvGS [41]): Ling et al. use volumetric NeRF emitters that incur prohibitive ray-marching costs and struggle with high-frequency lights. EnvGS focuses on view synthesis. By combining anisotropic 3D Gaussians with two-stage Otsu HDR optimization, this approach achieves a lightweight, spatially compact, and importance-sampleable near-field emitter representation.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneering paradigm shift transforming background pixels into an explicit 3DGS light emitter for uncalibrated NaPS]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Rigorous evaluation across 6 synthetic geometries, 3 real-world objects, view-count sensitivity, and component ablations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Compelling motivation, elegant mathematical formulation, and exceptionally clear pipeline diagrams]
  • Value: ⭐⭐⭐⭐⭐ [Offers a robust, laboratory-free inverse rendering foundation for high-fidelity 3D scanning in unconstrained environments]