Skip to content

DualResPS: Dual-Resolution Photometric Stereo Using a Frame-Event Hybrid Camera

Conference: ECCV 2026
Paper: ECCV Official
Area: 3D Vision
Keywords: Photometric Stereo, Event Camera, Neural Inverse Rendering, Hybrid Camera System, Surface Normal Estimation

TL;DR

DualResPS presents a dual-resolution photometric stereo pipeline utilizing a coaxial frame-event hybrid camera under a continuously rotating light source, integrating sparse long-exposure frames with high-temporal-resolution events via coordinate-based neural inverse rendering to recover high-resolution surface normals at only 18.1% of the data rate.

Background & Motivation

Photometric stereo (PS) is a cornerstone technique in computer vision for high-fidelity 3D geometric surface sensing, estimating fine surface normals from images captured under varying illumination directions. However, conventional frame-based photometric stereo relies heavily on capturing tens or hundreds of discrete images. This fundamental characteristic incurs prolonged acquisition time, heavy data bandwidth and storage consumption, and severely limits time-sensitive or dynamic applications. Recently, neuromorphic event cameras have emerged as an attractive alternative (e.g., EventPS) due to their high dynamic range, microsecond temporal resolution, and bandwidth-efficient sparse data streams. Nevertheless, owing to the circuit complexity of per-pixel event-triggering mechanisms, commercial event cameras are restricted to relatively low spatial resolution (typically around 720p), which fundamentally undermines fine-grained surface normal reconstruction compared to modern high-resolution frame sensors.

Hybrid camera systems combining conventional frame sensors with event sensors offer an intuitive path toward bridging spatial fidelity and temporal agility. However, effectively fusing these two disparate modalities in photometric stereo introduces three major challenges: i) Frame degradation: under directional light sweeps, frame pixels suffer severe signal loss from attached/cast shadows and highlight saturation; ii) Event resolution mismatch: the spatial resolution of event sensors is markedly coarser than high-frequency surface normals, sharp cast shadows, and specular highlights; iii) Temporal misalignment: global integrated exposures of frame cameras are temporally unaligned with continuous, asynchronous event streams.

This work addresses these challenges by fundamentally rethinking the illumination and dual-resolution capture paradigm: instead of taking many short-exposure static frames, it captures only three long-exposure frames concurrently with an event stream across one rotation period of a continuous point light. Core idea: aggregate multi-directional illumination over long exposures into equivalent directional lights, compensate non-ideal cast shadows and specular reflections via high-temporal-resolution events, and optimize surface normals through coordinate-based neural inverse rendering with asynchronous importance sampling.

Method

Overall Architecture

The input to DualResPS consists of three high-spatial-resolution global long-exposure RGB frames \(\{A_\text{f}^{(j)}\}_{j=1}^3\) captured across a rotation period \(T\) of a continuous light source (\(t_j - t_{j-1} = T/3\)), alongside an asynchronous stream of low-spatial-resolution events \(\mathcal{E}\) recorded through a beam splitter sharing a coaxial viewpoint. In the preprocessing phase, continuous cast shadow masks are identified from event profiles, and temporal importance sampling timestamps are determined. During neural optimization, spatial coordinates are fed into a coordinate-based multilayer perceptron (MLP) to output surface normals, diffuse albedos, and specular material parameters, which are then passed to an asynchronous differentiable physical renderer to synthesize integrated frames and logarithmic event changes under joint self-supervised losses.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Input: 3 Long-Exposure Frames + Asynchronous Events"] --> D1["Spatiotemporal Dual-Resolution Lighting Strategy<br/>Continuous sweeping illumination & equivalent light vectors"]
    D1 --> D2["Non-Ideal Shadow & Specular Decoupling<br/>Event interval profile shadow detection & BRDF parameterization"]
    D2 --> D3["Coordinate Neural Inverse Rendering & Importance Sampling<br/>Physics-decoupled MLP & pixel-wise asynchronous sampling"]
    D3 --> Out["Output: High-Resolution Surface Normals & Material Properties"]

Key Designs

1. Spatiotemporal Dual-Resolution Lighting Strategy: Overcoming the Bandwidth-Fidelity Bottleneck To reconcile the low sampling frequency of frame cameras with the spatial coarseness of event cameras, the authors propose a spatiotemporal dual-resolution illumination configuration. Over a rotation period \(T\), the light trajectory is divided into three equal exposure windows \(t_j - t_{j-1} = T/3\). Under the linearity of the time integral and dot product for Lambertian reflectance without shadowing, the temporally accumulated irradiance mathematically reduces to a single equivalent directional light \(\boldsymbol{l}^{*(j)} = \int_{t_{j-1}}^{t_j} \boldsymbol{l}(t) dt\). Three non-coplanar exposures yield three linearly independent equivalent lighting vectors, satisfying the theoretical minimum requirement for closed-form Lambertian normal and albedo recovery. This design replaces dozens of discrete strobed frames with merely three long-exposure captures, compressing the frame bandwidth footprint to an absolute minimum.

2. Non-Ideal Shadow & Specular Decoupling: Mitigating Observation Contamination Physical scene non-idealities—cast shadows, attached shadows, and specular highlights—violate the linear equivalent directional light model. For cast shadows, entering or exiting occlusion creates sharp logarithmic intensity transitions that are captured with microsecond fidelity by the event stream. The method detects cast shadow boundaries via an event interval profile metric \(f_{\boldsymbol{x}}^{(k)} = \frac{p_{\boldsymbol{x}}^{(k)} C}{t_{\boldsymbol{x}}^{(k)} - t_{\boldsymbol{x}}^{(k-1)}}\), inferring continuous binary visibility masks \(m_\text{c}(\boldsymbol{x}, t)\) that are assigned to adjacent high-resolution frame pixels leveraging shadow smoothness. For attached shadows, the thresholding mask \(m_\text{a}(\boldsymbol{x}, t; \boldsymbol{n}) = \mathbb{I}(\boldsymbol{l}(t) \cdot \boldsymbol{n} > 0)\) is explicitly parameterized by the surface normal, preserving normal solvability without extra degrees of freedom. For specular reflections, while three frames cannot separate diffuse and specular components, the dense temporal irradiance curves in adjacent event superpixels constrain the microfacet BRDF parameters \(\boldsymbol{\beta}(\boldsymbol{x}) = [F_0(\boldsymbol{x}), r(\boldsymbol{x})]\) (Fresnel specular color \(F_0\) and roughness \(r\)). This completely strips specular contamination from the frame intensities.

3. Coordinate Neural Inverse Rendering & Importance Sampling: Adaptive Cross-Resolution Bridging To bridge the disparate spatial resolutions without heuristic explicit interpolation, the surface parameters are modeled by a continuous coordinate-based MLP \(\mathcal{M}(\boldsymbol{x}; \Theta)\), outputting normal \(\boldsymbol{n}(\boldsymbol{x})\), diffuse albedo \(\boldsymbol{\rho}_\text{d}(\boldsymbol{x})\), and specular parameters \(\boldsymbol{\beta}(\boldsymbol{x})\). The implicit inductive bias naturally provides spatial smoothness and material sparsity, aligning high-frequency geometry with frame pixels while regularizing low-frequency material attributes through event supervision. To render long-exposure continuous integrals efficiently, a pixel-wise asynchronous importance sampling strategy determines discrete timestamps between adjacent events based on estimated instantaneous irradiance: $\(N_{\boldsymbol{x}, \text{e}}^{(k)} = \left\lceil \mathrm{round}\left( \sigma \sqrt{\Delta \hat{I}_{\boldsymbol{x}}^{(k)} \Delta t_{\boldsymbol{x}}^{(k)}} \right) \right\rceil\)$ Computational power is concentrated on high-intensity and high-fluctuation intervals, while dark, noise-prone event intervals are automatically downweighted by confidence weights \(w_{\boldsymbol{x}}^{(k)}\).

Loss & Training

The entire pipeline is trained self-supervised via gradient descent on the differentiable rendering loss: $\(\mathcal{L} = \lambda_\text{f} \mathcal{L}_\text{f} + \lambda_\text{e} \mathcal{L}_\text{e} + \lambda_\text{tv} \mathcal{L}_\text{tv}\)$ where the frame loss \(\mathcal{L}_\text{f} = \frac{1}{N_\text{f}} \sum |\tilde{A}_\text{f}^{(j)}(\boldsymbol{x}) - A_\text{f}^{(j)}(\boldsymbol{x})|\) penalizes reconstructed pixel exposures, and the event loss \(\mathcal{L}_\text{e} = \frac{1}{N_\text{e}} \sum w_{\boldsymbol{x}}^{(k)} \left| \log\left(\frac{\tilde{I}(\boldsymbol{x}, t_{\boldsymbol{x}}^{(k)}) + \epsilon}{\tilde{I}(\boldsymbol{x}, t_{\boldsymbol{x}}^{(k-1)}) + \epsilon}\right) - p_{\boldsymbol{x}}^{(k)} C \right|\) enforces fidelity on logarithmic irradiance changes across event intervals. A total variation term \(\mathcal{L}_\text{tv}\) regularizes normal and material fields during early warm-up stages. The network is optimized over 1200 epochs on an NVIDIA A6000 GPU.

Key Experimental Results

Main Results

Quantitative comparison (Mean Angular Error in degrees, lower is better) on the DiLiGenT semi-real benchmark dataset across 10 challenging objects. Data rate is normalized against the full 23-frame baseline (100%).

Method Type Input Modality Ball Buddha Cat Cow Goblet Harvest Pot1 Pot2 Reader Average MAE (°) Data Rate
Uni-MS-PS Supervised 4 frames 3.76 7.65 5.53 5.23 6.25 14.71 5.36 4.99 9.05 6.95 17.4%
SDM-UniPS Supervised 4 frames 1.83 7.91 6.60 5.73 6.53 13.29 7.28 5.72 9.29 7.13 17.4%
LINO Supervised 3 frames 2.41 7.58 6.79 6.45 7.97 11.33 6.72 5.69 10.37 7.26 13.0%
LL22a Self-supervised 23 frames 2.26 9.88 5.16 5.29 6.44 18.54 5.40 5.39 10.95 7.70 100.0%
DualResPS (Ours) Self-supervised 3 frames + Events 2.08 8.64 4.86 4.47 6.88 18.01 5.38 5.31 10.59 7.36 18.1%
EventPS Optimization/Self-sup Events only 11.00 19.96 12.10 26.14 18.94 38.29 11.54 15.79 26.18 19.99 5.0%

Real Prototype System & Complex Material Evaluation

Evaluation on real captured prototype data (Hikvision 5MP global-exposure frame camera + Prophesee EVK4 HD event camera via beam splitter) and the near-planar DiLiGenT-Pi dataset with diverse material properties:

Dataset / Test Target Surface Geometry & Material Characteristics DualResPS (Ours) EventPS-OP (Events) LL22a (Frame Baseline) Setup Note
Prototype: LION Moderate sculpture relief details 17.29 25.88 18.66 Frame baseline uses 8 frames
Prototype: KING Rich high-frequency relief structures 22.78 24.90 24.66 Frame baseline uses 8 frames
Prototype: STAR Simple geometry with severe cast shadows 15.12 28.06 26.70 Frame baseline uses 8 frames
DiLiGenT-Pi: COW Smooth surface with subtle specularity 4.47 26.14 5.29 Frame baseline uses 24 frames
DiLiGenT-Pi: READING Dense embossed characters and recesses 10.59 26.18 10.95 Frame baseline uses 24 frames
DiLiGenT-Pi: ASTRO Metallic specularity and roughness transitions 6.92 17.50 8.81 Frame baseline uses 24 frames
DiLiGenT-Pi: CLOUD-R Rough diffuse scattering 12.17 15.81 13.40 Frame baseline uses 24 frames

Key Findings

  • On DiLiGenT, DualResPS attains an average MAE of 7.36° using only 18.1% data bandwidth, outperforming the leading self-supervised frame-based method LL22a (7.70° with 23 frames) and rivaling heavily pre-trained supervised models.
  • Pure event-based EventPS suffers an average MAE of 19.99° due to spatial quantization and noise accumulation, frequently diverging around specularity and shadows; DualResPS anchors geometry with three global frames, effectively stabilizing convergence.
  • On the STAR prototype object, which exhibits severe cast shadows, traditional frame-based (26.70°) and event-based (28.06°) baselines degrade sharply, whereas DualResPS reduces MAE to 15.12° thanks to its fine temporal event shadow decoupling.

Highlights & Insights

  • Equivalent Directional Lighting Integration: Elegantly proves that continuous sweeping illumination over long-exposure windows collapses linearly into equivalent virtual directional lights, mathematically harmonizing sparse temporal frame capture with continuous lighting.
  • Physical Decoupling across Heterogeneous Resolutions: Leverages the insight that cast shadow boundaries and material parameters exhibit low-frequency spatial coherence suitable for event resolution, whereas surface normals and albedo demand high-resolution frame matching, synthesized harmoniously via neural representations.
  • Asynchronous Importance Rendering: Proposes adaptive temporal sampling based on estimated irradiance intervals, which drastically trims backward pass computation and naturally attenuates unreliable noisy dark events.

Limitations & Future Work

  • Absence of Inter-reflection Modeling: Secondary light bounces in non-convex concavities are not explicitly modeled, causing systematic residual errors in recessed regions.
  • Sensitivity to Extrinsic and Sensor Calibration: Self-supervised inverse rendering depends critically on precise beam-splitter alignment, light trajectory calibration, and accurate event contrast thresholds \(C\).
  • Future Directions: Exploring end-to-end feedforward foundation models for hybrid photometric stereo that generalize across complex geometries without requiring per-scene test-time optimization.
  • vs EventPS [Yu et al., CVPR 2024]: EventPS pioneered continuous rotating-light photometric stereo with event cameras but lacks spatial resolution and detail recovery; DualResPS introduces a hybrid beam-splitter camera with three long-exposure frames, eliminating spatial blurring.
  • vs LL22a [Li & Li, CVPR 2022]: LL22a employs neural inverse rendering on dense frame sequences with depth ray marching for shadows; DualResPS extracts explicit temporal shadow boundaries and specular profiles directly from events, surpassing LL22a's accuracy while slashing data requirements by over 80%.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering spatiotemporally dual-resolution photometric stereo with equivalent directional light theory and explicit event-driven non-ideality compensation.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous validation across DiLiGenT, DiLiGenT-Pi benchmarks and real-world 3D-printed prototype measurements.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical formulations, insightful motivations, and precise physical derivations.
  • Value: ⭐⭐⭐⭐⭐ Offers a practical, low-bandwidth, high-fidelity 3D surface sensing paradigm for high-speed industrial inspection and robotics.