Skip to content

SHINE-PPG: Non-Lambertian Intrinsic Decomposition for Illumination-Robust rPPG

Conference: ECCV 2026
Paper: ECCV Official Link
Code: https://github.com/Edmond-Yang/SHINE-PPG
Area: Human Understanding
Keywords: Remote Photoplethysmography (rPPG), Non-Lambertian Intrinsic Image Decomposition, Specular Highlight Separation, Self-Supervised Learning, Adversarial Illumination Enhancement

TL;DR

To overcome severe physiological signal corruption caused by environmental illumination variations and facial specular highlights in camera-based physiological monitoring, SHINE-PPG proposes a self-supervised non-Lambertian intrinsic decomposition framework that explicitly decouples facial videos into diffuse reflectance, environmental illumination, and sparse specular highlights, combined with a learnable AdaIN adversarial illumination enhancement strategy to achieve high-fidelity rPPG signal recovery under challenging out-of-distribution dynamic lighting.

Background & Motivation

Remote photoplethysmography (rPPG) enables non-contact cardiovascular monitoring by capturing subtle skin color variations induced by periodic blood volume changes within superficial facial capillaries using standard cameras. This technique exhibits promising potential across healthcare monitoring, affective computing, and face anti-spoofing. However, underlying rPPG signals are inherently weak and exceptionally sensitive to environmental disturbances. In uncontrolled real-world environments, lighting fluctuations—such as moving shadows, flickering indoor lamps, and sunlight patches during driving—induce intensity variations whose magnitudes exceed genuine blood volume modulations by orders of magnitude. Recent illumination-aware models attempt to mitigate this by either leveraging background regions as environmental priors (e.g., ND-DeeprPPG, Shao et al.) or utilizing classical Retinex theory to decouple illumination from reflectance.

Nevertheless, existing illumination-handling paradigms suffer from two fundamental physical and theoretical limitations. First, background-dependent strategies inherently rely on the specific ambient distributions present in the training set and exhibit severe performance collapse when confronting unseen, complex illumination during inference. Second, existing Retinex-based frameworks fundamentally rely on the Lambertian assumption, treating human facial skin as an ideal diffuse matte surface. In practice, facial skin inevitably exhibits non-Lambertian specular reflections ("shine" or highlights) caused by sweat, sebum secretion, and directional lighting. Under the Dichromatic Reflection Model (DRM), these specular highlights reflect the light source's chromaticity and high intensity, carrying zero pulsatile cardiovascular information. Enforcing a Lambertian decomposition forces these unmodeled specular highlights into the reflectance component, leaving localized white saturation artifacts that destroy fragile physiological waveforms.

To resolve this critical contradiction, this paper moves beyond the conventional Lambertian hypothesis to directly embrace non-Lambertian optical formation. Core idea: decompose facial videos into three physically distinct components—environmental illumination, intrinsic reflectance, and specular highlights—via self-supervised physics-based constraints (dark-channel priors and frequency-domain filtering) to completely isolate rPPG-rich reflectance, combined with a learnable AdaIN adversarial scheme that dynamically synthesizes challenging out-of-distribution lighting to achieve robust generalization.

Method

Overall Architecture

The overall architecture of SHINE-PPG centers on spatio-temporal non-Lambertian intrinsic decomposition and adversarial lighting expansion. Given an input facial video sequence \(V = \{V_t\}_{t=1}^T\) of length \(T\), the framework avoids direct pulse regression and instead decouples observed frames into three physical components via three parallel branches equipped with task-specific feature extractors and symmetric decoders with skip connections: a spatially smooth illumination component \(\bar{L} = \{\bar{L}_t\}_{t=1}^T\) modeling ambient lighting fields, an intrinsic diffuse reflectance component \(\bar{R} = \{\bar{R}_t\}_{t=1}^T\) preserving subcutaneous microvascular pulsations and skin textures, and a sparse specular highlight component \(\bar{H} = \{\bar{H}_t\}_{t=1}^T\) capturing surface oil and sweat reflections. Because cardiovascular pulses reside exclusively in diffuse reflections, the rPPG estimator \(E\) operates solely on the purified reflectance features \(F_R(V)\) to infer heart rate signals \(\bar{S}\).

To ensure stable decomposition without ground-truth physical annotations, color jitter is applied to synthesize paired samples \(V^{\text{aug}}\) to drive reflectance invariance. Concurrently, an adversarial illumination enhancement module \(A\) parameterized by learnable Adaptive Instance Normalization (AdaIN) dynamically maximizes pulse estimation errors to synthesize challenging out-of-distribution illumination components \(\bar{L}_{\text{adv}}\), which are recombined with \(\bar{R}\) to produce augmented training samples, comprehensively enforcing illumination-invariant feature representation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}, 'subGraphTitleMargin': {'top': 8, 'bottom': 16}}}%%
flowchart TD
    IN["Input facial video V<br/>and jittered pair V_aug"] --> DEC["Self-supervised 3-branch intrinsic decomposition<br/>extractors and decoders (F, D)"]
    subgraph S_DEC["Intrinsic Component Decoupling"]
        direction TB
        DEC --> L_COMP["Illumination component L<br/>spatio-temporal lighting"]
        DEC --> R_COMP["Reflectance component R<br/>intrinsic pulse & skin texture"]
        DEC --> H_COMP["Specular highlight H<br/>localized sparse shine"]
    end
    H_COMP --> SPEC_PHY["Physics-grounded specular highlight separation<br/>L1 sparsity + dark-channel KL + color consistency"]
    L_COMP --> FREQ_DEC["Frequency-constrained decomposition<br/>low-pass illumination + high-pass reflectance"]
    R_COMP --> FREQ_DEC
    R_COMP --> REF_INV["Reflectance invariance constraint<br/>L_reflect(R, R_aug)"]
    L_COMP --> ADAIN_ADV["Adversarial illumination enhancement<br/>learnable AdaIN synthesizes extreme OOD L_adv"]
    R_COMP --> REC_ADV["Recompose adversarial sample V_adv = R ◦ L_adv<br/>specular excluded to focus on lighting"]
    ADAIN_ADV --> REC_ADV
    REC_ADV --> PROG_TRAIN["Progressive 4-stage training strategy<br/>Lambertian init → specular isolation → joint → adversarial"]
    R_COMP --> EST_RPPG["Physiological pulse estimator E<br/>extracts clean rPPG S solely from reflectance"]
    EST_RPPG --> OUT_HR["Output: Illumination-robust high-fidelity rPPG signals"]

Key Designs

1. Physics-grounded specular highlight separation: dark-channel prior and color consistency to eliminate facial shine Under the non-Lambertian Dichromatic Reflection Model \(V_t = R_t \circ L_t + H_t\), specular highlights \(H_t\) exhibit high reflection intensity without penetrating the dermis, carrying zero physiological pulses. Because rPPG benchmarks lack ground-truth specular annotations, three physics-based self-supervised constraints are enforced. First, spatial sparsity is penalized via an \(L_1\) norm: $$ \mathcal{L}{\text{sparse}} = \sum}^T \left( |\bar{\mathbf{H}t|_1 + |\bar{\mathbf{H}}_t^{\text{aug}}|_1 \right) $$ Second, observing that highlights produce prominent peaks in the dark channel of RGB images, a normalized dark-channel prior is constructed as a spatial probabilistic distribution \(W_t(p) = \frac{\min_{c \in \{r,g,b\}} \mathbf{V}_{t,c}(p)}{\sum_{p=1}^P \min_{c} \mathbf{V}_{t,c}(p)}\), paired with the normalized magnitude of predicted highlights \(\bar{M}_t(p) = \frac{\|\bar{\mathbf{H}}_t(p)\|_2}{\sum_{p=1}^P \|\bar{\mathbf{H}}_t(p)\|_2}\). The Kullback-Leibler (KL) divergence is minimized to align the estimated highlight distribution with this physical prior: $$ \mathcal{L}) \right) $$ Third, based on the physical principle that specular reflections reflect the chromaticity of the illuminant rather than the surface albedo, a cosine similarity constraint weighted by the dark-channel prior is introduced: }} = \sum_{t=1}^T \left( \text{KL}(\mathbf{W}_t \parallel \bar{\mathbf{M}}_t) + \text{KL}(\mathbf{W}_t^{\text{aug}} \parallel \bar{\mathbf{M}}_t^{\text{aug}\(\mathcal{L}_{\text{color}} = \sum_{t,p} W_t(p) \cdot [1 - \cos(\bar{\mathbf{H}}_t(p), \bar{\mathbf{L}}_t(p))]\). Together, these terms form the total highlight loss \(\mathcal{L}_{\text{specular}}\), enabling the network to strip oily reflections and prevent highlight leakage into the reflectance component.

2. Frequency-constrained decomposition: complementary Fourier low-pass and high-pass filtering for stable physical separation Intrinsic image decomposition is an ill-posed inverse problem; relying solely on logarithmic reconstruction loss \(\mathcal{L}_{\text{rec}} = \sum_{t,p} \|\log(\mathbf{V}_t(p) - \bar{\mathbf{H}}_t(p)) - \log\bar{\mathbf{L}}_t(p) - \log\bar{\mathbf{R}}_t(p)\|_1\) risks trivial solutions or physiological leakage into the illumination component. To regularize this, the framework incorporates a frequency-domain prior: environmental illumination varies smoothly in space, whereas skin reflectance contains sharp anatomical boundaries and high-frequency capillary chromatic variations. A multi-scale frequency constraint is formulated: $$ \mathcal{L}{\text{freq}} = \sum; \sigma) \right) $$ where 2D Fast Fourier Transform (FFT) is paired with a Gaussian low-pass filter }^T \sum_{\sigma \in \mathcal{T}} \left( \mathcal{G}(\bar{\mathbf{R}}_t, \bar{\mathbf{R}}_t^{\text{aug}}; \sigma) + \hat{\mathcal{G}}(\bar{\mathbf{L}}_t, \bar{\mathbf{L}}_t^{\text{aug}\(M(u,v;\sigma) = \exp\left(-\frac{D^2(u,v)}{2\sigma^2}\right)\) with cutoff frequency \(\sigma\). The reflectance penalty \(\mathcal{G}\) applies \(M\) to penalize low-frequency energy in \(\bar{R}\), preserving texture and microvascular modulations. Conversely, the illumination penalty \(\hat{\mathcal{G}}\) applies the complementary high-pass filter \((1 - M)\) to suppress high-frequency components in \(\bar{L}\), enforcing spatial smoothness. Coupled with the cross-augmentation reflectance invariance loss \(\mathcal{L}_{\text{reflect}} = \sum_{t} \|\bar{\mathbf{R}}_t - \bar{\mathbf{R}}_t^{\text{aug}}\|_1\), physiological signals are strictly retained within the reflectance channel.

3. Adversarial illumination enhancement: learnable AdaIN to synthesize extreme out-of-distribution lighting Existing rPPG benchmarks feature relatively static laboratory settings, leading to poor generalization when testing models on dynamic in-car or outdoor recordings. Conventional data augmentations apply simple linear perturbations that fail to cover real-world non-linear lighting variations. This design introduces an adversarial illumination learning scheme using an Adaptive Instance Normalization (AdaIN) layer \(A\) parameterized by learnable scaling \(\alpha\) and shifting \(\beta\): \(\bar{\mathbf{L}}_{\text{adv}} = A(\bar{\mathbf{L}}) = \alpha \cdot \frac{\bar{\mathbf{L}} - \mu_{\bar{L}}}{\sigma_{\bar{L}}} + \beta\). In the adversarial optimization step, the reflectance extractor \(F_R\) and pulse estimator \(E\) are frozen while \(\alpha\) and \(\beta\) are optimized to maximize the negative Pearson correlation error: $$ \mathcal{L}{\text{adv}} = \max) $$ This objective forces } \text{NP}(\bar{\mathbf{S}}^{\text{adv}}, \mathbf{S\(A\) to synthesize the most challenging illumination distributions \(\bar{L}_{\text{adv}}\), which are recombined with intrinsic reflectance as \(V_t^{\text{adv}} = \bar{\mathbf{R}}_t \circ \bar{\mathbf{L}}_t^{\text{adv}}\) (where specular highlights are excluded to prevent synthetic highlight noise). In the subsequent step, \(F_R\) and \(E\) are fine-tuned to minimize estimation error under \(\bar{L}_{\text{adv}}\). This dynamic min-max loop forces the physiological estimator to develop strong invariance to extreme lighting shifts.

4. Progressive 4-stage training strategy: sequential component decoupling and lightweight inference Direct end-to-end joint training of all three decomposition branches and the adversarial module causes severe gradient conflict, causing the highlight and illumination branches to absorb each other's components. A four-stage progressive training strategy is implemented: - Stage 1 (Lambertian Baseline Initialization): Optimizes illumination, reflectance, and rPPG estimator modules using \(\mathcal{L}_{\text{reflect}} + \mathcal{L}_{\text{rec}} + \mathcal{L}_{\text{rPPG}} + \mathcal{L}_{\text{freq}}\) to establish a stable diffuse baseline. - Stage 2 (Specular Isolation): Freezes stage-1 parameters and trains only the specular branch \(\{F_H, D_H\}\) via \(\mathcal{L}_{\text{specular}} + \mathcal{L}_{\text{rec}}\), isolating surface highlights without corrupting the diffuse albedo. - Stage 3 (Joint Refinement): Unfreezes all branches and performs joint back-propagation across all physical loss terms to refine inter-component consistency. - Stage 4 (Adversarial Enhancement): Alternates between maximizing adversarial lighting loss and minimizing fine-tuning pulse estimation error. Crucially, during inference, all auxiliary branches (illumination decoder \(D_L\), highlight network \(\{F_H, D_H\}\), and AdaIN module \(A\)) are completely discarded. Only the reflectance extractor \(F_R\) and estimator \(E\) are executed, delivering the full robustness of non-Lambertian decomposition with zero additional runtime latency or memory overhead.

Loss & Training

The framework balances physical decomposition constraints and physiological regression objectives. Pulse waveform supervision is guided by Negative Pearson Correlation (NP): \(\mathcal{L}_{\text{rPPG}} = \text{NP}(\bar{\mathbf{S}}, \mathbf{S}) + \text{NP}(\bar{\mathbf{S}}^{\text{aug}}, \mathbf{S})\), where \(\text{NP}(x, y) = 1 - \frac{\sum (x - \mu_x)(y - \mu_y)}{\sqrt{\sum (x - \mu_x)^2 \sum (y - \mu_y)^2}}\). Loss balancing hyperparameters are set to \((\lambda_{\text{prior}}, \lambda_{\text{color}}) = 2\), \(\lambda_{\text{sparse}} = 1\), \(\lambda_{\text{adv}} = 1\), \(\lambda_{\text{enh}} = 1\), total highlight loss weight \(\lambda_{\text{specular}} = 5\), and \((\lambda_{\text{reflect}}, \lambda_{\text{rec}}, \lambda_{\text{rPPG}}, \lambda_{\text{freq}}) = 3\). Training is conducted on a single NVIDIA RTX 4090 GPU with AdamW (learning rate \(1 \times 10^{-4}\), batch size 1) for 100 epochs with clips of 320 consecutive frames cropped to \(64 \times 64\).

Key Experimental Results

Main Results

The model is evaluated across five benchmark datasets: UBFC, PURE, COHFACE, MR-NIRP Indoor, and the challenging MR-NIRP Car dataset featuring rapid illumination changes during driving. Performance is measured using Mean Absolute Error (MAE in bpm, lower is better), Root Mean Square Error (RMSE in bpm, lower is better), and Pearson correlation (\(R\), higher is better).

The table below reports intra-domain evaluation results under varying lighting conditions:

Method Type Method UBFC (MAE / RMSE / R) PURE (MAE / RMSE / R) COHFACE (MAE / RMSE / R) Indoor (MAE / RMSE / R) Car (MAE / RMSE / R)
General Supervised PhysFormer (CVPR 22) 1.88 / 4.47 / 0.96 0.43 / 0.57 / 0.99 3.24 / 6.46 / 0.81 1.57 / 3.58 / 0.93 4.72 / 7.92 / 0.45
Lightweight Supervised EfficientPhys (WACV 23) 0.72 / 1.77 / 0.99 0.30 / 0.36 / 0.99 2.75 / 6.99 / 0.81 1.16 / 3.06 / 0.96 4.79 / 9.76 / 0.48
Periodic Attention RhythmFormer (PR 25) 0.49 / 1.37 / 0.99 0.16 / 0.21 / 0.99 1.17 / 3.36 / 0.97 0.67 / 1.57 / 0.98 2.87 / 6.16 / 0.72
State-Space Model RhythmMamba (AAAI 25) 0.70 / 1.02 / 0.99 0.48 / 0.55 / 0.99 7.70 / 14.18 / 0.61 0.46 / 0.66 / 0.99 4.90 / 10.22 / 0.53
Background Prior ND-DeeprPPG (TIP 23) 0.78 / 1.94 / 0.98 0.24 / 0.32 / 0.99 1.56 / 3.95 / 0.93 0.52 / 0.93 / 0.99 4.28 / 8.76 / 0.54
Illumination-Aware Shao et al. (CVPR 25) 0.42 / 0.89 / 0.99 0.18 / 0.24 / 0.99 6.60 / 12.60 / 0.49 0.45 / 0.91 / 0.92 2.69 / 6.03 / 0.87
Lambertian Retinex Chen et al. (CVPRW 23) 0.62 / 2.20 / 0.99 0.19 / 0.23 / 0.99 2.23 / 5.81 / 0.85 0.55 / 1.33 / 0.99 2.16 / 4.90 / 0.78
Low-Light Retinex Yang et al. (TCE 25) 1.18 / 2.01 / 0.99 1.69 / 2.97 / 0.86 12.59 / 16.29 / 0.10 1.79 / 5.81 / 0.87 4.42 / 8.26 / 0.62
Non-Lambertian Intrinsic SHINE-PPG (Ours) 0.34 / 0.47 / 0.99 0.13 / 0.17 / 0.99 1.47 / 4.03 / 0.93 0.32 / 0.45 / 0.99 1.68 / 3.55 / 0.89

The table below reports cross-domain generalization from stable lighting to dynamic driving scenarios:

Method Type Method UBFC \(\to\) Indoor PURE \(\to\) Indoor UBFC \(\to\) Car PURE \(\to\) Car
General Supervised PhysFormer [41] 2.67 / 4.43 / 0.92 0.69 / 1.42 / 0.98 16.33 / 22.85 / 0.11 6.25 / 9.07 / 0.41
Lightweight Supervised EfficientPhys [22] 1.38 / 5.35 / 0.87 0.50 / 0.74 / 0.99 18.89 / 29.85 / 0.04 15.78 / 24.24 / 0.06
Periodic Attention RhythmFormer [43] 1.08 / 3.98 / 0.91 0.21 / 0.27 / 0.99 11.08 / 17.11 / 0.18 4.64 / 8.62 / 0.53
State-Space Model RhythmMamba [44] 0.85 / 1.74 / 0.98 0.56 / 0.77 / 0.99 9.80 / 16.30 / 0.24 9.37 / 16.62 / 0.28
Background Prior ND-DeeprPPG [21] 1.03 / 3.09 / 0.96 0.34 / 0.35 / 0.99 12.36 / 20.64 / 0.28 7.85 / 13.53 / 0.38
Contrastive Prior DD-rPPGNet [10] 1.22 / 2.67 / 0.96 0.63 / 1.50 / 0.98 14.92 / 21.33 / 0.15 7.47 / 11.78 / 0.43
Lambertian Retinex Chen et al. [4] 2.27 / 9.44 / 0.63 0.38 / 0.58 / 0.99 15.86 / 24.01 / 0.01 6.81 / 13.91 / 0.29
Non-Lambertian Intrinsic SHINE-PPG (Ours) 0.30 / 0.42 / 0.99 0.16 / 0.19 / 0.99 7.64 / 14.42 / 0.30 3.24 / 6.59 / 0.74

Ablation Study

Ablations are evaluated on the demanding cross-domain task PURE \(\to\) Car (where \(\sim\)11.9% of facial pixels are specular highlights).

Table 1: Ablation study of decomposition loss terms (evaluating physical constraints and non-Lambertian modeling) | \(\mathcal{L}_{\text{rPPG}}\) | \(\mathcal{L}_{\text{rec}}\) | \(\mathcal{L}_{\text{reflect}}\) | \(\mathcal{L}_{\text{freq}}\) | \(\mathcal{L}_{\text{specular}}\) | MAE (bpm) \(\downarrow\) | RMSE (bpm) \(\downarrow\) | \(R \uparrow\) | Note | |:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|---| | \(\checkmark\) | – | – | – | – | 8.65 | 14.03 | 0.32 | Supervised baseline: fails to isolate pulse under dynamic driving lighting | | \(\checkmark\) | \(\checkmark\) | – | – | – | 9.97 | 17.81 | 0.19 | Unconstrained reconstruction: ill-posedness destroys subtle pulsatile variations | | \(\checkmark\) | \(\checkmark\) | \(\checkmark\) | – | – | 6.23 | 11.52 | 0.39 | + Reflectance consistency: encourages invariant albedo across jittered frames | | \(\checkmark\) | \(\checkmark\) | – | \(\checkmark\) | – | 6.02 | 10.49 | 0.43 | + Frequency constraints: separates smooth illumination from high-frequency albedo | | \(\checkmark\) | \(\checkmark\) | \(\checkmark\) | \(\checkmark\) | – | 4.17 | 8.77 | 0.48 | Lambertian full baseline: unmodeled specular highlights cause residual noise | | \(\checkmark\) | \(\checkmark\) | \(\checkmark\) | \(\checkmark\) | \(\checkmark\) | 3.76 | 7.76 | 0.57 | Full non-Lambertian model: specular isolation yields significant correlation gain |

Table 2: Ablation study of illumination enhancement strategies (on full decomposition baseline) | Enhancement Strategy | MAE (bpm) \(\downarrow\) | RMSE (bpm) \(\downarrow\) | \(R \uparrow\) | Note | |---|---|---|---|---| | Baseline (Table 1 full decomposition, no extra enhancement) | 3.76 | 7.76 | 0.57 | Static training distribution limits OOD robustness | | Color Jitter | 3.56 | 7.74 | 0.58 | Global chromatic perturbation fails to capture complex lighting dynamics | | Noise Injection | 3.73 | 7.80 | 0.58 | Random noise lacks spatial coherence of lighting | | Standard AdaIN | 3.53 | 7.28 | 0.61 | Restricted to feature statistics from the training set | | Learnable Adversarial AdaIN (Ours) | 3.24 | 6.59 | 0.74 | Maximizes pulse error to synthesize worst-case OOD lighting distributions |

Table 3: Physiological signal content in decomposed physical components | Source Fed to Trained Estimator | MAE (bpm) \(\downarrow\) | RMSE (bpm) \(\downarrow\) | \(R \uparrow\) | Physical Finding | |---|---|---|---|---| | Decomposed Illumination (\(\bar{L}\)) | 15.91 | 21.32 | -0.04 | Near-zero correlation confirms illumination absorbs pure environmental noise | | Decomposed Specular Highlights (\(\bar{H}\)) | 18.46 | 23.64 | -0.03 | High error confirms highlights carry zero pulsatile information | | Original Video (\(V\)) | 1.68 | 3.55 | 0.89 | Untangled mixture of signal and noise |

Key Findings

  • The Lambertian assumption is the primary bottleneck under dynamic lighting: Table 1 demonstrates that even with full frequency and reflectance constraints, a Lambertian model achieves \(R = 0.48\); introducing \(\mathcal{L}_{\text{specular}}\) elevates \(R\) to \(0.57\). In real-world driving scenarios where \(11.9\%\) of face pixels are highlights, omitting specular modeling inevitably pollutes the reflectance channel.
  • Ill-posed decomposition requires rigorous physical priors: Blindly adding an unconstrained reconstruction loss \(\mathcal{L}_{\text{rec}}\) drops \(R\) from \(0.32\) to \(0.19\), proving that naive decomposition easily destroys delicate pulse signals. Spatial smoothness in the illumination channel and anatomical texture preservation in reflectance are indispensable.
  • Adversarial optimization discovers true out-of-distribution extremes: Standard augmentations yield modest gains (\(R \le 0.61\)), whereas dynamically optimizing AdaIN parameters via gradient ascent forces the network to learn invariant representations, boosting \(R\) to \(0.74\) and reducing MAE to \(3.24\) bpm.

Highlights & Insights

  • Seamless integration of classical optics with self-supervised deep learning: By formulating the Dichromatic Reflection Model into differentiable constraints (dark-channel prior, illumination chromaticity alignment, and Fourier filtering), SHINE-PPG achieves self-supervised physical decomposition without requiring costly ground-truth component labels.
  • Thoughtful highlight exclusion during adversarial synthesis: In generating adversarial samples \(V_t^{\text{adv}} = \bar{\mathbf{R}}_t \circ \bar{\mathbf{L}}_t^{\text{adv}}\), specular highlights are deliberately omitted. This prevents synthetic specular artifacts from contaminating training and focuses the adversarial generator entirely on modeling complex ambient illumination shifts.
  • Zero-overhead inference deployment: The multi-branch decomposition and adversarial modules are restricted to training. At inference, only the reflectance extractor and pulse estimator are executed, preserving high frame rates suitable for real-time edge devices and automotive driver monitoring systems.

Limitations & Future Work

  • Global AdaIN modulation lacks local spatial granularity: AdaIN operates on global channel statistics, effectively modeling overall brightening, dimming, and color temperature shifts. However, spatially non-uniform lighting—such as dynamic tree shadows and localized light streaks during driving—is not fully captured by global affine parameters. Future work should investigate spatially adaptive adversarial illumination maps.
  • Information loss under extreme saturation: When camera sensors experience severe overexposure (pixel intensities clipping at 255), underlying diffuse reflection is irretrievably lost at the hardware level. The model must rely on spatial neighborhood interpolation, which can introduce estimation variance during rapid head rotations.
  • Frequency cutoff sensitivity across camera resolutions: Fixed Gaussian frequency cutoffs \(\sigma \in \mathcal{T}\) may experience distribution shifts when video resolution or face scale changes dramatically. Adaptive multi-scale Fourier kernels conditioned on face bounding box scale could further improve cross-resolution stability.
  • vs. Background-prior methods (e.g., ND-DeeprPPG, DD-rPPGNet, Shao et al.): Background approaches assume strong correlation between background and facial lighting, but fail when backgrounds are pitch-black, overexposed, or dynamic (e.g., through car windows). SHINE-PPG operates purely on intrinsic facial optics, achieving superior cross-domain generalization.
  • vs. Conventional Retinex methods (e.g., Chen et al., Yang et al.): Existing Retinex pipelines neglect specular shine, allowing highlight residuals to contaminate reflectance. SHINE-PPG introduces a third specular branch to completely eliminate highlight corruption.
  • vs. General temporal architectures (e.g., PhysFormer, RhythmMamba): Relying solely on temporal self-attention or state-space models struggles when strong lighting shifts divert attention weights. SHINE-PPG demonstrates that enforcing physical invariance boundaries before feature extraction delivers superior robustness.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ First framework to introduce self-supervised non-Lambertian three-component intrinsic decomposition to rPPG, breaking free from Lambertian Retinex limitations with rigorous physical grounding.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluations across five benchmarks covering intra-domain, challenging cross-domain, illumination perturbation, ablation studies, and component information content verification.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical derivations, coherent physical reasoning, structured tables, and well-designed figures.
  • Value: ⭐⭐⭐⭐⭐ Highly impactful for automotive safety, mobile healthcare monitoring, and face anti-spoofing under unconstrained real-world illumination.