Skip to content

Tricam-rPPG: A Multimodal Multispectral Dataset for Remote Photoplethysmography

Conference: ECCV 2026
Paper: ECCV Official
Project Page: https://tricam-rppg.github.io
Area: AI Safety
Keywords: Remote Photoplethysmography (rPPG), Multispectral Imaging, Near-Infrared Sensing, Algorithmic Fairness, Cardiovascular Monitoring

TL;DR

To address severe systemic sensing bias where conventional visible-light rPPG degrades significantly on darker skin tones due to melanin absorption, this work introduces Tricam-rPPG, the first synchronized multimodal dataset containing high-resolution RGB, dual-band near-infrared (850 nm and 940 nm), and Meta Aria smart glasses streams, demonstrating that cross-spectral integration substantially improves signal-to-noise ratio and mitigates demographic disparity.

Background & Motivation

Cardiovascular disease (CVD) remains the leading cause of mortality worldwide, with more than three-quarters of deaths concentrated in low- and middle-income countries. Because conventional clinical evaluations are episodic and wearable adoption remains uneven, scalable, non-contact physiological monitoring is urgently needed. Remote photoplethysmography (rPPG) addresses this by extracting microscopic blood volume pulse (BVP) variations from subtle facial skin color fluctuations using standard cameras, showing great promise for telehealth, intensive care units, and driver monitoring systems. Nevertheless, almost all existing rPPG datasets and pipelines rely exclusively on visible-light RGB imaging, leaving systems susceptible to ambient illumination fluctuations, subject motion, and optical skin variations.

The most critical bottleneck lies in persistent racial and skin-tone bias. While commonly attributed to representation imbalance in training corpora, empirical evidence demonstrates that even when state-of-the-art deep learning architectures (such as TSCAN, DeepPhys, and EfficientPhys) are trained on skin-tone-balanced or dark-skin-predominant datasets like MMPD, heart rate estimation errors continue to deteriorate monotonically as skin pigmentation increases. The root cause is a fundamental physical sensing limitation rather than purely data imbalance: human melanin exhibits peak absorption in the visible green spectrum (500–600 nm)—the exact wavelength where blood volume pulsatile changes are most pronounced—thereby severely attenuating the pulsatile optical reflectance in darker skin types.

Overcoming this fairness barrier requires moving beyond the visible spectrum into near-infrared (NIR) wavelengths, where melanin absorption drops significantly. Core idea: establish the Tricam-rPPG benchmark dataset combining synchronized high-frame-rate RGB and dual-band NIR (850 nm and 940 nm) video with Meta Aria periocular/IMU streams and reference PPG ground truth, disentangling algorithmic bias from optical image formation physics and enabling robust multispectral physiological monitoring.

Method

Overall Architecture

Tricam-rPPG is designed as a controlled, physically aligned multispectral platform. The physical acquisition apparatus features three co-located industrial cameras mounted on a rigid support: one standard RGB camera and two monochrome NIR cameras equipped with bandpass filters centered at 850 nm and 940 nm, all operating synchronously under active NIR and controlled color-temperature illumination across 31 subjects. A BIOPAC MP46 clinical-grade finger PPG sensor simultaneously records ground-truth BVP waveforms at 125 Hz, while a subset of 14 participants also wear Meta Aria smart glasses recording 30 Hz periocular NIR video and IMU kinematics. The evaluation framework comprises hardware UTC alignment, facial landmark tracking, color-invariant skin segmentation, and cross-spectral temporal encoder-decoder benchmarking.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multimodal Multispectral Rig<br/>RGB + NIR 850nm + NIR 940nm + Aria Glasses"] --> B["Controlled Task Protocol<br/>Cool & Warm Illuminations + 8 Dynamic Tasks"]
    B --> C["Physical Alignment & Derived Annotations<br/>UTC microsecond sync + 468 landmarks + 760 facial patches"]
    C --> D["Cross-Spectral Temporal Feature Decoding<br/>Multichannel Demucs Encoder-Decoder + Bi-LSTM"]
    D --> E["Fair & Robust Cardiovascular Estimation<br/>Reconstructed BVP waveforms & accurate Heart Rate (HR)"]

Key Designs

1. Co-located Rigid Camera Array and Hardware Synchronization: Eliminating Parallax and Temporal Drift To enable rigorous cross-spectral comparative evaluation at both pixel and regional levels, the platform employs a rigid mounting frame hosting one color RGB camera and two monochrome NIR cameras (filtered at 850 nm and 940 nm). All three sensors record at \(2046 \times 2464\) high resolution and a continuous 60 fps with locked exposure, gain, and camera intrinsics, eliminating non-linear distortions introduced by commercial camera auto-exposure and ISP tone mapping. An in-scene X-Rite ColorChecker palette facilitates post-hoc photometric and color calibration. Video streams and the 125 Hz BIOPAC fingertip PPG reference signals are synchronized post-hoc via UTC timestamps, guaranteeing sub-frame microsecond temporal precision.

2. Dual-Temperature Illumination and Eight Physiological Induction Tasks: Capturing Dynamic Rate and Non-frontal Poses To transcend static frontal recordings, data collection is structured into two sessions under distinct illuminations: Session 1 (S1) under cool lighting (8000 K) and Session 2 (S2) under warm lighting (2700 K). Each recording segment lasts at least five minutes to support meaningful analysis of low-frequency heart rate variability (HRV). The eight distinct experimental tasks cover resting baseline with visual fixation (T1), cognitive load induction via mental arithmetic (T2), low-arousal calm video viewing (T3), sadness-inducing emotional stimulation (T4), static side-view profile posture (T5, enabling profile rPPG analysis), raised-hand baseline (T6, supporting pulse transit time (PTT) analysis across palm and face), reading aloud from Moby Dick (T1 in S2), and high-arousal excitement video viewing (T2 in S2).

3. Multi-tier Metadata and 760 Facial Patch Representations: Plug-and-Play Benchmarking In addition to lossless multi-stream video, the dataset provides comprehensive preprocessed metadata: 468 dense 3D facial landmarks per frame via MediaPipe across both RGB and NIR modalities; per-frame binary skin masks derived using color-invariant segmentation; and a spatial-temporal representation dividing each face into 760 anatomical patches. For each patch, mean channel intensities are computed over time, compressing multi-gigabyte video into compact spatiotemporal physiological response matrices that allow researchers to iterate rapidly on algorithm designs without repeated heavy decoding.

4. Cross-Spectral Channel Recombination and Demucs Temporal Decoder: Exploiting Deep NIR Tissue Penetration To evaluate multi-channel physiological recovery, the authors adapt a Demucs-based temporal encoder-decoder architecture. The network processes five input channels simultaneously (RGBIR: R, G, B, 850 nm, 940 nm) as well as various three-channel subsets (e.g., RG850, RG940, R+IR, G+IR, B+IR), specifically prioritizing the green (G) channel due to its high vascular response in visible light. Because NIR wavelengths undergo substantially lower absorption by epidermal melanin, near-infrared light penetrates deeper into subdermal vascular beds, providing the temporal encoder with clean pulsatile cues that bypass superficial melanin shielding.

Loss & Training

The cross-spectral Demucs model processes temporal sequences across the 760 facial patches. Multi-scale 1D convolutional downsampling layers extract temporal feature representations, followed by a bidirectional LSTM (Bi-LSTM) bottleneck to model long-term cardiac periodicity, and transposed convolutional layers that reconstruct continuous BVP waveforms. Training combines time-domain Mean Squared Error (MSE Loss) against ground-truth fingertip PPG signals with cross-entropy loss over power spectral density (PSD Loss) in the physiological band ($0.75\text}4.0\text{ Hz}$, corresponding to $45\text{240\text{ BPM}$), enforcing both temporal peak alignment and sharp spectral energy concentration.

Key Experimental Results

Main Results

To establish standardized reference baselines, representative supervised deep learning models implemented in the rPPG-Toolbox were evaluated in a cross-dataset zero-shot transfer configuration on Tricam-rPPG. Standard metrics include Mean Absolute Error (MAE, BPM), Mean Absolute Percentage Error (MAPE, %), Pearson correlation coefficient (\(\rho\)), and Signal-to-Noise Ratio (SNR, dB).

Method Training Dataset MAE (BPM) ↓ MAPE (%) ↓ Correlation \(\rho\) ↑ SNR (dB) ↑
TS-CAN [20] UBFC-rPPG 6.16 ± 1.00 7.63 ± 1.15 0.62 ± 0.07 -2.33 ± 0.38
TS-CAN [20] PURE 16.74 ± 2.89 21.09 ± 3.67 0.33 ± 0.09 -3.62 ± 0.31
TS-CAN [20] SCAMPS 37.31 ± 4.37 49.07 ± 5.96 0.09 ± 0.10 -4.90 ± 0.22
PhysNet [48] UBFC-rPPG 19.09 ± 2.02 26.75 ± 3.01 -0.14 ± 0.09 -5.70 ± 0.17
PhysNet [48] PURE 17.30 ± 1.76 23.76 ± 2.57 0.10 ± 0.09 -5.23 ± 0.16
PhysNet [48] SCAMPS 16.38 ± 1.40 22.45 ± 2.05 0.10 ± 0.09 -6.82 ± 0.08
PhysFormer [49] UBFC-rPPG 31.23 ± 2.77 43.88 ± 4.21 -0.29 ± 0.09 -5.46 ± 0.18
PhysFormer [49] PURE 12.81 ± 1.14 17.30 ± 1.69 0.23 ± 0.09 -5.01 ± 0.20
PhysFormer [49] SCAMPS 24.71 ± 2.01 34.03 ± 3.01 -0.04 ± 0.10 -6.54 ± 0.10
DeepPhys [7] UBFC-rPPG 3.85 ± 0.56 4.85 ± 0.71 0.77 ± 0.06 -1.28 ± 0.39
DeepPhys [7] PURE 5.21 ± 0.67 6.61 ± 0.85 0.67 ± 0.07 -1.89 ± 0.40
DeepPhys [7] SCAMPS 6.90 ± 0.77 8.82 ± 1.03 0.51 ± 0.09 -2.95 ± 0.32
EfficientPhys-C [21] UBFC-rPPG 6.30 ± 1.24 8.20 ± 1.70 0.40 ± 0.09 -2.47 ± 0.34
EfficientPhys-C [21] PURE 22.48 ± 3.46 29.26 ± 4.60 0.15 ± 0.09 -3.92 ± 0.27
EfficientPhys-C [21] SCAMPS 69.13 ± 4.24 92.14 ± 6.18 -0.00 ± 0.10 -5.80 ± 0.15

Ablation Study & Skin Tone Disparity Analysis

The study quantifies the severe degradation of visible-light models across skin pigmentation levels (Table 1) and verifies the restorative effect of incorporating NIR channels (Table 2).

Table 1: Performance degradation across skin tones for DeepPhys [7] (Test Set: Tricam-rPPG)

Training Set Skin Tone Group (Fitzpatrick) MAE (BPM) ↓ MAPE (%) ↓ Correlation \(\rho\) ↑ SNR (dB) ↑
UBFC-rPPG [4] Light (I–II) 1.72 ± 0.51 2.21 ± 0.63 0.94 ± 0.06 0.62 ± 0.63
UBFC-rPPG [4] Medium (III–IV) 4.42 ± 0.90 5.12 ± 0.99 0.78 ± 0.09 -1.39 ± 0.57
UBFC-rPPG [4] Dark (V–VI) 6.60 ± 1.58 9.43 ± 2.40 0.45 ± 0.22 -4.78 ± 0.20
PURE [41] Light (I–II) 1.50 ± 0.47 1.98 ± 0.63 0.95 ± 0.05 0.15 ± 0.65
PURE [41] Medium (III–IV) 4.92 ± 0.89 5.82 ± 1.03 0.76 ± 0.10 -1.95 ± 0.55
PURE [41] Dark (V–VI) 13.43 ± 1.65 17.99 ± 2.15 -0.01 ± 0.25 -5.79 ± 0.26
SCAMPS [?] Light (I–II) 3.26 ± 0.83 4.12 ± 1.07 0.82 ± 0.10 -1.61 ± 0.51
SCAMPS [?] Medium (III–IV) 6.45 ± 0.95 7.67 ± 1.11 0.67 ± 0.11 -2.90 ± 0.49
SCAMPS [?] Dark (V–VI) 15.38 ± 2.06 21.27 ± 3.06 -0.42 ± 0.23 -5.73 ± 0.16

Table 2: Relative performance gains of NIR configurations compared to RGB baseline (Demucs architecture)

Spectral Channel Setup Light Skin SNR Gain Light-Medium SNR Gain Medium-Dark SNR Gain Light Skin MAE Reduction Light-Med MAE Reduction (Absolute) Med-Dark MAE Reduction (Absolute)
RG850 (R+G+850nm) +50% +80% +70% -20% -35% (4.66 BPM) -25% (5.04 BPM)
RG940 (R+G+940nm) +10% +100% +85% moderate gain substantial reduction substantial reduction
RGBIR (all 5 channels) +50% +95% +100% -15% -45% (4.17 BPM) -30% (5.60 BPM)

Key Findings

  • Physical Sensing Barrier in Pure RGB Pipelines: Across multiple model paradigms—convolutional attention (DeepPhys), temporal shift attention (TS-CAN), or temporal difference transformers (PhysFormer)—testing on visible RGB alone exhibits severe failure on darker skin tones, with MAE worsening 3- to 10-fold and SNR dropping below -4.7 dB.
  • Pigmentation-Specific NIR Recovery: While NIR channels yield modest SNR improvements on light skin (10%–50%), they produce massive 70%–100% SNR gains on medium and dark skin tones, reducing MAE by 30%–45%.
  • Multispectral Synergy: The full RGBIR configuration delivers the lowest error and highest SNR across all demographic groups. Rather than replacing RGB, NIR channels complement visible bands by providing unattenuated pulsatile information under high melanin concentration without degrading lighter skin performance.

Highlights & Insights

  • Disentangling Systemic Physical Bias from Algorithmic Data Bias: The paper establishes that demographic disparities in rPPG are primarily rooted in sensor-level image formation physics (melanin absorption of green light) rather than sample imbalance alone, providing empirical proof that balancing RGB training sets cannot fundamentally resolve racial bias in rPPG.
  • Multimodal Multi-perspective Co-located Benchmark: Capturing synchronized \(2046 \times 2464\) 60 fps RGB and 850/940 nm NIR alongside Meta Aria glasses periocular views and IMU records, providing the community with a high-fidelity testbed for both cross-spectral fusion and wearable rPPG.
  • Structured 760-Patch Spatiotemporal Representation: Precomputed facial patches allow researchers to investigate cross-spectral physiological architectures directly without the prohibitive computational overhead of processing raw multi-gigabyte 2K videos.

Limitations & Future Work

  • Demographic Scale and Gender Imbalance: While the dataset features 31 subjects with significant representation in darker skin categories (14 subjects in group Dark), the demographic distribution remains male-dominant, especially in the dark-skin cohort (12 males vs. 2 females).
  • Subdermal Tissue Scattering in NIR: Because NIR wavelengths penetrate deeper into biological tissue, multiple forward scattering can induce spatial blurring of pulsatile signals, potentially attenuating high-frequency morphological landmarks such as dicrotic notches.
  • Unconstrained Dynamic Scenarios: Experiments are restricted to indoor seated protocols under controlled illumination, without evaluating extreme athletic exercise or unconstrained outdoor illumination.
  • vs MMPD / VIPL-HR: While MMPD provides mobile video across diverse environments and VIPL-HR offers RGB-D depth modalities, both operate entirely within visible spectra and suffer from melanin attenuation on darker skin. Tricam-rPPG fills this void by introducing dual-band NIR.
  • vs MR-NIRP: MR-NIRP pioneered vehicular NIR rPPG, but was limited to low resolution (\(640 \times 640\)) and 18 participants without 940 nm coverage or controlled color-temperature variation. Tricam-rPPG offers calibrated multi-spectral 2K resolution at 60 fps.
  • Insight: Future wearable and AR/VR systems can integrate lightweight dual-wavelength NIR/RGB micro-sensors paired with adaptive spectral attention mechanisms to achieve equitable, around-the-clock vital signs monitoring regardless of user skin tone.

Rating

  • Novelty: ⭐⭐⭐⭐☆ The first standardized dataset integrating co-located 2K RGB, dual-band NIR (850/940 nm), and Meta Aria glasses for remote cardiovascular monitoring.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarking with five SOTA methods from rPPG-Toolbox, cross-dataset transfer tests, and detailed skin-tone stratification.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous motivation from biological optics principles, detailed acquisition protocol documentation, and clearly structured findings.
  • Value: ⭐⭐⭐⭐⭐ Illuminates fundamental physical biases in computer vision healthcare tools and provides open infrastructure for equitable physiological AI.