Skip to content

RPM-Distill: Physiology-guided Adaptive Cross-modal Distillation for Robust Remote Physiological Measurement

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/WJULYW/RPM-Distill
Area: Others
Keywords: Remote Physiological Measurement, Cross-modal Knowledge Distillation, Radio Frequency Radar, Bilevel Meta-learning, Spectral Physiological Prior

TL;DR

Addressing the vulnerability of video-based remote physiological measurement under challenging lighting, skin tone, and motion, as well as the high deployment overhead of radar hardware at inference, RPM-Distill leverages synchronized radio-frequency radar as privileged training information, decouples spectral physiological constraints (fundamental peak, background noise, and morphology sharpness), and employs a bilevel meta-learned spectral policy network to adaptively gate and reweight distillation, achieving robust video-only physiological measurement.

Background & Motivation

Video-based remote photoplethysmography (rPPG) recovers vital signs, such as heart rate (HR) and blood volume pulse (BVP), in a non-contact, unobtrusive manner, holding enormous promise for daily healthcare monitoring and intelligent in-cabin driver sensing. However, camera-based rPPG fundamentally relies on detecting minute, quasi-periodic optical intensity variations reflected from facial capillary blood flow. In unconstrained real-world environments, motion-induced non-rigid facial deformations, severe ambient lighting variations, and optical attenuation in dark skin tones drastically deteriorate the signal-to-noise ratio (SNR), leading to severe temporal waveform distortion and widespread spurious peaks across the frequency spectrum. Conversely, radio-frequency (RF) radar senses cardio-respiratory activity by measuring nanometer-to-micrometer chest wall micro-displacements via phase modulation, rendering it intrinsically invariant to ambient illumination and skin pigmentation. To overcome optical fragility, recent studies predominantly resort to multimodal video-radar fusion; however, these approaches necessitate physical radar hardware during inference, severely hampering ubiquitous deployment on standard consumer cameras.

Treating radar as privileged information (LUPI) to guide a video student via cross-modal knowledge distillation (KD) during training is an attractive alternative, but directly adapting conventional distillation paradigms exposes three fundamental physical and physiological tensions: First, video surface reflectance and radar phase modulation stem from fundamentally disparate physical sensing mechanisms. Consequently, their intermediate representations and time-domain waveforms are incommensurate, meaning that naive feature- or time-domain alignment induces severe modality mismatch and negative transfer. Second, distillation must account for sample-level reliability; factors like radar multipath interference, transient body movements, and sensor clock drift introduce substantial variance into teacher quality, such that indiscriminate matching propagates corrupted guidance. Third, failure modes in remote physiological measurement are highly structured in the frequency domain—manifesting as fundamental peak drift, elevated broadband noise floors, and spectral smearing—which cannot be resolved by an unguided global distance metric.

This paper's core insight is that while RGB and RF signals exhibit distinct physical morphology in the time domain, they are modulated by the identical underlying cardiovascular rhythm and thus share a consistent latent periodic structure within the band-limited physiological frequency band (e.g., 45–180 bpm). Core idea: Anchor cross-modal privileged knowledge transfer within the band-limited spectral domain by decomposing supervision into three physiology-structured losses—fundamental peak alignment, off-peak background noise suppression, and spectral morphology sharpness consistency—while employing a bilevel meta-optimized spectral policy network to dynamically gate and reweight distillation per sample, thereby securing robust video-only inference.

Method

Overall Architecture

RPM-Distill introduces a privileged cross-modal distillation framework where a synchronized radar teacher guides a video student during training, while requiring only facial RGB video at inference. During training, the framework receives paired facial video clips \(v_i \in \mathbb{R}^{T \times H_v \times W_v \times C}\) and RF range-time phase representations \(r_i \in \mathbb{R}^{T \times 2H_r}\). A frozen, pretrained radar teacher network \(R(\cdot; \theta_R)\) outputs physiological pulse waveform predictions \(\hat{y}^r\), while the video student network \(S(\cdot; \theta_S)\) generates video pulse predictions \(\hat{y}^v\). Both predicted waveforms undergo Discrete Fourier Transform (DFT) to yield band-limited, z-score normalized log-power spectra \(\ell^r\) and \(\ell^v\).

To prevent negative transfer stemming from time-domain morphological mismatch, spectral supervision is decomposed into three complementary physiological constraints: fundamental peak anchoring loss \(\mathcal{L}_{\text{peak}}\), off-peak noise suppression loss \(\mathcal{L}_{\text{off}}\), and spectral morphology consistency loss \(\mathcal{L}_{\text{shape}}\). Concurrently, the student spectrum, teacher spectrum, and their element-wise absolute difference are stacked into a three-channel spectral relation map \(X_{\text{policy}}\), which is fed into a lightweight spectral policy network \(\Phi(\cdot; \phi)\) to dynamically predict sample-level distillation gates \(g\) and component weights \(\boldsymbol{\alpha}\). The policy network is optimized via a bilevel meta-learning scheme on a small labeled validation set, guaranteeing that teacher guidance is injected only when empirically beneficial.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400, 'subGraphTitleMargin': {'top': 8, 'bottom': 16}}}}%%
flowchart TD
    subgraph INPUT["Input & Feature Extraction Stage"]
        direction TB
        IN_V["Facial Video Input<br/>v ∈ R^{T×H×W×C}"] --> S_NET["Video Student Network<br/>S(·; θ_S)"]
        IN_R["Radar Time-Series Input<br/>r ∈ R^{T×2H_r}"] --> T_NET["Frozen Radar Teacher<br/>R(·; θ_R)"]
        S_NET --> SPEC_V["Student Band-limited Log-Spectrum<br/>ℓ^v = log(P^v + ε)"]
        T_NET --> SPEC_R["Teacher Band-limited Log-Spectrum<br/>ℓ^r = log(P^r + ε)"]
    end

    subgraph POLICY_STAGE["Spectral Relation Map & Adaptive Gating Policy"]
        direction TB
        SPEC_V & SPEC_R --> REL_MAP["Spectral Relation Map<br/>X_policy = [ℓ^v, ℓ^r, Δ]^T"]
        REL_MAP --> POLICY_NET["Spectral Policy Network<br/>1D Residual Conv + Matrix Decomposition"]
        POLICY_NET --> OUT_WEIGHTS["Predict Dynamic Coefficients<br/>Gate g ∈ (0,1) & Weights α ∈ Δ^2"]
    end

    subgraph DISTILL_STAGE["Physiology-Structured Spectral Distillation Constraints"]
        direction TB
        SPEC_V & SPEC_R --> PEAK_NODE["Fundamental Peak Alignment<br/>Gaussian soft window on teacher peak"]
        SPEC_V & SPEC_R --> OFF_NODE["Background Noise Suppression<br/>Out-of-window energy suppression"]
        SPEC_V & SPEC_R --> SHAPE_NODE["Spectral Morphology & Sharpness<br/>Centroid drift penalty + entropy matching"]
    end

    OUT_WEIGHTS --> WEIGHTED_LOSS["Adaptive Weighted Distillation Objective<br/>L_distill = g · Σ α_j L_j"]
    PEAK_NODE & OFF_NODE & SHAPE_NODE --> WEIGHTED_LOSS

    subgraph META_LOOP["Bilevel Meta-Optimization Closed Loop"]
        direction TB
        WEIGHTED_LOSS --> VIRTUAL_STEP["Virtual Student Forward Step on D_tr<br/>θ'_S = θ_S - η ∇ L_train"]
        VIRTUAL_STEP --> VAL_EVAL["Evaluate Physiological Fidelity on D_val<br/>Compute L_sup(θ'_S; D_val)"]
        VAL_EVAL --> META_UPDATE["Policy Backward Update<br/>φ* = φ - β ∇_φ L_sup"]
        META_UPDATE --> FINAL_STUDENT["Actual Student Parameter Update<br/>θ_S* = θ_S - η ∇ L_train(θ_S, φ*)"]
    end

Key Designs

1. Physiology-Structured Spectral Distillation: Overcoming Cross-modal Physical Incommensurability and Feature Gaps To circumvent the fundamental physical barrier where optical reflectance and radar phase micro-motion cannot align in the time domain or intermediate feature spaces, RPM-Distill grounds cross-modal knowledge transfer strictly within the discrete log-power spectrum across the physiological band \(\mathcal{B} = [45, 180]\text{ bpm}\). Defining the element-wise absolute spectral discrepancy as \(\Delta = |\boldsymbol{\ell}^v - \boldsymbol{\ell}^r| \in \mathbb{R}^K\), the method decomposes complex spectral corruption into distinct physiological terms. Using the teacher's normalized spectral distribution \(\mathbf{s}^r = \text{softmax}(\boldsymbol{\ell}^r)\), it computes the differentiable soft spectral centroid \(\hat{k}^r = \sum_{k \in \mathcal{B}} k \cdot \mathbf{s}^r[k]\) and establishes a Gaussian soft window mask centered at the teacher's peak: $$ \mathbf{m}{\mathrm{pk}}[k] = \exp\left(-\frac{(k - \hat{k}^r)^2}{2\sigma^2}\right), \quad \mathbf{m}}} = \max(0, 1 - \mathbf{m{\mathrm{pk}}) $$ Based on these complementary masks, three targeted physiological objectives are constructed: Fundamental peak alignment \(\mathcal{L}_{\text{peak}} = \frac{\|\Delta \odot \mathbf{m}_{\mathrm{pk}}\|_1}{\|\mathbf{m}_{\mathrm{pk}}\|_1 + \epsilon}\) rectifies student peak shifts caused by poor illumination or dark skin tones; background noise suppression \(\mathcal{L}_{\text{off}} = \frac{\|\Delta \odot \mathbf{m}_{\mathrm{oth}}\|_1}{\|\mathbf{m}_{\mathrm{oth}}\|_1 + \epsilon}\) penalizes out-of-band energy elevations induced by motion artifacts, preserving spectral contrast; and spectral morphology consistency \(\mathcal{L}_{\text{shape}}\) penalizes centroid drift while matching distribution sharpness via normalized entropy: $$ H(\mathbf{s}) = -\frac{1}{\log K} \sum}} \mathbf{s}[k] \log(\mathbf{s}[k] + \epsilon), \quad \mathcal{L{\text{shape}} = \text{Smooth}^r)) $$ This design explicitly maps distillation onto heart rate localization, background denoising, and rhythm concentration, completely avoiding destructive temporal waveform distortions.}(\hat{k}^v, \hat{k}^r) + \text{Smooth}_{L1}(H(\mathbf{s}^v), H(\mathbf{s

2. Matrix-Decomposition Spectral Policy Network: Sample-Adaptive Cross-modal Dependency and Quality Perception Because RF signals can suffer from multipath interference and cross-modal acquisition may exhibit temporal synchronization drift, uniform distillation risks enforcing corrupted teacher signals onto the student. To resolve this, RPM-Distill devises an adaptive gating network taking as input the localized spectral relation map \(X_{\text{policy}} = [\boldsymbol{\ell}^v, \boldsymbol{\ell}^r, \Delta]^T \in \mathbb{R}^{3 \times K}\). This tensor explicitly retains frequency-aligned cross-modal correlations. The network encoder applies 1D convolutions with Group Normalization residual blocks to extract multi-scale spectral representations. Subsequently, a matrix-decomposition decoder models basis and coefficient attentions across the frequency axis to distill low-rank signal quality tokens.

These decoded tokens are fused with globally pooled features and projected via dual linear heads to generate a global distillation gate \(g = \text{Sigmoid}(h_g) \in (0, 1)\) and component weights \(\boldsymbol{\alpha} = \text{Softmax}(h_\alpha) \in \Delta^2\). The overall distillation loss is formulated as: $$ \mathcal{L}{\text{distill}}(\hat{\mathbf{y}}^v, \hat{\mathbf{y}}^r; \phi) = g \cdot \sum_j $$ When the teacher waveform is degraded by severe subject movement or sensor misregistration, the policy outputs a gate value near zero, proactively severing erroneous gradient flow. Conversely, under dim lighting where radar remains pristine, the network increases }, \text{off}, \text{shape}}} \alpha_j \mathcal{L\(g\) and heavily weights \(\alpha_{\text{peak}}\), precisely guiding the student model back to the true cardiac rhythm.

3. Bilevel Meta-Optimization: Closed-Loop Distillation Tuning Without Explicit Gating Ground Truth During training, explicit ground-truth supervision for optimal sample gates \(g\) and component weights \(\boldsymbol{\alpha}\) does not exist. Manual heuristic tuning cannot accommodate dynamic driving scenarios or fluctuating lighting. Hence, RPM-Distill casts policy parameter learning into a bilevel meta-optimization formulation. The outer objective minimizes the student's supervised physiological waveform loss \(\mathcal{L}_{\text{sup}}\) (combining negative Pearson correlation and SNR loss) on a held-out labeled validation split \(\mathcal{D}_{\text{val}}\), subject to the inner objective optimizing student parameters \(\theta_S\) on training set \(\mathcal{D}_{\text{tr}}\).

To maintain computational feasibility, an online first-order alternating approximation is employed: First, a virtual gradient descent step on a training batch under current policy \(\phi\) produces temporary student parameters \(\theta'_S = \theta_S - \eta \nabla_{\theta_S} \mathcal{L}_{\text{train}}(\theta_S, \phi; \mathcal{D}_{\text{tr}})\). Second, the virtual student \(\theta'_S\) is evaluated on a validation batch to compute \(\mathcal{L}_{\text{sup}}\), enabling meta-gradient backpropagation into the policy: \(\phi^* \leftarrow \phi - \beta \nabla_\phi \mathcal{L}_{\text{sup}}(\theta'_S(\phi); \mathcal{D}_{\text{val}})\). Finally, the updated policy \(\phi^*\) directs the true physical update of student network \(\theta_S\) on \(\mathcal{D}_{\text{tr}}\). This meta-learning loop aligns distillation incentives directly with empirical physiological generalization.

Loss & Training

The overall training objective for the student model operates in a label-optional configuration: $$ \mathcal{L}{\text{train}} = \mathbb{I}} \mathcal{L{\text{sup}}(\hat{\mathbf{y}}^v, \mathbf{y}) + \lambda^r; \phi) $$ where }} \mathcal{L}_{\text{distill}}(\hat{\mathbf{y}}^v, \hat{\mathbf{y}\(\mathbb{I}_{\{y\}}\) indicates the availability of ground-truth physiological labels. Unlabeled pairs are supervised purely via self-guided \(\mathcal{L}_{\text{distill}}\). Models are trained at 30 Hz sampling rate over clip length \(T=256\), with physiological band \(\mathcal{B} = [45, 180]\text{ bpm}\). The student network is instantiated with FactorizePhys, and the teacher is a pretrained, frozen RF pulse predictor. Optimization uses Adam with learning rate \(\eta = 5 \times 10^{-5}\) and weight decay \(10^{-2}\) for the student, and Adam with \(\beta = 10^{-4}\) for the policy network, performing meta-updates every 10 student iterations.

Key Experimental Results

Main Results

Cross-dataset performance was evaluated by training on one dataset and testing on the other across PhysDrive and EquiPleth. Evaluation metrics include standard deviation of heart rate error (STD), mean absolute error (MAE), root mean square error (RMSE), and Pearson correlation coefficient (\(r\)).

Modality Method PhysDrive STD↓ PhysDrive MAE↓ PhysDrive RMSE↓ PhysDrive \(r\)↑ EquiPleth STD↓ EquiPleth MAE↓ EquiPleth RMSE↓ EquiPleth \(r\)↑
Video CHROM 17.17 16.72 23.97 0.19 10.77 7.31 17.42 0.54
Video POS 17.33 17.38 24.38 0.20 11.51 4.64 15.07 0.68
Video PhysNet 12.73 18.45 25.45 0.15 11.10 7.93 13.64 0.55
Video PhysFormer 17.10 18.87 21.36 0.16 12.56 9.47 13.20 0.34
Video FactorizePhys 14.74 15.67 20.55 0.19 11.83 8.21 14.40 0.78
Video BeatFormer 16.10 17.16 21.80 0.17 13.82 9.08 13.98 0.62
RF Vilesov et al. 13.60 15.21 20.42 0.21 7.79 8.46 9.44 0.78
Video-RF Fusion FusionPhys 15.91 15.35 22.48 0.23* 5.24 4.88 6.48 0.89
Video-RF Fusion SATM 15.22 13.79* 21.99 0.23* 5.75 3.92 6.62 0.88
Cross-modal KD FitNets 15.26 20.20 28.31 0.09 6.60 3.09 7.23 0.85
Cross-modal KD C2KD 13.21 18.25 25.37 0.12 6.92 2.45 6.88 0.90
Cross-modal KD MST-Distill 14.69 16.55 23.47 0.17 5.58 2.99 7.31 0.87
Ours (Video-only) RPM-Distill 12.67* 14.15 20.24* 0.22 4.14* 1.57* 4.64* 0.94*

(Note: Asterisks * indicate statistically significant improvements with \(p < 0.005\) via paired t-tests; best results in bold.)

On EquiPleth, RPM-Distill slashes the MAE of the leading video-only baseline FactorizePhys from 8.21 bpm down to 1.57 bpm (an 81% error reduction) and boosts Pearson correlation \(r\) from 0.78 to 0.94. Crucially, it outperforms multimodal fusion baselines like SATM (MAE 3.92 bpm), despite operating under video-only inference. On the highly challenging PhysDrive dataset, it achieves top-tier STD (12.67) and RMSE (20.24).

Ablation Study

A systematic ablation study was conducted under the cross-dataset protocol on PhysDrive and EquiPleth to isolate the contributions of spectral loss terms, adaptive weighting components, and bilevel optimization.

Model Variant PhysDrive MAE↓ PhysDrive RMSE↓ PhysDrive \(r\)↑ EquiPleth MAE↓ EquiPleth RMSE↓ EquiPleth \(r\)↑ Note
Full Model 14.15 20.24 0.22 1.57 4.64 0.94 Complete RPM-Distill framework
w/o \(\mathcal{L}_{\text{peak}}\) 16.28 22.51 0.17 4.30 9.92 0.75 Removes fundamental peak anchoring; severe drop
w/o \(\mathcal{L}_{\text{off}}\) 16.10 22.37 0.16 3.40 8.29 0.82 Removes off-peak background suppression
w/o \(\mathcal{L}_{\text{shape}}\) 15.95 22.29 0.16 3.13 7.93 0.83 Removes centroid and entropy sharpness matching
w/o weight 15.83 22.06 0.17 3.25 8.19 0.83 Replaces predicted \(\boldsymbol{\alpha}\) with fixed equal weights
w/o gate 15.47 21.66 0.18 3.45 8.53 0.81 Removes gate \(g\); noisy teachers cause negative transfer
w/o \(\Phi(\cdot; \phi)\) 15.68 21.89 0.18 2.78 7.35 0.85 Eliminates policy network and meta-learning entirely
w/o bilevel 16.15 22.54 0.16 3.07 7.74 0.84 Optimizes policy jointly on training data without meta loop

Under synthetic cross-modal synchronization drift (shifting the RF window by 1s / 30 frames), stress tests show that as the shift probability reaches 1.0, the ungated model (w/o gate) degrades markedly from 15.47 to 18.08 bpm MAE (RMSE reaching 25.64 bpm). In contrast, RPM-Distill dynamically downweights corrupted pairs, maintaining a stable MAE of 14.80 bpm.

Key Findings

  • Fundamental peak alignment is paramount: Ablating \(\mathcal{L}_{\text{peak}}\) inflicts the most catastrophic performance collapse (EquiPleth MAE deteriorates from 1.57 to 4.30 bpm, and RMSE jumps from 4.64 to 9.92 bpm). This confirms that under low optical SNR, the primary failure mode of video models is rhythm mis-localization, making the radar peak anchor indispensable.
  • Dynamic gating safeguards against unsynchronized and corrupted supervision: Under cross-modal synchronization shifts, omitting gating accumulates corrupted gradients and destabilizes learning. The learned gate effectively attenuates unaligned sample pairs.
  • Moderate supervision yields optimal distillation balance: In label-scarce experiments, utilizing 40% labeled training pairs delivered better overall performance than 100% full supervision. An over-reliance on source-domain hard labels predisposes models to domain-specific appearance bias, whereas balanced spectral distillation enforces genuine physiological dynamics.

Highlights & Insights

  • Bridging incommensurate modalities via shared spectral physiology: Instead of forcing mismatched video reflectance and radar phase signals into unnatural time-domain or intermediate feature alignment, RPM-Distill insightfully identifies the band-limited frequency spectrum as the true physical substrate where cross-modal periodicity coincides.
  • Overcoming unlabelled adaptive distillation via bilevel meta-learning: By treating the distillation policy as a differentiable inductive bias evaluated against empirical downstream physiological recovery on a validation split, the framework elegantly sidesteps the absence of explicit gating ground truth.
  • Privileged learning delivers deployment efficiency: By confining radar hardware exclusively to the training phase, the model matches or exceeds multimodal fusion models while preserving lightweight, single-camera inference suitable for real-world devices.

Limitations & Future Work

  • Fixed physiological band limits detection of extreme arrhythmias: The strict bandpass constraint of \([45, 180]\text{ bpm}\) limits utility for patients suffering from severe pathological bradycardia (<45 bpm) or extreme tachycardia (>180 bpm).
  • Dependency on synchronized training pairs: The training protocol still requires co-recorded video and radar data from the same subject. Developing unsupervised distillation across unpaired in-the-wild video repositories and unaligned radar databases remains an open frontier.
  • Extension to multi-vital sign monitoring: The current formulation focuses on heart rate; incorporating respiration rate (RR) and pulse transit time (PTT) into a unified multi-scale spectral distillation scheme represents a natural next step.
  • vs Multimodal Fusion (e.g., SATM [Liang et al., 2025], FusionPhys [Ying et al., 2025]): Multimodal fusion methods aggregate video and radar features during inference, imposing heavy hardware burdens at deployment; RPM-Distill adopts the learning using privileged information (LUPI) paradigm, keeping inference strictly video-only while achieving superior cross-domain generalization.
  • vs Generic Cross-modal KD (e.g., FitNets [Romero et al., 2015], C2KD [Huo et al., 2024], MST-Distill [Li et al., 2025]): Standard cross-modal KD targets intermediate semantic features or broad category logits, which fails on periodic physiological signals; RPM-Distill decomposes supervision into three physically grounded spectral constraints alongside meta-learned dynamic gating, successfully preventing negative transfer.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ First framework to introduce RF radar as privileged training information for RPM; the spectral decomposition and bilevel meta-policy are conceptually elegant and physically principled.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive cross-dataset evaluations, label-scarcity curves, synchronization drift stress tests, and qualitative waveform analyses provide solid empirical backing.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear problem formulation, rigorous mathematical exposition, and convincing ablation justifications.
  • Value: ⭐⭐⭐⭐⭐ Bridges the gap between robust multi-sensor academic research and practical, low-cost camera deployment in digital health and smart mobility.