Revisiting the Volumetric Data of 4DME: Compression, Extension and Benchmarking for Micro-Expression Analysis¶
Conference: ECCV 2026
Paper: ECCV 2026 Official
Code: https://github.com/1nO1SB/4DMEplus
Area: Human Understanding
Keywords: Micro-Expression Analysis, Action Unit Detection, 3D Volumetric Compression, V-PCC, Dual-Stream Optical Flow
TL;DR¶
To tackle extreme volumetric data size, stereo reconstruction artifacts, and the lack of standardized benchmarks in 4D facial micro-expression action unit (ME-AU) detection, this work systematically cleans and doubles 4DME into a culturally balanced release, introduces the dual-stream VoluME network, and validates that MPEG V-PCC compression achieves an approximately 3000× data reduction with only ~1% F1 loss.
Background & Motivation¶
Micro-expressions (MEs) are involuntary, rapid facial movements that manifest under emotional concealment, typically lasting only 40 to 500 milliseconds. While traditional micro-expression recognition (MER) maps video clips to coarse, discrete emotion classes, micro-expression action unit (ME-AU) detection grounds analysis directly in the Facial Action Coding System (FACS), providing a more objective, fine-grained, and compositional taxonomy of local muscle group activations. However, micro-expression movements are inherently sparse in both space and time, and conventional 2D video analysis struggles with lighting variations, head pose changes, and ambiguous out-of-plane facial deformations. Although recent volumetric (3D+t) datasets like 4DME and CAS(ME)3 provide temporally coherent 3D surface reconstructions that capture subtle depth deformations, their real-world adoption has been severely limited by massive storage footprints, heterogeneous file formats, and the lack of standardized experimental benchmarks.
A closer inspection of the original 4DME release reveals that due to stereo reconstruction and multi-camera capture artifacts, geometric flaws are prevalent: 40% of sequences suffer from temporally persistent reconstruction noise, 55% contain isolated peripheral artifacts, and 45% are corrupted by sensor-level acquisition outliers, leaving only 33% of the dataset completely artifact-free. Such noise disrupts temporal registration, point-to-point correspondence, and voxelization schemes, which previously caused roughly half of the captured sequences to be discarded entirely. Furthermore, while 3D compression standards are essential for practical distribution and storage of high-rate 4D meshes, standard codecs are calibrated for global human visual perception or general geometric fidelity; whether lossy compression preserves the subtle, high-frequency spatial-temporal signals critical for ME-AUs has remained an open question.
This paper addresses these gaps with a unified, end-to-end framework covering geometric denoising, data extension, task-aware compression benchmarking, and dual-stream dynamic modeling. The core idea is to design a three-stage geometric denoising pipeline that recovers previously discarded sequences into an ethnically balanced 4DME extension, propose a dual-stream architecture (VoluME) that decouples visual and geometric surface flows with optical strain, and demonstrate that MPEG V-PCC achieves ~3000× compression with negligible degradation in downstream ME-AU detection.
Method¶
Overall Architecture¶
The proposed framework for volumetric micro-expression action unit analysis and compression benchmarking consists of three main components: 1. Three-stage geometric denoising and dataset extension: Raw stereo-reconstructed meshes are filtered sequentially via statistical GMM outlier pruning, sequence-wide occupancy-mask filtering, and manual interactive refinement, expanding usable sequences from 130 to 265 while achieving an equal Asian/European cultural balance. 2. Task-aware volumetric compression benchmarking: Draco (geometry-based) and MPEG V-PCC (video-based point cloud compression) are evaluated across multiple quantization levels, combining objective surface metrics, human perceptual AU blind studies, and downstream model performance. 3. VoluME dual-stream ME-AU detection: Projecting the 4D facial surface into orthogonal RGB appearance and depth streams, computing onset-to-apex optical flow alongside optical strain, fusing high-level convolutional features, and classifying individual AUs via decoupled single-task heads.
The overall data processing pipeline and model inference flow are depicted below:
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Raw captured 4DME mesh sequence<br/>(contains stereo artifacts & outliers)"] --> B["Three-stage mesh geometric denoising<br/>Statistical GMM + Occupancy mask + Refinement"]
B --> C["Extended 4DME dataset<br/>265 sequences / Balanced culture (23:20)"]
C --> D["MPEG V-PCC / Draco volumetric compression<br/>(fidelity & downstream task evaluation)"]
D --> E["Surface projection into dynamic RGB-D streams<br/>Appearance C(x,y,t) and depth D(x,y,t)"]
E --> F["Dual-stream optical flow & strain computation<br/>OF_rgb and OF_depth (each with u, v, |ε|)"]
F --> G["VoluME convolutional feature extraction & fusion<br/>(CNN encoders + feature-level concatenation)"]
G --> H["AU-specific decoupled classifiers<br/>(independent MLP heads for each AU)"]
Key Designs¶
1. Three-stage geometric denoising and cross-cultural extension: salvaging corrupted sequences Standard mesh filtering tools often blur high-curvature facial contours or fail against large phantom surfaces that share texture and vertex density with the face. To address this, the authors introduce a coarse-to-fine three-stage denoising strategy: first, acquisition outliers are detected by fitting a Gaussian Mixture Model (GMM) to edge-length distributions within each frame, removing abnormal edges and isolated vertices; second, to eliminate persistent stereo matching ghosts across frames, all temporal frames are aggregated into a unified sequence occupancy mask, where consistent facial surfaces form dense clusters while transient artifacts exhibit lower spatio-temporal coherence, allowing neighborhood outlier filtering to isolate the true facial envelope; third, residual dense artifacts adjacent to facial boundaries are manually trimmed. This pipeline reduces normal divergence by 86%, cuts occupancy instability by 52%, and boosts curvature consistency by 47%. Consequently, the usable 4DME sequence count increases from 130 to 265 (AU6 samples increase by +233% and AU12 by +224%), establishing the first volumetric micro-expression benchmark with an explicit cross-cultural balance (23 Asian vs. 20 European subjects).
2. Decoupled surface projection and dual-stream flow: resolving 3D volume flow scale mixing and strain omission Prior 3D micro-expression methods attempted to compute full volumetric optical flow by extending optical flow constraints directly to 3D voxel space: \(\frac{\partial I}{\partial x}u + \frac{\partial I}{\partial y}v + \frac{\partial I}{\partial z}w + \frac{\partial I}{\partial t} = 0\). However, this treats planar in-plane movement and normal depth displacements on an identical scale, mixing appearance dynamics with surface geometry into a single vector while discarding optical strain. VoluME reformulates motion modeling under the principle that the human face acts as a single visible surface \(P^*(x, y, t) := \{C(x, y, t), D(x, y, t)\}\). By separating color appearance \(C\) and depth geometry \(D\), optical flow assumptions are applied independently: $$ \frac{\partial C}{\partial t} = -\left(\frac{\partial C}{\partial x}u_{\text{rgb}} + \frac{\partial C}{\partial y}v_{\text{rgb}}\right), \quad \frac{\partial D}{\partial t} = -\left(\frac{\partial D}{\partial x}u_{\text{depth}} + \frac{\partial D}{\partial y}v_{\text{depth}}\right) $$ Between onset and apex frames, horizontal displacement \(u\), vertical displacement \(v\), and normal optical strain \(|\epsilon_{uv}| = \sqrt{(\frac{\partial u}{\partial x} - \frac{\partial v}{\partial y})^2 + (\frac{\partial u}{\partial y} + \frac{\partial v}{\partial x})^2}\) are extracted for both streams. Feeding the three-channel descriptor \((u, v, |\epsilon_{uv}|)\) directly preserves physical deformation priors for subtle skin stretching.
3. Feature-level dual-stream fusion and decoupled AU-specific classifiers: preventing overfitting and multi-label interference Given the limited sample scale typical of micro-expression benchmarks, heavy 3D vision transformers are prone to catastrophic overfitting. VoluME uses lightweight CNN backbones (such as SSSNet, STSTNet, or Off-ApexNet) to process each stream independently. Rather than stacking RGB and depth at the input layer (early fusion), VoluME keeps the visual and geometric representations separated through intermediate convolutions, merging them only at the late feature representation stage. This preserves subtle texture signals (e.g., eye crinkles) and geometric height shifts (e.g., brow raises) before joint modeling. Furthermore, because adjacent facial action units frequently co-occur in close spatial proximity, standard shared multi-label heads suffer from cross-AU interference and competition. VoluME deploys dedicated multi-layer perceptron (MLP) binary heads for each action unit in a One-vs-Rest fashion, significantly boosting sensitivity in high-density facial regions.
Loss & Training¶
The framework is evaluated under a Leave-One-Subject-Out (LOSO) cross-validation protocol on 4DME. Because positive activations are sparse across the dataset, each independent AU classification head is trained using weighted binary cross-entropy loss: $$ \mathcal{L}_{\text{BCE}}^{(k)} = - \left( w_k y_k \log(\hat{y}_k) + (1 - y_k) \log(1 - \hat{y}_k) \right) $$ where \(y_k \in \{0, 1\}\) is the ground-truth label for the \(k\)-th action unit, \(\hat{y}_k\) is the predicted probability, and \(w_k\) denotes the positive class reweighting coefficient calculated from the negative-to-positive ratio in the training fold. In Leave-One-Dataset-Out (LODO) cross-dataset experiments, models trained on 4DME are directly evaluated on CAS(ME)3 Part C without target domain fine-tuning.
Key Experimental Results¶
Main Results¶
The table below reports subject-independent Leave-One-Subject-Out (LOSO) results on the extended, denoised 4DME dataset using pooled binary F1-score (BF1) and per-AU F1-scores. VoluME variants consistently surpass previous 3D baselines.
| Model | AU1 | AU2 | AU4 | AU6 | AU7 | AU12 | AU17 | AU45 | Overall BF1 |
|---|---|---|---|---|---|---|---|---|---|
| RGBD AlexNet [Li et al., 2023] | 0.111 | 0.182 | 0.264 | 0.200 | 0.371 | 0.452 | 0.000 | 0.000 | 0.323 |
| CCDN [Li et al., 2023] | 0.130 | 0.255 | 0.361 | 0.000 | 0.323 | 0.304 | 0.000 | 0.085 | 0.336 |
| VoluME (Off-ApexNet) | 0.283 | 0.291 | 0.247 | 0.071 | 0.398 | 0.457 | 0.039 | 0.210 | 0.397 |
| VoluME (STSTNet) | 0.279 | 0.311 | 0.511 | 0.117 | 0.402 | 0.474 | 0.073 | 0.145 | 0.424 |
| VoluME (SSSNet, Ours) | 0.344 | 0.327 | 0.467 | 0.189 | 0.383 | 0.424 | 0.000 | 0.452 | 0.475 |
Under the cross-dataset (LODO) setting across 4DME and CAS(ME)3 Part C, VoluME (SSSNet) achieves the highest domain generalization BF1 of 0.387, substantially outperforming RGBD AlexNet (0.318) and CCDN (0.274).
Ablation Study¶
To validate each design component, ablations were conducted under both within-dataset (LOSO) and cross-dataset (LODO) settings, measured by unweighted macro-averaged F1 (Macro-F1):
| Config / Variant | LOSO Macro-F1 | LODO Macro-F1 | Note |
|---|---|---|---|
| VoluME-SSSNet (full model) | 0.485 | 0.389 | Dual-stream feature fusion + optical strain (u, v, |ε|) |
| Single stream: RGB-only | 0.405 | 0.300 | Drops 8.0% on LOSO without 3D geometric motion |
| Single stream: Depth-only | 0.364 | 0.312 | Drops 12.1% on LOSO without RGB appearance flow |
| Input Fusion (early fusion) | 0.396 | 0.217 | Direct RGB+Depth channel stacking; LODO drops by 17.2% |
| No Strain (w/o |ε|) | 0.364 | 0.175 | Removing strain collapses cross-dataset generalization |
Volumetric Compression Benchmark¶
The impact of volumetric compression on downstream ME-AU detection accuracy across varying compression ratios on 4DME is summarized below:
| Codec & Quality Level | Geom. Ratio (CRgeom) | Total Ratio (CRtotal) | Mean Macro-F1 | Degradation |
|---|---|---|---|---|
| Uncompressed baseline | 1.0× | 1.0× | 0.356 | 0.0% (reference) |
| V-PCC R3 | 640× | 1575× | 0.356 | 0.0% (lossless retention) |
| V-PCC R2 | 1665× | 3155× | 0.344 | -1.2% (negligible decline) |
| V-PCC R1 | 2955× | 5780× | 0.306 | -5.0% (severe geometric loss) |
| Draco QP9 | 10.7× | - | 0.352 | -0.4% |
| Draco QP8 | 12.0× | - | 0.321 | -3.5% |
| Draco QP6 | 16.0× | - | 0.283 | -7.3% |
Key Findings¶
- Optical strain is crucial for domain generalization: Removing optical strain \(|\epsilon_{uv}|\) causes cross-dataset (LODO) performance to drop sharply from 0.389 to 0.175 (a 55% relative decline), proving that local surface strain is far more invariant across camera sensors and subject populations than raw displacement vectors.
- Feature-level fusion outclasses early input concatenation: Early channel-wise stacking degrades LODO score to 0.217, whereas late feature fusion retains 0.389, verifying that independent representation learning for texture and depth avoids destructive interference.
- V-PCC achieves ~3000× compression with negligible downstream penalty: Under the V-PCC R2 configuration (3155× total compression), the downstream detection Macro-F1 drops by only 1.2% (0.356 to 0.344), and human perceptual AU retention remains above 95.6%, demonstrating that volumetric ME datasets can be compressed drastically without compromising research utility.
Highlights & Insights¶
- Pragmatic 2.5D surface flow formulation: Instead of brute-forcing computationally heavy 3D voxel convolutions or point cloud attention, the authors observe that human facial deformation can be treated as a single visible surface manifold, decoupling motion into 2D appearance flow and 2D depth flow while retaining optical strain advantages.
- Task-specific and perception-driven compression evaluation: By going beyond conventional distortion metrics (PSNR, Chamfer distance) to conduct human AU perceptual studies and downstream detector performance checks, the authors establish an empirical upper bound on volumetric compression for subtle facial dynamics.
- Reusable data restoration SOP: Combining edge-length GMM outlier rejection with sequence-wide spatio-temporal occupancy consistency offers a practical recipe for rehabilitating historical 3D/4D facial capture datasets plagued by stereo reconstruction noise.
Limitations & Future Work¶
- Long-tail sparsity in rare action units: Extreme data scarcity persists for certain AUs like AU17 (chin raiser), which appears in only 6 subjects across the extended dataset, resulting in near-zero F1 scores under subject-independent evaluation.
- Apex frame dependency: The current pipeline relies on known onset and apex frame boundaries; evaluating continuous, in-the-wild micro-expression spotting and detection directly from compressed 4D streams remains unaddressed.
- Future directions: Integrating V-PCC compressed representations directly into continuous 4D Gaussian Splatting (4D-GS) or neural volumetric representations to facilitate streaming and real-time inference.
Related Work & Insights¶
- vs. CCDN and RGBD AlexNet: Prior 3D methods either process static single-frame RGB-D maps or coarse 3D convolutions, losing critical onset-to-apex temporal motion; VoluME explicitly injects displacement and strain dynamics, improving 4DME LOSO BF1 from 0.336 to 0.475.
- vs. Heavy 3D Vision Transformers: In low-data 3D micro-expression domains, complex attention backbones quickly overfit; VoluME demonstrates that lightweight CNN encoders coupled with strong inductive priors (optical strain, decoupled AU heads) deliver superior performance and generalization.
Rating¶
- Novelty: ⭐⭐⭐⭐ [First systematic study of 4D micro-expression data compression limits combined with decoupled surface optical strain modeling]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Spans mesh denoising metrics, human perceptual studies, within- and cross-dataset benchmarks, codec comparisons, and component ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear structural progression from data engineering challenges to architectural solutions and compression analysis]
- Value: ⭐⭐⭐⭐⭐ [Salvages unusable 4DME sequences into an open, balanced benchmark and demonstrates 3000× compression viability for community adoption]