MeanTalker: Efficient and Expressive Speech-Driven 3D Facial Animation via Geometric-Aware Mean Flow¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: [To be released]
Area: 3D Vision
Keywords: speech-driven facial animation / mean flow / one-step generation / 3D facial animation / geometric-aware manifold constraints
TL;DR¶
Tackling the severe latency of diffusion sampling and the over-smoothing of deterministic regression, MeanTalker introduces Mean Flow into speech-driven 3D facial animation via a Dual-Time Denoising Transformer (DTDT) and Geometric-Aware Trajectory Learning (GATL), achieving single-step inference (NFE=1, RTF≈0.004) with superior lip-sync precision and expressive facial dynamics over diffusion SOTA.
Background & Motivation¶
Speech-driven 3D facial animation plays a foundational role in virtual reality, gaming, and digital avatar interaction, aiming to synthesize synchronized, natural, and expressive time-varying 3D mesh sequences from continuous acoustic inputs and reference speaking styles. Early deterministic regression paradigms—whether predicting vertex displacements like VOCA and FaceFormer or regressing 3DMM coefficients like SadTalker and EmoTalk—formulate the task as a direct mapping from speech to facial geometry. However, because speech-to-facial motion is inherently an ill-posed many-to-many problem (a single phoneme can correspond to diverse valid facial expressions and subtle muscular tensions), deterministic models tend to collapse to conditional expectations, resulting in severely over-smoothed animations that lack vivid emotional nuances.
To overcome over-smoothing, recent probabilistic generative approaches adopt Denoising Diffusion Probabilistic Models (DDPMs), such as FaceDiffuser, DiffPoseTalk, and MSMD. By capturing the underlying conditional distribution of facial motions, diffusion models significantly improve motion variance and expressive dynamics. Nonetheless, this expressiveness comes at the cost of prohibitive computational latency: the iterative denoising process typically requires tens or hundreds of sequential sampling steps (NFE > 50), yielding real-time factors around 0.9 to 1.0 that prevent deployment in real-time interactive systems. While some works have investigated model distillation or shortcut paths to accelerate diffusion sampling, they frequently suffer from noticeable visual degradation, lip desynchronization, and geometric jitter.
Adapting modern single-step flow paradigms (e.g., Flow Matching or Mean Flow) to 3D facial motion presents unique challenges. Directly training average velocity fields in complex cross-modal latent spaces leads to severe numerical instability and gradient explosion. Furthermore, unconstrained latent-space optimization lacks physical manifold grounding: large one-step discrete jumps readily cause trajectories to deviate from the valid 3D facial manifold, inducing severe geometric artifacts such as incomplete lip closure and unnatural cheek collapse. Core idea: formulate speech-driven 3D motion generation via Mean Flow, employing an interleaved dual-time encoding to progressively transition from instantaneous to average velocities, while enforcing implicit manifold recovery and semantic-aware geometric velocity consistency to strictly straighten the trajectory within the physical domain.
Method¶
Overall Architecture¶
MeanTalker completely bypasses the iterative numerical integration of standard diffusion models by directly learning a straight-line vector field from a Gaussian noise distribution to the target motion distribution, synthesizing the entire 3D facial sequence in a single forward evaluation. As shown in the framework diagram below, the system comprises two synergistic components: the Dual-Time Denoising Transformer (DTDT), which ingests acoustic features, identity geometry, and reference speaking style to predict generative velocity fields under dual temporal anchors; and Geometric-Aware Trajectory Learning (GATL), which analytically recovers clean motion endpoints, maps them to the FLAME mesh manifold, and simultaneously regularizes vertex physics and semantic velocity consistency to eliminate trajectory curvature.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input Data<br/>Speech Waveform + Reference Video + Noisy Motion zt"] --> DTDT["Dual-Time Denoising Transformer<br/>Interleaved dual temporal anchors + masked cross-attention velocity field"]
DTDT --> Proj["Implicit Manifold Projection<br/>Analytic recovery of clean endpoint xθ decoded via FLAME"]
Proj --> GVC["Semantic-Aware Manifold-GVC<br/>Jacobian linearization + region-specific lip and structural trajectory rectification"]
GVC --> Out["Output Results<br/>High-precision lip-sync and expressive 3D facial meshes (NFE=1)"]
Key Designs¶
1. Dual-Time Denoising Transformer: Interleaved Anchoring and Progressive Velocity Modeling Directly training an average velocity field from scratch on high-dimensional cross-modal motion representations typically leads to severe numerical instability. DTDT resolves this by introducing an interleaved dual-time encoding mechanism with two temporal anchors: current simulation time \(t \sim \mathcal{U}[0, 1]\) and target reference time \(r\). Both scalars are projected through sinusoidal positional encodings (PE) and dedicated MLPs to form composite time embeddings \(E_{\text{time}} = \text{MLP}_t(\text{PE}(t)) + \text{MLP}_r(\text{PE}(r))\), which are added to the static identity shape \(\beta\) and motion style code \(S\) to form a time-aware prefix \(C_{\text{prefix}}\). In the Transformer decoder stack, bi-directional self-attention captures long-range temporal motion context, while a masked cross-attention module with a diagonal alignment mask restricts each motion frame to attend strictly to its local phonetic window extracted by HuBERT. To enable stable curriculum learning, DTDT alternates the reference time \(r\) via an interleaved switch: $\(r = \begin{cases} t, & \text{Phase I (Instantaneous Velocity } v_\theta) \\ u \sim \mathcal{U}[0, 1], & \text{Phase II (Average Velocity } u_\theta) \end{cases}\)$ This mechanism allows the network to first master fundamental acoustic-lip correlations under standard Conditional Flow Matching (CFM) before expanding into global average velocity prediction across arbitrary horizons.
2. Implicit Manifold Projection: Analytic Recovery and Physical Geometric Regularization In single-step generation, latent feature regression without physical guidance easily drifts off the authentic 3D facial manifold, causing distorted topologies and lip tearing. GATL exploits the explicit displacement property of the average velocity \(u_\theta\) to analytically extrapolate the clean motion endpoint \(x_\theta\) directly from any intermediate noisy state \(z_t\) without iterative integration: $\(x_\theta = z_t + (1 - t) u_\theta(z_t, 1, t)\)$ The recovered motion vector \(x_\theta\) is immediately passed into the differentiable FLAME 3DMM decoder \(\mathcal{V}\) to reconstruct the 5,023-vertex surface mesh. GATL then imposes explicit physical geometric constraints \(\mathcal{L}_{\text{geo}} = \lambda_{\text{vert}}\mathcal{L}_{\text{vert}} + \lambda_{\text{vel}}\mathcal{L}_{\text{vel}} + \lambda_{\text{coef}}\mathcal{L}_{\text{coef}}\), where \(\mathcal{L}_{\text{vert}}\) penalizes absolute lip-vertex Euclidean errors, \(\mathcal{L}_{\text{vel}}\) penalizes inter-frame vertex velocity jitters to guarantee temporal smoothness, and \(\mathcal{L}_{\text{coef}}\) enforces boundary motion continuity, anchoring the latent trajectory to the physical surface manifold.
3. Semantic-Aware Manifold-GVC: Jacobian Linearization and Region-Adaptive Trajectory Rectification To eliminate trajectory curvature in one-step inference, the predicted average velocity \(u_\theta\) is supervised against a JVP-rectified target \(u_t = (x_1 - x_0) - (t - r)\frac{d}{dt}u(z_t, r, t)\) via an adaptive flow loss \(\mathcal{L}_{\text{flow}} = w \cdot \|u_\theta - \text{sg}(u_t)\|^2\). Because latent perturbations propagate non-linearly to 3D facial surfaces, GATL introduces Geometric Velocity Consistency (Manifold-GVC). Given latent velocity error \(\delta u = u_\theta - \text{sg}(u_t)\), it approximates mesh deformation \(\delta V\) around the detached ground-truth anchor \(x_{\text{base}} = \text{sg}(x_1)\) using a localized Jacobian-aware gradient step: $\(\delta V \approx \frac{\mathcal{V}(x_{\text{base}} + \epsilon \cdot \delta u) - \mathcal{V}(x_{\text{base}})}{\epsilon}\)$ Furthermore, recognizing that facial dynamics are decoupled—the mouth region governs phonetic articulation, while the forehead, brows, and cheeks carry emotional style and blinking—GATL deploys category-wise semantic weighting: $\(\mathcal{L}_{\text{gvc}} = \lambda_{\text{lip}}\mathcal{L}_{\text{gvc}}^{\text{lip}} + \lambda_{\text{str}}\mathcal{L}_{\text{gvc}}^{\text{str}}\)$ By allocating stringent constraints to the articulation-critical lip region (\(\mathcal{L}_{\text{gvc}}^{\text{lip}}\)) while applying moderate structural regularization to the rest of the face (\(\mathcal{L}_{\text{gvc}}^{\text{str}}\)), GATL prevents stochastic trajectory variations from compromising speech articulation, ensuring sharp and expressive single-step synthesis.
Loss & Training¶
MeanTalker is trained in a progressive two-stage curriculum: - Stage I (Instantaneous Velocity Pretraining): Reference time is fixed as \(r \equiv t\), reducing the model to standard Conditional Flow Matching. The objective is \(\mathcal{L}_{\text{StageI}} = \mathcal{L}_{\text{cfm}} + \mathcal{L}_{\text{geo}}\), with \(\mathcal{L}_{\text{cfm}} = \|v_\theta - (x_1 - x_0)\|^2\), establishing a robust cross-modal alignment baseline. - Stage II (Mean Flow Trajectory Rectification): Initialized from Stage I weights with \(\text{MLP}_r\)'s final layer zero-initialized. Training switches dynamically: in each iteration, \(r = t\) with 50% probability (reinforcing instantaneous dynamics) and \(r \sim \mathcal{U}[0, 1]\) with 50% probability (learning the average velocity field \(u_\theta\)). The total objective is: $\(\mathcal{L}_{\text{StageII}} = \mathcal{L}_{\text{flow}} + \mathcal{L}_{\text{gvc}} + \mathcal{L}_{\text{geo}}\)$ At inference, setting \(t=0, r=1\) enables direct single-step generation from Gaussian noise \(z_0 \sim \mathcal{N}(0, I)\) via \(x_1 = z_0 + u_\theta(z_0, 1, 0)\) with NFE=1.
Key Experimental Results¶
Main Results¶
Quantitative evaluations are conducted on the standard TFHP benchmark (1,052 high-definition clips across 588 subjects; test set with 64 unseen subjects and 120 video sequences). Inference latency (RTF) is measured on a single NVIDIA RTX 3090 GPU.
| Methods | Paradigm | LVE (mm) ↓ | FDD (\(10^{-5}\) m) ↓ | MOD (mm) ↓ | MVE (mm) ↓ | RTF ↓ |
|---|---|---|---|---|---|---|
| MultiTalk (arXiv 2024) | Autoregressive | 11.099 | 8.348 | 3.853 | 1.494 | 0.288 |
| ScanTalk (ECCV 2024) | Regression | 13.145 | 12.550 | 4.365 | 2.254 | 0.049 |
| DiffPoseTalk (SIGGRAPH 2024) | Diffusion | 9.119 | 5.458 | 3.225 | 0.938 | 0.919 |
| ARTalk (arXiv 2025) | Autoregressive | 11.783 | 12.049 | 4.057 | 1.643 | 0.020 |
| MSMD (SIGGRAPH 2025) | Diffusion | 8.390 | 5.763 | 2.982 | 0.829 | 1.015 |
| MeanTalker (Ours) | Mean Flow | 7.879 | 5.189 | 2.830 | 0.788 | 0.004 |
Ablation Study¶
To investigate the contribution of dual-time curriculum learning, physical manifold regularization, and style conditioning, rigorous ablation studies are conducted on the TFHP benchmark.
| Variant | Paradigm / Setting | LVE (mm) ↓ | FDD (\(10^{-5}\) m) ↓ | MOD (mm) ↓ | MVE (mm) ↓ |
|---|---|---|---|---|---|
| Pre-T (Standard FM) | Pretrained Standard FM (NFE=25, coupled \(t \equiv r\)) | 8.906 | 4.778 | 3.155 | 0.918 |
| w/o Rectification | Forced one step (NFE=1, coupled \(t \equiv r\)) | 9.189 | 6.511 | 3.231 | 0.940 |
| w/o Style Code | Full stage without temporal style code | 9.144 | 5.919 | 3.076 | 0.930 |
| w/o Progressive Init | Direct Stage-II training without FM pretraining | 8.272 | 5.977 | 3.133 | 0.810 |
| w/o GATL | Latent MSE loss only (no manifold projection & GVC) | 8.735 | 5.152 | 3.187 | 0.874 |
| MeanTalker (Full) | Full model (DTDT + GATL progressive optimization) | 7.879 | 5.189 | 2.830 | 0.788 |
Key Findings¶
- 200× Acceleration in Latency: MeanTalker achieves an exceptional Real-Time Factor (RTF) of 0.004, running more than 200× faster than state-of-the-art diffusion models such as DiffPoseTalk (0.919) and MSMD (1.015). Even when including HuBERT audio feature extraction and FLAME mesh decoding, the full end-to-end pipeline processes 1 second of audio in just 0.0048 seconds, fully unblocking real-time interactive avatars.
- Physical Manifold Grounding is Indispensable: Discarding GATL and training solely via latent MSE (w/o GATL) causes LVE to jump from 7.879 mm to 8.735 mm. Error heatmaps confirm that unconstrained latent regression leads to severe under-articulation in mouth opening and closure, proving that explicit 3D mesh projection is vital for maintaining geometric validity along straight trajectories.
- Training Stability from Progressive Initialization: Training the average velocity model from scratch without Stage I FM pretraining (w/o Progressive Init) degrades LVE to 8.272 mm and more than doubles convergence iterations (65K vs. 30K), verifying that instantaneous velocity fields provide essential guidance for trajectory straightening.
- Catastrophic Distortions in Forced Single-Step FM: Forcing the standard Flow Matching model to run in one step (w/o Rectification) results in severe global geometric distortions across the cheeks and brow ridge (LVE=9.189 mm), demonstrating that curved probability flows cannot be traversed in one jump without Mean Flow velocity correction.
Highlights & Insights¶
- First Successful Deployment of Mean Flow to Cross-Modal 3D Motion: While Mean Flow was originally developed for 2D image synthesis, MeanTalker demonstrates its capability in continuous 3D facial animation, demonstrating that straight-line average velocity fields can resolve the classic speed-quality trade-off in speech-driven graphics.
- Jacobian-Linearized Manifold Rectification (Manifold-GVC): Calculating higher-order gradients directly through the non-linear 3DMM decoder during trajectory optimization is computationally prohibitive. The finite-difference gradient approximation combined with articulation-aware semantic masking provides an elegant, scalable solution for physical manifold regularization.
- Generalizable Curriculum Framework: The dual-time encoding paradigm (alternating simulation/reference time with zero-initialized transfer) provides a clear blueprint for adapting one-step flow matching to other continuous motion tasks, such as full-body avatar dance and gesture generation.
Limitations & Future Work¶
- Limitations Admitted by Authors: The current model focuses on general speaking styles and lip-sync synchronization, lacking fine-grained emotional category labels and explicit emotion intensity control. Moreover, the topological modeling is confined to the head and facial mesh, omitting gaze dynamics and full-body skeletal coordination.
- Future Directions: Extending Manifold-GVC to the full-body SMPL-X parametric manifold to enable real-time one-step synthesis of body postures, arm gestures, and facial expressions from continuous speech; integrating streaming audio encoders (e.g., DAC/EnCodec) for ultra-low-latency microsecond interactions.
Related Work & Insights¶
- vs DiffPoseTalk / MSMD: DiffPoseTalk and MSMD rely on iterative diffusion sampling requiring dozens of forward passes to synthesize natural facial expressions; MeanTalker leverages the Mean Flow identity to compress inference into a single step, boosting speed by over 200× while attaining higher lip-sync precision (LVE 7.879 mm vs. 8.390 mm) and superior expression naturalness (FDD 5.189 vs. 5.763).
- vs ARTalk / MultiTalk: Autoregressive approaches suffer from cumulative drift over extended sequences and discrete code quantization loss; MeanTalker operates on continuous deformation fields in a single step, bypassing error accumulation and lowering LVE significantly (7.879 mm vs. 11.099 mm and 11.783 mm).
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First framework introducing Mean Flow to speech-driven 3D facial animation, with elegant dual-time encoding and manifold-aware trajectory rectification.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous benchmark evaluations on TFHP, comprehensive ablation studies, qualitative error heatmaps, and a blind user study across 5 baselines.
- Writing Quality: ⭐⭐⭐⭐⭐ Mathematically rigorous, clearly structured, with well-motivated design choices and self-contained formulations.
- Value: ⭐⭐⭐⭐⭐ Drastically cuts inference latency to 4 ms per audio second, bridging the gap between high-fidelity 3D generative avatars and real-time deployment.