RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction¶
Conference: ECCV 2026
Paper: ECCV Official
Project: https://RoboGesture.github.io
Code: https://RoboGesture.github.io
Area: Human Understanding
Keywords: Humanoid Robot, Co-Speech Gestures, Semantic Alignment, Conditional Flow Matching, Embodied Safety
TL;DR¶
RoboGesture addresses the trilemma of semantic gesture data scarcity, modality eclipse from motion inertia, and physical self-collision in humanoid co-speech generation by unifying semi-synthetic data synthesis, hierarchical acoustic-semantic alignment, anti-inertia flow matching, and real-time MPC filtering into an end-to-end policy on a 41-DoF Unitree G1 humanoid.
Background & Motivation¶
In social human-robot interaction and humanoid embodied intelligence, empowering humanoid robots to actively listen to human speech and synthesize synchronized, semantically expressive body and hand gestures is a cornerstone for elevating machines from rigid tools to intuitive social companions. However, transferring prior co-speech gesture generation methodsโwhich are largely built for virtual avatars or rely on offline kinematic retargetingโto physical humanoid platforms exposes several fundamental bottlenecks. Existing audio-motion benchmarks (such as BEAT) overwhelmingly comprise repetitive rhythmic beat motions, whereas gestures with explicit semantic or iconic meaning remain sparse in the long tail. Prior efforts attempting to bridge speech and gestures via intermediate text representations (such as ASR transcription followed by LLM retrieval) incur significant pipeline latency, erase vital prosodic nuances like intonation and emphasis, and cannot capture subtle pre-phonetic anticipatory signals where the human body instinctively prepares physical energy prior to vocalizing an emphatic word.
A more insidious architectural challenge in online streaming generation is the "modality eclipse" phenomenon: when autoregressive or temporal conditioning models generate the current motion chunk given past kinematics, they tend to discover a degenerate shortcutโrelying excessively on kinematic inertia from preceding frames while discounting weaker acoustic cues. Consequently, robots collapse into repetitive, monotonous motion loops regardless of incoming speech. Furthermore, unconstrained avatar motion spaces mapped onto a physical 41-degree-of-freedom humanoid (17 DoFs for upper torso and arms, 24 DoFs for dexterous hands) inevitably suffer from severe high-frequency motor jitter, velocity exceedance, and hazardous self-collisions (such as arms penetrating the torso or entangled fingers).
To overcome these obstacles, this work discards traditional post-hoc retargeting and adopts a robot-centric data-model-control co-design. Core idea: unify multi-granular acoustic-semantic disentanglement, anti-inertia conditional flow matching, and online MPC safety filtering to generate real-time, semantically aligned, and collision-free co-speech gestures directly within the humanoid robot's native action space.
Method¶
Overall Architecture¶
RoboGesture is a streaming-to-streaming co-speech generation framework tailored for a 41-DoF humanoid robot (instantiated on a Unitree G1 equipped with dual BrainCo dexterous hands). The system processes continuous incoming streaming audio chunks and produces motion trajectories in chunk-level increments of \(T = 30\) frames (1 second of motion at 30 FPS), conditioned on the past \(T_{\text{hist}} = 2T = 60\) frames of historical motion to ensure temporal smoothness.
The architecture comprises three tightly coupled modules: first, a Hierarchical Semantic-Acoustic Aligner tokenizes raw streaming audio using a neural audio codec and decouples it into low-level rhythmic pulses and high-level discrete semantic classes through multi-task auxiliary supervision; second, a Streaming Conditional Motion Generator based on a Diffusion Transformer (DiT) uses Conditional Flow Matching (CFM) to synthesize continuous robot joint velocities, guided by Cross-Attention for frame-level acoustic/motion alignment and FiLM for global semantic tone modulation; within this generator, Anti-Inertia CFG Masking and Kinematic Supervision breaks the modality eclipse via random history context masking while constraining dexterous hand kinematics and motion smoothness; finally, an MPC-based Kinematic Safety Filter executes per-frame quadratic programming in 5.6 ms during inference to eliminate self-collision risks before sending trajectory commands to low-level PD motor controllers.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Streaming Raw Audio Chunks<br/>Mimi RVQ Tokens"] --> B["Hierarchical Semantic-Acoustic Aligner<br/>Shallow Beat & Deep Semantic Disentanglement"]
B --> C["Dual-Path Conditioning & Flow Matching<br/>FiLM Global Modulation + Cross-Attention Alignment"]
C --> D["Anti-Inertia CFG & Spatiotemporal Kinematics<br/>Context Masking + Hand & Jitter Regularization"]
D --> E["MPC-based Kinematic Safety Filter<br/>Real-Time OSQP Collision Avoidance"]
E --> F["Humanoid Robot Execution<br/>Unitree G1 + BrainCo 41-DoF"]
Key Designs¶
1. Hierarchical Semantic-Acoustic Aligner: Text-Free Multi-Granular Acoustic Disentanglement Targeting the latency overhead and prosodic erasure inherent in text-based ASR pipelines, this module directly mines multi-granular acoustic cues from raw audio tokens. The system employs the Mimi neural audio codec to convert streaming audio into Residual Vector Quantization (RVQ) tokens: the primary quantizer captures high-level semantic representations, while subsequent residual quantizers preserve acoustic nuances such as pitch, volume, and intonation. A hierarchical transformer backbone processes these tokens with multi-task auxiliary supervision. Shallow layers extract rapid energy transients routed to a Beat Head, supervised by a Mean Squared Error (MSE) target \(B_{\text{target}}\) combining normalized joint velocity magnitude and acoustic onset strength. Deep layers compress contextual semantics into a Semantic Head predicting 300 discrete gesture categories via Cross-Entropy loss over 1,000 hours of semi-synthetic data. The shallow rhythmic pulses and deep semantic tokens provide complementary guidance, capturing pre-phonetic anticipation (e.g., physical preparation 0.4s prior to emphatic words) that text-based representations fail to express.
2. Dual-Path Conditioning & Flow Matching: Decoupled Micro-Alignment and Macro-Modulation Traditional discrete-token autoregressive models suffer from quantization coarseness on high-DoF dexterous hands and recursive error accumulation over extended horizons. RoboGesture adopts a continuous Diffusion Transformer (DiT) parameterized by Conditional Flow Matching (CFM) directly in the 41-dimensional robot joint space. To balance millisecond-level rhythmic responsiveness with long-term semantic stability, a dual-path condition injection mechanism is introduced: (1) Framewise Micro-alignment: cross-attention projects shallow rhythmic features \(\mathcal{H}_{\text{low}}\), deep local representations \(\mathcal{H}_{\text{high}}\), and historical motion chunks into Key-Value pairs, enforcing tight frame-by-frame temporal synchronization across chunk boundaries; (2) Global Macro-modulation: compressed semantic classifications from the aligner act as overarching behavioral anchors, injected via Feature-wise Linear Modulation (FiLM) to apply affine transformations to intermediate DiT feature maps, thereby governing global emotional posture without corrupting local fine-grained kinematic rhythms.
3. Anti-Inertia CFG & Spatiotemporal Kinematics: Breaking Modality Eclipse and Hand Regularization In online conditional motion synthesis, generators readily exploit historical kinematic inertia as a shortcut, ignoring incoming speech variations (the "modality eclipse"). To force the model to proactively harvest control signals from the audio stream, an Anti-Inertia Classifier-Free Guidance (CFG) masking strategy randomly masks the past motion context with a 15% probability during training, compelling the network to perform audio-driven "cold-start" generations. Furthermore, because subtle 24-DoF dexterous hand movements are easily overwhelmed by larger torso-arm dynamics, a spatial weight matrix \(W_s\) (\(w_{\text{hand}} = 4.0\) for hand joints) and a temporal weight vector \(W_t\) (\(w_{\text{frame}} \in \{5, 10\}\) for semantic intervals) are applied to the velocity matching loss. To suppress high-frequency hardware shuddering, a single-step Euler approximation yields a Kinetic Consistency Loss (\(\mathcal{L}_{\text{kin}}\)) penalizing inter-frame displacement differences: $$ \mathcal{L}{\text{Stage 2}} = \mathcal{L}}} + \lambda_{\text{kin}}\mathcal{L{\text{kin}} + \lambda) $$}}\text{CE}(\hat{Y}_{\text{target}
4. MPC-based Kinematic Safety Filter: Real-Time Convex Optimization for Collision-Free Execution While generating motions directly within robot joint space circumvents standard retargeting distortions, generative flow stochasticity cannot theoretically guarantee zero hardware collisions. RoboGesture incorporates an online Model Predictive Control (MPC) safety filter operating as a per-frame Convex Quadratic Program (QP) solved via OSQP. The optimizer minimizes tracking error and acceleration roughness subject to joint velocity limits and collision-avoidance constraints among robot mesh bounding geometries. Operating with an average latency of only 5.6 ms per frame, this filter reduces the self-collision frame ratio of generated motions from 4.16% to 0.13%, establishing strict hardware safety before driving physical motors.
Loss & Training¶
The framework is optimized in two consecutive stages: - Stage 1 (Representation Disentanglement Pre-training): The DiT generator remains frozen. The aligner is trained on 1,000 hours of chunk-level semi-synthetic data using a \(3T\)-frame audio window (current chunk plus history). The shallow layers optimize beat regression while deep layers optimize 300-class gesture classification: \(\mathcal{L}_{\text{Stage 1}} = \lambda_{\text{beat}}\text{MSE}(\hat{B}_{\text{target}}) + \text{CE}(\hat{Y}_{\text{target}})\). - Stage 2 (Joint Generative Training): The DiT generator is unfrozen and trained jointly on the 1,000-hour semi-synthetic dataset blended with a \(3\times\) upsampled BEAT dataset (76 hours). Supervision combines the spatiotemporally weighted velocity matching loss \(\mathcal{L}_{\text{vel}}\), kinetic consistency loss \(\mathcal{L}_{\text{kin}}\), and semantic cross-entropy loss.
Key Experimental Results¶
Main Results¶
Evaluation is conducted on both the standard BEAT benchmark and the proposed SemanticBEAT test set (comprising 1,000 real-world speech videos with manually annotated semantic intervals). Baselines include LivelySpeaker, DiffSHEG, SemTalk, and Semantic Gesticulator (SG). Metrics include Frรฉchet Gesture Distance (FGD), Beat Consistency (BC), structural Mean Squared Error (MSE), Diversity (DIV), and physical Collision Rate (Col.).
| Dataset | Method | FGD โ | BC โ | MSE โ | DIV โ | Col. (%) โ |
|---|---|---|---|---|---|---|
| BEAT | LivelySpeaker (ICCV 2023) | 3.0350 | 0.1811 | 0.1861 | 0.1485 | 21.41 |
| DiffSHEG (CVPR 2024) | 2.2316 | 0.1851 | 0.1753 | 0.1218 | 0.85 | |
| SemTalk (ICCV 2025) | 7.9328 | 0.1828 | 0.3931 | 0.1781 | 52.82 | |
| Semantic Gesticulator (SIGGRAPH 2024) | 3.0147 | 0.1771 | 0.2375 | 0.2818 | 13.11 | |
| RoboGesture (Ours) | 0.8452 | 0.1866 | 0.1347 | 0.2075 | 0.88 | |
| SemanticBEAT | LivelySpeaker (ICCV 2023) | โ | 0.2852 | โ | 0.1384 | 13.36 |
| DiffSHEG (CVPR 2024) | โ | 0.2859 | โ | 0.1209 | 1.20 | |
| SemTalk (ICCV 2025) | โ | 0.2913 | โ | 0.1440 | 42.65 | |
| Semantic Gesticulator (SIGGRAPH 2024) | โ | 0.2906 | โ | 0.2831 | 13.84 | |
| RoboGesture (Ours) | โ | 0.2950 | โ | 0.2041 | 0.13 |
Note: SemanticBEAT is an out-of-domain evaluation set lacking ground-truth motion capture data; thus, FGD and MSE are omitted.
Ablation Study¶
A subjective user study involving 200 participants evaluates key design choices across four dimensions: Semantic Action Score (SA), Hand Detail Score (HD), Human-likeness & Naturalness (HN), and Beat Matching Score (BM).
| Category | Configuration | SA โ | HD โ | HN โ | BM โ | Note |
|---|---|---|---|---|---|---|
| Data Scale | Ours w/o Semi-data | 4.281 | 2.321 | 4.945 | 4.628 | Collapses into rhythmic beat waving without semantic hand gestures |
| Ours (1/4 Semi-data, ~250h) | 6.316 | 5.980 | 6.882 | 6.892 | Sub-optimal semantic expressiveness and hand detail | |
| Architecture & Strategy | Ours w/o Context Motion | 4.237 | 4.263 | 2.192 | 1.928 | Severe inter-chunk discontinuities; drastic drop in naturalness |
| Ours w/o FiLM Injection | 5.181 | 4.389 | 4.506 | 4.589 | Loss of global semantic anchoring weakens intent delivery | |
| Ours w/o Semantic Classification | 5.007 | 4.747 | 4.750 | 5.268 | Aligner fails to distill explicit semantic cues from audio tokens | |
| Ours w/o CFG Masking | 4.628 | 4.885 | 6.453 | 6.212 | Suffers from modality eclipse; over-relies on kinematic inertia | |
| Autoregressive (AR) Strategy | 4.843 | 5.031 | 5.358 | 6.850 | Error accumulation causes long-term semantic drift | |
| Loss & Refinement | Ours w/o Kinetic-Aware (KA) Loss | 5.573 | 3.947 | 4.763 | 5.587 | Severe high-frequency joint jittering degrades hand clarity |
| Ours w/o MPC Safety Filter | 6.541 | 6.913 | 6.891 | 7.387 | Occasional self-collisions harm perceived kinematic smoothness | |
| Full Model | Ours (Full Model) | 7.175 | 7.203 | 7.394 | 7.529 | Achieves superior performance across all dimensions |
Key Findings¶
- Semi-synthetic data is the critical enabler for semantic gestures: Completely omitting semi-synthetic data causes Hand Detail (HD) to plummet from 7.203 to 2.321, turning the model into a purely rhythmic waving policy and demonstrating the indispensable value of large-scale annotated gesture data.
- Anti-inertia CFG masking directly resolves modality eclipse: Removing CFG masking drops the Semantic Action score (SA) from 7.175 to 4.628, proving that randomly forcing "cold-start" generations compels the model to actively mine conditioning cues from audio rather than passively riding historical inertia.
- Inference speed substantially surpasses real-time constraints: On the Unitree G1 platform, the streaming motion generator achieves \(\approx 120\) FPS (\(\approx 0.25\) s per 1-second chunk) and the MPC safety filter adds only 5.6 ms per frame, operating well above the 30 Hz hardware control loop. The dominant system latency stems from upstream speech pipeline components (ASR, LLM, TTS), where first-chunk audio duration is typically \(<1.5\) s.
Highlights & Insights¶
- Native robot-space generation circumvents kinematic retargeting pitfalls: Conventional approaches generate human motions before retargeting to humanoids, which frequently causes joint singularities and self-penetration; RoboGesture synthesizes directly in the 41-DoF robot action space, ensuring both kinematic expressiveness and mechanical executability.
- Text-free hierarchical audio decoupling captures pre-phonetic anticipation: By operating directly on neural audio codec tokens rather than intermediate ASR text, the model preserves rich prosody and successfully models the 0.4s pre-phonetic bodily preparation prior to verbal emphasis.
- Anti-inertia CFG masking is broadly applicable to temporal generation: The strategy of randomly masking temporal historical conditioning to counter modality eclipse offers an elegant, lightweight solution for diverse sequential multimodal diffusion and flow-matching tasks.
Limitations & Future Work¶
- Author-admitted limitations: The current implementation restricts active generation to the 41-DoF upper body and dexterous hands while the lower body maintains a static standing pose; it lacks full-body coordination integrating locomotion and gesticulation (Whole-Body Loco-gesticulation).
- Observed limitations: While the motion stack runs at 120 FPS, the overall human-robot conversation turnaround is gated by the upstream ASR-LLM-TTS speech pipeline (first-chunk audio latency \(\approx 1.5\) s), leaving slight turn-taking delays in fast-paced dialogues.
- Future directions: Integrating whole-body locomotion control with dynamic balance via reinforcement learning, and exploring tighter joint fine-tuning between speech foundation models and motion heads to compress first-packet latency into the sub-second regime.
Related Work & Insights¶
- vs Semantic Gesticulator (SIGGRAPH 2024): SG uses offline LLM retrieval and discrete motion concatenation, incurring high latency, failing streaming support, and exhibiting a 13.11% collision rate; RoboGesture achieves 120 FPS continuous flow generation and suppresses collision rates to 0.13%-0.88%.
- vs SemTalk (ICCV 2025): SemTalk relies on text features that miss vocal prosody and pre-phonetic anticipation, while showing over 40% collision rates when transferred to humanoids; RoboGesture's text-free pipeline and MPC filter provide significantly higher rhythmic fidelity and physical safety.
Rating¶
- Novelty: โญโญโญโญโ [Pioneering data-model-control co-design for humanoid co-speech gestures solving modality eclipse and self-collision]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive benchmark evaluations, out-of-domain tests, 200-subject ablation study, and complete physical deployment on Unitree G1]
- Writing Quality: โญโญโญโญโญ [Structured, rigorous formulation of flow matching, loss objectives, and real-time optimization filters]
- Value: โญโญโญโญโญ [Provides a production-grade, collision-free, high-FPS motion foundation for humanoid social interaction]