EgoGVAE: Ego-body Mesh Reconstruction via Guided Variational Autoencoder¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/DCVL-3D/EgoGVAE_release
Area: 3D Vision
Keywords: Egocentric Human Mesh Reconstruction, Variational Autoencoder, Latent Space Guidance, One-Step Sampling, Real-Time Inference
TL;DR¶
Addressing the severe latency and computational cost of diffusion-based head-to-motion generation, EgoGVAE leverages a full-body motion variational autoencoder as a training-phase prior to guide the latent distribution of a head-conditioned network, achieving superior full-body mesh recovery with single-step sampling at over 50x faster inference speed.
Background & Motivation¶
With the advent of physical AI and interactive spatial computing, smart glasses (such as Project Aria) and lightweight head-mounted displays (HMDs) are becoming indispensable daily wearable interfaces. In these interactive egocentric scenarios, estimating full-body 3D motions of the wearer is critical for avatar animatronics, assistive robotics, and immersive communication. However, because egocentric cameras look outward and cannot view the user's torso or limbs, inferring unobserved full-body kinematics solely from a single joint trajectory—the 3D head pose derived from visual SLAM—presents an acutely ill-posed challenge.
Early attempts mitigated this under-constrained formulation by incorporating auxiliary hand sensors or wristbands, which severely hindered consumer ergonomics and general accessibility. More recently, several studies turned to diffusion-based probabilistic generation models (e.g., EgoEgo and EgoAllo) to model full-body distributions conditioned on head poses. While diffusion models capture multi-modal distributions reasonably well, their iterative denoising process incurs formidable computational costs (taking several seconds and hundreds of GFLOPs per 128 frames); aggressively reducing denoising steps invariably triggers severe temporal jitter and unstable limb configurations.
This paper approaches the problem from a distinct angle: instead of paying the heavy runtime penalty of iterative diffusion, one can harvest the expressive power of a full-body motion manifold via variational latent guidance. By constructing an auxiliary variational autoencoder that takes ground-truth full-body motions during training and aligning the head-conditioned network's latent distribution with this prior, the generator can reliably synthesize natural full-body poses in a single sampling step during inference. The core idea is to establish a guided variational autoencoder framework (EgoGVAE) where learnable body tokens complement the unobserved limbs, aligning the head-to-motion latent distribution with a full-body motion prior during training to enable high-fidelity, single-step mesh decoding at test time.
Method¶
Overall Architecture¶
EgoGVAE comprises two transformer-based variational autoencoder networks: an auxiliary Motion-to-Motion Network that encodes ground-truth full-body motion sequences into a target latent Gaussian distribution, and a Head-to-Motion Network that encodes normalized head trajectories augmented with learnable limb tokens. During training, the two networks are trained jointly with symmetric Kullback–Leibler divergence to enforce latent alignment. At inference time, the motion-to-motion guidance branch is discarded entirely; the head-to-motion network encodes the wearer's head trajectory, samples a latent vector via one-step reparameterization, and decodes realistic SMPL-H body meshes and contact probabilities through a shared transformer decoder.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Head Trajectory<br/>(16D normalized relative pose)"] --> B["Head Embedding & Learnable Body Token Concatenation"]
B --> C["Head-to-Motion Variational Encoder<br/>Predicts (μ^H, Σ^H)"]
D["Training Full-body Pose Sequence<br/>(Local joint rotations)"] --> E["Motion-to-Motion Variational Encoder<br/>Predicts (μ^M, Σ^M)"]
C & E --> F["Variational Latent Space Alignment<br/>(Symmetric KL & Standard Normal Regularization)"]
F --> G["One-step Reparameterized Sampling z^H"]
G --> H["Shared Transformer Motion Decoder"]
H --> I["Predicted SMPL-H Parameters & Contact Probabilities"]
Key Designs¶
1. Variational Latent Guidance: Full-Body Motion Prior Aligning Head Generation Predicting full-body human motion from only the head trajectory is heavily ill-posed and prone to generating static arms or floating lower bodies. EgoGVAE designs a Motion-to-Motion auxiliary branch as a variational guide. Appending specialized mean and variance tokens (\(\boldsymbol{\mu}_{token}^M, \boldsymbol{\Sigma}_{token}^M\)) to the input full-body joint rotations, the transformer encoder produces sequence-level Gaussian parameters \(\mathcal{N}(\boldsymbol{\mu}^M, \boldsymbol{\Sigma}^M)\). The Head-to-Motion network mirrors this variational formulation to predict \(\boldsymbol{\mu}^H, \boldsymbol{\Sigma}^H\). During training, the model minimizes the symmetric Kullback–Leibler divergence between the two distributions:
Simultaneously, both distributions are regularized toward a standard normal distribution \(\mathcal{N}(\mathbf{0}, \mathbf{I})\). This enforces the head-driven latent space to faithfully mirror the manifold of natural human movements, enabling single-step forward sampling without multi-step denoising.
2. Learnable Body Tokens: Bridging the Input Asymmetry Gap A 16-dimensional head pose vector per frame (comprising relative translation, canonicalized orientation, and floor height) has a profound dimensional and semantic mismatch with 21-joint full-body rotation inputs. Feeding solitary head tokens directly into the transformer encoder struggles to establish self-attention patterns compatible with full-body sequences. EgoGVAE introduces temporal learnable tokens \(L_{token} \in \mathbb{R}^{T \times \text{dim}}\) to serve as explicit structural proxies for unobserved torso and limb kinematics. A lightweight transformer projects head poses into embeddings \(E_W \in \mathbb{R}^{T \times \text{dim}}\), which are concatenated channel-wise with \(L_{token}\) prior to variational encoding. This provides the self-attention layers with structural handles for the missing limbs, effectively closing the representational gap between head inputs and full-body priors.
3. Shared Temporal Transformer Decoder with Decoupled Queries To translate sampled latent representations back into coherent frame-by-frame kinematics, the framework utilizes a weight-shared transformer motion decoder. In the head-to-motion stream, the globally sampled latent vector \(\mathbf{z}^H \sim \mathcal{N}(\boldsymbol{\mu}^H, \boldsymbol{\Sigma}^H)\) is projected as the Key and Value in cross-attention blocks, while the temporal token embeddings (head plus learnable limb representations) serve as Query vectors. Consequently, the latent vector injects global motion style, joint coordination, and physical plausibility, while the temporal queries anchor the per-frame alignment against head motion. The decoder directly outputs SMPL-H local joint rotations \(\hat{\theta}_t\), invariant shape parameters \(\hat{\beta}\), and binary joint contact probabilities \(\hat{\psi}_t\).
Loss & Training¶
The entire architecture is optimized end-to-end using a joint objective balancing reconstruction fidelity, distributional regularization, and kinematic smoothness:
The reconstruction objective \(\mathcal{L}_\text{rec}\) incorporates joint rotation error \(\mathcal{L}_\text{rot}\), forward-kinematic 3D joint position error \(\mathcal{L}_\text{pos}\), SMPL-H shape error \(\mathcal{L}_\text{shape}\), and binary cross-entropy foot-floor contact error \(\mathcal{L}_\text{contact}\). The velocity loss \(\mathcal{L}_\text{vel}\) penalizes the first-order difference between predicted and ground-truth joint velocities to eliminate high-frequency jitter. Balancing hyperparameters are set to \(\lambda_\text{shape}=0.003, \lambda_\text{contact}=0.003, \lambda_\text{KL}=0.0004\), and \(\lambda_\text{vel}=0.003\). For arbitrary-length test sequences, a 128-frame sliding window processes online motion at 26 ms per frame.
Key Experimental Results¶
Main Results¶
The authors benchmarked EgoGVAE across 128-frame motion sequences on the AMASS dataset and performed cross-dataset generalization on the RICH benchmark (trained on AMASS, tested on RICH without fine-tuning), comparing against diffusion-based (EgoEgo, EgoAllo) and adapted regression baselines (AvatarPoser, EgoPoser).
| Dataset | Method | MPJPE (mm) ↓ | PA-MPJPE (mm) ↓ | Ground (mm) ↓ | \(T_{head}\) (mm) ↓ | Jitter (\(10^2 m/s^3\)) ↓ | Foot sliding (mm) ↓ |
|---|---|---|---|---|---|---|---|
| AMASS | AvatarPoser† | 142.2 | 127.5 | 48.1 | 42.9 | 4.74 | 13.5 |
| AMASS | EgoPoser† | 143.9 | 121.0 | 49.3 | 40.8 | 5.27 | 10.6 |
| AMASS | EgoEgo | 167.4 | 145.8 | 47.1 | 54.9 | 4.22 | 11.7 |
| AMASS | EgoAllo (SOTA) | 119.7 | 101.1 | 26.3 | 6.2 | 4.06 | 10.2 |
| AMASS | EgoGVAE (Ours) | 106.7 | 89.9 | 21.6 | 5.5 | 3.08 | 8.2 |
| RICH (Cross-eval) | AvatarPoser† | 297.5 | 287.4 | 281.8 | 75.6 | 4.43 | 12.4 |
| RICH (Cross-eval) | EgoPoser† | 314.4 | 305.9 | 161.3 | 108.5 | 5.12 | 9.5 |
| RICH (Cross-eval) | EgoEgo | 210.7 | 180.1 | 93.6 | 134.5 | 3.61 | 12.7 |
| RICH (Cross-eval) | EgoAllo (SOTA) | 193.1 | 172.9 | 71.9 | 7.7 | 5.07 | 15.9 |
| RICH (Cross-eval) | EgoGVAE (Ours) | 179.9 | 165.9 | 59.0 | 5.4 | 3.87 | 11.9 |
Note: † indicates baselines retrained using solely head pose inputs. On the real-world headset tracking dataset EgoBody, EgoGVAE also achieved an MPJPE of 162.9 mm, outperforming EgoAllo's 172.9 mm.
Efficiency & Ablation Study¶
The table below compares computational complexity and latency against diffusion models (evaluated on an RTX 3090 Ti for 128 frames) alongside systematic ablations of core architectural modules on AMASS.
| Model / Configuration | Params (M) ↓ | FLOPs (G) ↓ | Time (s) ↓ | MPJPE (mm) ↓ | PA-MPJPE (mm) ↓ | Description |
|---|---|---|---|---|---|---|
| EgoEgo (Diffusion) | 10.97 | 3025.13 | 7.222 | 167.4 | 145.8 | 1000-step iterative sampling, intractable latency |
| EgoAllo (Diffusion) | 50.45 | 378.16 | 1.510 | 119.7 | 101.1 | 30-step denoising, heavy parameters and >1.5s delay |
| EgoGVAE (Full Model) | 13.88 | 3.45 | 0.026 | 106.7 | 89.9 | One-step sampling, 58x speedup, lowest error |
| w/o Guidance | - | - | - | 125.6 | 108.6 | Discards motion-to-motion prior alignment |
| w/o Learnable Tokens | - | - | - | 113.7 | 94.5 | Encodes head token alone without limb proxies |
| Baseline (w/o Guidance & Tokens) | - | - | - | 128.3 | 110.3 | Simple regression VAE; highly degraded |
| Latent-to-Latent Mapping | - | - | - | 128.1 | 111.9 | Deterministic MLP mapping between latents fails |
| Fixed Pretrained Motion Prior | - | - | - | 111.2 | 94.5 | Freezing the guide restricts joint manifold learning |
Key Findings¶
- Latent guidance resolves head-only under-determination: Ablating the variational guide causes MPJPE to deteriorate from 106.7 mm to 125.6 mm (+18.9 mm), confirming that aligning the head-conditioned space with full-body dynamics prevents limb collapse.
- Drastic reduction in runtime complexity: In contrast to EgoAllo's 30 diffusion steps (1.51 s, 378.16 GFLOPs), EgoGVAE requires only 3.45 GFLOPs and 0.026 s—yielding a 58x speedup while simultaneously cutting parameters by 72.5%.
- Joint training outperforms frozen motion priors: Aligning the two variational models interactively through symmetric KL divergence outperforms transferring from a frozen motion prior (106.7 mm vs. 111.2 mm), ensuring the learned latent manifold is jointly conditioned.
Highlights & Insights¶
- Training-time guidance with zero test-time overhead: By treating the complete motion autoencoder as an implicit teacher in latent space, full-body priors are absorbed entirely into network weights, eliminating teacher computation or iterative steps during deployment.
- Learnable tokens as virtual kinematic anchors: Introducing trainable temporal tokens provides structural slots for missing bodily extremities, allowing self-attention layers to propagate head kinematics into leg and arm motions without hand-crafted heuristic links.
- Challenging diffusion model hegemony in wearable tracking: Demonstrates that for constrained conditional synthesis tasks like ego-body tracking, properly regularized variational autoencoders can outperform truncated diffusion models in both accuracy and physical plausibility.
Limitations & Future Work¶
- Reduced motion diversity in unconstrained dynamics: Variational Gaussian assumptions favor modal, smooth poses. For non-periodic, highly dynamic acrobatics or erratic movements, the reconstruction may tend slightly toward mean trajectories compared to unconstrained diffusion sampling.
- Absence of 3D scene geometry awareness: The model currently infers poses purely from head kinematics and implicit contact loss without explicit 3D scene point clouds, which occasionally causes penetration when interacting with furniture (e.g., chairs or tables).
- Future directions: Integrating sparse depth readings or environmental normal cues into the learnable token sequence could enforce collision-free contact constraints in complex scenes.
Related Work & Insights¶
- vs. EgoAllo (SIGGRAPH Asia 2024): EgoAllo adapts diffusion sampling on floor-aligned head poses; while accurate, it requires 1.5 seconds per 128 frames and 50M parameters. EgoGVAE's single-step variational formulation reduces latency by 58x and delivers better accuracy across joint error, jitter, and foot sliding.
- vs. AvatarPoser / EgoPoser: Conventional regression baselines rely heavily on auxiliary hand tracking. When restricted solely to head trajectories, their performance degrades drastically (MPJPE exceeds 142 mm). EgoGVAE proves that full-body latent distribution guidance provides sufficient structural regularities to compensate for missing extremities.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ [Presents an elegant, highly effective variational latent guidance strategy and learnable token design for head-only ego-pose estimation]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across AMASS, RICH cross-dataset testing, EgoBody real sensor data, 256-frame sequences, online streaming, and thorough architectural ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Well-structured narrative with crisp motivation, explicit mathematical formulations, and clear comparative figures]
- Value: ⭐⭐⭐⭐⭐ [Provides an immediate real-time solution for lightweight smart glasses and consumer VR headsets with open-sourced code]