Skip to content

Ego-Human Motion Prediction with 3D-Aware LLM

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://jaewoo97.github.io/Ego3DLM/
Area: 3D Vision
Keywords: Egocentric Motion Prediction, Multimodal LLM, 3D Scene Awareness, Reinforcement Learning, Cross-Modal Alignment

TL;DR

Ego3DLM casts egocentric motion forecasting into a 3D scene-aware language model that simultaneously decodes past and future 3D poses alongside natural language narrations in a single autoregressive pass, refined by multi-modal GRPO to ensure physical plausibility and semantic fidelity.

Background & Motivation

Predicting human motion from an egocentric perspective is foundational for proactive assistance in AR/VR, human-robot collaboration, and embodied AI. In contrast to third-person video, egocentric observations directly mirror the wearer's authentic field of view, capturing immediate intent, environmental affordances, and near-field object interactions. Nevertheless, egocentric forecasting is inherently ill-posed: the wearer's camera captures little of their own body, proprioceptive input is restricted to sparse wearable signals like head and hand poses, and identical partial observations are compatible with numerous physically viable future trajectories.

Recent motion-language models attempt to alleviate this ambiguity by injecting text as a semantic prior or discretizing human motion into motion tokens for unified modeling. However, the majority of existing techniques focus on motion reconstruction, captioning, or editing from third-person perspectives, largely neglecting the explicit 3D spatial and semantic context that governs environmental affordances, obstacles, and physical movement constraints. Even when 3D scene geometry is incorporated, prior works typically inject scene representations as passive conditioning without instilling genuine spatial reasoning; furthermore, they decouple pose forecasting and language generation into disjoint inference streams, discarding the mutual regularization between physical actions and underlying semantic goals.

The core premise of this work is that accurate motion anticipation requires explicit spatial and semantic comprehension of the surrounding 3D environment, and that 3D pose and language narration must be predicted holistically in a single pass to anchor physical dynamics to semantic intent. Core idea: develop Ego3DLM, a 3D-aware multimodal LLM that sequentially generates spatial scene reasoning, past pose, future pose, past narration, and future description in a single autoregressive sequence, optimized via Group Relative Policy Optimization (GRPO) with intra- and inter-modal alignment rewards.

Method

Overall Architecture

Ego3DLM consumes three modalities: sparse three-point tracking data (6D poses of head and hands), continuous egocentric video frames, and reconstructed 3D point cloud scene features. The model produces a structured autoregressive sequence that unifies spatial scene reasoning, discrete human motion tokens (past and future), and corresponding natural language descriptions.

To bridge 3D geometry with temporal kinematics and high-level semantics, the training scheme is structured into three progressive stages: 1. 3D Scene Feature Extraction and Spatial-Semantic Scene Awareness Pretraining (Stage I): lifts dense 2D semantic representations onto the 3D point cloud and pretrains the language model on directional clearance and scene semantic QA; 2. Multi-Modal Multi-Task Instruction Tuning (Stage II): trains holistic single-pass autoregressive generation across spatial reasoning, past/future poses, and past/future narrations; 3. Multi-Modal Reward GRPO Reinforcement Finetuning (Stage III): samples candidate trajectories and applies group-relative policy optimization with pose, language, and inter-modal distance rewards to maximize physical plausibility and cross-modal fidelity.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Inputs: Three-point Tracking + Video + 3D Point Cloud"] --> B["3D Scene Feature Extraction & Alignment<br/>2D feature lifting + Q-Former compression"]
    B --> C["Spatial-Semantic Scene Awareness Pretraining<br/>Directional clearance & semantic QA"]
    C --> D["Multi-Modal Multi-Task Instruction Tuning<br/>Spatial reasoning → poses → text narration chain"]
    D --> E["Multi-Modal Reward GRPO<br/>Inter-modal text-motion distance optimization"]
    E --> F["Outputs: Coherent 3D Motion & Grounded Descriptions"]

Key Designs

1. 3D Scene Feature Extraction: Lifting 2D open-vocabulary semantics to point clouds with compact query compression

To address the limited field of view in egocentric video and the sparse semantics in raw point clouds, the model extracts dense 2D instance masks and visual features using Mask2Former and EVA-ViT-G. Using per-frame camera poses and a closest-viewpoint selection strategy, these features are lifted onto the 3D point cloud while suppressing occlusion artifacts. The enriched point cloud is aligned to the egocentric initial coordinate frame, voxelized, augmented with 3D sinusoidal positional encodings, and compressed into \(K=32\) query embeddings \(\mathbf{S}\) via a Q-Former. Concurrently, frame-wise CLIP visual features \(\mathbf{V}\) are linearly mapped into the language model's latent space, providing unified visual-spatial conditioning.

2. Spatial-Semantic Scene Awareness Pretraining: Internalizing environmental layout and navigable free space

Before introducing human kinematic motion, the model is pretrained on a large-scale automatically generated QA corpus to establish a spatial-semantic prior. The semantic awareness set encompasses six categories (object identity, location, attributes, color, counting, and presence). The spatial awareness set partitions the forward space into three directional sectors (front, left, right), querying directional clearance levels (low, mid, high), obstruction status (blocked, clear), and the optimal navigable direction. Both objectives are optimized via joint autoregressive negative log-likelihood: $\(\mathcal{L}_{\text{Pre}} = - \sum_i \left( \log p_\theta(a_i^{\text{sem}} \mid a_{<i}^{\text{sem}}, q^{\text{sem}}, \mathbf{E}) + \log p_\theta(a_i^{\text{spa}} \mid a_{<i}^{\text{spa}}, q^{\text{spa}}, \mathbf{E}) \right)\)$ where \(\mathbf{E} = (\mathbf{S}, \mathbf{V})\). This stage equips the LLM with obstacle layout understanding and spatial affordances prior to motion synthesis.

3. Multi-Modal Multi-Task Instruction Tuning: Chain-of-thought spatial grounding to holistic cross-modal decoding

Unlike prior pipelines that separate tracking from prediction or isolate pose from text, Stage II executes unified instruction tuning. Given tracking features \(\mathbf{P}\), video \(\mathbf{V}\), and scene embeddings \(\mathbf{S}\), the model decodes in a strict sequence: [spatial reasoning a^spa] → [past pose x^past] → [future pose x^fut] → [past narration y^past] → [future description y^fut]. Human poses are parameterized by a 23-joint kinematic tree and discretized into 4096 codebook tokens via PQ-VAE, extending the LLM's vocabulary. The prepended spatial reasoning acts as a chain-of-thought buffer that steers kinematic trajectory generation away from obstacles, while the generated motion states subsequently guide semantic narration.

4. Multi-Modal Reward GRPO: Reinforcement learning for joint motion-language consistency

Supervised likelihood training alone cannot guarantee mutual semantic coherence between simultaneously generated poses and text descriptions. Stage III applies Group Relative Policy Optimization (GRPO) without a separate value critic, sampling \(G\) candidate outputs per prompt. The scalar reward incorporates intra-modal accuracy (\(R_{\text{motion}} = \max(0, 1 - \text{JPE})\) and \(R_{\text{text}} = \text{BLEU-4}\)) alongside format penalties. Critically, it incorporates an inter-modal matching reward \(R_{\text{matching}}\) computed via pretrained cross-modal encoders in a shared embedding space: $\(R_{\text{matching}} = - \left[ d_{\text{gp}}(\mathbf{e}_{t_{\text{gt}}}, \mathbf{e}_{m_{\text{pred}}}) + d_{\text{pg}}(\mathbf{e}_{t_{\text{pred}}}, \mathbf{e}_{m_{\text{gt}}}) + d_{\text{pp}}(\mathbf{e}_{t_{\text{pred}}}, \mathbf{e}_{m_{\text{pred}}}) \right]\)$ Here, \(d_{\text{pp}}\) directly penalizes discrepancies between the simultaneously predicted text and predicted motion trajectory, steering the model toward outputs that are physically and semantically concordant.

Key Experimental Results

Main Results

Evaluation on the in-the-wild Nymeria benchmark comparing future motion forecasting (5-second duration) and past motion tracking. The backbone model is GPT-2 Medium.

Method Pred APE ↓ Pred JPE ↓ Pred ADE2s ↓ Pred FDE2s ↓ Pred FID ↓ Track APE ↓ Track Upper ↓ Track Lower ↓ Text Bleu-4 ↑ Alignment \(d_{\text{pp}}\) ↓
FIction 206.2 564.7 416.7 494.9 0.3275 181.1 114.0 282.3 - -
EgoLM (SFT) 184.9 579.4 329.9 519.2 0.2137 161.9 95.6 265.6 0.0649* 9.8686
UniEgoMotion 151.5 424.3 223.9 360.4 0.1530 152.2 79.4 233.3 - -
Ego3DLM (Ours) 147.9 364.5 205.9 312.6 0.0160 96.4 53.1 152.7 0.1039 4.2571

(Note: EgoLM only produces text narration for past tracking with Bleu-4 of 0.0649; Ego3DLM attains 0.1107 on past narration Bleu-4.)

Ablation Study

Ablation of training stages and structural components under instruction tuning (evaluated relative to the full SFT model):

| Config | Pred APE ↓ | Pred JPE ↓ | Pred ADE2s ↓ | Track APE ↓ | Text R@3 ↑ | Alignment \(d_{\text{pp}}\) ↓ | Note | |---|---|---|---|---|---|---| | Full SFT model | 148.5 | 368.3 | 208.8 | 98.3 | 0.2706 | 5.9170 | Full pretraining & spatial reasoning | | w/o 3D Scene (w/o S) | 156.2 | 388.5 | 220.8 | 105.8 | 0.2574 | 6.3012 | Lacks physical scene boundaries; errors spike | | w/o Spatial QA (w/o Qspa) | 152.9 | 378.3 | 214.3 | 99.3 | 0.2453 | 6.1612 | Lacks obstacle clearance prior; alignment drops | | w/o Semantic QA (w/o Qsem) | 153.5 | 381.1 | 221.9 | 122.1 | 0.2558 | 5.9190 | Tracking degrades sharply (+24.2% APE) | | w/o Spatial Reasoning (w/o SSR) | 151.5 | 376.6 | 213.4 | 99.4 | 0.2638 | 6.3930 | Unbuffered CoT leads to severe alignment degradation |

Key Findings

  • 3D Scene Conditioning is Vital for Physical Boundary Adherence: Removing explicit 3D scene inputs increases global joint error (JPE) from 368.3 to 388.5, and qualitative visualizations confirm severe collisions with walls and furniture without 3D geometry.
  • Explicit Spatial Scene Reasoning Acts as Chain-of-Thought: Prepending obstacle reasoning before trajectory generation noticeably boosts motion-language alignment (\(d_{\text{pp}}\) drops from 6.3930 to 5.9170), indicating that environmental geometry serves as a crucial semantic stepping stone.
  • Mutual Synergies in Cross-Modal GRPO: Introducing \(R_{\text{matching}}\) in Stage III contracts the predicted text-motion distance \(d_{\text{pp}}\) from 5.9170 to 4.2571 (a 28% reduction) while concurrently lifting pose accuracy and BLEU metrics, demonstrating that multimodal alignment reinforces both individual modalities.

Highlights & Insights

  • Discretized Motion in Autoregressive LLM Vocabulary: Mapping 4096 PQ-VAE motion codes directly into the LLM token dictionary enables standard causal transformers to jointly synthesize continuous kinematics and discrete text tokens.
  • Critic-Free Multimodal GRPO: Converting cross-modal embedding distances into group-relative policy rewards bypasses the instability and overhead of multi-modal value networks, directly enforcing physical-semantic correspondence.
  • Structured Reasoning Pipeline: The progressive sequence—spatial environment understanding, kinematic motion planning, and high-level linguistic interpretation—provides an effective blueprint for embodied agent planning.

Limitations & Future Work

  • Reliance on Pre-Reconstructed 3D Point Clouds: The framework assumes access to a pre-existing 3D scene representation, limiting deployment in unmapped, dynamic environments.
  • Coarse Handling of Fine-Grained Hand-Object Interactions: While effective for whole-body navigation and obstacle avoidance, prediction accuracy remains limited for dexterous finger manipulation and small tool use.
  • Future Directions: Integrating online incremental 3D reconstruction (e.g., 3D Gaussian Splatting) into the LLM pipeline for real-time exploratory embodied navigation.
  • vs EgoLM: EgoLM operates solely on video and sparse tracking without 3D scene topology, separating pose and text generation into isolated streams. Ego3DLM incorporates explicit 3D grounding and single-pass joint decoding with GRPO, reducing global joint error JPE by 37.1% and cutting cross-modal alignment distance by more than half.
  • vs UniEgoMotion: UniEgoMotion utilizes 3D scenes for motion prediction but completely lacks language understanding and intention interpretation. Ego3DLM outperforms UniEgoMotion across all kinematic metrics while providing synchronized natural language explanations.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering framework combining explicit 3D scene context, single-pass four-output autoregressive decoding, and multi-modal GRPO for egocentric human motion forecasting.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation on the in-the-wild Nymeria benchmark covering motion, text, and cross-modal embedding alignment with detailed ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Highly coherent presentation with clear methodological motivations, detailed equations, and well-structured empirical validation.
  • Value: ⭐⭐⭐⭐⭐ Establishes a compelling paradigm for anticipatory AR/VR systems, proactive assistive robotics, and embodied AI.