EgoTraj: Real-World Egocentric Human Trajectory¶
Conference: ECCV 2026
Paper: ECCV Paper
Code: https://github.com/yehiahmad/EgoTraj
Area: Autonomous Driving
Keywords: Egocentric Trajectory Prediction / Multimodal Dataset / Gaze Tracking / Embodied Navigation / Consumer AR Headset
TL;DR¶
EgoTraj is the first open multimodal egocentric human trajectory dataset and benchmark collected with consumer AR headsets (Meta Quest Pro) in real-world outdoor urban environments, capturing synchronized 6DoF head poses, 3D gaze vectors, RGB videos, and VLM scene descriptions across 75 subjects navigating self-chosen routes, where cascaded cross-attention demonstrates the decisive predictive gain of gaze and scene priors.
Background & Motivation¶
Egocentric perception constitutes the fundamental perceptual substrate for human spatial navigation and obstacle avoidance in the physical world. When navigating dynamic urban environments, human pedestrians continuously perceive traffic flows, identify crossing opportunities, and anticipate turns; crucially, visual neuroscience demonstrates that eye gaze naturally fixates on navigation waypoints and obstacles 1 to 2 seconds before motor execution. Translating this intuitive perceptual ability into machines is central to embodied artificial intelligence, humanoid robotics, smart wheelchairs, and assistive navigation systems for blind and visually impaired individuals. Nevertheless, existing human trajectory prediction methods are predominantly trained on third-person benchmarks captured from static surveillance cameras or bird's-eye views (such as ETH/UCY, Stanford Drone Dataset, inD, and nuScenes). While these external perspectives accurately record past coordinate displacements, they inherently fail to capture how humans perceive their immediate surroundings and how internal visual attention governs future motor actions.
Although recent first-person computer vision efforts have introduced extensive datasets like Ego4D and Ego-Exo4D, their scopes concentrate heavily on action recognition and procedural video understanding rather than continuous trajectory forecasting and navigation telemetry. A limited number of pioneering egocentric trajectory datasets (such as KrishnaCam, EgoMotion, LookOut, and EgoCogNav) suffer from critical limitations: they are often restricted to indoor corridors or single-participant recordings, rely solely on chest-mounted cameras or offline SfM approximations without real-time 6DoF ground truth, or omit synchronized gaze tracking entirely. Most crucially, almost none leverage commercial AR headsets, leaving open questions regarding how trajectory models cope with rapid head ego-motion, dynamic view shifts, and low-power wearable edge sensors in naturalistic environments.
To bridge this gap between external motion forecasting and first-person intention-aware perception, the authors collect a rich multimodal navigation corpus using a commodity AR headset. Core idea: construct EgoTraj, the first 10.7-hour multimodal outdoor egocentric trajectory benchmark encompassing 75 participants walking unscripted urban routes with synchronized 6DoF head pose, 3D gaze telemetry, RGB video, and structured VLM scene annotations, demonstrating that gaze and scene context fundamentally alleviate trajectory prediction bottlenecks.
Method¶
Overall Architecture¶
The construction and benchmarking of EgoTraj span an integrated pipeline from hardware-software data acquisition and multimodal spatiotemporal synchronization to VLM-based scene annotation and cascaded cross-attention trajectory forecasting. Participants wear a Meta Quest Pro (MQPro) operating in full-color passthrough mode and walk between designated landmark pairs in open urban settings (sidewalks, crosswalks, and busy intersections) following their own unscripted route decisions. A custom Unity application interfaces with the headset's visual-inertial SLAM and infrared eye-tracking cameras to record 6DoF head poses, 3D gaze vectors, and RGB video at a synchronized 30 Hz rate, packaged into structured HDF5 files. A quadratic calibration mapping projects 3D gaze rays into image-plane pixel coordinates, while Qwen2.5-VL-32B generates structured scene and behavioral intent annotations. For downstream benchmarking, forecasting models ingest past ego-motion, social poses, scene segmentation, and gaze coordinates to predict future 3.5-second trajectories and 6DoF head orientations.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["MQPro Hardware Sensing<br/>6DoF Pose + 3D Gaze + RGB Video"] --> B["Multimodal Spatiotemporal Calibration<br/>30Hz Timestamp Sync + Gaze Pixel Reprojection"]
B --> C["VLM Scene Understanding & Anonymization<br/>EgoBlur Face/Plate Blurring + Qwen2.5-VL Annotation"]
C --> D["EgoViz Dashboard Quality Control<br/>Interactive Pose-Gaze-Video-Semantics Inspection"]
D --> E["Multimodal Cascaded Cross-Attention Benchmark<br/>CXA-Transformer / EgoCast Intent Modeling"]
E --> F["Future Trajectory & Rotation Output<br/>Future 3.5s Displacement ADE/FDE + Head Rotation Error"]
Key Designs¶
1. Naturalistic Outdoor Data Collection and Quadratic Gaze-to-Pixel Calibration To overcome the constraints of rigid scripted paths and cumbersome research rigs, data collection employs the standalone Meta Quest Pro headset in full-color passthrough mode. Using two infrared eye-tracking cameras, four tracking cameras, and a 6-axis IMU, the system streams 6DoF head poses (translation, orientation quaternion, linear velocity, angular velocity) and binocular 3D gaze vectors alongside H.264 RGB video at 30 Hz. Rather than following fixed tracks, 75 diverse participants navigate between 7 outdoor landmark waypoints (forming 21 origin-destination pairs) choosing their own routes. Pairwise Dynamic Time Warping (DTW) distances across identical waypoint pairs yield a median of 122.8 m (range 19.3 to 379.4 m), confirming substantial path diversity. To project 3D gaze rays into the egocentric camera frame, a per-session quadratic polynomial calibration model maps gaze yaw and pitch angles to 2D pixel coordinates \((u, v)\) on the image plane, providing direct visual grounding of spatial attention.
2. Vision-Language Model Pipeline for Navigation-Relevant Scene Annotation Raw kinematic data lacks high-level behavioral and semantic context. The authors establish an automated annotation pipeline on 253 privacy-blurred video segments (anonymized via EgoBlur for faces and license plates), sampling frames at 1 fps to generate 38,606 annotated frames using Qwen2.5-VL-32B-Instruct. Using few-shot exemplars and chain-of-thought (CoT) prompting, the model extracts egocentric environmental elements, traffic conditions, gaze fixation targets, and inferred navigation intent (such as following pedestrian flows or waiting at crosswalks). Across stratified sample audits, structural compliance reached 96% after up to two retries, and independent human review achieved Cohen's \(\kappa\) values of 0.83 to 0.96 across annotation fields, providing high-quality sidecar JSON annotations synchronized with the primary HDF5 files.
3. Multimodal Cascaded Cross-Attention Trajectory Forecasting Benchmark To benchmark trajectory prediction under severe egocentric motion and frequent turning, the evaluation adapts CXA-Transformer using an observation window of \(T_{\text{obs}} = 1.5\text{ s}\) (45 frames) to forecast a future horizon of \(T_{\text{pred}} = 3.5\text{ s}\) (105 frames). In addition to ego-motion history (\(Y\)), complementary spatial and social representations are extracted: nearby pedestrian poses (\(P\)), bounding boxes (\(B\)), and torso centers (\(C\)) via YOLOv8-Pose; semantic scene segmentation (\(S\)) via OneFormer; metric relative depth (\(D\)) via Depth Anything V2; and calibrated 2D gaze coordinates (\(G\)). Rather than simple early feature concatenation, the CXA-Transformer architecture sequentially routes ego-motion representations through cascaded cross-attention blocks conditioned on social poses, scene geometry, and gaze vectors. This allows the model to leverage gaze fixations as dynamic spatial priors that guide predicted trajectory curvature toward attended turning points and passable routes.
Loss & Training¶
Trajectory displacement is evaluated using Average Displacement Error (ADE) and Final Displacement Error (FDE) across the prediction horizon: $\(\text{ADE} = \frac{1}{T_{\text{pred}}} \sum_{i=1}^{T_{\text{pred}}} \|\hat{p}_{t+i} - p_{t+i}\|_2, \quad \text{FDE} = \|\hat{p}_{t+T_{\text{pred}}} - p_{t+T_{\text{pred}}}\|_2\)$ Head orientation prediction is evaluated in the rotation matrix manifold via mean \(L_1\) rotation distance (\(L_1^{\text{head}}\)): $\(\mathcal{L}_{\text{rot}} = \frac{1}{T_{\text{pred}}} \sum_{i=1}^{T_{\text{pred}}} \|\hat{R}_{t+i} R_{t+i}^{\top} - I\|_1\)$ The dataset is partitioned into subject-disjoint training (80%), validation (10%), and test (10%) splits. To assess out-of-distribution robustness, the benchmark also defines two stricter evaluation protocols: a Waypoint Held-Out split (holding out 3 landmark pairs across 10 sessions) and an Unfamiliar split (holding out 8 participants unfamiliar with the area), with 95% bootstrap confidence intervals computed over 1,000 resamples.
Key Experimental Results¶
Main Results¶
The benchmark compares standard kinematic baselines against neural sequence architectures on the held-out EgoTraj test split for a 1.5 s observation and 3.5 s forecasting window.
| Model Category | Model Name | Modalities | ADE (m) ↓ | FDE (m) ↓ | \(L_1^{\text{head}}\) ↓ |
|---|---|---|---|---|---|
| Kinematic Baseline | Const_Vel (Constant Velocity) | Ego Translation + Rotation | 0.24 | 0.35 | 0.82 |
| Kinematic Baseline | Lin_Ext (Linear Extrapolation) | Ego Translation + Rotation | 0.26 | 0.39 | 1.39 |
| Neural Model | M_Transformer (Early Fusion) | Ego Translation + Rotation | 0.20 | 0.32 | 0.74 |
| Neural Model | CXA-Transformer (Cascaded Attention) | Ego Translation + Rotation | 0.19 | 0.29 | 0.69 |
| Neural Model | EgoCast (Adapted Forecasting) | Ego Translation + Rotation | 0.16 | 0.28 | 0.78 |
Ablation Study¶
Using CXA-Transformer as the backbone, the ablation examines the individual and combined impact of ego-motion (\(Y\)), social cues (\(C, B, P\)), scene segmentation (\(S\)), relative depth (\(D\)), and projected gaze (\(G\)):
| Modality Configuration | Feature Description | ADE (m) ↓ | FDE (m) ↓ | \(L_1^{\text{head}}\) ↓ |
|---|---|---|---|---|
| \(Y\) | Ego-motion only (translation + rotation) | 0.19 | 0.29 | 0.69 |
| \(Y + C\) | Ego-motion + Pedestrian center points | 0.18 | 0.29 | 0.79 |
| \(Y + B\) | Ego-motion + Pedestrian 2D bounding boxes | 0.18 | 0.30 | 0.81 |
| \(Y + P\) | Ego-motion + Pedestrian keypoint poses (YOLOv8-Pose) | 0.17 | 0.27 | 0.77 |
| \(Y + S\) | Ego-motion + Scene semantic segmentation (OneFormer) | 0.16 | 0.26 | 0.74 |
| \(Y + D\) | Ego-motion + Relative depth (Depth Anything V2) | 0.18 | 0.29 | 0.78 |
| \(Y + G\) | Ego-motion + Projected 2D gaze coordinates \((u, v)\) | 0.15 | 0.26 | 0.69 |
| \(Y + C + G\) | Pedestrian centers + Gaze | 0.14 | 0.25 | 0.67 |
| \(Y + B + G\) | Pedestrian bounding boxes + Gaze | 0.16 | 0.26 | 0.70 |
| \(Y + P + G\) | Pedestrian keypoint poses + Gaze | 0.12 | 0.24 | 0.63 |
| \(Y + S + G\) | Scene segmentation + Gaze | 0.12 | 0.25 | 0.65 |
| \(Y + D + G\) | Relative depth + Gaze | 0.15 | 0.27 | 0.71 |
| \(Y + P + S + G\) | Full Multimodal (Motion + Pose + Segmentation + Gaze) | 0.12 | 0.23 | 0.58 |
Generalization Across Stricter Splits¶
To ensure that models are learning transferable multimodal navigational priors rather than memorizing landmark paths, performance is evaluated across three distinct splits (values in meters with 95% bootstrap confidence intervals):
| Modality Configuration | Random Participant (\(n=8\)) ADE / FDE (m) | Waypoint Held-Out (\(n=10\)) ADE / FDE (m) | Unfamiliar Subject (\(n=8\)) ADE / FDE (m) |
|---|---|---|---|
| \(Y\) (Motion only) | 0.19±.014 / 0.29±.021 | 0.21±.018 / 0.32±.024 | 0.23±.019 / 0.34±.027 |
| \(Y + P\) (Motion + Social Pose) | 0.17±.011 / 0.27±.019 | 0.19±.015 / 0.29±.022 | 0.20±.013 / 0.31±.020 |
| \(Y + S\) (Motion + Scene Seg) | 0.16±.013 / 0.25±.014 | 0.18±.012 / 0.28±.018 | 0.18±.016 / 0.29±.023 |
| \(Y + G\) (Motion + Gaze) | 0.15±.009 / 0.26±.017 | 0.16±.014 / 0.26±.013 | 0.16±.010 / 0.29±.018 |
| \(Y + P + S + G\) (Full Multimodal) | 0.12±.008 / 0.23±.012 | 0.14±.010 / 0.25±.011 | 0.14±.012 / 0.26±.014 |
Key Findings¶
- Gaze serves as the single strongest intent signal: Incorporating gaze into pure ego-motion (\(Y+G\)) reduces ADE from 0.19 m to 0.15 m (a 21.1% error reduction), outperforming standalone scene segmentation (0.16 m) and social bounding boxes (0.18 m). This aligns with the visual neuroscience finding that eye gaze fixates on turning targets and clearance corridors 1–2 seconds prior to physical displacement.
- Fine-grained human pose markedly surpasses bounding boxes: Among social representations, skeletal pose (\(P\)) delivers lower trajectory errors than coarse bounding boxes (\(B\)) or torso centroids (\(C\)). Body orientation and gait dynamics convey essential interaction cues that indicate whether approaching pedestrians are yielding or crossing.
- Full multimodal fusion minimizes trajectory and rotation errors simultaneously: Combining motion, human pose, scene parsing, and gaze (\(Y+P+S+G\)) achieves the lowest ADE (0.12 m), FDE (0.23 m), and head orientation error (\(L_1^{\text{head}} = 0.58\)), proving the mutual complementarity of visual attention and environmental geometry.
- Robust transfer across unseen routes and unfamiliar subjects: Across held-out landmark routes and subjects unfamiliar with the environment, the full multimodal model incurs only a minor degradation (ADE 0.12 m \(\to\) 0.14 m), verifying that cross-attention modules learn generalizable navigation cues rather than overfitting to route geometries.
Highlights & Insights¶
- Scalable Real-World Capture with Commodity AR Headsets: Demonstrates that consumer-grade hardware (Meta Quest Pro) with built-in VIO-SLAM and infrared eye tracking provides academic-grade multimodal telemetry, substantially lowering hardware deployment barriers compared to specialized research rigs.
- Interactive Multimodal Quality Control via EgoViz Dashboard: Accompanied by an open-source multi-stream inspection dashboard synchronizing 2D trajectory maps, BEV paths, first-person RGB frames, gaze overlays, and VLM descriptions, setting a strong engineering benchmark for synchronization verification.
- Illuminating the Core Dilemma of Sharp Abrupt Turns: Qualitative failure analysis highlights that sharp \(\sim 90^\circ\) intersection turns after signal changes cause all baselines to underestimate path curvature. Prior to turning, walking speed and head angular velocity remain entirely within normal steady-state bounds, highlighting that deterministic regression must ultimately yield to multimodal probabilistic intent modeling.
Limitations & Future Work¶
- Deterministic Forecasting Under Multimodal Intention Branching: All evaluated models operate deterministically, struggling to represent multiple distinct future paths at ambiguous intersection corners.
- Environmental Diversity and Adverse Weather Conditions: The current 75 sessions predominantly cover daytime fair-weather urban scenes, leaving low-light night conditions, glare, rain, and snow unrepresented.
- One-Way Decoupling of VLM Reasoning: High-level semantic annotations are generated offline by Qwen2.5-VL-32B rather than serving as an end-to-end differentiable multimodal foundation backbone for real-time robotic dialogue and control.
Related Work & Insights¶
- vs LookOut (ICCV 2025): LookOut recorded 4 hours of pedestrian motion using Project Aria glasses without high-level scene annotations; EgoTraj scales up to 10.7 hours across 75 subjects and incorporates structured VLM intent annotations alongside public benchmarking code.
- vs EgoCogNav (2025): EgoCogNav targeted cognitive uncertainty across 6 hours of mixed indoor-outdoor environments; EgoTraj focuses on open urban pedestrian transit (sidewalks, crosswalks, intersections), directly supporting assistive AR guidance and autonomous delivery robotics.
- vs Third-Person Trajectory Benchmarks (Social-LSTM, TUTR): Conventional surveillance benchmarks observe pedestrians externally, missing internal perceptual and intent cues; EgoTraj unlocks first-person modeling of how visual attention guides locomotion.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ First large-scale open multimodal outdoor egocentric trajectory dataset leveraging commercial AR headsets.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive baseline evaluations, 13-variant ablation matrix, and 3 cross-split generalization evaluations with bootstrap confidence intervals.
- Writing Quality: ⭐⭐⭐⭐⭐ Well-structured, transparently detailing hardware pipelines, quadratic gaze calibration, and qualitative failure cases.
- Value: ⭐⭐⭐⭐⭐ Highly valuable for humanoid robotics, embodied AR navigation, assistive technology for visually impaired pedestrians, and autonomous driving.