Skip to content

E-MOTION: A Dataset for Event-Based Scene Flow Estimation with Independent Moving Objects

Conference: ECCV 2026
Paper: ECCV Official
Code: https://emotion.hds.utc.fr/
Area: Segmentation
Keywords: event camera / scene flow estimation / independent moving objects / dense depth estimation / dynamic dataset

TL;DR

E-MOTION is the first large-scale benchmark dataset captured with real high-resolution stereo event cameras, delivering 200 Hz sub-millimeter dense 3D scene flow, depth, optical flow, and instance segmentation ground truth in scenes featuring independent moving objects and complex handheld motions.

Background & Motivation

Three-dimensional scene flow characterizes the instantaneous 3D Cartesian velocity vectors of all visible surface points in the physical world, serving as a core foundational perceptual representation for autonomous mobile robot navigation, non-rigid motion analysis, and dynamic human-robot interaction. Compared to conventional frame-based cameras that suffer from motion blur and dynamic range constraints, and LiDAR systems restricted by sparse point clouds and low temporal update rates, event cameras offer microsecond-level temporal resolution, high dynamic range, and extremely low latency. In principle, they represent the ideal sensor modality for high-speed dynamic scene flow estimation. However, progress in event-based 3D scene flow has long lagged behind, with very few works addressing this challenging task.

The fundamental bottleneck behind this stall is the critical scarcity of high-quality real-world scene flow benchmark datasets. Generating accurate 3D scene flow ground truth requires not only precise 6DoF sensor trajectories, but also sub-millimeter 3D geometric models, accurate poses, and instantaneous velocities of all static surfaces and independent moving objects (IMOs) at microsecond time scales. To bypass these difficulties, existing event camera datasets (such as MVSEC, DSEC, and M3ED) universally adopt a static scene assumption, relying solely on static LiDAR re-projection or camera visual odometry. Whenever autonomous moving objects like vehicles, pedestrians, or quadrotors appear in the field of view, the static assumption breaks down, leading to catastrophic motion discontinuities, missing object flows, and projection artifacts. The only two prior event datasets featuring IMOs and scene flow labels (EKubric and BlinkVision) are synthetic, failing to emulate realistic sensor circuit noise, non-ideal bias characteristics, and dynamic optical artifacts, thus severely impeding sim-to-real transfer.

To bridge this crucial gap, this paper introduces and open-sources the E-MOTION dataset. Core idea: by combining hardware-synchronized high-resolution stereo event cameras (Prophesee EVK4) with a 33-camera motion capture system, millimeter-level scene laser scans, and foreground object meshes, and applying joint spatiotemporal optimization along with hierarchical Delaunay triangulation, the authors establish the first real high-resolution 200 Hz dense 3D scene flow benchmark covering complex independent moving objects and diverse handheld camera dynamics.

Method

Overall Architecture

The data generation pipeline of E-MOTION is designed to reconstruct spatiotemporally continuous physical ground truth from multimodal physical measurements. The system ingests four primary inputs: asynchronous stereo event streams from handheld EVK4 event cameras, high-precision rigid-body trajectories from an Optitrack optical motion capture system, millimeter-accuracy 3D static point clouds from a FARO Focus3D laser scanner, and precise 3D geometric models of quadrotors and YCB objects. The pipeline consists of three sequential phases: hardware-level temporal synchronization and full-chain spatial calibration, high-fidelity geometric reconstruction with foreground hierarchical Delaunay densification, and analytical scene flow / optical flow derivation via Lie algebra differentiation. It yields dense depth maps, 2D optical flow fields, 3D scene flow vectors, 6DoF object poses, and instance segmentation masks at 200 Hz.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multimodal Raw Inputs<br/>Stereo events + Optitrack poses + FARO laser scans + 3D meshes"] --> B["Hardware-level Temporal Sync & Spatial Calibration<br/>eSync 2 square-wave trigger + Quasi-Newton hand-eye optimization"]
    B --> C["Geometric Reconstruction & Hierarchical Delaunay Densification<br/>Background inpainting + Local mesh triangulation + Z-buffering"]
    C --> D["Analytical Flow Derivation via Lie Algebra Differentiation<br/>20ms pose smoothing + 40ms Lie twist + Projected time derivative"]
    D --> E["High-Precision Dense Benchmark Outputs<br/>200 Hz scene flow / Dense depth / 2D optical flow / 6DoF poses / IMO masks"]

Key Designs

1. Hardware-level temporal synchronization and spatial calibration: eliminating sub-millisecond clock drifts and rigid-body offset errors

Due to the asynchronous and continuous nature of event streams, microsecond-level clock jitter can induce substantial geometric misalignments in dynamic re-projections. To resolve this, an Optitrack eSync 2 synchronization hub delivers periodic square-wave pulses to the EVK4 cameras. Each electrical pulse edge captured during optitrack frame sampling is injected directly into the event stream as a special hardware event timestamped by the camera's internal oscillator, enabling strict clock alignment across the motion capture system and both cameras. Spatially, the optitrack tracks rigid bodies attached to the camera rig (\(H\)) and object marker shells (\(I_i\)), which do not inherently coincide with the camera optical centers (\(L, R\)), the static scenario model origin (\(M\)), or the mesh origins (\(M_i\)).

To resolve the hand-eye extrinsic transformation \(^{H}T_L\) and camera intrinsic parameters \(\theta_L\), the pipeline reconstructs high-frame-rate intensity images from event streams to detect AprilGrid corners, formulating a joint non-linear optimization objective:

\[f(^{O}T_P, {}^{H}T_L, \theta_L) = \sum_{i=0}^{n-1} \|\pi(P_i(^{O}T_P, {}^{H}T_L), \theta_L) - p_i\|^2\]

where \(\pi(P, \theta)\) projects 3D points onto the camera image plane with radial-tangential distortion models, and \(^{O}T_P\) denotes the grid pose in the global optitrack frame. By iteratively minimizing \(f\) using a quasi-Newton algorithm, the re-projection error distribution is compressed from an initial 6–8 pixels down to below 1 pixel. Similarly, the scenario extrinsic \(^{O}T_M\) and object extrinsics \(^{I_i}T_{M_i}\) are aligned with an average error of only 1.1 cm by cross-identifying the precise physical locations of optitrack cameras and reflective tape markers in the high-resolution dense point clouds.

2. High-fidelity geometric reconstruction and hierarchical Delaunay densification: mitigating background bleeding and surface pixel gaps

Projecting raw environment and object point clouds directly onto the image plane results in two severe artifacts caused by perspective ray divergence: background points penetrating through sparse foreground object points (background bleeding) and irregular unlabeled gaps on object surfaces. E-MOTION introduces a hierarchical rendering and local meshing strategy to eliminate these artifacts.

The pipeline first isolates the static environment point cloud by removing all foreground objects (both static props and dynamic IMOs), projecting the background points onto the left camera view and applying nearest-neighbor depth clustering to inpaint interior perspective gaps. Next, each foreground object is projected independently to produce an isolated sparse 2D point set. To recover a gap-free manifold, the algorithm constructs a 2D Delaunay triangulation on the projected points, discarding low-quality stretched triangles whose side lengths exceed 4 pixels, and interpolating continuous depth across the interior pixels. This operation produces sharp foreground depth patches and crisp binary instance segmentation masks. Finally, foreground depth layers are composited over the background map using standard Z-buffering (\(Z_{\text{final}} = \min(Z_{\text{fg}}, Z_{\text{bg}})\)), completely resolving edge penetration and depth ambiguities.

3. Analytical scene flow and optical flow derivation via Lie algebra differentiation: suppressing finite-difference noise and ensuring continuous physical mapping

Scene flow is defined as the 3D velocity vector of physical points relative to the camera frame. Applying naive discrete finite-differencing directly to high-rate optitrack poses amplifies high-frequency tracking noise, corrupting the estimated velocity field. E-MOTION designs a two-tier differentiation scheme: 6DoF poses from the 200 Hz optitrack are first processed with a 20 ms sliding-window mean filter to remove tracking jitter; next, over a 40 ms temporal window, the relative motion from the start to the end pose is parameterized as a Lie twist on the SE(3) Lie group, directly yielding smoothed linear velocity \(v \in \mathbb{R}^3\) and angular velocity \(\omega \in \mathbb{R}^3\).

Based on rigid-body kinematics, the instantaneous 3D velocity of any surface point \(P = [X, Y, Z]^T\) relative to the camera center defines the 3D scene flow vector \(\dot{P} = [\dot{X}, \dot{Y}, \dot{Z}]^T\). Applying the chain rule to the time derivative of the perspective projection equation \(p = [u, v]^T = \frac{1}{Z} K [X, Y, Z]^T\), the pipeline derives an analytical closed-form relationship linking 2D image optical flow directly to 3D scene flow:

\[\begin{bmatrix} \dot{u} \\ \dot{v} \end{bmatrix} = \frac{1}{Z^2} K \begin{bmatrix} Z\dot{X} - X\dot{Z} \\ Z\dot{Y} - Y\dot{Z} \end{bmatrix}\]

where \(K\) is the camera calibration matrix. This formulation yields dense 2D optical flow fields \([\dot{u}, \dot{v}]^T\) and axial motion maps \(\dot{Z}\) on rectified images, while also extending to raw distorted image coordinates, ensuring mathematically grounded ground truth for both event-based and frame-based models.

A Worked Example

Consider the highly challenging dynamic sequence drone_fast_1: an operator holds the stereo EVK4 setup with vigorous translational velocities up to \(1.67\text{ m/s}\) and angular velocities up to \(202^{\circ}/\text{s}\), while a Parrot AR 2.0 Elite quadrotor maneuvers rapidly in the foreground, inducing compound camera-object motion and rotating rotor dynamics. 1. Pulse synchronization and pose smoothing: The eSync 2 unit emits square-wave clock edges at 200 Hz, recorded as special event packets with camera timestamps \(t_k\). The system filters optitrack poses within a 20 ms window around \(t_k\), and computes the Lie twist over a 40 ms baseline, yielding instantaneous camera velocity \(v_C = [0.82, -0.45, 1.21]^T\text{ m/s}\) and quadrotor velocity \(v_D = [-1.10, 0.60, -0.30]^T\text{ m/s}\). 2. Foreground isolation and Delaunay meshing: The static background point cloud is rendered and densified onto the \(1280 \times 720\) canvas. The quadrotor model is projected separately into approximately 15,000 sparse screen points; a 2D Delaunay triangulation filters triangles with edge lengths \(>4\text{ px}\), densely interpolating foreground depth (mean depth \(1.35\text{ m}\)) and generating an accurate binary silhouette mask. 3. Compositing and flow projection: In the quadrotor region, \(Z_D (1.35\text{ m}) < Z_{\text{bg}} (3.80\text{ m})\), overriding background pixels via Z-buffering. The relative 3D velocity vector field is calculated per pixel and analytically projected via the differential camera matrix, yielding sharp motion discontinuities at object boundaries across both 3D scene flow and 2D optical flow fields.

Key Experimental Results

Main Results

The table below presents a comprehensive benchmark comparison between E-MOTION and prominent existing event camera datasets for depth, optical flow, and scene flow estimation. The comparison highlights camera resolution, stereo configuration, real-world data capture, support for independent moving objects (IMOs), dense ground truth generation, and scene flow support.

Dataset Sensor Model Resolution Stereo Rig Real Data IMOs Support Dense GT Frequency [Hz] Duration [min] Depth GT Optical Flow GT Scene Flow GT
MVSEC [42] DAVIS 346 \(346 \times 260\) ✓ ✓ ✗ ✗ 20 43 ✓ ✓ ✗
DSEC [10] Prophesee Gen3 \(640 \times 480\) ✓ ✓ ✗ ✗ 20 53 ✓ ✓ ✗
SHEF [38] Prophesee Gen3 \(640 \times 480\) ✓ ✓ ✗ ✓ 125 23 ✓ ✗ ✗
VECtor [9] Prophesee Gen3 \(640 \times 480\) ✓ ✓ ✗ ✗ 120 21 ✓ ✗ ✗
EVIMO2 [2] DVS-Gen3 / Sam. \(640 \times 480\) ✓ ✓ ✓ ✓ 200 41 ✓ ✓ ✗
M3ED [4] Prophesee IMX636 \(1280 \times 720\) ✓ ✓ ✗ ✗ 10 203 ✓ ✓ ✗
CoSEC [29] Prophesee IMX636 \(1280 \times 720\) ✓ ✓ ✗ ✗ \(\le 20\) 60 ✓ ✓ ✗
EKubric [37] ESIM Simulator \(960 \times 540\) ✗ ✗ ✓ ✓ NA 15K frames ✓ ✓ ✓
BlinkVision [22] DVS-Voltmeter \(640 \times 480\) ✗ ✗ ✓ ✓ NA 15K frames ✓ ✓ ✓
E-MOTION (Ours) Prophesee IMX636 \(1280 \times 720\) ✓ ✓ ✓ ✓ 200 42 ✓ ✓ ✓

Ablation Study & Sequence Analysis

E-MOTION comprises 43 diverse sequences categorized into static scenes without dynamic objects (Scene), quadrotor flight sequences (Drones), and handheld YCB object interactions (YCB). The table below summarizes representative sequences, their split designations, moving object counts, event counts, and motion dynamics.

Category Sequence Name Split # IMOs Duration [s] Trajectory [m] Events [M] Velocity Avg / Max [m/s] Angular Vel Avg / Max [°/s] Motion Dynamics & Challenges
Scene transl_slow_1 Train (TR) 0 100 10.57 25 0.11 / 0.27 3 / 21 Slow handheld translation, smooth optical flow
Scene transl_fast_1 Train (TR) 0 31 19.93 82 0.90 / 2.09 20 / 962 High-speed translation, intense event rate
Scene rot_fast_1 Train (TR) 0 24 3.28 184 0.26 / 0.95 69 / 1140 Rapid rotation exceeding \(1000^{\circ}/\text{s}\)
Scene rand_normal_1 Test (TE) 0 77 20.65 71 0.30 / 0.78 26 / 2422 Handheld random motion across multiple axes
Drones drone_static Train (TR) 1 38 0.00 0.02 0.00 / 0.00 0 / 20 Static stereo rig observing free quadrotor flight
Drones drone_slow_1 Test (TE) 1 39 4.83 60 0.14 / 0.73 18 / 860 Moderate motion, evaluates baseline separation
Drones drone_fast_1 Test (TE) 1 40 17.43 202 0.64 / 1.67 73 / 1825 Fast flight and camera motion, high complexity
Drones two_drones_1 Test (TE) 2 73 14.34 67 0.23 / 0.79 22 / 1890 Dual quadrotors inducing mutual dynamic occlusions
YCB mustard_bottle Train (TR) 1 136 10.44 246 0.36 / 1.84 30 / 3304 Handheld fishing line manipulation with random swing
YCB mult_2 Test (TE) 2 60 17.28 105 0.35 / 1.46 19 / 1199 Multi-object interaction, benchmark evaluation

To investigate the performance of established event-based algorithms on E-MOTION, the authors benchmarked the top-performing multi-temporal stereo depth method MCEMVS alongside a contrast maximization optical flow baseline:

Evaluated Method Benchmark Task Performance on Slow Motion (drone_slow_1) Performance on Fast Motion (drone_fast_1) Performance on IMO Regions (Drone / YCB) Root Failure Analysis
MCEMVS [12] Stereo Depth Estimation Dense, accurate depth in static background regions Severe drop in density; sparse reconstructed points Completely fails on IMOs (zero depth output) Relies on static world assumption; non-epipolar events from IMOs are treated as outliers
Contrast Maximization [32] Dense 2D Optical Flow Background flow aligns well with camera ego-motion Image warp blurring under multi-axis rotations Severe velocity divergence and artifacts on moving objects Single motion warp model cannot simultaneously compensate ego-motion and independent motion

Key Findings

  • Violation of the static scene assumption fundamentally degrades existing methods: The empirical benchmark demonstrates that state-of-the-art stereo depth (MCEMVS) and optical flow methods fail on moving objects because their underlying models assume all events stem exclusively from camera ego-motion. When an independent quadrotor navigates the scene, its independent velocity violates epipolar constraints and single-warp compensation models, causing methods to discard the object or output distorted flow.
  • Graded motion difficulty spans several orders of magnitude: Ranging from gentle translations (\(<0.3\text{ m/s}\)) to aggressive random maneuvers (peak angular speed \(>2000^{\circ}/\text{s}\) and velocity \(>2\text{ m/s}\)), the sequences provide a rigorous evaluation continuum that reveals the operational limits of current algorithms under high dynamic conditions.
  • 200 Hz analytical Lie algebra differentiation is essential for event ground truth: Naive finite differencing of raw 6DoF poses introduces substantial high-frequency noise at high sampling rates. Combining a 20 ms sliding-window filter with 40 ms Lie twist integration successfully recovers smooth, physically consistent 3D velocity vectors.

Highlights & Insights

  • Pioneering real high-resolution stereo event scene flow benchmark: Replaces synthetic approximations (EKubric / BlinkVision) with real-world Sony IMX636 sensor measurements at \(1280 \times 720\), providing an indispensable physical benchmark to address the sim-to-real domain gap in event vision.
  • Elegant hierarchical Delaunay rendering for occlusion handling: Instead of relying on heuristic bilateral filtering or point splatting, the pipeline decouples foreground and background projections, using edge-constrained Delaunay triangulation and Z-buffering to cleanly eliminate background bleeding artifacts.
  • Reusable Lie algebra kinematics and analytical flow projection pipeline: The open-sourced tools provide an exact mathematical mapping from SE(3) continuous-time body velocities to pixel-level scene flow and optical flow, establishing a methodological template for future robotic vision benchmarks.

Limitations & Future Work

  • Visual artifacts from fishing line suspension in YCB sequences: Suspending YCB objects via fishing lines occasionally produces minor stray event bursts caused by specular glints, and manual operator movements are briefly noticeable along scene borders in select sequences.
  • Absence of onboard physical IMU measurements: The sensor rig lacks a dedicated hardware IMU. While software scripts generate synthetic IMU data with configurable Gaussian noise from optitrack poses, small discrepancies remain compared to real physical IMU bias drift and thermal noise in tightly coupled VIO research.
  • Future directions: Subsequent iterations could incorporate robotic manipulators or magnetic levitation for fully autonomous object trajectory control, mount industrial-grade synchronized IMU hardware, and extend data collection to diverse outdoor illumination environments.
  • vs DSEC [10] & M3ED [4]: DSEC and M3ED provide high-resolution stereo event streams for automotive and robotic platforms, but derive labels from static point clouds, causing ground truth to break down around moving targets; E-MOTION explicitly models independent moving objects with accurate dynamic scene flow vectors.
  • vs EVIMO2 [2]: EVIMO2 investigated independent object motion segmentation and 6DoF tracking, but was confined to older \(640 \times 480\) sensors and did not formalize full 3D Cartesian scene flow ground truth; E-MOTION scales to megapixel resolution and completes the 3D-to-2D velocity representation.
  • vs BlinkVision [22]: BlinkVision uses Blender to synthesize event streams and scene flow labels, but cannot fully replicate sensor-level analog noise, dark current, and latency fluctuations; E-MOTION serves as the indispensable real-world counterpart.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Establishes the first real high-resolution stereo event scene flow benchmark with rigorous multimodal synchronization.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprises 43 sequences (42 minutes), extensive multi-rate dynamics, and detailed empirical failure analysis of established baselines.
  • Writing Quality: ⭐⭐⭐⭐⭐ Features clean mathematical formulations, thorough geometric derivation, and transparent documentation of calibration procedures.
  • Value: ⭐⭐⭐⭐⭐ Delivers vital benchmark infrastructure for event-based 3D scene flow, high-speed robot navigation, and dynamic perception.