Cooking beyond Frames: A Stereo Event Camera Dataset in the Kitchen¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://chengmingf.github.io/EventKitchen.github.io/
Area: Autonomous Driving
Keywords: Event Camera, Stereo Vision, Egocentric Vision, Action Recognition, Stereo Depth Estimation
TL;DR¶
EventKitchen is the first large-scale, real-world, unscripted egocentric stereo event camera benchmark dataset capturing daily cooking activities across 13 diverse kitchens, featuring 5.5 hours of synchronized multimodal recordings, 10,762 action segments, 13,482 bounding boxes, and dense ground-truth depth maps across three core perception tasks.
Background & Motivation¶
Event cameras (neuromorphic sensors) operate asynchronously with microsecond temporal resolution, high dynamic range (HDR), and low power consumption, making them remarkably resilient to fast motion blur and abrupt illumination transitions. Despite their profound theoretical advantages, research and benchmark datasets in event-based vision remain heavily skewed toward automotive perception (such as 1Mpx, GEN1, and DSEC) and drone racing. In sharp contrast, human-centric daily-life scenarios remain largely underexplored, creating an acute bottleneck for transferring neuromorphic perception algorithms to embodied robotics, indoor human-computer interaction, and wearable edge intelligence.
Existing event datasets for human activity understanding exhibit severe limitations. Most prior benchmarks were recorded in sterile laboratory settings under rigid, scripted protocols, resulting in repetitive motions and unnatural behaviors. Furthermore, static viewpoints dominate these setups, leaving backgrounds virtually devoid of the rich event streams generated by natural head and body movements. Meanwhile, synthetic workarounds such as simulator-generated events (e.g., N-EPIC-Kitchens via ESIM) and monitor replay captures (e.g., UCF-Crime-DVS) suffer from substantial Sim-to-Real gaps, refresh rate bottlenecks, and distorted sensor noise profiles. Real indoor kitchens involve high-speed hand-object interactions, severe occlusions, small cutlery manipulation, and complex spatial depth variations, creating an urgent demand for a native, unscripted, multimodal stereo event benchmark.
To overcome the challenges of wearable sensor payload, natural unconstrained behavior capture, and ground-truth alignment on noisy event streams, this work develops a helmet-mounted multisensory recording rig and records unscripted cooking sessions with diverse participants across 13 real kitchens. Core idea: build the first large-scale, unscripted, egocentric stereo event camera benchmark dataset, EventKitchen, leveraging a cross-modal 3D point cloud back-projection pipeline to provide high-fidelity annotations for action recognition, object detection, and stereo depth estimation beyond the automotive domain.
Method¶
Overall Architecture¶
The EventKitchen dataset pipeline integrates four core stages: multimodal wearable sensor acquisition, multisensor joint calibration and temporal synchronization, cross-modal 3D geometric back-projection for ground-truth generation, and a rigorous kitchen-level cross-environment evaluation protocol. The wearable helmet integrates stereo HD event cameras, stereo CMOS RGB cameras, and an RGB-D sensor. Synchronized by nanosecond ROS timestamps, high-confidence annotations from the RGB-D domain are projected into the stereo event coordinate frame via 3D point clouds, establishing baselines across three fundamental neuromorphic perception tasks.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Wearable Multimodal Rig<br/>Stereo Event + Stereo RGB + Depth/IMU"] --> B["Joint Calibration & Synchronization<br/>E2Calib Rebuilding + Unified ROS Timestamps"]
B --> C["Cross-Modal 3D Back-Projection<br/>D-RGB Annotations Smoothly Mapped to Event Domain"]
C --> D["Kitchen-Level Split Protocol<br/>9 Training Kitchens vs 4 Unseen Test Kitchens"]
D --> E["Multitask Benchmark Evaluation<br/>Action Recognition / Object Detection / Stereo Depth"]
Key Designs¶
1. Multimodal Helmet-Mounted Synchronized Acquisition and Joint Calibration: Overcoming Daily-Life Embodied Sensing Bottlenecks
Capturing fine-grained manipulation in natural kitchens demands a wide field-of-view, unconstrained bimanual motion, and accurate spatial depth. The custom wearable rig mounts two Prophesee Gen4 HD event cameras (\(1280 \times 720\) resolution) configured as a stereo baseline, flanked by two CMOS RGB cameras (30 fps), and topped by an Intel RealSense depth camera supplying D-RGB frames and 16-bit depth maps (15 fps) alongside a 6-axis IMU (200 fps). All data streams connect via USB 3.0 to a recording laptop carried in a backpack, managed under ROS to ensure sub-millisecond timestamp alignment. To calibrate event sensors where conventional checkerboard corner detectors fail, E2Calib first reconstructs asynchronous event streams into grayscale video frames, after which OpenCV standard corner solvers compute camera intrinsics. Extrinsics are calculated using the depth camera as the global frame of reference, with the left event camera designated as the reference for stereo and depth alignment. The resulting intrinsic and extrinsic reprojection errors remain strictly under two pixels across all sensors.
2. Cross-Modal 3D Point Cloud Back-Projection and Stereo Rectification: Solving Event Ground-Truth Annotation Barriers
Directly annotating bounding boxes and precise boundaries on sparse, polarity-coded event frames is notoriously error-prone due to high sensor noise and missing static edges. EventKitchen sidesteps this challenge through a cross-modal 3D back-projection pipeline. Human annotators first label rigid kitchenware across 12 categories (such as bowl, knife, chopping board, pan, and spatula) on D-RGB frames sampled at 1 frame per 4 seconds. Exploiting the pixel-to-pixel correspondence between D-RGB and depth maps, pixels inside each 2D box are lifted into 3D camera coordinates. These 3D point clouds are then mapped onto left and right event camera planes via calibrated transformation matrices. To counter boundary jitter and projection distortion, a smoothing step averages y-coordinates along horizontal edges and x-coordinates along vertical edges, while discarding boxes outside the event sensor field-of-view (FoV), producing 13,482 accurate bounding boxes in the event domain. For stereo depth estimation, depth maps from RealSense are projected into the left event camera FoV and stereo-rectified against the event pair, yielding 297,457 aligned dense ground-truth depth frames.
3. Unscripted Long-Horizon Sequences and Kitchen-Level Generalization Protocol: Confronting Real-World Long-Tail Complexity
To capture genuine human behavioral diversity rather than robotic compliance, EventKitchen recruited 10 volunteers of diverse nationalities, ages, and genders across 8 private and 5 public kitchens. Participants freely chose among 13 daily cooking activities (such as cutting bread, frying eggs, preparing salad, and washing dishes) with zero procedural constraints or step-by-step scripts. The resulting recordings average 180 seconds per sequence, preserving realistic hesitation, variable speeds, repetitions, and natural pauses. Annotations span 32 verbs, 268 fine-grained action classes, and 10,762 action segments. Crucially, the benchmark splits data at the kitchen level: 9 kitchens (82 sequences, \(\approx 80\%\)) form the training split, while 4 entirely unseen kitchens (28 sequences, \(\approx 20\%\)) constitute the test split. This protocol strictly prevents visual leakage of countertop geometries, background clutter, and ambient illumination, imposing a rigorous test on cross-environment domain generalization.
Loss & Training¶
All baseline models ingest fused binocular event streams obtained by summing left and right event representations over a temporal slice \(\Delta t\): - Action Recognition: Evaluates Temporal Shift Module (TSM, ResNet-50 backbone) and Video Swin Transformer (Swin-Base backbone) pretrained on Kinetics-400. Models are fine-tuned on the 69 action classes \(C_{69}^a\) with at least 5 test instances and the 18 shared verb classes \(C_{18}^v\) using cross-entropy classification loss \(\mathcal{L}_{cls} = - \sum_{k} y_k \log \hat{y}_k\). - Object Detection: Evaluates frame-based YOLOv10-x (COCO pretrained) alongside event-native RVT-Base and EvRT-DETR-B (1Mpx automotive pretrained). Models are optimized using the joint detection loss \(\mathcal{L}_{det} = \lambda_{cls} \mathcal{L}_{cls} + \lambda_{box} \mathcal{L}_{box} + \lambda_{dfl} \mathcal{L}_{dfl}\). - Stereo Depth Estimation: SE-CFF is trained from scratch in an Event-only setting using Stacking by Number (SBN) representations within the valid depth range \([200\,\text{mm}, 1500\,\text{mm}]\), supervised by smooth \(L_1\) disparity loss. A zero-shot baseline using FoundationStereo evaluates disparity estimation on E2VID-reconstructed grayscale frames.
Key Experimental Results¶
Main Results¶
Evaluating baseline models across the three benchmark tasks under the unseen test kitchens demonstrates the substantial difficulty of unscripted egocentric neuromorphic perception.
| Task | Model | Backbone / Pretraining | Key Metric (Top-1 / AP) | Secondary Metric (Top-5 / AP50) | Relaxed Metric (AP05) |
|---|---|---|---|---|---|
| Action Recognition (\(C_{69}^{a}\)) | TSM | ResNet-50 (Kinetics-400) | 19.23% | 42.26% | โ |
| Action Recognition (\(C_{69}^{a}\)) | Video Swin | Swin-Base (Kinetics-400) | 24.69% | 56.48% | โ |
| Verb Classification (\(C_{18}^{v}\)) | TSM | ResNet-50 (Kinetics-400) | 38.17% | 87.20% | โ |
| Verb Classification (\(C_{18}^{v}\)) | Video Swin | Swin-Base (Kinetics-400) | 46.08% | 90.67% | โ |
| Object Detection (12 classes) | YOLOv10-x | YOLOv10-x (COCO) | 16.2% | 29.9% | 38.1% |
| Object Detection (12 classes) | RVT-Base | RVT-B (1Mpx) | 7.6% | 16.1% | 22.7% |
| Object Detection (12 classes) | EvRT-DETR | RT-DETR-B (1Mpx) | 8.5% | 16.7% | 22.7% |
Ablation Study¶
For stereo depth estimation, SE-CFF was ablated across event stacking limits and sampling frequencies against zero-shot FoundationStereo, compared directly with the inherent standard deviation and dispersion of the ground truth.
| Method / Configuration | Stacking Budget & Rate | Input Modality | RMSE (mm) โ | MAE (mm) โ | Status vs. GT Dispersion |
|---|---|---|---|---|---|
| SE-CFF | 5M @ 3Hz | Event-only | 88.19 | 59.32 | Lower than GT STD (effective) |
| SE-CFF | 5M @ 1Hz | Event-only | 88.21 | 57.38 | Lower than GT STD (effective) |
| SE-CFF | 15M @ 1Hz | Event-only | 84.91 | 54.84 | Best convergence |
| FoundationStereo | Zero-shot | E2VID Reconstructed Grayscale | 155.06 | 123.85 | Exceeds GT STD (degraded) |
| Ground Truth Reference | โ | RealSense Depth Ground Truth | STD: 106.26 | MAD: 82.00 | Ground-truth benchmark variance |
Fine-grained object detection analysis reveals a drastic performance divide across object scales and interaction dynamics: - Large, tabletop-anchored utensils perform substantially better: Pan reaches 36.5% AP, Chopping board reaches 32.2% AP, and Bowl reaches 26.0% AP under YOLOv10-x. - Small, handheld cutlery subject to rapid motion and severe hand occlusions almost completely fail: Fork achieves only 1.1% AP (0.0% on RVT and EvRT-DETR), and Spoon achieves only 1.0% AP (0.1% on RVT and 0.2% on EvRT-DETR).
Key Findings¶
- Stack size dominates temporal sampling rate for high-resolution event depth: Increasing the event stack budget from 5M to 15M at 1Hz reduces SE-CFF RMSE from 88.21 mm to 84.91 mm, whereas varying the sampling rate between 1Hz and 3Hz yields virtually identical error (88.21 mm vs 88.19 mm). This demonstrates that spatial edge density accumulated from millions of microsecond spikes is the primary driver of stereoscopic correspondence.
- Negative transfer of automotive event detectors to egocentric domains: RVT and EvRT-DETR achieve 47.4% and 50.1% AP on automotive highway data (1Mpx), but plunge to 7.6% and 8.5% on EventKitchen, falling well behind frame-based YOLOv10-x (16.2%). Automotive models assume distant, rigid, translational motion and fail when confronted with close-up bimanual manipulation, non-rigid occlusions, and sudden egocentric head turns.
- Reconstruct-then-match paradigm suffers acute domain gaps: FoundationStereo applied to E2VID-reconstructed grayscale frames incurs an MAE of 123.85 mm, significantly worse than the ground truth mean absolute deviation (82.00 mm). Video reconstruction artifacts and residual polarity noise disrupt epipolar consistency, underscoring the necessity of native neuromorphic stereo learning.
Highlights & Insights¶
- Geometric Cross-Modal Projection Pipeline: By transforming 2D RGB bounding boxes into 3D depth point clouds and reprojecting them onto stereo event sensors, the authors bypass the extreme difficulty of directly delineating fuzzy, sparse event boundaries, guaranteeing sub-pixel temporal and spatial fidelity.
- First Truly Unscripted Daily Living Event Corpus: The dataset moves beyond synthetic video recycling and rigid laboratory poses, capturing authentic cooking workflows. High-frequency action verbs closely replicate the natural distribution of EPIC-KITCHENS (led by 'Put' and 'Take'), proving unscripted behavioral fidelity.
- Unveiling the Egocentric Neuromorphic Domain Gap: The benchmark exposes the sharp collapse of specialized automotive event architectures when transferred to indoor bimanual environments, providing an indispensable catalyst for embodied neuromorphic AI.
Limitations & Future Work¶
- Limitations Admitted by Authors: As a naturalistic real-world capture, the dataset exhibits strong long-tail imbalance across action classes and object categories; additionally, the 13k human-annotated bounding boxes remain modest in comparison to automated million-scale automotive datasets.
- Deeper Limitations Identified: Object annotations are strictly restricted to 12 rigid kitchenware categories, omitting non-rigid food ingredients (e.g., vegetables, meats, dough, and liquids) whose geometric and state transitions are central to procedural cooking intelligence.
- Future Directions: Integrating 3D hand pose tracking (such as MANO mesh estimation) with stereo event streams to study continuous bimanual grasping, as well as developing foundation event models capable of open-vocabulary localization and zero-shot ingredient state tracking.
Related Work & Insights¶
- vs N-EPIC-Kitchens / UCF-Crime-DVS: Synthetic datasets produced by video simulators or screen playback suffer from monitor refresh limits, artificial brightness clipping, and missing neuromorphic noise; EventKitchen captures genuine physical events with dual Prophesee Gen4 sensors.
- vs HARDVS / DailyDVS-200 / THUMV-EACT-50: Prior real-world event activity datasets rely on third-person static views and choreographed laboratory scripts; EventKitchen is the first egocentric, stereo, unscripted, multitask kitchen benchmark.
- vs DSEC / 1Mpx: While DSEC and 1Mpx dominate outdoor autonomous driving benchmarks, EventKitchen pioneers near-field indoor stereo depth and human-object manipulation, opening a new frontier for neuromorphic vision in everyday living.
Rating¶
- Novelty: โญโญโญโญโญ (Pioneering unscripted egocentric stereo event dataset with multimodal ground-truth alignment)
- Experimental Thoroughness: โญโญโญโญโญ (Systematic evaluation across action recognition, object detection, and stereo depth estimation over 7 baseline architectures)
- Writing Quality: โญโญโญโญโญ (Exemplary clarity in sensor rig construction, calibration mathematics, and error analyses)
- Value: โญโญโญโญโญ (Essential cornerstone benchmark transitioning neuromorphic perception from automotive settings to embodied indoor robotics)