NutriBench-Kitchen: Benchmarking Embodied AI for Nutrition Management¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/V1ol1n/NutriBench-Kitchen
Area: Robotics & Embodied AI
Keywords: embodied AI, nutrition management, long-form video understanding, structured memory, food state tracking
TL;DR¶
This paper introduces NutriBench-Kitchen, a benchmark covering 160 long cooking videos and 1,500 human-verified multi-task QA pairs to evaluate Embodied Nutrition Management, and proposes Nutri-Vgent, an agent featuring decoupled episodic, food-state, and recipe memories that substantially bridges the performance gap in long-horizon state maintenance and constraint-aware planning.
Background & Motivation¶
Embodied assistants deployed in real-world kitchens cannot rely solely on isolated video frames to make informed decisions. Cooking procedures naturally span several minutes or tens of minutes, during which ingredients undergo a series of irreversible physical transformations including pouring, weighing, cutting, mixing, transferring, and partial consumption. When a user asks whether enough ingredients remain to prepare another serving or whether the current preparation complies with specific dietary restrictions, the relevant items may have already left the camera view or been sealed inside opaque cookware. Static food understanding models operating on single frames cannot capture this temporal history, while conventional embodied planning benchmarks typically assume that the current ground-truth environment state is fully observable and provided, overlooking the fundamental challenge of continuously inferring and maintaining latent physical states from uncurated visual streams.
Prior food-related benchmarks predominantly target single-image categorization, calorie estimation, or cross-modal recipe retrieval (e.g., Nutrition5k, Recipe1M+, FoodKG). Meanwhile, egocentric datasets such as EPIC-KITCHENS and Ego4D prioritize action recognition, spatial-temporal detection, and procedural step segmentation. None of these benchmarks close the loop from visual interaction events to continuous ingredient inventory updates, and finally to constraint-grounded downstream decision-making. If an embodied assistant fails to translate temporally distributed visual observations into persistent state representations, even if it accurately retrieves the exact timestamp where flour was initially added, it remains incapable of determining how many grams of flour remain after multiple subsequent baking and consumption stages.
To address this gap, this paper formalizes the problem of Embodied Nutrition Management and presents NutriBench-Kitchen, a multi-dimensional benchmark spanning five core task families: Ingredient Entry, Memory Management, Recipe Query, Long-Term Planning, and Short-Term Planning. The core idea is to decouple transient perceptual observations from persistent physical state ledgers and recipe domain rules through a three-tiered structured memory architecture (Nutri-Vgent), converting ephemeral video events into persistent state tracking that supports long-horizon, constraint-aware embodied decision-making.
Method¶
Overall Architecture¶
Embodied Nutrition Management requires an agent to process a continuous long-form cooking video \(V=\{v_t\}_{t=1}^T\), a user query \(q\), dialogue history \(h\), and external recipe/nutrition knowledge \(\mathcal{K}\), to infer the latent kitchen state \(S_t\) and output a constraint-compliant response \(\hat{y} = \mathcal{F}(V, q, h, r; \mathcal{K})\). To overcome the severe temporal forgetting and context confusion exhibited by standard vision-language models, the authors propose Nutri-Vgent. Rather than concatenating raw observations into an unstructured retrieval pool, Nutri-Vgent organizes its memory into three specialized tiers: Episodic Memory for timestamped action observations, Food Memory for an explicit physical inventory ledger, and Recipe Memory for domain constraints and substitution rules.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Long Cooking Video Stream<br/>Sliding Temporal Window"] --> B["LVLM Event Perception<br/>Action & Entity Proposals"]
B --> C["Episodic Memory<br/>Timestamped Actions & Spatial Relations"]
B --> D["Food Memory Ledger<br/>Quantitative Tracking & State Updates"]
E["Recipe Knowledge Graph<br/>Formulas / Dietary Constraints / Substitutions"] --> F["Function Calling Tools<br/>DishSpec / FoodByCalories"]
C --> G["Constraint-Aware Reasoning Engine<br/>State-Rule Reconciliation"]
D --> G
F --> G
G --> H["Constraint-Grounded Decision & QA Output"]
Key Designs¶
1. Multi-Dimensional Task Families and Dual Evaluation Regimes: Full-Lifecycle Assessment
NutriBench-Kitchen structures its evaluation around five interconnected task families representing the entire workflow of nutrition management: ① Ingredient Entry (IE), which tests fine-grained visual grounding and initial quantity/calorie estimation under severe occlusions; ② Memory Management (MM), which evaluates the agent's ability to maintain and reconstruct the updated inventory state after compound actions such as adding, mixing, transferring, and consuming; ③ Recipe Query (RQ), which assesses recipe matching conditioned on dynamic pantry availability and dietary restrictions; ④ Long-Term Planning (LTP), which probes high-level understanding of procedural dependencies and multi-stage execution order; and ⑤ Short-Term Planning (STP), which tests immediate next-action selection conditioned on local object states and active subgoals. The benchmark further introduces two disjoint evaluation regimes: the Easy split provides single-turn queries with a localized visual clue (frame_hint), whereas the Hard split presents multi-turn dialogues without hints, placing full pressure on autonomous long-range visual memory and state reconstruction.
2. Dual-Source Grounding and Consensus Verification: Combining Diversity with Quantitative Precision The benchmark collects 160 long-form cooking videos (yielding 1,500 meticulously annotated QA pairs) by combining two complementary data sources. First, 145 sequences selected from the HD-EPIC dataset provide natural egocentric kitchen interactions, realistic clutter, and authentic visual occlusions. Second, 15 self-recorded long-form cooking videos incorporate physical kitchen scale measurements and precise nutritional calculations, establishing accurate ground truth for ingredient gram weights that cannot be recovered from public footage. During annotation, candidate proposals generated by Qwen3-VL-max are manually corrected and temporal spans are aligned frame-by-frame. Furthermore, an external Recipe Knowledge Graph containing over 10,000 recipes is constructed to canonicalize ingredient aliases and verify nutritional constraints, ensuring strict validity for every question and plausible distractor.
3. Decoupled Structured Memory Architecture (Nutri-Vgent): Isolating Observations from State Ledgers
Generic visual RAG frameworks frequently suffer from context pollution and hallucinations when directly concatenating raw retrieval snippets into prompts. Nutri-Vgent addresses this failure by decomposing memory into three distinct functional modules:
- Episodic Memory: Houses time-indexed procedural records detailing when and where actions occurred (e.g., slicing carrots, pouring milk into a bowl), providing causal evidence for how states evolved;
- Food Memory: Acts as an explicit, persistent inventory ledger tracking canonical food identities, estimated remaining quantities (in grams or milliliters), spatial containers (e.g., bowl, frying pan, refrigerator), and freshness. Any state-changing event directly triggers incremental mathematical updates to this ledger, retaining numerical states even after items disappear from camera sight;
- Recipe Memory: Normalizes culinary knowledge into structured ingredient-step-tool graphs and exposes specialized prompt-level function calling tools—DishSpec (retrieving standard ingredient proportions for a target dish) and FoodByCalories (querying ingredient candidates within specified caloric bounds). By reconciling physical inventories against recipe rules prior to response generation, Nutri-Vgent eliminates hallucinations during complex multi-step reasoning.
Loss & Training¶
Nutri-Vgent is instantiated with pretrained large vision-language backbones (such as Qwen3-VL-4B and Qwen3-VL-8B) using zero-shot chain-of-thought (CoT) prompting alongside external tool execution, without requiring end-to-end parameter updates. For evaluation, multiple-choice questions are scored using deterministic string parsing accuracy, while continuous numerical estimation questions (primarily in Ingredient Entry and Memory Management) are evaluated via Mean Relative Accuracy (MRA):
where \(\mathcal{C} = \{0.50, 0.55, \dots, 0.95\}\) denotes a series of increasingly strict relative error tolerances, and \(\epsilon\) prevents division by zero. This metric comprehensively penalizes inaccurate numerical quantity tracking across different error bounds.
Key Experimental Results¶
Main Results¶
The authors evaluate both proprietary commercial frontier models (Gemini 2.5 Pro, GPT-4o) and representative open-source vision-language backbones against Nutri-Vgent across the five task families under Easy (E.) and Hard (H.) regimes.
| Models | IE (E./H.) | LTP (E./H.) | MM (E./H.) | RQ (E./H.) | STP (E./H.) | Avg. |
|---|---|---|---|---|---|---|
| Random Chance | 16.7 / 25.0 | 25.0 / 25.0 | 25.0 / 25.0 | 25.0 / 25.0 | 25.0 / 25.0 | 24.2 |
| Frequency Baseline | 12.6 / 6.2 | 28.5 / 22.4 | 5.3 / 2.1 | 29.8 / 23.5 | 34.2 / 26.8 | 19.1 |
| Human Performance | 92.5 / 88.0 | 95.0 / 91.5 | 89.0 / 84.5 | 96.0 / 92.0 | 97.5 / 94.0 | 92.0 |
| Proprietary LVLMs | ||||||
| GPT-4o | 30.7 / 16.7 | 69.3 / 66.7 | 26.7 / 61.3 | 46.7 / 50.5 | 49.3 / 59.3 | 47.7 |
| Gemini 2.5 Pro | 46.7 / 14.7 | 95.3 / 78.0 | 47.3 / 66.7 | 65.3 / 69.3 | 64.7 / 63.3 | 61.1 |
| Open-Source LVLMs | ||||||
| InternVL3.5-8B-HF | 26.7 / 18.0 | 81.3 / 45.3 | 36.7 / 58.7 | 50.7 / 53.5 | 48.7 / 49.3 | 46.9 |
| GLM-4.1V-9B-Base | 38.0 / 5.3 | 90.7 / 57.3 | 34.7 / 64.0 | 58.0 / 60.4 | 50.0 / 50.0 | 50.8 |
| Qwen3-VL-4B | 32.7 / 16.7 | 93.3 / 40.0 | 44.0 / 70.7 | 70.0 / 69.3 | 60.7 / 57.3 | 55.5 |
| + Vgent (Flat RAG) | 18.0 / 15.3 | 83.6 / 34.6 | 21.5 / 14.6 | 21.7 / 27.7 | 33.3 / 17.8 | 28.8 |
| + Nutri-Vgent (Ours) | 22.7 / 73.3 | 72.7 / 35.1 | 42.7 / 68.7 | 61.3 / 95.0 | 76.0 / 73.0 | 62.1 |
| Qwen3-VL-8B | 30.7 / 14.0 | 95.3 / 48.7 | 38.7 / 66.0 | 70.7 / 63.4 | 53.3 / 52.7 | 53.4 |
| + Vgent (Flat RAG) | 16.9 / 20.7 | 77.6 / 58.8 | 20.8 / 26.2 | 25.5 / 21.3 | 30.4 / 18.5 | 31.7 |
| + Nutri-Vgent (Ours) | 20.0 / 73.3 | 68.0 / 53.5 | 42.0 / 69.3 | 58.0 / 94.1 | 73.3 / 71.3 | 62.3 |
Ablation Study¶
On both Qwen3-VL-4B and Qwen3-VL-8B backbones, the contribution of individual memory modules was systematically investigated across Visual-only (Vis.), Episodic-only (Epi.), Food-only (Food), and combined structured memory (Epi. + Food).
| Model Backbone | Vis. | Epi. | Food | IE | LTP | MM | RQ | STP | Overall |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3-VL-4B | ✓ | 24.70 | 66.65 | 57.35 | 69.65 | 59.00 | 55.5 | ||
| ✓ | 50.00 | 44.30 | 46.30 | 44.55 | 51.00 | 47.3 | |||
| ✓ | 47.35 | 57.70 | 43.00 | 26.20 | 53.00 | 45.5 | |||
| ✓ | ✓ | 48.00 | 53.90 | 55.70 | 78.15 | 74.50 | 62.1 | ||
| Qwen3-VL-8B | ✓ | 22.35 | 72.00 | 52.35 | 67.05 | 53.00 | 53.4 | ||
| ✓ | 57.00 | 53.65 | 44.30 | 50.05 | 53.70 | 51.8 | |||
| ✓ | 43.65 | 69.00 | 42.65 | 26.85 | 52.35 | 46.9 | |||
| ✓ | ✓ | 46.65 | 60.75 | 55.65 | 76.05 | 72.30 | 62.3 |
Key Findings¶
- Standard LVLMs Suffer from Catastrophic Temporal Forgetting: Monolithic vision-language models fail dramatically when temporal spans widen. For example, Qwen2.5-VL-7B drops from 30.7% on Easy IE to 3.3% on Hard IE, and GPT-4o scores only 26.7% on Easy Memory Management. This underscores that scale alone cannot resolve persistent physical quantity tracking.
- Unstructured Flat RAG Triggers Context Pollution: Naively prepending retrieved video captions (as done by generic Vgent) causes substantial performance degradation (Qwen3-VL-4B drops from 55.5% to 28.8%), as unorganized video fragments dilute the model's focus on active constraints.
- Nutri-Vgent Excels under Complex Hard Constraints: Decoupled state tracking empowers Nutri-Vgent to achieve massive gains on Hard tasks: on Hard Ingredient Entry it jumps from 14.0% to 73.3%, and on Hard Recipe Query it surges from 63.4% to 94.1%, surpassing Gemini 2.5 Pro (61.1% overall) and GPT-4o (47.7% overall) with a compact 4B/8B footprint.
Highlights & Insights¶
- Persistent State Ledger vs. Ephemeral Visual Saliency: Nutri-Vgent demonstrates that embodied kitchen reasoning fundamentally requires maintaining an explicit symbolic state ledger rather than treating long videos as homogeneous retrieval pools. This converts complex implicit multi-step deductions into clean, incremental ledger lookups.
- Complementary Dual-Source Benchmark Design: By pairing uncurated in-the-wild interactions from HD-EPIC with strictly measured physical ground truths from laboratory cooking sessions, NutriBench-Kitchen achieves both natural visual complexity and exact numerical precision.
- Broad Transferability for Embodied AI: The decoupled tri-memory paradigm (episodic trajectory, physical asset inventory, and task domain rules) provides a generalizable blueprint applicable to other long-horizon embodied tasks, such as robotic assembly, laboratory experimentation, and warehouse logistics.
Limitations & Future Work¶
- Retrieval Overhead in Trivial Visual Scenes: In simple scenarios where target objects remain continuously visible, enforcing external knowledge graph calls can occasionally introduce over-correction, leading to slight performance drops on selected Easy splits.
- Heuristic Physical Quantity Grounding: The current pipeline estimates ingredient grams from 2D bounding boxes and heuristic density priors, rather than tightly integrating continuous 3D fluid/deformable physical simulation engines.
- Future Directions: The authors highlight extending the single-agent pipeline to multi-agent human-robot collaboration, and embedding evolving user-specific health profiles (e.g., glycemic targets, allergic histories) into the dynamic memory constraint loop.
Related Work & Insights¶
- vs HD-EPIC / Ego4D: While prior egocentric datasets benchmark atomic action recognition and temporal localization, NutriBench-Kitchen introduces evaluation of cumulative ingredient inventory tracking and constraint-aware decision-making over extended cooking procedures.
- vs Nutrition5k / Recipe1M+: Static image nutrition estimation ignores the sequential transformations, cutting, mixing, and transferring common in real cooking; this work formalizes nutrition management in dynamic, sequential video streams.
- vs CookBench / ET-Plan-Bench: Existing embodied planning benchmarks assume complete, structured state inputs provided directly by the environment simulator; NutriBench-Kitchen requires models to autonomously extract, update, and retain these physical states from raw visual observations.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneers the formal definition of Embodied Nutrition Management and state-tracking benchmarks]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [160 long cooking videos, 1,500 verified QA pairs, thorough cross-model evaluations and ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Well-structured narrative, mathematically rigorous formulations, and clear system design]
- Value: ⭐⭐⭐⭐⭐ [Establishes a critical bridge between long-video understanding and practical embodied dietary assistance]