Skip to content

Self-Evolving Just-In-Time Memory for Proactive Embodied Safety

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/DyMessi/JIT-Memory
Area: Robotics & Embodied AI
Keywords: Embodied Planning, Interactive Safety, Agent Memory, Test-Time Evolution, Vision-Language Models

TL;DR

Addressing the dilemma where embodied agents frequently cause dynamic physical risks during benign-goal household execution, this paper proposes the Self-Evolving Just-In-Time Memory (JIT-Memory) framework, combining a topological belief graph, intent-triggered factual rules, procedural meta-skills, and an annotation-free test-time verification loop to achieve proactive risk mitigation without stalling task progress.

Background & Motivation

Vision-Language Models (VLMs) have dramatically enhanced the planning abilities of embodied agents, empowering them to interpret high-level natural language instructions and carry out multi-step interactions in complex physical households. However, most existing safety research centers on "malicious-goal safety"โ€”settings where the prompt itself specifies a harmful command (e.g., pouring water onto an active electrical appliance). In those scenarios, prevailing solutions deploy runtime guardrails or temporal logic verifiers that simply intercept the dangerous behavior or abort execution. While effective at stopping deliberate threats, treating safety purely as command refusal fails in real-world deployments.

In benign-goal interactive safety, the target instruction is completely harmless (e.g., "place the apple on the plate"), yet latent physical hazards emerge dynamically as the agent interacts with its surroundings (e.g., placing food on a dusty surface, leaving a stove burner on, or walking across a wet spill). Under these conditions, halting the task is unacceptable, while blind task execution leads to dangerous hazard accumulation. The fundamental tension arises because interactive safety is not an intrinsic property of the physical environment alone; rather, it is a bipartite function of the current environmental state and the agent's imminent actionable intent. Existing agents suffer from perceptual forgetting under partial observability, while static step-wise safety prompting (such as Safe-CoT) induces severe over-caution, causing planning deadlocks and task aborts.

To break this safetyโ€“progress trade-off, an agent requires persistent state tracking across partial observations, intent-conditioned hazard anticipation before executing actions, and executable, progress-preserving mitigation strategies. Core idea: formalize interactive safety as a state-agency bipartite function, develop a Just-In-Time Memory framework coordinating a topological belief graph, factual triggering rules, and procedural meta-skills, and continually refine mitigation strategies via an automated, annotation-free Test-Verify-Write loop at test time.

Method

Overall Architecture

The JIT-Memory framework equips closed-loop embodied planning under POMDPs with lightweight, proactive safety awareness. At each timestep \(t\), the planner takes the visual observation and generates an immediate action \(a_t\) alongside a predictive one-step intent \(\tilde{a}_{t+1}\). Once \(a_t\) is executed, an auxiliary perception VLM applies an action-conditioned patch update to maintain a Risk-Sufficient Topological Belief Graph (RSG), retaining safety-critical states across partial views without full-scene reconstruction overhead. The Agency-Grounded Factual Memory then uses the predictive intent \(\tilde{a}_{t+1}\) as a lookahead probe against the RSG: if a situational or temporal hazard is detected, the Experience Memory injects structured procedural Meta-Skills into the planner prompt, directing the agent to resolve the risk while preserving the overarching task goal. Upon episode completion, the system programmatically audits the trajectory using historical RSG snapshots, mining verified causal case patches to self-evolve the stored Meta-Skills.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Local Observation + Predictive Intent"] --> B["1. Risk-Sufficient Topological Belief Graph: Patch Update Tracking"]
    B --> C["2. Agency-Grounded Factual Memory: Just-In-Time Lookahead Triggering"]
    C -->|No risk triggered| D["Standard Task Execution"]
    C -->|Risk matched| E["3. Procedural Experience Memory: Meta-Skill Progress-Preserving Guidance"]
    E --> F["Safe & Executable Mitigation Action"]
    F --> G["4. Test-Verify-Write Loop: Annotation-Free Trace Verification & Evolution"]
    G -.->|Update buffers & evolve meta-skills| E

Key Designs

1. Risk-Sufficient Topological Belief Graph: Action-Conditioned Patch Tracking Under POMDPs
Full-scene 3D reconstruction is computationally prohibitive and prone to perceptual hallucinations in dynamic home environments, whereas single-frame visual representations lose track of hazardous objects once the camera pans away (such as forgetting a running faucet or an open flame behind the robot). The working memory addresses this by instantiating the belief state as a compact Risk-Sufficient Topological Belief Graph \(G_t = (V_t, E_t, \sigma_t, \tau)\). Here, \(V_t\) and \(E_t\) maintain task-relevant entities and sparse topological links, \(\sigma_t\) tracks dynamic unary states (e.g., open, toggled_on, dusty), and \(\tau\) maps instance details into generalized semantic/functional tags (e.g., abstracting sponges and rags into FUNCTION_CLEANING_TOOL, and paper towels into SAFETY_FLAMMABLE). To update the graph under partial observations, the agent defines focus seeds \(S_{t+1} = \text{Args}(a_t) \cup \text{Args}(\tilde{a}_{t+1})\) and extracts its local \(k\)-hop subgraph \(F_{t+1}\). An auxiliary VLM compares the new observation \(o_{t+1}\) against this local context to produce a structural patch \(\Delta_{t+1}\), which is merged into the global graph via \(G_{t+1} = \text{Merge}(G_t, \Delta_{t+1})\), ensuring unmitigated hazards remain tracked indefinitely.

2. Agency-Grounded Factual Memory: Deterministic Just-In-Time Triggering via Lookahead
Traditional semantic memory retrieval relies on embedding similarity across broad scene captions, which frequently causes false alarms due to irrelevant background objects. The Factual Memory compiles human safety norms into structured rule schemas \(r = \langle \text{mode}, \phi_r, \psi_r, \mathcal{V}_r \rangle\), categorized into Situational and Temporal modes. Situational risks represent negative affordancesโ€”environmental states that are benign until coupled with an incompatible action intent:

\[h^{\text{sit}}(G_{t+1}, \tilde{a}_{t+1}) = \mathbf{1}[\exists r \in \mathcal{R}_{\text{sit}}, \phi_r(G_{t+1}) \wedge \psi_r(\tilde{a}_{t+1})]\]

For instance, a dusty plate is harmless on its own, but when the lookahead intent is \(\tilde{a}_{t+1} = \text{PLACE\_ON\_TOP(apple, plate)}\), the joint state-intent condition triggers a situational intervention. Temporal risks \(h^{\text{tem}}(G_t) = \mathbf{1}[\exists u \in \mathcal{U}, u(G_t)]\) monitor ongoing hazardous states (such as an open gas valve). The verification condition \(\mathcal{V}_r\) is expressed as a computable logical formula over graph predicates (e.g., \(\neg\text{NEXTTO}(\text{SAFETY\_FLAMMABLE}, \text{FUNCTION\_HEAT\_SOURCE})\)), enabling deterministic downstream verification.

3. Procedural Experience Memory: Decontextualized Meta-Skills for Actionable Mitigation
Simply alerting the agent to an impending hazard does not tell it how to act, and passing raw episodic trajectories fails to generalize across different floor plans. Experience Memory abstracts verified execution traces into modular, decontextualized Meta-Skills \(m_r = \langle \textsc{Purpose}, \textsc{When}, \textsc{How}, \textsc{Grounding}, \textsc{Timing}, \textsc{Constraints} \rangle\). Specifically, \(\textsc{How}\) formalizes the standard mitigation sequence (e.g., wiping a surface prior to placing edible items); \(\textsc{Grounding}\) instructs the planner on how to bind alternative tools from the current RSG (e.g., using a rag or searching for another cleaning implement); and \(\textsc{Timing}\) schedules risk resolution so that ongoing workflows are not disrupted prematurely (e.g., delaying stove shutdown until boiling completes). At planning time, injecting only the active \(m_r\) and a single verified exemplar maintains a lean prompt context.

4. Test-Verify-Write Loop: Causal Patch Mining and Test-Time Evolution
To eliminate manual demonstration labeling, the framework incorporates an automated test-time self-evolution loop. At episode termination, the system queries the recorded RSG snapshots against the formal verification target \(\mathcal{V}_r\). If the hazard was resolved successfully, the system extracts the minimal causal trajectory segment between the trigger timestep \(t_r^{\text{on}}\) and resolution timestep \(t_r^{\text{off}}\):

\[c_r = \{(a_t, G_t)\}_{t=t_r^{\text{on}}}^{t_r^{\text{off}}}\]

This patch is routed to the positive buffer \(\mathcal{B}_r^+\), while safety violations or task lockups are logged in the negative buffer \(\mathcal{B}_r^-\). When a buffer reaches capacity or repeated failures occur, the agent updates the procedural skill by contrasting successful and failed traces:

\[m_r \leftarrow \mathrm{Evolve}(m_r, r, \mathcal{B}_r^+, \mathcal{B}_r^-)\]

This empirical contrastive loop systematically refines execution timing (preventing premature task aborts) and grounding precision (adapting tool choices to novel layouts).

A Worked Example

Consider a household scenario: "Clean an apple and place it on a plate inside the cabinet": 1. Belief Tracking & Lookahead: The agent retrieves the apple from the refrigerator and emits a lookahead intent \(\tilde{a}_{t+1} = \text{PLACE\_ON\_TOP(apple, plate)}\). 2. Rule Matching: In the RSG, the target plate node carries the unary state dusty. The factual rule FOOD_ONLY_ON_CLEAN_PLACE evaluates \(\phi_r(\text{dusty}(y)) \wedge \psi_r(\text{PLACE\_*}(\text{food}, y))\) to true, triggering a situational risk. 3. Meta-Skill Mitigation: The memory injects Meta-Skill \(m_r\), which specifies: "Cleaning must precede food placement; locate an available cleaning tool from the RSG." The planner schedules \(a_{t+1} = \text{WIPE(plate, rag)}\), updates the plate to clean, and seamlessly proceeds with food placement without aborting the main task. 4. Temporal Obligation Tracking: When the faucet is turned on to rinse the apple, the temporal risk monitor flags the persistent obligation. The Meta-Skill's timing instructions prevent shutting the faucet off prematurely, allowing rinsing to finish before turning off the tap. Upon completion, the RSG confirms the faucet is closed, and the trace is cataloged into \(\mathcal{B}_r^+\).

Key Experimental Results

Main Results

Evaluation was conducted on IS-Bench within the OmniGibson simulation environment, spanning 161 benign household tasks containing 388 latent physical risks. Performance was measured via Task Success (TS), Safe Success (SS, requiring both task completion and zero safety violations), Pre-mitigation rate (Pre, for situational risks), and Post-mitigation rate (Post, for temporal risks).

VLM Backbone Method TS (%) โ†‘ SS (%) โ†‘ Pre (%) โ†‘ Post (%) โ†‘
Qwen3-VL-8B Vanilla 69.3 18.4 17.2 46.2
Qwen3-VL-8B Safe-CoT 57.6 30.2 40.0 61.1
Qwen3-VL-8B JIT-Memory (Ours) 71.3 48.7 46.8 89.4
Qwen3-VL-32B Vanilla 70.5 25.2 16.7 76.0
Qwen3-VL-32B Safe-CoT 71.2 33.3 32.8 63.8
Qwen3-VL-32B JIT-Memory (Ours) 76.9 58.7 61.5 94.1
GPT-4o Vanilla 77.7 35.3 15.4 72.0
GPT-4o Safe-CoT 71.7 31.2 29.5 72.3
GPT-4o JIT-Memory (Ours) 81.3 66.2 70.9 96.8
GPT-5.2 Vanilla (Ref) 78.0 28.0 20.6 71.7
GPT-5.2 Safe-CoT (Ref) 74.7 45.3 39.3 77.2
Gemini-3.1-Pro-Preview Vanilla (Ref) 88.0 25.8 21.3 61.3
Gemini-3.1-Pro-Preview Safe-CoT (Ref) 82.0 46.0 38.7 81.3

Ablation Study

Component ablations evaluated on the Qwen3-VL-8B backbone illustrate the necessity of each architectural component:

Method Variant TS (%) โ†‘ SS (%) โ†‘ Pre (%) โ†‘ Post (%) โ†‘ Analysis
Full System (Ours) 71.3 48.7 46.8 89.4 All modules coordinated; resolves hazards without task stalling
Replace with Dense Retrieval 62.0 24.0 26.8 76.0 Global semantic similarity introduces false prompts, distracting planner
Replace with Per-step Re-parse 68.7 40.0 45.2 72.3 Full re-parsing forgets out-of-view hazards, dropping Post mitigation by 17.1%
Only Episodic Cases 68.0 21.3 28.0 42.0 Raw trajectories overfit specific scenes and generalize poorly
Only Factual Principles 71.3 28.1 35.6 55.4 Abstract safety rules lack grounding on how and when to mitigate
Perception-CoT Baseline 65.2 20.0 17.2 42.8 Verbal scene descriptions fail to prompt counterfactual hazard anticipation
RSG-Augmented Baseline 70.0 28.0 19.2 57.5 Belief graph alone lacks procedural meta-skills, yielding low Pre-risk mitigation

Key Findings

  1. Breaking the Safetyโ€“Progress Trade-Off: Safe-CoT severely impairs Task Success on Qwen3-VL-8B (dropping from 69.3% to 57.6%) because indiscriminate safety tips lead to excessive caution and stalled progress. In contrast, JIT-Memory elevates Safe Success by +30.3% (from 18.4% to 48.7%) while maintaining high overall task completion (71.3%).
  2. Small Open-Source Models Outperforming Frontier LLMs: An 8B model equipped with JIT-Memory achieves a higher Safe Success rate (48.7%) than frontier proprietary models using Safe-CoT, such as GPT-5.2 (45.3%) and Gemini-3.1-Pro-Preview (46.0%), highlighting the power of structured memory over brute-force prompting.
  3. Autonomous Timing Calibration via Test-Time Evolution: In test-time evolution runs across 60 training episodes, initial uncalibrated meta-skills caused an initial dip in TS (from 63.3% to 60.0%) due to overly hasty mitigation. Through closed-loop trace verification, the agent learned correct temporal scheduling, ultimately lifting TS to 64.4% and boosting SS from 20.0% to 42.2%.

Highlights & Insights

  • Interactive Safety as State-Agency Coupling: Reframing safety from a static environment property into an intent-conditioned negative affordance offers a mathematically sound and practical lens for embodied hazard prevention.
  • Just-In-Time Lookahead Triggering: Employing one-step predictive intents as lookahead probes activates safety rules deterministically while keeping prompts clean and dormant during standard operation.
  • Annotation-Free Test-Time Evolution: Utilizing graph-based logical verifiers to mine causal trajectory patches creates an automated self-healing loop for continuous procedural skill refinement without human intervention.

Limitations & Future Work

  • Perceptual Bottlenecks in Auxiliary VLMs: The accuracy of the RSG depends on the auxiliary model's ability to detect subtle state changes (such as minor oil spills or fine cracks). Future work could integrate specialized embodied 3D scene-graph estimators.
  • Predefined Rule Coverage: Factual rules are currently pre-compiled from human safety principles. Extending this to open-world scenarios will require methods for autonomous online rule discovery.
  • Sim-to-Real Transfer: The current benchmark relies on simulated physics in OmniGibson. Validating this architecture on physical robotic platforms will require managing real-world sensor noise and low-latency motor control constraints.
  • vs Safe-CoT / HomeGuard: While prior methods apply continuous safety reflection prompts that induce hallucinations and task abandonment, JIT-Memory uses targeted lookahead triggering to intervene only when intent and state collide.
  • vs AgentSpec / Plug in the Safety Chip: Rather than relying on rigid temporal logic guardrails that freeze the agent, JIT-Memory translates formal constraints into actionable procedural Meta-Skills that resolve hazards actively.
  • vs Voyager / RoboMemory: Unlike traditional embodied memory architectures focused on task decomposition and scene mapping, JIT-Memory introduces an explicit risk-sufficient representation with verified test-time procedural learning.

Rating

  • Novelty: โญโญโญโญโญ Formulates interactive safety as a state-intent coupling and introduces an elegant just-in-time triggering and self-evolving memory loop.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluation across open-source and proprietary backbones on IS-Bench, supported by extensive ablations and evolution trajectories.
  • Writing Quality: โญโญโญโญโญ Rigorous mathematical framing, clean architectural design, and clear, structured narrative.
  • Value: โญโญโญโญโญ Solves the critical safetyโ€“progress trade-off in interactive embodied planning, providing a valuable blueprint for reliable real-world deployment.