Skip to content

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/hananshafi/RoboTALES
Area: Robotics & Embodied AI
Keywords: Video World Models, Embodied Control, Hierarchical Planning, Diffusion Policy, Representation Alignment

TL;DR

RoboTALES introduces a single-stage robot learning framework that steers a video diffusion world model via hierarchical sub-goal plans and aligns its latent representations with a frozen VLM critic using differentiable policy optimization (DDPO), directly conditioning an action diffusion policy for robust long-horizon manipulation.

Background & Motivation

Humans rarely act purely reactively when performing complex physical tasks; instead, they decompose high-level objectives into ordered subtasks, mentally simulate possible future outcomes, and commit to actions only after rejecting unfavorable trajectories. In recent embodied AI research, diffusion-based video generation models have emerged as promising predictive backbones for robot control, offering rich visual dynamics and physical interaction priors learned from large-scale video distributions. However, standard video generators are trained primarily on perceptual reconstruction or aesthetic objectives, meaning their imagined futures frequently drift away from the actual task specification and struggle to remain faithfully conditioned on low-level motor actions over extended horizons.

Concurrently, large language models (LLMs) and vision-language models (VLMs) have shown remarkable strengths in high-level reasoning and abstract task decomposition. Yet, most prior robot systems treat language planning and physical prediction as decoupled stages: language either specifies high-level discrete sub-goals or selects among reactive primitives, never directly shaping the predictive latent dynamics inside the visual world model. This structural gap leaves the agent's abstract reasoning, visual imagination, and motor execution poorly aligned, causing catastrophic task drift in multi-stage manipulation.

The critical insight of this work is that abstract reasoning must serve as an explicit structural scaffold embedded directly into the video model's latent space, coupled with reward-driven feedback to strictly discipline simulated futures. Core idea: build a single-stage joint framework, RoboTALES, where an LLM planner decomposes goals into structured sub-goals to condition video diffusion, a frozen VLM critic aligns internal representations via DDPO, and action diffusion loss backpropagates end-to-end into the video decoder to tightly couple imagination with control.

Method

Overall Architecture

RoboTALES organizes long-horizon visuomotor control into four tightly integrated components: a high-level LLM planner \(F_P\), a video diffusion generator \(G_\theta\) (adapted from Stable Video Diffusion), a frozen VLM critic \(F_R\), and a 1D action diffusion policy \(\pi_\phi\). Given current multi-view visual observations \(s_t\) and a natural-language task instruction \(\tau\), the system generates continuous motor action chunks \(a_{t:t+H}\). During training, RoboTALES performs single-stage joint optimization where action-level gradients flow back without stop-gradient into the video decoder layers, ensuring that internal representations are simultaneously optimized for visual predictive fidelity and downstream physical control.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Task instruction and multi-view observation"] --> B["Hierarchical Planning & Goal Anchoring<br/>LLM planner decomposes task into sub-goal chain"]
    B --> C["Conditioned Video Diffusion Simulation<br/>Cross-attention injects text plan & visual states"]
    C --> D["Critic-Guided Representation Steering<br/>Frozen VLM scores rollouts & backprops DDPO gradient"]
    D --> E["Coupled Gradient Action Diffusion<br/>Action loss backpropagates directly into video decoder"]
    E --> F["Executable continuous robot trajectory"]

Key Designs

1. Hierarchical Planning & Goal Anchoring: removing ambiguity through structured temporal scaffolds Raw task descriptions \(\tau\) (such as "prepare ingredients for zucchini curry") are typically too compressed and under-specified to steer a generative diffusion model across extended manipulation horizons. The LLM planner \(F_P\) decomposes \(\tau\) conditioned on observation \(s_t\) into \(K\) ordered, physically grounded sub-goals \(C = \{c^{(1)}, \dots, c^{(K)}\}\) (\(K \in [2, 5]\)), concatenating them into an augmented plan: $\(C^* = [\,\tau\,;\;c^{(1)}\,;\;c^{(2)}\,;\;\dots\,;\;c^{(K)}\,]\)$ Encoded by a CLIP text encoder, \(C^*\) is injected into the cross-attention layers of the video generator \(G_\theta\). These sub-goal tokens serve as temporal anchors that transform unconstrained visual rollout generation into milestone-driven visual simulation.

2. Critic-Guided Representation Steering: enforcing semantic alignment via differentiable policy optimization While standard video diffusion loss guarantees low-level visual fidelity, it does not penalize task-irrelevant visual hallucinations. RoboTALES formulates the iterative video denoising chain as a multi-step MDP and leverages a frozen VLM critic \(F_R\) to evaluate decoded rollout keyframes \(\hat{s}_{t:t+\Delta}\) against goal \(\tau\), producing a scalar reward \(r = F_R(\hat{s}_{t:t+\Delta}, \tau)\). With baseline \(b(\tau)\) and advantage \(A = r - b(\tau)\), the trainable parameters \(\theta^*\) are optimized via a REINFORCE-style DDPO objective: $\(L_{\mathrm{DDPO}}(\theta^*) = - \mathbb{E}\left[ A \cdot \frac{1}{|\mathcal{T}_K|} \sum_{t \in \mathcal{T}_K} \nabla_{\theta^*} \log p_{\mathcal{G}_{\theta^*}}(x_{t-1} \mid x_t, C^*) \right]\)$ This policy gradient update directly steers the latent transition kernel toward goal-consistent trajectories, suppressing non-functional visual branches.

3. Coupled Gradient Action Diffusion: closing the imagination-action loop with end-to-end backpropagation Conventional two-stage pipelines freeze the video generator before training an action head, yielding visual features agnostic to control-critical physical cues such as contact gaps and gripper orientation. In RoboTALES, a 1D action diffusion UNet \(\pi_\phi\) ingests hidden features \(f_{\mathcal{G}_{\theta^*}, C^*}^{\{l\}}\) from selected video decoder blocks to predict continuous actions. Crucially, no stop-gradient is applied between \(\pi_\phi\) and \(G_{\theta^*}\), allowing action diffusion loss gradients to backpropagate directly into the video decoder layers \(\theta^*\): $\(L_{\mathrm{total}}(\theta^*, \phi) = L_{\mathrm{video}}(\theta^*) + \beta L_{\mathrm{DDPO}}(\theta^*) + \gamma L_{\mathrm{action}}(\phi, \theta^*)\)$ This co-adaptation encourages the video generator's latent space to preserve representations that are not only visually and semantically coherent but also maximally informative for precise, low-jerk motor actuation.

Loss & Training

The entire architecture is trained end-to-end in a single stage combining three balanced objectives: standard video diffusion loss \(L_{\mathrm{video}}\) for visual realism, DDPO reward loss \(L_{\mathrm{DDPO}}\) for semantic task compliance, and action diffusion loss \(L_{\mathrm{action}}\) for trajectory accuracy. To prevent catastrophic forgetting and maintain computational tractability, updates to the video backbone are strictly targeted to the text cross-attention layers and designated decoder feature blocks \(\theta^*\), keeping the remaining video UNet layers frozen.

Key Experimental Results

Main Results

Evaluations were conducted on 24 challenging manipulation tasks in the physics-based RoboCasa simulation suite (spanning pick-and-place, door and drawer actuation, knobs, levers, and buttons) and the LIBERO10 benchmark, with 50 rollout evaluations per task.

Table 1: Success rate comparison on the RoboCasa simulation benchmark validation set (trained on 50 human demonstrations)

Category / Task DP-3D DP-ResNet DP-CLIP GR00T FPV OpenVLA UVA VideoPolicy Ours (RoboTALES)
Pick and Place
PnPCabToCounter 0.04 0.06 0.00 0.20 0.10 0.10 0.26 0.16 0.44
PnPCounterToCab 0.02 0.06 0.02 0.36 0.14 0.32 0.18 0.38 0.44
PnPCounterToSink 0.00 0.14 0.08 0.10 0.08 0.30 0.16 0.32 0.50
PnPSinkToCounter 0.00 0.10 0.22 0.33 0.30 0.56 0.38 0.52 0.56
PnPStoveToCounter 0.00 0.02 0.06 0.29 0.26 0.62 0.24 0.64 0.66
Doors
OpenSingleDoor 0.24 0.42 0.32 0.59 0.74 0.42 0.54 0.78 0.80
OpenDoubleDoor 0.20 0.70 0.82 0.15 0.92 0.80 0.90 0.92 0.94
CloseDoubleDoor 0.56 0.78 0.84 0.75 0.78 0.84 0.76 0.90 0.94
CloseSingleDoor 0.62 0.78 0.48 0.83 0.84 1.00 0.88 0.98 0.98
Drawers
OpenDrawer 0.36 0.64 0.60 0.79 0.72 0.66 0.28 0.40 0.80
CloseDrawer 0.48 0.82 0.96 0.99 0.94 1.00 0.72 0.88 1.00
Knobs & Levers
TurnOnStove 0.24 0.38 0.28 0.56 0.66 0.64 0.50 0.34 0.40
TurnOnSinkFaucet 0.32 0.66 0.66 0.63 0.70 0.56 0.62 0.80 0.84
TurnOffSinkFaucet 0.42 0.68 0.70 0.73 0.78 0.72 0.64 0.68 0.78
Buttons
CoffeePressButton 0.16 0.76 0.68 0.85 0.90 0.86 0.84 0.92 0.96
TurnOnMicrowave 0.38 0.68 0.88 0.78 0.68 0.84 0.94 0.86 0.96
TurnOffMicrowave 0.54 0.62 1.00 0.71 0.96 0.86 0.96 0.94 0.96
Average Success (All 24 Tasks) 0.23 0.41 0.43 0.50 0.51 0.57 0.50 0.575 0.64

Table 2: Average success rates on the LIBERO10 manipulation benchmark

Model DP-C DP-T OpenVLA \(\pi_0\) \(\pi_0\)-FAST UVA VideoPolicy Ours (RoboTALES)
Average Success Rate (%) 53% 58% 54% 85% 60% 90% 94% 97%

Ablation Study

Systematic ablations were performed across core algorithmic components to isolate individual contributions.

Table 3: Ablation study of core RoboTALES components on RoboCasa benchmark tasks

Configuration Planning Mechanism VLM Critic Feedback End-to-End Action Gradients Success Rate / Observations
VideoPolicy Baseline None (raw instruction \(\tau\)) None Decoupled feature extraction 0.575 (suffers from semantic drift)
Zero-shot Plan Appending Injected \(C^*\) at inference None Decoupled feature extraction Lowest performance (unadapted attention)
Planner-Conditioned Video Finetuned cross-attention None Decoupled feature extraction Solid gain over baseline (structured milestones)
Full Model (RoboTALES) Finetuned cross-attention Active DDPO reward steering Single-stage backprop 0.640 (lowest jerk, highest noise robustness)

Key Findings

  • Co-adaptation is essential: Directly feeding an LLM sub-goal string into an unadapted video baseline degrades performance, demonstrating that video diffusion layers must be explicitly tuned to parse and ground structured sub-goal tokens.
  • Superior sample efficiency: Trained on only 50 demonstrations, RoboTALES achieves 64% average success across RoboCasa, outperforming methods like OpenVLA (57%) and UVA (50%) which utilize 300 demonstrations. Even with as few as 10 demonstrations, RoboTALES maintains a competitive 43% success rate.
  • Physical motion smoothness & robustness: Under observation noise injection (\(\sigma = 0.06\)), baseline policies collapse toward near-zero completion, while RoboTALES retains steady execution. Critic-aligned representations also yield significantly lower latent and actuator jerk, ensuring smooth and purposeful trajectories.

Highlights & Insights

  • Diffusion sampling as an RL policy MDP: Framing the iterative multi-step video denoising process as a policy trajectory allows DDPO to propagate VLM rewards directly into continuous diffusion latents, bypassing non-differentiable rendering barriers.
  • Action-grounded world modeling: By allowing action diffusion gradients to shape video decoder blocks, the model resolves the classical trade-off between visual realism and control utility, producing representations tailored for manipulation.
  • Negligible inference overhead: Because the LLM planner is queried only once at task initiation, the execution overhead is merely ~1 second per episode, preserving real-time control feasibility.

Limitations & Future Work

  • Coarse semantic reward granularity: The frozen VLM critic yields global scalar alignment scores that lack fine-grained sensitivity to millimetric contact dynamics, causing occasional failures in tight peg-in-hole insertions.
  • High training compute footprint: Joint training of video diffusion decoder layers and action UNet requires 6โ€“7 days on dual NVIDIA A100 80GB GPUs, which poses challenges for scaling to massive multi-robot datasets.
  • Future directions: Integrating a learnable multi-modal critic that co-evolves with policy capabilities, and transferring the learned policies to physical dual-arm robotic systems.
  • vs VideoPolicy (Liang et al., 2025): VideoPolicy freezes SVD features and feeds them unidirectionally into an action head, lacking high-level reasoning and semantic reinforcement. RoboTALES introduces hierarchical sub-goal conditioning and DDPO critic steering with end-to-end action gradient backpropagation, boosting success by +6.5%.
  • vs Diffusion Policy (Chi et al., 2024): Diffusion Policy acts as a reactive imitation learner without internal visual foresight. RoboTALES integrates a predictive video world model that anticipates intermediate states, dramatically reducing action variance and execution failures.
  • vs OpenVLA / \(\pi_0\): Generalist VLAs generate discrete action tokens autoregressively with substantial inference latency. RoboTALES pairs high-level reasoning with continuous action diffusion, achieving higher sample efficiency under low-data regimes.

Rating

  • Novelty: โญโญโญโญโญ [Pioneering integration of hierarchical LLM planning, video diffusion simulation, VLM critic DDPO alignment, and end-to-end action diffusion]
  • Experimental Thoroughness: โญโญโญโญโญ [Evaluated on 24 diverse RoboCasa tasks and LIBERO10, supplemented with motion jerk, noise robustness, and demonstration scaling analyses]
  • Writing Quality: โญโญโญโญโญ [Clear problem formulation, elegant mathematical modeling of diffusion MDPs, and comprehensive empirical presentation]
  • Value: โญโญโญโญโญ [Provides an impactful blueprint for transforming generative video models into reliable, goal-aligned world models for robotic manipulation]