Skip to content

TDSR-VLA: Transition-aware Denoising Sequence Representations for Vision-Language-Action

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Robotics & Embodied AI
Keywords: Vision-Language-Action Models, Subgoal Planning, Flow Matching, Denoising Sequence Representations, Bimanual Manipulation

TL;DR

TDSR-VLA proposes a decoupled VLA architecture that reuses early sequence representations (SR) and iterative updated sequence representations (USR) produced during the denoising process of a Stable Diffusion 3 Vision Planner, injecting them via layer-wise Prefix-KV conditioning to guide a flow-matching Action Expert for continuous action chunk generation, substantially boosting small-data long-horizon bimanual manipulation.

Background & Motivation

Vision-Language-Action (VLA) models adapt large pretrained vision-language backbones on multi-robot demonstrations, holding great promise for semantic instruction grounding and zero-shot physical generalization. However, monolithic end-to-end VLA policies face severe challenges when deployed under distribution shift with scarce target-domain supervision: when provided with only dozens to a few hundred demonstration trajectories, a unified network must simultaneously learn high-level semantic planning ("what to do next") and low-level continuous control ("how to transition smoothly"). This parameter entanglement makes the policy fragile, frequently leading to error accumulation, drifting, or idle execution over extended task horizons.

To alleviate this issue, subgoal-image planning interfaces have gained widespread adoption, employing generative diffusion models to synthesize a visual observation of a future state \(\Delta\) steps ahead as an explicit target. Nevertheless, a static subgoal image merely specifies the final target appearance to reach, but severely underconstrains the physical transitions and corrective maneuvers required to transition toward it. Recent works attempt to supplement subgoal images with optical flow motion tokens, video features, or generated human videos; however, these approaches rely on auxiliary networks and discard the rich internal dynamic trajectories already formed inside the subgoal generator itself.

The key insight of this paper is that as a diffusion visual planner iteratively denoises a future subgoal image, its cross-modal attention mechanisms naturally compute and refine the contextual evolution of the scene over time. Core idea: decouple VLA into a visual planner and an action expert, directly reusing the internal denoising-time sequence representationsโ€”an early current-grounded Sequence Representation (SR) and a final Updated Sequence Representation (USR)โ€”to guide a continuous flow-matching Action Expert through layer-wise Prefix-KV conditioning, achieving robust long-horizon manipulation without robot-demonstration pretraining.

Method

Overall Architecture

TDSR-VLA decouples visual reasoning with internet-scale generative priors from action execution requiring precise, small-data motor tuning. At each execution cycle, the system takes current multi-view RGB observations, proprioceptive robot states, and language instructions. A fine-tuned Vision Planner first predicts a future subgoal image at a fixed horizon alongside two conditioning hidden sequences extracted across denoising steps. Subsequently, the Action Expert treats these planner-side representations and vision features as layer-wise Prefix Key-Value pairs, generating a continuous chunk of action vectors via Conditional Flow Matching (CFM).

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Current Multi-view Observations & Language<br/>$o_t = [I_t^1, \dots, I_t^n, q_t], \ell$"] --> DGF["Decoupled Generative Factorization<br/>Planning & Control split: $P_\psi(S_{t+\Delta}, R_t \mid I_t^1, q_t, \ell) \cdot P_\varphi(A_{t:t+H-1} \mid S_{t+\Delta}, R_t, o_t, \ell)$"]
    DGF --> DSR["Diffusion Denoising Sequence Representations<br/>MM-DiT extracts grounded SR & refined transition USR"]
    DSR --> PKV["Layer-wise Prefix-KV Conditioning<br/>Layers 0-11 inject SR / current views; Layers 12-17 inject USR / subgoal"]
    PKV --> CFM["Continuous Flow-Matching Action Generation<br/>Gemma backbone + AdaRMS modulation + 10-step ODE integration"]
    CFM --> Out["Continuous Action Chunk Execution<br/>$A_{t:t+H-1} \in \mathbb{R}^{H \times d_a}, H=50$"]

Key Designs

1. Decoupled Generative Factorization: Disentangling Planning from Control To overcome the failure mode where monolithic policies overfit or collapse when jointly optimizing high-level spatial planning and low-level torque commands on scarce data, TDSR-VLA explicitly factorizes the joint probability of action chunks \(A_{t:t+H-1}\) and planning signals conditioned on observations \(o_t\) and instruction \(\ell\): $\(P_\theta(A_{t:t+H-1}, S_{t+\Delta}, R_t \mid o_t, \ell) = P_\psi(S_{t+\Delta}, R_t \mid I_t^1, q_t, \ell) \cdot P_\varphi(A_{t:t+H-1} \mid S_{t+\Delta}, R_t, o_t, \ell)\)$ Here, \(S_{t+\Delta}\) is the reachable subgoal image at prediction horizon \(\Delta\) (default offset of 10 frames), and \(R_t = (SR_t, USR_t)\) denotes internal planner transition representations. This decoupling allows the Vision Planner \(P_\psi\) to inherit general visual and physical priors from large-scale generative models, while the low-level Action Expert \(P_\varphi\) concentrates entirely on action trajectory generation under explicit, local target guidance.

2. Diffusion Denoising Sequence Representations: Extracting Grounded and Transition Cues Static subgoal images display reachable appearances but lack step-by-step transition and error-correction information. TDSR-VLA adapts Stable Diffusion 3 with an MM-DiT backbone. In addition to text instructions, the current primary image \(I_t^1\) is encoded via SigLIP into visual tokens and concatenated with fixed-point serialized robot state vectors \(q_t\). Throughout the multi-step rectified flow denoising process, conditioning tokens and noisy image latents undergo joint cross-attention. Rather than discarding intermediate conditioning streams upon decoding the RGB subgoal, TDSR-VLA extracts two conditioning hidden sequences of dimension \(L_p \times d_p^{\text{SD3}}\) (\(333 \times 1536\)): - Sequence Representation (SR): An early-step conditioning hidden sequence anchoring the policy to the current grounded scene \(I_t^1\), proprioceptive state \(q_t\), and text instruction; - Updated Sequence Representation (USR): The final-step conditioning hidden sequence obtained after full denoising iterations, having repeatedly attended to the evolving image latent, thus encoding dynamic transition cues from the current state toward the predicted subgoal.

3. Layer-wise Prefix-KV Conditioning: Lower-layer Grounding and Upper-layer Subgoal Guidance To allow the policy to balance current physical feasibility with prospective target reaching, TDSR-VLA introduces a phased Prefix-KV conditioning scheme. In the 18-layer Gemma-based transformer Action Expert, layer prefix keys and values are partitioned across functional roles: $\(\begin{aligned} \text{Layer}^{0:11}_{\text{current}}: \quad & K^{\text{pre}}_i = \text{MLP}(SR_t), \quad & V^{\text{pre}}_i = \text{MLP}(\text{SigLIP}(I_t^{1:n})) \\ \text{Layer}^{12:17}_{\text{subgoal}}: \quad & K^{\text{pre}}_i = \text{MLP}(USR_t), \quad & V^{\text{pre}}_i = \text{MLP}(\text{SigLIP}(S_{t+\Delta})) \end{aligned}\)$ The lower 12 layers (Layers 0โ€“11) serve as current-grounding layers, utilizing \(SR_t\) as keys and SigLIP features from multi-view current cameras (top-down and wrist views) as values to anchor actions in the actual physical geometry. The upper 6 layers (Layers 12โ€“17) act as subgoal-guidance layers, employing \(USR_t\) as keys and SigLIP embeddings of the synthesized subgoal \(S_{t+\Delta}\) as values to steer action trajectories. Suffix action tokens attend bidirectionally across the full \(H\)-step chunk and cross-attend to prefix tokens, ensuring seamless global temporal coordination.

4. Continuous Flow-Matching Action Generation: Bidirectional Chunking without Simulators Departing from autoregressive discrete tokenization that suffers from compounding errors and slow sequential sampling, TDSR-VLA builds on Conditional Flow Matching (CFM) for continuous action chunk prediction. Action vectors are zero-padded to \(d_{\max}=32\) and linearly projected to latent action tokens over horizon \(H=50\). Flow time \(\tau \sim \text{Beta}(1.5, 1.0)\) is encoded via sinusoidal embeddings and an MLP, modulating transformer blocks through Adaptive RMSNorm (AdaRMS). At test time, a standard Euler ODE solver samples smooth, coordinated action chunks within just \(K=10\) integration steps, eliminating the unidirectional constraint of causal masks.

Loss & Training

The architecture is trained in two decoupled stages: 1. Vision Planner Fine-tuning: Freezing the VAE and text/vision encoders, MM-DiT is fine-tuned with LoRA under the rectified-flow mean squared error loss: $\(\mathcal{L}_{\text{RF}} = \mathbb{E}_{\tau, x_0^S, \epsilon} \left[ \| v_\psi(x_\tau^S, \tau \mid I_t^1, q_t, \ell) - u_\tau^S \|_2^2 \right]\)$ where the target velocity field is \(u_\tau^S = \epsilon - x_0^S\). 2. Precomputation Caching & Action Expert Training: To eliminate repetitive and costly diffusion rollouts during action learning, the fine-tuned Vision Planner precomputes and caches \((S_{t+\Delta}, SR_t, USR_t)\) across the entire dataset. The Action Expert is subsequently trained with a frozen planner directly loading cached representations under the CFM objective: $\(\mathcal{L}_{\text{CFM}} = \mathbb{E}_{\tau, x_0, z} \left[ \| v_\varphi(x_\tau, \tau \mid \{K_i^{\text{pre}}, V_i^{\text{pre}}\}_{i=0}^{17}) - u_\tau \|_2^2 \right]\)$ where \(x_\tau = (1-\tau)x_0 + \tau z\) and \(u_\tau = z - x_0\). This caching prevents optimization instabilities between planning and control.

Key Experimental Results

Main Results

On the LIBERO simulation benchmark, TDSR-VLA is evaluated on 10 tasks in the Goal suite and 10 long-horizon tasks in the Long suite (50 demonstrations per task, 500 evaluation episodes per suite). Without any robot-demonstration pretraining, TDSR-VLA outperforms the fully pretrained OpenVLA baseline on both suites, and surpasses \(\pi_0\) on the long-horizon benchmark.

Model Demonstration Pretraining LIBERO-Goal Success Rate (%) LIBERO-Long Success Rate (%) Summary / Gain
OpenVLA Yes (Open X-Embodiment) 79.2 53.7 Autoregressive baseline
\(\pi_0\) No 89.0 48.0 Monolithic flow matching, degrades on long horizons
TDSR-VLA (Ours) No 82.0 56.0 Outperforms OpenVLA (+2.8% / +2.3%); outperforms \(\pi_0\) on Long (+8.0%)

In real-world bimanual manipulation using two 7-DOF OpenArm v0.3 arms (16-DOF total control) across 6 tasks (4 sequential, 2 object-cooperative, 100 demonstrations per task, fine-tuned for 10 epochs, 20 trials per task, scored 0.5 per arm):

Task Suite SmolVLA (Score) \(\pi_0\) (Score) GR00T N1.5 (Score) TDSR-VLA (Ours Score) Gain over Top Baseline
Sequential Tasks Mean (Tasks a, b, d, e) 0.58 0.76 0.76 0.80 +0.04
Object-Cooperative Tasks Mean (Tasks c, f) 0.59 0.63 0.70 0.88 +0.18
Overall Mean Score 0.58 0.71 0.74 0.825 (95% CI: [0.767, 0.879]) +0.083 (Paired diff CI: [0.004, 0.163])

Ablation Study

Ablations investigate the depth of subgoal injection layers and the indispensability of the Updated Sequence Representation (USR).

Configuration Subgoal Layers / Key Setting OpenArm Score Findings & Observations
0 Subgoal Layers All 18 layers conditioned on current views & SR ~0.71 Lacks prospective visual guidance, prone to stagnation
3 Subgoal Layers 15 current layers + 3 subgoal layers ~0.81 Captures target guidance, closely trailing 6 layers
6 Subgoal Layers (Full Model) 12 current (SR) + 6 subgoal (USR) 0.83 Optimal balance between current grounding and prospective guidance
9 Subgoal Layers 9 current layers + 9 subgoal layers ~0.73 Excessive subgoal injection weakens grounding to current physical state
Replace with SR (w/o USR) Subgoal layer keys use initial SR instead of USR 0.70 Drops by 0.13, demonstrating USR encodes crucial dynamic transition cues

Key Findings

  • USR provides vital dynamic transition and error-correction cues: Linear ridge probing on unseen LIBERO trajectories shows that frozen USR representations achieve \(R^2 = 0.48\) for 10-step future end-effector displacement, \(R^2 = 0.40\) for mean future actions, and a Macro-F1 of 0.74 for progress bin classification. USR substantially outperforms proprioceptive state, subgoal image tokens, and shuffled USR, proving it captures temporal transition manifolds rather than static appearance.
  • Substantial advantages on coordinated bimanual tasks: In tightly coupled cooperative tasks such as bimanual cube transfer (Task c) and wine glass handover (Task f), TDSR-VLA achieves an impressive score of 0.88, vastly outstripping GR00T N1.5 (0.70) and \(\pi_0\) (0.63), demonstrating the value of explicit subgoal cues in multi-arm coordination.

Highlights & Insights

  • Turning diffusion denoising byproducts into policy guidance: While standard pipelines view generative models as black-box image renderers, TDSR-VLA reuses the internal condition-stream representations (USR) shaped during cross-attention denoising, obtaining rich temporal transition signals without training auxiliary optical flow or video models.
  • Precomputed caching decouples heavy generative training from control: By caching visual subgoals and intermediate states offline, Action Expert training avoids expensive multi-step diffusion sampling during backpropagation, reducing training wall-clock time and preventing adversarial gradient coupling.

Limitations & Future Work

  • Inference latency from synchronous planning: End-to-end latency per action chunk is 3453.4 ms, heavily dominated by the Vision Planner (3014.8 ms, ~87%). In a 30 Hz control loop, this causes an estimated 1.79 s idle pause between chunks, leading to a slight performance dip on high-tempo tasks like Dual Cup Sorting (0.55 vs. GR00T's 0.58).
  • Future directions: Developing asynchronous subgoal planning pipelines that overlap planner inference with action chunk execution, and leveraging few-step consistency distillation to compress diffusion latency.
  • vs FlowVLA / VideoVLA: FlowVLA requires explicit optical flow generators, and VideoVLA relies on heavy video diffusion models. TDSR-VLA achieves state transition awareness directly from the internal denoising states of a single-image planner, offering a cleaner modular architecture.
  • vs \(\pi_0\) / GR00T: Monolithic architectures map observations straight to continuous actions, which degrade over long horizons when trained without vast multi-robot datasets. TDSR-VLA proves that decoupling visual planning from action execution significantly boosts sample efficiency on scarce real-world demonstration budgets.

Rating

  • Novelty: โญโญโญโญโญ Reusing internal diffusion conditioning states (SR/USR) as action transition cues is highly innovative.
  • Experimental Thoroughness: โญโญโญโญโญ Evaluated across LIBERO simulation and 6 real-world 14-DOF bimanual tasks with layer ablations, feature probing, and latency analysis.
  • Writing Quality: โญโญโญโญโญ Well-structured, clear mathematical formulation of decoupled factorizations, and comprehensive tables.
  • Value: โญโญโญโญโญ Offers a robust blueprint for data-efficient, long-horizon multi-arm robotic manipulation.