Practice Makes Perfect: From Explicit Decomposition to Reinforced Latent Planning in Text-to-Human Motion¶
Conference: ECCV 2026
Paper: ECCV Paper
Code: https://github.com/Erwin2233/LaCT-Motion
Area: Human Understanding
Keywords: text-to-motion generation, latent chain-of-thought, latent planning, reinforcement learning, GRPO
TL;DR¶
Addressing the severe inference latency of explicit Chain-of-Thought (CoT) in text-to-motion generation, LaCT-Motion mirrors human skill acquisition ("practice, internalization, perfection") by progressively compressing variable-length explicit reasoning steps into continuous latent tokens via a multi-pass curriculum, subsequently refined by latent GRPO to surpass explicit CoT generation quality with zero textual reasoning overhead at test time.
Background & Motivation¶
Text-to-motion generation represents a foundational capability for embodied robotics, 3D animation, and human-computer interaction. Recently, quantizing motion into discrete tokens (e.g., via VQ-VAE) and autoregressively modeling them using Large Language Models (LLMs) has emerged as a dominant paradigm. To tackle complex, multi-stage compositional motion instructions involving multi-joint coordination and fine-grained temporal dynamics, recent pioneering methods such as Motion-R1 and UniMo equip LLMs with explicit Chain-of-Thought (CoT) reasoning. In these explicit planning systems, the model sequentially predicts linguistic thought traces (e.g., body part decomposition, kinetic progression, and physical constraints) before generating discrete motion tokens, yielding remarkable enhancements in semantic fidelity and physical plausibility.
However, the explicit planning paradigm suffers from an unavoidable inference latency bottleneck: autoregressively generating lengthy, variable-length text reasoning tokens at test time incurs dramatic computational delays, making real-time deployment impractical. Conversely, the traditional "No Planning" paradigm directly maps input text to motion tokens; while fast, it lacks compositional reasoning capacity and frequently collapses when executing compound, long-horizon motion prompts. Furthermore, existing latent reasoning techniques developed in the NLP literature primarily target symbolic and mathematical domains with fixed-length reasoning steps and deterministic verifiable answers. Directly applying them to open-ended spatio-temporal motion synthesis presents unique challenges due to highly heterogeneous reasoning steps (1 to 9 atomic stages) and continuous multi-joint coordination dynamics.
The core insight of this work is that explicit reasoning should act as a pedagogical scaffold during training rather than a mandatory deliverable at inference time. Just as a novice pianist begins by deliberately reading sheet music note by note and gradually internalizes the passages into fluid, automatic muscle memory, a motion model should advance from conscious step-by-step reasoning to internalized continuous latent planning. Core idea: LaCT-Motion establishes a three-stage "practice-to-mastery" training recipe—comprising explicit CoT practice, progressive latent curriculum internalization via multi-pass hidden-state injection, and fully latent GRPO perfection with verifiable physical/semantic rewards—enabling superior spatio-temporal motion planning with zero explicit reasoning token overhead at inference.
Method¶
Overall Architecture¶
LaCT-Motion builds upon a pre-trained language model backbone (Qwen2.5-3B-Instruct) and a pre-trained motion discrete tokenizer (T2M-GPT VQ-VAE). The learning pipeline systematically transitions across three sequential stages: Stage I (Practice) performs supervised fine-tuning (SFT) on paired motion data with explicit CoT traces, instilling fundamental spatio-temporal reasoning priors; Stage II (Internalization) applies a stage-based curriculum schedule coupled with an iterative multi-pass forward mechanism, progressively compressing heterogeneous explicit reasoning steps into compact continuous latent tokens; and Stage III (Perfection) refines the policy under the fully latent regime using Group Relative Policy Optimization (GRPO) driven by format compliance, motion fidelity, and semantic alignment rewards. During inference, the model executes a single multi-pass injection over latent tokens and immediately reuses the KV cache for autoregressive greedy decoding of motion tokens, completely bypassing textual reasoning latency.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Motion Description q"] --> B["Explicit CoT Practice & Vocabulary Extension<br/>Incorporate motion and latent special tokens to establish structured reasoning foundation"]
B --> C["Progressive Latent Curriculum Compression & Stochastic Scheduling<br/>Excise explicit steps across curriculum stages and allocate continuous latent tokens"]
C --> D["Multi-Pass Forward Hidden-State Injection & Differentiable Chain<br/>Overwrite latent token embeddings with preceding hidden states and reuse KV cache"]
D --> E["Fully Latent Shared-Cache Reinforcement Perfection<br/>GRPO refinement with format/motion/semantic rewards and single latent forward rollout"]
E --> F["Output Discrete Motion Code Sequence a"]
Key Designs¶
1. Explicit CoT Practice & Vocabulary Extension: Building Structured Discrete Foundations
Because contemporary foundation LLMs lack native motion representations and latent reasoning boundaries, the model cannot initially generate or process motion tokens. LaCT-Motion expands the tokenizer with 512 discrete motion code tokens (<Motion_0> to <Motion_511>) matching the VQ-VAE codebook, 2 boundary markers (<Motion>, </Motion>), 2 explicit CoT delimiters (<think>, </think>), and 3 latent markers (<|start-latent|>, <|end-latent|>, <|latent|>). The input embedding matrix and language modeling head are resized accordingly, initializing the <|latent|> embedding from the delimiter "«" to supply an informative starting prior. In Phase A, the model undergoes supervised fine-tuning over 66,515 paired instances containing \(N_i\) atomic reasoning steps \(s_1, \ldots, s_{N_i}\) wrapped inside <think>...</think>, followed by target motion sequence \(\mathbf{a}\), establishing a solid foundation for physical decomposition before any latent compression begins.
2. Progressive Latent Curriculum Compression & Stochastic Scheduling: Smoothly Bridging Heterogeneous Spatial Steps
Directly forcing a model to reason in continuous latent space from scratch leads to optimization collapse due to uninformative random initializations. To mitigate this, Phase B introduces a progressive curriculum schedule. For a sample \(i\) with \(N_i\) explicit reasoning steps (ranging dynamically from 1 to 9 steps) and maximum curriculum stage \(K_{\max}\), stage \(k\) removes the first \(\min(k, N_i)\) explicit steps and replaces them with \(\min(k, N_i) \times c\) continuous latent tokens (where \(c\) is the latent token allocation budget per step): $$ \mathbf{x}^{(k)} = \big(q, \; t_{\mathrm{s}}, \; \underbrace{l_1, \ldots, l_{\min(k, N_i) \cdot c}}{\text{latent}}, \; t\big) $$ When }}, \; \underbrace{s_{\min(k, N_i)+1}, \ldots, s_{N_i}}_{\text{explicit CoT}}, \; \mathbf{a\(k \ge N_i\), all explicit reasoning text vanishes. To prevent the model from overfitting to rigid stage transitions, stochastic stage sampling is introduced: with probability \(p_u\), a random stage \(k' \sim \text{Uniform}(0, N_i)\) is sampled independently per training instance. Finally, Phase C consolidates learning by setting \(k=K_{\max}\) with \(p_u=0\), padding the latent region to a fixed budget of \(K_{\max} \cdot c\) tokens and enforcing pure continuous latent conditioning without residual explicit text.
3. Multi-Pass Forward Hidden-State Injection & Differentiable Chain: Maintaining Autoregressive Dependence
To ensure continuous latent tokens carry dynamic spatio-temporal representations rather than collapsing into static vectors, the framework implements an iterative multi-pass forward injection procedure. Given an input sequence with \(M\) latent tokens, computation is partitioned into \(M\) iterative passes followed by 1 final forward pass. At the \(i\)-th pass, the model processes up to latent token \(l_i\), extracts the last-layer hidden state \(\mathbf{h}_{i-1}\) from the immediately preceding position, and performs an in-place differentiable write: $$ \mathbf{e}(l_i) \leftarrow \mathbf{h}_{i-1} $$ Because \(\mathbf{h}_{i-1}\) resides at the top of the residual stream, its dimensionality naturally matches the input embedding space without requiring auxiliary projection layers. Crucially, \(\mathbf{h}_{i-1}\) remains attached to the computational graph (no gradient detachment), establishing a recurrent differentiable chain \(\mathbf{e}(l_i) = f_\theta(\mathbf{x}_{<\mathrm{pos}(l_i)}; \mathbf{e}(l_1), \ldots, \mathbf{e}(l_{i-1}))\) where downstream gradients back-propagate smoothly through all \(M\) injection steps. Furthermore, accumulated key-value (KV) caches are reused across passes to eliminate redundant computation. During loss calculation, question tokens \(q\), boundary markers \(t_{\mathrm{s}}, t_{\mathrm{e}}\), and latent tokens \(l_1, \ldots, l_M\) are masked with label -100, focusing cross-entropy supervision strictly on explicit reasoning tokens (when present) and motion tokens \(\mathbf{a}\).
4. Fully Latent Shared-Cache Reinforcement Perfection: Zero-Latency Policy Alignment
While curriculum SFT internalizes reasoning capabilities, fine-grained spatio-temporal fidelity requires further policy refinement. Stage III implements GRPO in the fully latent regime (\(k=K_{\max}\)). Bypassing value network estimation, the model optimizes against group-normalized advantages using three verifiable reward functions:
$$
\mathcal{R} = \lambda_1 r_{\mathrm{format}} + \lambda_2 r_{\mathrm{motion}} + \lambda_3 r_{\mathrm{semantic}}
$$
where format reward \(r_{\mathrm{format}} \in \{0, 1\}\) enforces <Motion> syntax compliance; motion similarity reward \(r_{\mathrm{motion}} = \cos(f_{\mathrm{motion}}(\hat{\mathbf{m}}), f_{\mathrm{motion}}(\mathbf{m}))\) computes cosine similarity between generated motion \(\hat{\mathbf{m}}\) and ground truth \(\mathbf{m}\) using a frozen pre-trained motion encoder; and semantic similarity reward \(r_{\mathrm{semantic}} = \cos(f_{\mathrm{motion}}(\hat{\mathbf{m}}), f_{\mathrm{text}}(\mathbf{q}))\) assesses cross-modal text-motion alignment. Importantly, because the latent multi-pass injection is completely deterministic under fixed model weights, sampling \(G\) rollouts per prompt requires executing the latent forward procedure only once; all \(G\) candidate trajectories share the resulting KV cache, and diversity arises purely from temperature-based autoregressive motion decoding. This eliminates redundant multi-pass computations and substantially accelerates RL training.
Loss & Training¶
During the supervised fine-tuning phases (Phase A, B, and C), the model minimizes sequence-level cross-entropy loss: $$ \mathcal{L}{\mathrm{SFT}} = - \sum_j) $$ where active target index set }} \log p_\theta(x_{j+1} \mid \hat{\mathbf{y}\(\mathcal{A}\) contains only unmasked explicit reasoning tokens and answer motion tokens. In Stage III (GRPO), the policy samples \(G=8\) rollouts per prompt and incorporates a KL divergence penalty (\(\beta=0.001\)) against the reference policy \(\pi_{\mathrm{ref}}\). Fine-tuning runs for 7 epochs on 2 NVIDIA H200 GPUs using the Qwen2.5-3B-Instruct base model, with curriculum hyper-parameters set to \(K_{\max}=8\) and \(c=2\).
Key Experimental Results¶
Main Results¶
On the benchmark HumanML3D dataset (22 joints, 263-dimensional representation), LaCT-Motion is evaluated against leading diffusion models, discrete auto-regressive LLM models, and explicit CoT frameworks. All metrics report the mean and 95% confidence intervals across 20 independent runs.
| Method Category | Model | R-Precision (Top-1) ↑ | R-Precision (Top-2) ↑ | R-Precision (Top-3) ↑ | FID ↓ | MM-Dist ↓ | Diversity ↑ |
|---|---|---|---|---|---|---|---|
| Diffusion | MDM | 0.320 ± .005 | 0.498 ± .004 | 0.611 ± .007 | 0.544 ± .044 | 5.566 ± .027 | 9.559 ± .086 |
| Diffusion | MLD | 0.481 ± .003 | 0.673 ± .003 | 0.772 ± .002 | 0.473 ± .013 | 3.196 ± .010 | 9.724 ± .082 |
| Diffusion | MotionDiffuse | 0.491 ± .001 | 0.681 ± .001 | 0.782 ± .001 | 0.630 ± .001 | 3.113 ± .001 | 9.410 ± .049 |
| Discrete / LLM | T2M | 0.457 ± .002 | 0.559 ± .007 | 0.740 ± .003 | 1.067 ± .002 | 3.340 ± .008 | 9.188 ± .002 |
| Discrete / LLM | TM2T | 0.424 ± .003 | 0.618 ± .003 | 0.729 ± .002 | 1.501 ± .017 | 3.467 ± .011 | 8.589 ± .076 |
| Discrete / LLM | T2M-GPT | 0.491 ± .003 | 0.680 ± .003 | 0.775 ± .002 | 0.116 ± .004 | 3.118 ± .011 | 9.761 ± .081 |
| Discrete / LLM | MotionGPT | 0.492 ± .003 | 0.681 ± .003 | 0.778 ± .002 | 0.232 ± .008 | 3.096 ± .008 | 9.528 ± .071 |
| Discrete / LLM | MoMask | 0.521 ± .002 | 0.713 ± .002 | 0.807 ± .002 | 0.045 ± .002 | 2.958 ± .008 | 9.620 ± .064 |
| Discrete / LLM | MotionAgent | 0.515 ± .004 | 0.691 ± .003 | 0.801 ± .004 | 0.230 ± .009 | 2.967 ± .020 | 9.908 ± .102 |
| Discrete / LLM | MotionGPT-3 | 0.542 ± .003 | 0.741 ± .003 | 0.839 ± .002 | 0.065 ± .003 | 2.768 ± .008 | 9.744 ± .082 |
| Discrete / LLM | MG-MotionLLM | 0.523 ± .003 | 0.718 ± .003 | 0.814 ± .002 | 0.090 ± .002 | 2.926 ± .009 | 9.632 ± .076 |
| Explicit CoT + RL | Motion-R1 | 0.515 ± .003 | 0.719 ± .002 | 0.818 ± .002 | 0.201 ± .004 | 2.854 ± .010 | 10.026 ± .075 |
| Explicit CoT SFT | UniMo | 0.539 ± .003 | 0.738 ± .002 | 0.831 ± .002 | 0.177 ± .004 | 2.768 ± .010 | 10.042 ± .076 |
| Latent Reasoning | TBYM | 0.537 ± .005 | 0.721 ± .004 | 0.810 ± .005 | 0.040 ± .002 | 2.895 ± .017 | 9.668 ± .077 |
| Ours | LaCT-Motion | 0.577 ± .002 | 0.783 ± .002 | 0.873 ± .002 | 0.167 ± .004 | 2.506 ± .005 | 9.507 ± .098 |
Ablation Study¶
1. Ablation on Training Stages and Curriculum Mechanisms¶
This ablation examines the progressive contribution of each training stage (Stage I: Practice; Stage II: Internalization; Stage III: Perfection) and compares against single-pass distillation.
| Experimental Configuration | Stage I (SFT) | Stage II (Curriculum) | Stage III (GRPO) | R Top-1 ↑ | R Top-2 ↑ | R Top-3 ↑ | FID ↓ | MM-Dist ↓ | Diversity ↑ |
|---|---|---|---|---|---|---|---|---|---|
| Pipeline Checkpoint 1 | ✓ | 0.167 | 0.240 | 0.289 | 24.119 | 7.069 | 5.655 | ||
| Pipeline Checkpoint 2 | ✓ | ✓ | 0.453 | 0.640 | 0.736 | 0.466 | 3.344 | 10.216 | |
| Pipeline Checkpoint 3 (full model) | ✓ | ✓ | ✓ | 0.577 | 0.783 | 0.873 | 0.167 | 2.506 | 9.507 |
| Independent baseline UniMo* (pure SFT) | ✓* | 0.438 | 0.613 | 0.702 | 0.201 | 3.588 | 9.748 | ||
| Independent baseline (latent curriculum conv.) | ✓ | ✓ | 0.486 | 0.660 | 0.756 | 0.244 | 3.190 | 9.915 | |
| Independent baseline UniMo† (explicit CoT+GRPO) | ✓† | ✓ | 0.529 | 0.739 | 0.832 | 0.203 | 2.743 | 9.780 | |
| Single-pass latent distillation baseline ‡ | ✓ | ✓‡ | ✓ | 0.520 | 0.716 | 0.822 | 0.221 | 2.838 | 9.924 |
2. Ablation on Latent Token Budget \(c\) and GRPO Rewards¶
The table evaluates the number of latent tokens allocated per replaced step \(c\) (evaluated after Stage II on all 4,646 test set samples) alongside the decomposition of GRPO reward signals in Stage III.
| Variable | Configuration / Component | R Top-1 ↑ | R Top-2 ↑ | R Top-3 ↑ | FID ↓ | MM-Dist ↓ | Diversity ↑ | Evaluation Time (s) ↓ |
|---|---|---|---|---|---|---|---|---|
| Latent Token Budget \(c\) | \(c = 1\) | 0.458 | 0.629 | 0.724 | 0.536 | 3.357 | 9.802 | 369 |
| Latent Token Budget \(c\) | \(c = 2\) (default) | 0.453 | 0.640 | 0.736 | 0.466 | 3.344 | 10.216 | 387 |
| Latent Token Budget \(c\) | \(c = 4\) | 0.451 | 0.630 | 0.729 | 0.529 | 3.381 | 9.873 | 455 |
| GRPO Rewards | \(r_f + r_m\) | 0.581 | 0.768 | 0.861 | 0.217 | 2.543 | 9.745 | - |
| GRPO Rewards | \(r_f + r_s\) | 0.576 | 0.771 | 0.859 | 0.261 | 2.529 | 9.641 | - |
| GRPO Rewards | \(r_f + r_m + r_s\) (full model) | 0.577 | 0.783 | 0.873 | 0.167 | 2.506 | 9.507 | - |
Key Findings¶
- Latent Continuous Representations Outperform Explicit Discrete Text: LaCT-Motion attains an R-Precision Top-1 of 0.577, outperforming SFT-based models (UniMo at 0.539, MotionGPT-3 at 0.542) and surpassing explicit CoT + GRPO (Motion-R1 at 0.515) by +0.062 (+12.0% relative improvement). This indicates that continuous latent embeddings represent complex spatial coordination and kinematic constraints more effectively than discrete natural language tokens.
- Multi-Pass Forward Injection is Crucial for Internalization: In contrast to the single-pass distillation baseline (Top-1: 0.520, FID: 0.221), the progressive curriculum with multi-pass hidden-state injection boosts Top-1 by +0.057 and lowers FID by 0.054 under identical capacity and supervision, verifying that recursive differentiable hidden injection is vital for preserving reasoning capacity.
- Orders of Magnitude Speedup in Inference: While explicit CoT planning (Motion-R1) requires approximately 1 hour to evaluate the full HumanML3D test set due to auto-regressive textual decoding, LaCT-Motion completes evaluation in only 7 minutes (an 88%+ reduction in wall-clock time). Setting \(c=2\) provides the optimal trade-off between generation diversity (10.216) and physical fidelity (FID 0.466).
- Complementary Synergy Among Verifiable Rewards: In GRPO reward ablation, omitting the motion similarity reward \(r_m\) severely impairs FID (degrading from 0.167 to 0.261). Incorporating semantic similarity \(r_s\) lowers multimodal distance (MM-Dist to 2.506) and boosts Top-3 R-Precision to 0.873, demonstrating the distinct value of physical and semantic supervision.
Highlights & Insights¶
- Decoupling Pedagogical Scaffolding from Deployment Inference: Treating explicit CoT as a training-time cognitive scaffold and progressively internalizing it into continuous hidden states resolves the dilemma between high-level reasoning capability and real-time generation latency.
- Zero-Cost Rollout Acceleration via KV Cache Sharing: Because multi-pass latent injection is strictly deterministic given fixed model parameters, all \(G\) candidate rollouts per prompt in GRPO share a single precomputed KV cache, reducing the multi-pass injection overhead from \(G\) passes to 1 pass.
- Continuous Latents Breaking Discrete Linguistic Bottlenecks: Empirical findings demonstrate that continuous vectors in residual stream space avoid the quantization and syntactic rigidity of natural language, providing a more natural substrate for physical and temporal motion representations.
Limitations & Future Work¶
- Dependency on Underlying Discrete Tokenizer: Generation fidelity remains bounded by the codebook capacity of the pre-trained VQ-VAE; high-frequency limb vibrations or rapid contacts occasionally show quantization artifacts.
- Static Curriculum Hyper-parameters: The number of latent tokens per step (\(c=2\)) and maximum stages (\(K_{\max}=8\)) are currently fixed heuristics, lacking dynamic allocation based on input prompt complexity.
- Future Directions: Exploring input-dependent dynamic latent token budgeting, and extending latent planning to human-object interaction (HOI) and multi-agent coordination scenarios.
Related Work & Insights¶
- vs Motion-R1: Motion-R1 pioneers explicit CoT and GRPO for motion generation but requires full auto-regressive generation of explicit reasoning tokens during test time (1-hour test set latency); LaCT-Motion compresses reasoning into latent space, cutting inference time to 7 minutes while boosting R-Precision Top-1 by +0.062.
- vs UniMo: UniMo relies on explicit CoT SFT for unified motion understanding and generation; LaCT-Motion demonstrates that internalized latent tokens coupled with GRPO achieve substantially tighter text-motion alignment without requiring explicit textual traces.
- vs TBYM (Think-Before-You-Move): Concurrent work TBYM explores latent motion reasoning by training an auto-regressive Transformer from scratch without general LLM foundations or RL refinement; LaCT-Motion leverages a 3B LLM backbone and a practice-to-mastery recipe, achieving markedly higher R-Precision across all settings.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First framework to adapt latent CoT reasoning to continuous multimodal human motion generation via a complete "practice-internalize-perfect" paradigm.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarking across diffusion/discrete baselines, multi-stage checkpoints, reward ablations, token budget studies, and qualitative evaluations.
- Writing Quality: ⭐⭐⭐⭐⭐ Compelling cognitive analogy of skill acquisition, accompanied by mathematically rigorous formulations of differentiable multi-pass injection.
- Value: ⭐⭐⭐⭐⭐ Provides a viable and effective blueprint for resolving the inference latency bottleneck of reasoning-augmented generative models.