BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion¶
Conference: NeurIPS2026
arXiv: 2609.35407
Paper: Project page
Area: Human Understanding
Keywords: bidirectional motion-text generation, masked discrete diffusion, unified vocabulary, two-stage training, self-correction
TL;DR¶
BiMoGen combines a unified motion-text vocabulary with a bidirectional masked diffusion model for motion generation and captioning, using correspondence learning before conditional generation and generation-aware self-correction to achieve T2M R@1 of 0.555 and M2T CIDEr of 60.2 on HumanML3D, without outperforming existing unified models on every metric.
Background & Motivation¶
Text-to-motion generation (T2M) and motion-to-text captioning (M2T) rely on the same semantic correspondence: actions, directions, and temporal order expressed in a sentence should be grounded in human motion. MotionGPT and related methods discretize motion into tokens and model them alongside text autoregressively, enabling model sharing. However, a fixed left-to-right generation order is poorly matched to this correspondence. An incorrect early motion token becomes an immutable prefix for subsequent predictions, even when later context could help correct it. Masked generation methods such as MoMask can exploit bidirectional context, but mainly serve unidirectional T2M rather than replacing a unified bidirectional model.
Masked discrete diffusion appears to resolve the ordering problem, but still faces two difficulties. Early in training, the model is asked to generate one modality from the other before learning reliable language-motion correspondence. Differences in sequence length can also bias unified prediction toward motion reconstruction. Less visibly, ordinary masked training retains ground-truth tokens at visible positions, whereas sampling retains the model's own predictions as context. In early sampling states with high mask ratios, errors committed from sparse context can still affect later predictions. Parallel generation alone does not establish an ability to correct mistakes.
The paper therefore separates learning correspondence from learning to handle self-generated errors. Pretraining corrupts both motion and text to establish shared semantic grounding; fine-tuning then preserves the condition and predicts only the target. Self-correction trains the same model to recover ground truth from its own complete predictions and revises committed tokens during sampling, rather than merely filling unresolved masks. Core Idea: separate unified bidirectional masked generation into cross-modal correspondence pretraining and conditional-generation fine-tuning, then use generation-aware self-correction to explicitly reject the assumption that generated tokens are always correct.
Method¶
Overall Architecture¶
BiMoGen uses one bidirectional Transformer rather than separate T2M and M2T generators. A pretrained VQ-VAE compresses continuous human motion into discrete tokens, and its motion vocabulary is appended to the text vocabulary. Both modalities enter a single sequence, with the same model predicting clean tokens at any position. T2M preserves text and generates motion tokens that the VQ-VAE decodes into motion; M2T preserves motion tokens and generates text.
The pipeline contains four designs: Unified Token Space, Decoupled Training, GASC Training, and GASC Sampling. The first two define representation and learning stages; the latter two introduce self-generated context into both training and sampling. Solid arrows indicate data or model handoffs, while dashed arrows specifically identify training supervision. Ground-truth target supervision is absent at inference time.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Paired text + motion"] --> B["Unified Token Space"]
B --> C["Decoupled Training<br/>PT then SFT"]
C --> D["GASC Training<br/>in PT and SFT"]
G["Ground-truth tokens"] -.->|training supervision| C
G -.->|all-region supervision| D
D -->|trained model| E["GASC Sampling"]
H["Visible condition + masked target"] --> E
E -->|iterative refinement| E
E --> F["Motion decoding or text output"]
The motion tokenizer is an existing single-layer VQ-VAE, not a new high-capacity residual quantization architecture. The model uses a LLaDA-style bidirectional backbone, but the appendix states that BiMoGen is trained from scratch. Adopting the LLaDA architecture should therefore not be interpreted as inheriting a pretrained language model's capabilities. Self-correction introduces no separate teacher, discriminator, or correction network: the same Transformer generates and revises predictions.
Key Designs¶
1. Unified Token Space: share a discrete prediction interface across both directions
Continuous motion and text originally occupy different output spaces. BiMoGen uses a pretrained VQ-VAE encoder to downsample motion temporally and map it to codebook indices. The experimental downsampling rate is 4, so a 120-frame motion corresponds to 30 motion tokens. This example illustrates compression only: tokens are not independent action labels, and an individual codebook index cannot directly be interpreted as โwalkingโ or โturning.โ After generation, the VQ-VAE decoder reconstructs continuous motion from the complete token sequence.
Appending the motion codebook to the text vocabulary establishes a shared token-ID space and prediction interface, not identical semantics for motion and text tokens. After concatenation, bidirectional attention allows prediction at a masked position to access visible tokens on both sides and visible content from the other modality. Unlike causal attention restricted to a left prefix, this interface better supports mutual constraints between motion chronology and caption semantics.
Forward corruption replaces tokens with [M] independently at a shared sampled mask ratio, rather than adding Gaussian noise to skeletal coordinates. Training spans nearly complete to nearly fully masked contexts; reverse generation starts from a fully masked target and progressively recovers confident tokens. Switching direction primarily changes which segment serves as condition and which as target, without changing the architecture. Motion completion can use the same interface.
2. Decoupled Training: learn correspondence before specializing in conditional generation
The first stage is masked pretraining (PT). Each paired sample is arranged as either textโmotion or motionโtext with equal probability, preventing a fixed modality order from becoming a positional shortcut. Both segments undergo independent token corruption with the same mask ratio, and all masked positions in both modalities receive supervision. โSymmetricโ means that either modality can be predicted; it does not force equal numbers of masked tokens in the two segments. The model learns both intra-modal structure and semantic recovery through the other modality.
The second stage is supervised fine-tuning (SFT). For T2M, complete text is the condition and motion is the target; M2T reverses these roles. Ordinarily, the condition remains fully visible, only the target is corrupted, and the standard masked loss covers only masked target positions. This distinction matters: pretraining asks what is missing from a joint sequence, whereas fine-tuning asks how to generate the other modality from a complete condition. PT alone does not guarantee generation quality, as the ablation below directly demonstrates.
Fine-tuning also includes auxiliary motion-to-motion (M2M) supervision: randomly selected motion positions form the condition and the remaining positions form the target. This provides denser intra-motion supervision to supplement the limited constraints supplied by text. Fine-tuning batches mix T2M, M2T, and M2M at 8:1:1, so the two cross-modal directions are not trained in equal proportions. Condition dropout replaces the entire condition with masks in 10% of cases, allowing unconditional prediction for classifier-free guidance (CFG) in T2M. This exception does not change the ordinary SFT rule of corrupting only the target.
3. GASC Training: learn all-region correction from complete self-generated predictions
Ordinary masked training uses ground-truth visible context, whereas inference can contain erroneous generated tokens. GASC Training first performs a standard prediction pass on corrupted input, then replaces every token in the current generation region with the model's argmax prediction and stops gradient propagation through these discrete predictions. The generation region is the entire motion-text sequence in PT and only the target in SFT; tokens outside it remain unchanged. Crucially, replacement is not limited to originally masked positions: it covers the entire generation region, including originally visible ground-truth tokens.
The same Transformer then reads this self-generated context and predicts the ground-truth tokens again. Correction loss covers the entire generation region rather than only originally masked positions, because a visible, apparently confident token can also be wrong during sampling. The model therefore learns not just to fill blanks but to recover a true sequence from its own potentially inconsistent predictions. Detachment trains correction through the second pass without propagating through the first pass's nondifferentiable argmax; the standard masked loss still trains the first prediction pass.
Correction training uses a complete candidate constructed from one prediction pass, rather than reproducing the full 20-step sampling trajectory. It reduces the mismatch between training and inference contexts but does not establish complete elimination of exposure bias. PT generates both modalities, whereas conditional sampling generates only the target, so their error states are not identical either. The contribution is to learn revision using a shared model and additional supervision, not to guarantee that every revision is more accurate.
4. GASC Sampling: revise committed tokens while leaving unresolved positions for later steps
Standard sampling preserves the condition and initializes the target entirely as [M]. Each step predicts masked positions in parallel, uses the maximum predicted probability at each position as confidence, progressively accepts confident tokens, and leaves low-confidence positions masked. T2M receives its target motion length externally; M2T generates up to a maximum text length and truncates at the first [EOS]. The paper does not solve automatic motion-duration inference from unrestricted text.
The default schedule uses 20 sampling steps and triggers correction at steps 5 and 10. At a trigger, visible tokens are retained and masked positions are temporarily filled with the current predictions to construct a complete target candidate. The model performs an additional correction pass on this candidate. Corrected predictions are written back only to target positions that were already unmasked. Uncommitted positions are not permanently filled by correction and remain subject to subsequent ordinary sampling. The appendix algorithm reuses the current masked predictions to complete the candidate, so each trigger explicitly adds one correction-model forward pass.
This is not an ordinary fill-in-the-blank trick that remasks erroneous positions: the essential change is that visible, committed tokens can also be replaced. The complete candidate supplies provisional semantic evidence from uncommitted positions, helping the correction model reconsider early commitments. Restricted writeback prevents all low-confidence provisional predictions from being fixed prematurely. Corrections occur in the first half of sampling because sparse early context is more error-prone; experiments also show that denser or later revision is not uniformly better.
T2M uses CFG with a default scale of 4; M2T does not use CFG. The paper expresses guidance as a linear extrapolation of conditional and unconditional probabilities. With a scale above 1, the unconditional term has a negative coefficient, but nonnegativity handling and normalization are not explicitly specified. This note therefore preserves the conditional/unconditional extrapolation and confidence-based sampling mechanism without inventing a valid probability formula or treating logits-based guidance as a confirmed implementation detail.
Twenty steps means ordinary sampling iterations, not exactly 20 Transformer executions. T2M CFG requires conditional and unconditional predictions, and the two correction triggers add forward passes. Whether these calls are batched is an implementation detail. The reported 0.67-second T2M latency measures the complete default setting, not a bare 20-forward-pass configuration excluding CFG or correction overhead.
A Worked Example¶
Consider the illustrative prompt โa person walks forward, then turns left,โ with an externally specified motion length of 120 frames. Text remains visible, and the target starts as 30 masked motion tokens. The model begins accepting confident predictions in parallel. Suppose an accepted token before step 5 resembles a right turn. This hypothetical error illustrates the evolving state and is not a documented example from the paper.
At step 5, current predictions temporarily complete the remaining masks into a 30-token candidate, which is fed back into the Transformer together with the original text. Correction can replace the accepted right-turn-related token with a prediction more consistent with turning left, while unresolved positions remain uncommitted. Step 10 checks early commitments again, and only after all 20 steps are the final motion tokens passed to the VQ-VAE decoder.
For M2T, the same model reads visible motion tokens, generates text, and truncates at [EOS]. Correction acts on accepted text tokens rather than changing the input motion. A shared bidirectional generator therefore does not mean that inference simultaneously rewrites both condition and target.
Loss & Training¶
One expression summarizes ordinary masked supervision in PT and SFT: the prediction region is the entire sequence in PT and the target region in SFT, with only masked positions contributing to this loss. The following juxtaposes ordinary supervision with all-region correction, retaining their essential difference in supervision scope:
This consolidates the notation of original equations (1), (2), (4), and (5), rather than introducing another objective. \(t\) is sampled uniformly from \((0,1]\) and denotes the token masking probability; \(\mathcal{G}\) is the stage-specific generation region, and \(\tilde{\mathbf{x}}\) is the detached self-generated input. The ordinary term supervises only masked positions, the correction term removes the mask indicator, and \(1/t\) retains the original mask-ratio reweighting.
The model has 20 layers, hidden size 1024, and 334M parameters. PT runs for 100 epochs at learning rate 2e-4; SFT runs for 200 epochs at 8e-5, both using AdamW. The correction weight is \(\lambda=1\), and experiments use one NVIDIA H100. The appendix compares pretrained BERT backbones with BiMoGen trained from scratch, which does not support an explanation requiring pretrained language weights.
Key Experimental Results¶
Main Results¶
HumanML3D contains 14,616 motions and 44,970 descriptions; KIT-ML contains 3,911 motions and 6,278 descriptions. R@k is the fraction of paired items appearing among the top k candidates in the shared evaluator space. FID measures generated-versus-real motion feature distributions, and MM Dist measures average distance between paired text and motion features. CIDEr measures agreement with reference captions, while BERTScore measures embedding-based semantic similarity. The table preserves the original reporting scales.
The following selects unified models from original Tables 1 and 2, rather than mixing single-task or separate-model systems into claims about the best unified model.
| Dataset | Model | T2M R@1 โ | T2M FID โ | T2M MM Dist โ | M2T BLEU@4 โ | M2T CIDEr โ | M2T BERTScore โ |
|---|---|---|---|---|---|---|---|
| HumanML3D | MotionGPT | 0.492 | 0.232 | 3.096 | 12.5 | 29.2 | 32.4 |
| HumanML3D | MotionGPT3 | 0.553 | 0.208 | 2.725 | 19.4 | 28.7 | 35.2 |
| HumanML3D | DiMo | 0.528 | 0.047 | 2.862 | 22.7 | 58.1 | 37.7 |
| HumanML3D | BiMoGen | 0.555 | 0.069 | 2.733 | 20.1 | 60.2 | 38.5 |
| KIT-ML | MoTe | 0.419 | 0.256 | 3.216 | 14.1 | 55.6 | 35.9 |
| KIT-ML | DiMo | 0.406 | 0.206 | 2.983 | 17.8 | 68.7 | 37.7 |
| KIT-ML | BiMoGen | 0.421 | 0.222 | 2.763 | 17.3 | 62.2 | 40.4 |
On HumanML3D, BiMoGen leads in T2M R@1 and M2T CIDEr and BERTScore, but DiMo has better FID and BLEU@4, and MotionGPT3 has slightly lower MM Dist. On KIT-ML, the principal advantages are MM Dist and BERTScore, not every retrieval or captioning metric. DiMo uses an RVQ-VAE and an additional residual Transformer, whereas BiMoGen uses a single-layer VQ-VAE. Their FID comparison is not a controlled experiment changing only self-correction.
Ablation Study¶
Original Table 3 compares training strategies on HumanML3D. SC denotes the training self-correction objective, not merely enabling or disabling correction at sampling time.
| Config | T2M R@1 โ | T2M FID โ | T2M MM Dist โ | M2T R@1 โ | M2T CIDEr โ | M2T BERTScore โ |
|---|---|---|---|---|---|---|
| PT only | 0.416 | 1.219 | 3.650 | 0.434 | 6.7 | 16.8 |
| SFT only | 0.466 | 0.325 | 3.259 | 0.474 | 41.3 | 26.4 |
| PT + SFT | 0.514 | 0.102 | 3.017 | 0.547 | 44.0 | 36.7 |
| PT + SFT + SC | 0.555 | 0.069 | 2.733 | 0.573 | 60.2 | 38.5 |
Original Table 5 keeps the trained model and ordinary sampling-step count fixed, changing only correction timing during sampling. T2M uses default CFG, while M2T does not use CFG.
| Sampling correction schedule | T2M R@1 โ | T2M FID โ | T2M MM Dist โ | T2M latency (seconds) โ | M2T CIDEr โ | M2T BERTScore โ |
|---|---|---|---|---|---|---|
| Sampling correction disabled | 0.551 | 0.070 | 2.745 | 0.55 | 58.8 | 39.6 |
| Step 15 only | 0.542 | 0.056 | 2.833 | 0.61 | 59.6 | 38.8 |
| Every step | 0.529 | 0.066 | 2.948 | 0.78 | 61.2 | 41.0 |
| Steps 5 and 10 (default) | 0.555 | 0.069 | 2.733 | 0.67 | 60.2 | 38.5 |
Key Findings¶
- PT and SFT have distinct roles: adding PT to SFT reduces MM Dist from 3.259 to 3.017; training SC further reduces it to 2.733 and increases CIDEr from 44.0 to 60.2. Training correction produces a much larger improvement than default sampling correction relative to disabling sampling correction.
- More frequent revision is not uniformly better. Every-step correction reaches M2T CIDEr of 61.2 but reduces T2M R@1 to 0.529 and increases latency to 0.78 seconds. Default BERTScore of 38.5 is also below the 39.6 obtained with sampling correction disabled; this trade-off should be preserved.
- Table 6 defines modification rate as the fraction of committed tokens revised by correction and stability as the fraction of revised tokens retaining that prediction at the next step. At steps 5 and 10, these are 31.82%/84.70% and 29.48%/83.72%, respectively. Stability does not establish correctness against ground truth.
- Table 4 reports FID of 0.118/0.069/0.071, latency of 0.20/0.67/0.95 seconds, and FLOPs of 0.34T/1.08T/1.57T for 5/20/30 steps. Thirty steps do not uniformly outperform 20, so additional iterations should not be described as monotonically improving quality.
- Original Table 1 reports DiMo FID of 0.047 and M2T BERTScore of 37.7, whereas Table 4 reports 0.050 and 35.4 for DiMo at 20 steps. The main table above retains Table 1 values; baseline rows from different tables should not be merged into one experimental setting.
Highlights & Insights¶
- Correspondence learning and conditional generation are different objectives. Symmetric corruption followed by target-only corruption provides a clear curriculum-like transition for a unified model without requiring another model.
- Self-correction makes committed tokens revisable again. The transferable mechanism is all-region supervision on self-generated inputs with restricted writeback, rather than simply adding more fill-in-the-blank iterations.
- Correction can improve semantic alignment but is not a physical-constraint verifier. Late correction improves FID while weakening retrieval in Table 5, showing that motion distribution quality and text fidelity are not interchangeable.
Limitations & Future Work¶
- The authors note that VQ-VAE reconstruction fidelity and codebook expressiveness still limit subtle motions, while the sequence-level representation lacks explicit independent control of hands, arms, and legs. Stronger part-aware representations merit investigation, but the paper's Part VQ-VAE comparison degrades performance; replacing the tokenizer cannot be assumed to improve results.
- T2M requires a given motion length. Two standard datasets do not establish open-domain generalization to long sequences, complex hand interactions, or scene constraints. Motion in-betweening is primarily demonstrated qualitatively; the user study uses 30 prompts and 15 participants.
- Probability-based CFG extrapolation lacks explicit nonnegativity and normalization details, and correction's additional forward-pass cost should be measured separately. Reproduction should verify these implementation boundaries rather than hide them with reconstructed formulas.
Related Work & Insights¶
- vs MotionGPT / MotionGPT3: both pursue unified motion-language modeling, but BiMoGen replaces fixed causal ordering with bidirectional masked prediction and allows early commitments to be revised. Differences in motion decoding heads still affect cross-model quality comparisons.
- vs DiMo: both use bidirectional discrete diffusion. DiMo combines RVQ prediction and reinforcement-learning fine-tuning, whereas BiMoGen emphasizes two-stage training and generation-aware correction. BiMoGen is not the first application of discrete diffusion to unified motion-text tasks.
- vs MoMask / MaskGIT: BiMoGen inherits parallel masked prediction and confidence-based selection, adding bidirectional motion-text tasks and all-region correction supervision on self-generated context. It does not merely supply text as an external condition for motion generation.
Rating¶
- Novelty: 4/5. Combining shared-model correction during training and sampling addresses a specific weakness, but concurrent unified discrete diffusion already exists.
- Experimental Thoroughness: 4/5. Two datasets, training and sampling ablations, computational analysis, and a user study provide broad coverage, without a controlled DiMo comparison using the same tokenizer.
- Writing Quality: 4/5. Generation regions and writeback scope are clear, but CFG formulation and cross-table baseline settings require verification during reproduction.
- Value: 4/5. A practical correction mechanism for unified human motion generation and understanding, with fine-grained control and automatic length modeling unresolved.