K-Mask: Kinematic-Aware Masked Modeling for Controllable Text-to-Motion Synthesis¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Human Understanding
Keywords: text-to-motion synthesis, discrete tokenization, kinematic grouping, masked modeling, localized motion control
TL;DR¶
K-Mask introduces a kinematic-aware motion generation framework that couples an anatomically grounded residual VQ-VAE (KG-RVQ) with a latent-aware kinematic dropout (LAKD) loss and spatiotemporal axis-factorized masked transformers, achieving competitive generation fidelity and precise limb-level localized control without separate limb-specific encoders.
Background & Motivation¶
Text-to-motion generation aims to synthesize physically plausible and temporally coherent 3D human movements aligned with natural language instructions, serving as a core foundation for character animation, virtual reality, filmmaking, and human-robot interaction. While recent continuous diffusion models and discrete masked generative models have drastically pushed the boundary of generation realism and text-motion alignment, real-world interactive creation demands fine-grained, localized control—such as modifying only the right arm to wave while preserving the underlying walking trajectory and lower-body locomotion gait.
Existing generative paradigms face an acute trade-off between control granularity and architectural complexity. Continuous diffusion-based frameworks offer rich sample diversity but incur high inference costs, and localized edits typically require complex sampling-time attention guidance or per-sample latent optimization. Conversely, discrete masked frameworks enable rapid parallel sampling and native temporal inpainting; however, standard models rely on whole-body, part-agnostic tokenization that collapses all joints into a single token stream, causing collateral corruption and unnatural trajectory drifts to non-target limbs during editing. Previous part-aware attempts either deploy separate local generator branches and fusion modules, heavily inflating parameter counts and coordination overhead, or fragment the representation to individual joint levels, which dilutes codebook capacity and imposes prohibitive spatiotemporal attention complexity.
The key insight of this paper is that rigorous anatomical disentanglement does not require isolated limb subnetworks. By partitioning the latent channel space of a unified full-body tokenizer into kinematic groups and imposing an explicit anti-leakage regularization objective during training, a single masked generator can directly address and manipulate an anatomically structured token lattice. Core idea: integrate kinematic-group residual vector quantization (KG-RVQ) with a latent-aware kinematic dropout (LAKD) loss to suppress inter-group leakage, combined with axis-factorized spatiotemporal transformers and structured kinematic masking for decoupled yet globally coordinated motion generation and editing.
Method¶
Overall Architecture¶
The K-Mask framework operates over a structured time-group spatiotemporal lattice \((t, g)\). It comprises three sequential components: First, a unified Kinematic-Group Residual VQ-VAE (KG-RVQ) encodes the motion sequence into 6 anatomically grounded sub-streams with proportional channel allocation and independent RVQ codebooks, trained under the LAKD objective to prevent cross-limb latent entanglement; second, a Base Transformer (Base-T) uses axis-factorized temporal-then-spatial attention to predict the coarse base-layer motion tokens (\(q=0\)) under a hybrid masking strategy; third, a Residual Transformer (Res-T) conditions on preceding dequantized history latents to autoregressively synthesize the remaining residual levels (\(q>0\)), which are subsequently decoded back into 3D continuous motion.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Motion Sequence & Text Prompt"] --> B["Kinematic-Group Residual Quantization & Allocation<br/>Full-body encoding into 6 sub-streams with independent RVQ"]
B --> C["Latent-Aware Kinematic Dropout LAKD<br/>Within-group reconstruction & inter-group leakage suppression"]
C --> D["Axis-Factorized Spatiotemporal Attention<br/>Temporal-then-spatial self-attention for reduced complexity"]
D --> E["Hybrid Structured Kinematic Group Masking<br/>Joint training with random token and body subtree occlusion"]
E --> F["Two-Stage Synthesis & Localized Editing<br/>Freeze non-target groups and resample target limb tokens"]
Key Designs¶
1. Kinematic-Group Residual Quantization & Allocation: Anatomical Biomechanics with Adaptive Capacity Distribution
Prior part-aware systems frequently employ independent encoders for distinct limbs, sacrificing the unified receptive field of whole-body dynamics. Conversely, uniformly splitting a shared latent representation ignores the fact that different body parts possess vastly disparate kinematic degrees of freedom and complexity. K-Mask processes the full-body motion sequence with a shared 1D temporal convolutional encoder to obtain a unified latent representation \(Z_t \in \mathbb{R}^d\), which is explicitly partitioned along the channel dimension into \(G=6\) kinematic groups (root, spine, left leg, right leg, left arm, right arm): \(Z_t = [Z_t^{(1)} \parallel \cdots \parallel Z_t^{(G)}]\). Drawing inspiration from perceptual coding bit-allocation principles, the model abandons uniform splits (\(d_g = d/G\)) and dynamically assigns latent capacity proportional to the input feature dimension \(D_g\) of each group:
Each group \(g\) is assigned an independent stack of \(Q\) residual codebooks \(\mathcal{C}_{g,q} = \{e_{g,q,k}\}_{k=1}^K\). The resulting quantized codes are concatenated and passed through a shared convolutional decoder, organizing motion representations into a structured tensor \(I \in \{1, \dots, K\}^{T' \times G \times Q}\).
2. Latent-Aware Kinematic Dropout (LAKD): Eliminating Cross-Group Information Leakage
Merely slicing the channel dimension does not guarantee semantic disentanglement: shared parameters in the encoder and decoder can easily exploit non-target channels to store auxiliary whole-body context, leading to noticeable motion spillovers when editing an individual limb. To enforce true anatomical independence, K-Mask introduces the Latent-Aware Kinematic Dropout (LAKD) loss. During training, a target group \(g\) is randomly sampled to construct two complementary representations using a binary channel mask: an "only group \(g\)" latent \(\hat{Z}_{\text{only } g}\) and an "except group \(g\)" latent \(\hat{Z}_{\text{except } g}\). The LAKD loss combines positive fidelity with bidirectional negative leakage suppression:
Here, the positive term \(\mathcal{L}_{\text{pos}}(g) = \|P_g(g_\phi(\hat{Z}_{\text{only } g})) - P_g(X)\|_1\) requires the group's dedicated latent to be completely self-sufficient for reconstructing its own motion; the in-group suppression term \(\mathcal{L}_{\text{neg-in}}(g) = \|P_g(g_\phi(\hat{Z}_{\text{except } g}))\|_1\) penalizes latents from other groups from predicting motion in group \(g\); and the out-group suppression term \(\mathcal{L}_{\text{neg-out}}(g) = \|P_{\bar{g}}(g_\phi(\hat{Z}_{\text{only } g}))\|_1\) penalizes group \(g\)'s latents from contributing to any non-target parts \(\bar{g}\). Incorporating LAKD cuts the cross-group leakage ratio from 2.484 down to 0.088, providing clean isolation for downstream editing.
3. Axis-Factorized Spatiotemporal Attention: Balancing Cross-Limb Coordination with Linear Temporal Scalability
Flattening the \(T' \times G\) discrete token grid into a standard self-attention module incurs a prohibitive quadratic complexity of \(\mathcal{O}((T'G)^2)\), which degrades training efficiency and dilutes limb-specific temporal nuances. K-Mask implements a factorized temporal-then-spatial (T\(\to\)S) attention architecture across its 8-layer transformer blocks. The token input combines shared token embeddings, temporal positional encodings \(P_t\), learnable group embeddings \(P_g\), and projected text condition embeddings from CLIP. Each block first processes individual group dynamics independently along the temporal axis, followed by spatial attention across the \(G\) group tokens at each timestep. This factorized structure scales as \(\mathcal{O}(T'^2 G + T' G^2)\) and empirically outperforms both spatial-then-temporal (S\(\to\)T) factorization and standard full self-attention.
4. Hybrid Structured Kinematic Group Masking: Training Inter-Group Dynamics for Native Inpainting
To empower the base generator with robust whole-body coordination during localized edits, Base-T is trained using a hybrid masking strategy. For each sequence, with probability \(p_{\text{kin}} = 0.25\), the model applies structured kinematic group masking, fully occluding entire anatomical subtrees across the entire temporal window (e.g., masking symmetric limb pairs such as both legs, upper or lower body segments, or arbitrary random subsets). This compels the model to infer missing limb dynamics purely from unmasked limbs and text semantics. With probability \(1 - p_{\text{kin}} = 0.75\), the model executes standard random token masking sampled from a cosine schedule. During localized inference, non-target tokens are strictly frozen while target tokens are iteratively resampled; users can also smoothly expand the edit scope to the spine or pelvis to accommodate biomechanical compensation.
Loss & Training¶
The framework is optimized in two disjoint stages: 1. Tokenizer Optimization: The KG-RVQ objective integrates a smooth-\(\ell_1\) reconstruction loss \(\mathcal{L}_{\text{rec}}\), a commitment loss \(\mathcal{L}_{\text{commit}}\), and the expected LAKD loss with weights \(\lambda_{\text{com}}=0.02\) and \(\lambda_{\text{LAKD}}=2.0\). It is trained for 50 epochs using AdamW with cosine learning rate decay starting from \(2 \times 10^{-4}\). 2. Transformer Optimization: Base-T optimizes cross-entropy loss over masked base tokens with classifier-free guidance text dropping; Res-T randomly samples residual target level \(q\) and injects Gaussian noise into dequantized history latents to mitigate exposure bias, training for 2,000 epochs (batch size = 64). Base-T utilizes 18 refinement steps during inference.
Key Experimental Results¶
Main Results¶
On the standard HumanML3D and KIT-ML benchmarks, K-Mask achieves competitive generation fidelity and text-motion retrieval performance across all metrics against diffusion and discrete baselines.
| Dataset | Method | FID ↓ | R-Precision (Top-1) ↑ | R-Precision (Top-3) ↑ | Multimodal Dist ↓ | Diversity → |
|---|---|---|---|---|---|---|
| HumanML3D | T2M-GPT [56] | 0.141±.005 | 0.492±.003 | 0.775±.002 | 3.121±.009 | 9.761±.081 |
| HumanML3D | MotionDiffuse [57] | 0.630±.001 | 0.491±.001 | 0.782±.001 | 3.113±.001 | 9.410±.049 |
| HumanML3D | MoMask [16] | 0.045±.002 | 0.521±.002 | 0.807±.002 | 2.958±.008 | - |
| HumanML3D | BAMM [35] | 0.055±.002 | 0.525±.002 | 0.814±.003 | 2.919±.008 | 9.717±.089 |
| HumanML3D | ParCo [62] | 0.109±.005 | 0.515±.003 | 0.801±.002 | 2.927±.008 | 9.576±.088 |
| HumanML3D | K-Mask (Ours) | 0.041±.003 | 0.531±.003 | 0.824±.002 | 2.917±.006 | 9.671±.074 |
| KIT-ML | T2M-GPT [56] | 0.514±.029 | 0.416±.006 | 0.745±.006 | 3.007±.023 | 10.86±.094 |
| KIT-ML | MoMask [16] | 0.204±.011 | 0.433±.007 | 0.781±.005 | 2.779±.022 | - |
| KIT-ML | BAMM [35] | 0.183±.013 | 0.438±.009 | 0.788±.005 | 2.723±.026 | 11.008±.094 |
| KIT-ML | ParCo [62] | 0.453±.027 | 0.430±.004 | 0.772±.006 | 2.820±.028 | 10.95±.094 |
| KIT-ML | K-Mask (Ours) | 0.154±.025 | 0.455±.006 | 0.793±.006 | 2.671±.020 | 10.993±.096 |
On the 200-case HumanML3D localized editing evaluation set across 8 representative action classes, K-Mask significantly outperforms prior models in target edit success while preserving surrounding motions with minimal drift:
| Method | FID ↓ | Target Success Rate (TSR ↑) | Retrieval R@3 ↑ | Non-target Spillover ↓ | Pelvis Drift ↓ | Foot Skate ↓ |
|---|---|---|---|---|---|---|
| MoMask [16] | 0.045 | 20.0 | 26.6 | 0.164 | 0.74 | 0.009 |
| BAMM [35] | 0.055 | 7.0 | 21.9 | 0.136 | 1.33 | 0.009 |
| ParCo [62] | 0.109 | 11.5 | 27.6 | 0.190 | 1.15 | 0.016 |
| MoGenTS [54] | 0.028 | 19.0 | 25.5 | 0.168 | 0.92 | 0.018 |
| K-Mask (Ours) | 0.041 | 22.5 | 28.4 | 0.108 | 0.41 | 0.002 |
Ablation Study¶
Systematic ablations on grouping granularity, capacity allocation, and LAKD regularization weighting under a controlled 30-epoch short-run schedule are summarized below:
| Setting | Group Count \(G\) | Allocation Strategy | \(\lambda_{\text{LAKD}}\) | Recon FIDrec ↓ | MPJPE (mm) ↓ | Cross-group Leakage Ratio ↓ |
|---|---|---|---|---|---|---|
| Whole-body | 1 | - | - | 0.033 | 30.6 | - |
| Coarse | 3 | Proportional | 2.0 | 0.010 | 17.9 | 0.071 |
| No LAKD | 6 | Proportional | 0.0 | 0.008 | 16.4 | 2.484 |
| Low \(\lambda\) | 6 | Proportional | 0.5 | 0.009 | 18.7 | 0.184 |
| High \(\lambda\) | 6 | Proportional | 4.0 | 0.015 | 19.5 | 0.104 |
| Uniform | 6 | Uniform | 2.0 | 0.010 | 18.5 | 0.106 |
| Joint-level | 22 | Proportional | 2.0 | 0.013 | 20.3 | 1.489 |
| K-Mask (Full) | 6 | Proportional | 2.0 | 0.010 | 17.2 | 0.088 |
Key Findings¶
- LAKD is an Essential Anti-Leakage Shield: Disabling LAKD yields marginally better unconstrained reconstruction (MPJPE 16.4mm vs 17.2mm) due to free cross-channel shortcuts, but cross-group leakage explodes from 0.088 to 2.484 (nearly a 30-fold increase). LAKD successfully eliminates inter-group interference with negligible reconstruction penalty.
- Anatomical Granularity Sweet Spot: Fragmenting tokens to joint-level (\(G=22\)) disperses codebook capacity, degrading MPJPE to 20.3mm and leakage to 1.489. In contrast, \(G=6\) strikes the ideal balance between localized addressability and kinematic continuity.
- Kinematic Masking Promotes Inter-Limb Dynamics: Increasing kinematic masking probability from \(p_{\text{kin}}=0\) to \(0.25\) boosts generation FID from 0.108 to 0.041, confirming that structured subtree occlusion forces the model to learn organic cross-limb coordination.
Highlights & Insights¶
- Unified Architecture with Decoupled Latents: Rather than training fragmented limb-specific encoders and complex fusion layers, K-Mask achieves rigorous anatomical disentanglement within a single compact network via channel allocation and LAKD suppression.
- True In-Place Editing: In strict target editing mode, non-target tokens are strictly frozen, reducing spillover displacement from 0.164 to 0.108 and reducing pelvis trajectory drift by more than 40% compared to MoMask.
- Flexible Biomechanical Scope: The tokenized kinematic lattice allows users to either perform strict localized edits or expand the sampling window to the spine and pelvis to absorb physical balance shifts.
Limitations & Future Work¶
- High-Dynamic Whole-Body Coupling: For intensely coupled acrobatic or martial arts maneuvers (e.g., spinning jumps, flips), strictly editing a single limb can violate global momentum conservation and introduce visual stiffness.
- Fixed Subtree Granularity: The current 6-group partition cannot address fine-grained finger articulation or facial cues; extending to hierarchical kinematic trees remains an open direction.
- Multi-Turn Temporal Sequencing: In multi-step narrative prompts with sharp temporal transitions, localized smoothing at boundary frames still requires global latent diffusion assistance.
Related Work & Insights¶
- vs MoMask [16] / BAMM [35]: Whole-body tokenizers treat localized edits as full temporal window inpainting, inevitably causing inadvertent motion drift in unrelated limbs; K-Mask introduces an anatomically addressable lattice that protects non-target body parts.
- vs ParCo [62]: ParCo deploys separate local generators and a dedicated fusion module, adding substantial architectural complexity; K-Mask retains a single generator with intrinsic kinematic factorization.
- vs MoGenTS [54]: MoGenTS partitions motion down to 22 joints, suffering from codebook capacity fragmentation and complex attention; K-Mask validates that 6 anatomically grounded groups provide superior trade-offs.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Elegant LAKD formulation solving cross-group leakage in a unified motion tokenizer.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous evaluations on two benchmarks, matched editing protocols, and extensive ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Well-structured narrative with crisp architectural diagrams and biomechanically grounded motivations.
- Value: ⭐⭐⭐⭐⭐ Sets a new practical benchmark for localized 3D character animation and controllable motion generation.