RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://hidream-ai.github.io/RASA/
Area: Human Understanding
Keywords: Cross-Identity Character Animation, Diffusion Transformer, Spatial Pose Alignment, 3D Motion Priors, Disentangled Representations
TL;DR¶
Addressing severe structural mismatches in cross-identity character animation, RASA introduces a hierarchical framework that injects disentangled spatial-motional priors into a Diffusion Transformer, utilizing a Spatial Prior Calibrator (SPC) at the generative onset for 2D geometric alignment and an Inherent Motional Guider (IMG) in intermediate layers for 3D anatomical consistency.
Background & Motivation¶
Driven by advances in Latent Diffusion Models (LDMs) and Diffusion Transformers (DiTs), character image animation has emerged as a foundational technology for digital avatar synthesis and interactive media production. However, existing animation pipelines encounter significant performance bottlenecks when applied to real-world cross-identity driving scenarios. In practical settings, the driving subject and the reference character typically exhibit pronounced discrepancies in global spatial position, bounding scale, and skeletal limb proportionsโa fundamental phenomenon termed poseโreference misalignment. Prevailing diffusion frameworks implicitly assume geometric congruence between the driving skeleton and the reference image; consequently, structural discrepancies trigger severe visual distortions, unnatural limb warping, and catastrophic identity corruption.
Conventional remedies primarily rely on explicit pose retargeting preprocessing, heuristic pose augmentation, or late-stage multi-modal feature fusion in deep denoising layers. Nonetheless, preprocessing retargeting strategies often discard subtle motion dynamics and peripheral trajectories, while late-stage feature fusion inadvertently entangles structural calibration with appearance texture rendering, causing irreversible accumulation of visual artifacts. The core tension lies in the fact that cross-identity animation fundamentally demands two orthogonal capabilities: resolving 2D spatial discrepancies (scale, position, and skeletal proportions) and providing viewpoint-consistent, physically plausible motion control. Conflating these two tasks prevents models from achieving robust spatial alignment while preserving fine-grained dynamic fidelity.
RASA resolves this dilemma by systematically disentangling spatial geometric mapping from fine-grained motion control: neutralizing geometric discrepancies at the generative onset while continuously enforcing shape-agnostic 3D articulation dynamics within deeper network layers. Core idea: propose RASA, a cross-identity character animation framework that systematically disentangles spatial and motional priors, deploying a Spatial Prior Calibrator (SPC) at the generative onset to resolve 2D spatial mismatches and an Inherent Motional Guider (IMG) across intermediate DiT blocks to inject shape-agnostic SMPL articulation semantics for anatomically consistent animation.
Method¶
Overall Architecture¶
RASA builds upon the open-source Wan2.1 (1.3B) Diffusion Transformer (DiT) backbone, utilizing Flow Matching to model the generative velocity field between noise and clean latent distributions. The overall inputs comprise a reference character image \(I_r\) with its 2D pose \(P_r\), a driving pose sequence \(P_d\) spanning \(T\) frames, and corresponding 3D articulation parameters \(\theta\). During training, stochastic spatial transformations are applied to simulate extreme cross-identity structural mismatches. The generation process is orchestrated via a dual-stream injection paradigm: the Spatial Prior Calibrator (SPC) takes the driving and reference poses and performs frame-wise 2D RoPE cross-attention to produce an aligned spatial latent that is element-wise added to the noisy video latent at the start of each denoising step; concurrently, the Inherent Motional Guider (IMG) extracts shape-agnostic articulation vectors and hierarchically injects high-level motion semantics into intermediate DiT blocks (Blocks 2 through 15).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Inputs: Reference Image $I_r$ & Driving Pose $P_d$"] --> B["Misalignment Simulation & Feature Extraction<br/>Random Spatial Perturbation + Shared Encoder"]
B --> C["Spatial Prior Calibrator (SPC)<br/>Driving-query 2D RoPE Cross-Attention"]
C --> D["Onset Geometric Rectification<br/>Element-wise addition to initial noisy latents"]
B --> E["Inherent Motional Guider (IMG)<br/>Shape-agnostic SMPL articulation vector"]
E --> F["Intermediate High-Level Motion Injection<br/>Multi-depth projection into DiT Blocks 2-15"]
D --> G["DiT Backbone Velocity Prediction & ODE Solver"]
F --> G
G --> H["Output: Structurally Consistent & Identity-Preserved Video"]
Key Designs¶
1. Spatial Prior Calibrator (SPC): Resolving 2D Geometric Misalignment at Generative Onset
Prior feature-fusion approaches entangle structural guidance with texture generation, leading to skeletal deformation. SPC reframes alignment as a learnable structural calibration task in latent space prior to DiT processing. Given perturbed driving poses \(\tilde{P}_d\) and the reference pose \(P_r\), a shared lightweight encoder extracts a static reference latent \(z_r \in \mathbb{R}^{C \times H \times W}\) and a temporal driving latent \(z_d \in \mathbb{R}^{T \times C \times H \times W}\). To capture non-local spatial correspondences across disparate skeletons, SPC establishes a frame-wise spatial cross-attention mechanism equipped with 2D Rotary Positional Encoding (RoPE) to safeguard 2D topological integrity:
where \(Q_t = z_{d,t}\) is flattened into spatial sequence tokens, and \(K = z_r, V = z_r\). Crucially, using the dynamic driving features as queries allows the model to actively retrieve and adapt the geometric scale of the reference character; swapping query and key roles suppresses dynamic motion cues. The calibrated spatial tensor \(\hat{z}_d\) is element-wise added to the noisy latent at the beginning of every reverse diffusion step, anchoring skeletal proportions and global positions from the generative onset and preventing error accumulation.
2. Inherent Motional Guider (IMG): Intermediate Layer Anatomical and Volumetric Guidance
While SPC resolves 2D planar discrepancies, 2D representations suffer from projection ambiguity and lack volumetric consistency during complex joint rotations. IMG introduces an explicit 3D parametric motion prior decoupled from subject appearance. Using ScoreHMR, 3D mesh parameters are estimated from the driving sequence. To guarantee absolute geometry invariance, subject-specific shape parameters (\(\beta\)) are discarded, preserving exclusively the articulation parameters (\(\theta\)). The dynamic sequence is formulated into a motion vector \(m \in \mathbb{R}^{T \times 263}\), encoding root-relative joint rotations, velocities, and foot-contact flags, which is then mapped to latent sequence \(z_{\text{3Dm}} \in \mathbb{R}^{T \times D}\) via a pretrained motion encoder.
Because 3D motion dynamics represent high-level semantic signals, RASA deploys \(L = 14\) learnable linear projection layers to inject \(z_{\text{3Dm}}\) into the intermediate blocks (specifically Blocks 2 through 15) of the DiT backbone. This multi-depth conditioning leaves initial spatial layouts to SPC while providing continuous, viewpoint-consistent anatomical refinement throughout latent feature evolution, ensuring physical plausibility under extreme cross-identity variations.
3. Reference Identity Conditioning & CIM-Bench Benchmark
To isolate invariant identity from temporal dynamics, the reference image latent \(z_r\) is concatenated with the video latents along the temporal axis with a large Positional Encoding Offset, enabling global self-attention to attend to fine textures without conflating static identity with dynamic motion flows. Furthermore, recognizing that existing synthetic perturbation datasets fail to capture real cross-identity discrepancies, the authors introduce CIM-Bench. By pairing diverse stylized characters generated by modern T2I models with real human portraits, and executing identical driving motions via commercial video synthesis backbones, CIM-Bench delivers 656 rigorously curated, high-quality test pairs featuring authentic skeletal and proportional discrepancies.
Loss & Training¶
The DiT backbone is optimized end-to-end using the standard Flow Matching conditional velocity objective:
where condition vector \(c\) integrates text prompts, reference latent \(z_r\), calibrated spatial features \(\hat{z}_d\), and 3D motion embeddings \(z_{\text{3Dm}}\). The system is trained on 8 NVIDIA A100 GPUs using the AdamW optimizer with a learning rate of \(1 \times 10^{-5}\). Videos are uniformly sampled into 41-frame clips at \(832 \times 480\) resolution, with total training completing in approximately 48 hours.
Key Experimental Results¶
Main Results¶
Quantitative comparisons on the self-driven TikTok benchmark and the cross-identity CIM-Bench demonstrate that RASA (1.3B) consistently outperforms existing state-of-the-art models, showing superior robustness under severe structural misalignment.
| Dataset | Method | PSNR โ | SSIM โ | LPIPS โ | FID โ | FVD โ | FID-VID โ | Sim-Arc โ |
|---|---|---|---|---|---|---|---|---|
| TikTok | MTVCrafter (5B) | 19.37 | 0.784 | 0.217 | 19.46 | 140.60 | 6.98 | - |
| TikTok | One-to-All (14B) | 18.07 | 0.812 | 0.254 | 50.49 | 297.94 | 13.93 | - |
| TikTok | RASA (Ours, 1.3B) | 22.03 | 0.830 | 0.189 | 18.93 | 190.82 | 3.32 | - |
| CIM-Bench | Animate-X | 14.05 | 0.426 | 0.474 | 105.14 | 1312.60 | 75.47 | 0.51 |
| CIM-Bench | MTVCrafter | 14.19 | 0.431 | 0.457 | 85.21 | 1723.89 | 85.21 | 0.54 |
| CIM-Bench | Wan-Animate | 11.29 | 0.372 | 0.585 | 178.58 | 1920.88 | 137.05 | 0.35 |
| CIM-Bench | One-to-All (1.3B) | 13.75 | 0.383 | 0.471 | 96.70 | 1324.74 | 76.79 | 0.51 |
| CIM-Bench | RASA (Ours, 1.3B) | 15.93 | 0.472 | 0.372 | 60.96 | 667.96 | 17.63 | 0.72 |
Ablation Study¶
Component-wise ablation experiments across TikTok and CIM-Bench validate the complementary necessity of both SPC and IMG:
| Config | PSNR โ (TikTok/CIM) | SSIM โ (TikTok/CIM) | LPIPS โ (TikTok/CIM) | FID โ (TikTok/CIM) | FVD โ (TikTok/CIM) | FID-VID โ (TikTok/CIM) | Note |
|---|---|---|---|---|---|---|---|
| SMS only (random pose perturbations) | 18.90 / 14.60 | 0.799 / 0.453 | 0.209 / 0.393 | 23.04 / 80.96 | 316.45 / 737.48 | 10.77 / 30.93 | Lacks explicit 2D calibration; poor on CIM-Bench |
| IMG only (3D motion guider only) | 15.47 / 13.89 | 0.707 / 0.416 | 0.356 / 0.445 | 41.57 / 91.17 | 1567.03 / 1625.69 | 15.90 / 37.46 | Lacks 2D geometric grounding; visual quality collapses |
| SMS + IMG | 18.51 / 14.14 | 0.811 / 0.470 | 0.238 / 0.399 | 28.83 / 77.34 | 487.66 / 706.22 | 22.06 / 34.51 | No frame-wise cross-attention; structural drift persists |
| SMS + FCA (Full SPC) | 20.86 / 15.26 | 0.820 / 0.471 | 0.202 / 0.391 | 22.01 / 71.16 | 289.16 / 686.79 | 6.56 / 27.05 | Robust 2D alignment, but lacks 3D temporal dynamics |
| Full Model (SPC + IMG) | 22.03 / 15.93 | 0.830 / 0.472 | 0.189 / 0.372 | 18.93 / 60.96 | 190.82 / 667.96 | 3.32 / 17.63 | Disentangled synergy yields superior across-the-board results |
Key Findings¶
- SPC resolves spatial deformation while IMG drives temporal motion stability: Adding SPC (SMS+FCA) elevates TikTok PSNR from 18.90 to 20.86. Layering IMG onto SPC triggers an unprecedented drop in FVD (from 289.16 to 190.82 on TikTok, and from 686.79 to 667.96 on CIM-Bench), with FID-VID halving from 6.56 to 3.32, demonstrating that 3D joint articulation priors are essential for preventing temporal flickering and limb collapse.
- Minimal parameter and computational overhead: SPC and IMG collectively add only 19.78M parameters (< 1.52% of the 1.3B DiT backbone) and 90.46G FLOPs. On a single NVIDIA H100 GPU, generating a 41-frame video at \(832 \times 480\) resolution requires 61 seconds (only a 2-second increase over the 59-second baseline) with 11.95 GB VRAM usage, outperforming Wan-Animate (42.29 GB VRAM, 327 seconds) by a large margin.
Highlights & Insights¶
- Decoupled hierarchical injection paradigm: Grounding 2D skeletal scale at the generative onset while infusing 3D kinematic dynamics into intermediate DiT blocks effectively resolves the long-standing coupling between geometric adaptation and appearance texture synthesis.
- Shape-agnostic motion descriptor: Discarding SMPL shape parameters \(\beta\) and preserving solely articulation features (\(m \in \mathbb{R}^{T \times 263}\)) yields an intrinsically identity-agnostic motion prior suitable for diverse body proportions.
- Driving-as-query structural retrieval: Utilizing driving poses as queries in cross-attention allows dynamic motions to query and adopt the static reference character's skeletal proportions without losing motion vitality.
Limitations & Future Work¶
- Strict 1:1 humanoid topology assumption: RASA presumes topological correspondence between source and target bodies, failing on non-humanoid entities with disparate topologies (e.g., multi-limbed, quadruped, or limb-less creatures).
- Reliance on 2D poses for hands and facial expressions: Because SMPL omits detailed hand articulations and expressive facial dynamics, fine facial expressions and finger movements remain constrained by 2D DWPose detections.
- Future directions: The authors plan to integrate whole-body parametric models (SMPL-X / MANO) and extend the framework toward general, non-humanoid topology-aware animation.
Related Work & Insights¶
- vs Champ / UniAnimate: Champ aggregates five disparate motion inputs (depth, normal, skeletons), incurring excessive complexity; UniAnimate relies on crude bounding-box normalization. RASA uses only 2D poses and compact motion vectors with latent cross-attention, offering superior flexibility.
- vs Animate-X / One-to-All: Animate-X degrades significantly under large structural discrepancies; One-to-All exhibits substantial identity drift due to unconstrained generation. RASA achieves state-of-the-art identity retention (0.72 Sim-Arc on CIM-Bench) alongside precise motion fidelity.
Rating¶
- Novelty: โญโญโญโญโ Elegant decoupling of 2D spatial calibration and 3D motion dynamics within DiTs.
- Experimental Thoroughness: โญโญโญโญโญ Rigorous evaluation across TikTok and the newly introduced CIM-Bench with comprehensive ablations.
- Writing Quality: โญโญโญโญโญ Well-structured narrative with crisp technical explanations and clear motivation.
- Value: โญโญโญโญโญ High practical relevance and engineering value for virtual character animation and digital human synthesis.