ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos¶
Conference: ECCV 2026
Paper: ECCV Official Page
Code: https://vision.cs.utexas.edu/projects/expert_edit/
Area: Human Understanding
Keywords: motion editing, motion generation, skill assessment, expert motion prior, masked sequence modeling
TL;DR¶
Addressing the lack of paired novice-expert data and explicit editing prompts, ExpertEdit learns an expert motion prior strictly from unpaired expert videos via masked motion infilling, selectively projecting novice executions at kinematic peaks onto the learned expert manifold to enable guidance-free skill refinement.
Background & Motivation¶
Motor skill acquisition is vital in competitive athletic training and physical rehabilitation. Psychological research on video self-modeling demonstrates that observing near-perfect versions of one's own execution accelerates motor learning and boosts retention far more effectively than merely watching expert demonstrations, as it engages the observer's perceptual-motor system directly. Consequently, the ability to "auto-tune" an amateur's performanceโinfusing expert-level form and cadence while preserving their unique spatial trajectory, viewpoint, and body identityโholds immense promise for AI coaching, stunt visual effects, and synthesizing expert-quality demonstrations for embodied robotics.
However, existing 3D human motion editing methods cannot be readily applied to skill refinement. Conventional motion editing frameworks rely on two restrictive assumptions: first, they require explicit edit conditioning such as text instructions or temporally synchronized reference clips, which breaks down in skill correction where the semantic action category remains identical and low-level kinematic instructions (e.g., "tuck the right elbow by 15 degrees") demand substantial domain expertise; second, supervised approaches depend on large-scale paired motion triplets ("before-edit, text prompt, after-edit"), which are virtually nonexistent for fine-grained athletic skills due to the prohibitive cost of expert diagnosis and frame-level alignment.
This paper tackles this problem from a fresh perspective: skill variations are concentrated in biomechanically critical phases of goal-directed actions (e.g., takeoff in a basketball layup or foot contact in a soccer kick), meaning that global motion reconstruction is unnecessary when one can locally project critical phases onto an expert motion manifold. Core Idea: cast skill-driven motion refinement as contextual motion infilling, learning an expert motion manifold exclusively from unpaired expert demonstrations via a masked language modeling (MLM) objective, and using kinematic peak criteria at inference time to automatically mask and infill novice action phases with expert-like refinements without requiring paired supervision or inference-time edit guidance.
Method¶
Overall Architecture¶
ExpertEdit takes as input an unaligned novice 3D motion sequence \(\mathbf{X}^{\text{nov}} = \{(\mathbf{r}_t, \mathbf{o}_t, \mathbf{p}_t)\}_{t=1}^{T}\), where \(\mathbf{r}_t \in \mathbb{R}^3\) denotes global root translation, \(\mathbf{o}_t \in \mathbb{R}^3\) denotes root orientation, and \(\mathbf{p}_t \in \mathbb{R}^{3J}\) represents the axis-angle rotations of \(J\) skeletal joints. The objective is to synthesize an edited sequence \(\mathbf{X}^{\text{edit}}\) of identical duration that exhibits expert-level execution quality. To preserve the performer's body shape, spatial path, and execution rhythm, the global translation \(\{\mathbf{r}_t\}_{t=1}^T\) and orientation \(\{\mathbf{o}_t\}_{t=1}^T\) are preserved, restricting modifications strictly to joint rotations \(\mathbf{p}_t\) during selected skill-critical action windows.
The framework operates in three tightly integrated stages:
1. Causal Pose Tokenization (Pose Tokenizer): Trains a causal Transformer-based VQ-VAE on expert motion sequences to discretize continuous kinematics into a compact vocabulary of skilled motion primitives.
2. Kinematic Phase Masking & Motion Infilling (MotionInfiller): Automatically identifies skill-critical moments using scalar kinematic statistics, applies symmetric temporal masks around peak frames, and trains a bidirectional Transformer with a masked language modeling (MLM) objective to learn expert motion priors.
3. Inference Mask-and-Infill & Pose Decoding: Extracts kinematic peaks from an input novice motion, replaces the critical window with [MASK] tokens, predicts expert-like discrete tokens via bidirectional infilling, and reconstructs refined joint rotations with the decoder while maintaining unedited surrounding poses.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Motion Sequence<br/>Extract joint rotations and root trajectory"] --> B["Causal Pose Tokenization<br/>Temporal causal Transformer VQ-VAE"]
B --> C["Kinematic Phase Masking<br/>Locate critical moments via velocity/acceleration"]
C --> D["Motion Infilling Training & Inference<br/>Bidirectional Transformer masked token infilling"]
D --> E["Pose Decoding & Local Blending<br/>Retain root trajectory, output edited motion"]
Key Designs¶
1. Causal Pose Tokenization: Constructing a Dynamically Consistent Expert Motion Vocabulary
Frame-wise quantization neglects temporal velocity and acceleration constraints, frequently leading to physical jitter and unnatural transitions in the decoded motion. ExpertEdit adopts a causal Transformer VQ-VAE as its Pose Tokenizer. For each frame, it constructs a joint feature vector combining root translation, root orientation, and joint rotations \(\mathbf{x}_t^{\text{exp}} = [\mathbf{r}_t^{\text{exp}}, \mathbf{o}_t^{\text{exp}}, \mathbf{p}_t^{\text{exp}}] \in \mathbb{R}^{6+3J}\). The encoder \(E_\phi\) enforces causal attention masking so that the latent state \(\mathbf{z}_t\) attends strictly to past and current frames: $\(\mathbf{z}_t = E_\phi(\mathbf{x}_{\leq t}^{\text{exp}})\)$ Each latent vector is quantized to its nearest neighbor in a learned discrete codebook \(\mathcal{C} = \{\mathbf{e}_k\}_{k=1}^K\) yielding \(\mathbf{e}_{k_t^*}\). A causal Transformer decoder \(D_\phi\) reconstructs each frame conditioned on past and present tokens \(\hat{\mathbf{x}}_t = D_\phi(\mathbf{e}_{\leq t})\). The tokenizer is trained end-to-end using the standard VQ-VAE objective: $\(\mathcal{L}_{\text{VQ}} = \|\mathbf{x}_t^{\text{exp}} - \hat{\mathbf{x}}_t\|_2^2 + \|\text{sg}[E_\phi(\mathbf{x}_{\leq t}^{\text{exp}})] - \mathbf{e}_{k_t^*}\|_2^2 + \beta \|E_\phi(\mathbf{x}_{\leq t}^{\text{exp}}) - \text{sg}[\mathbf{e}_{k_t^*}]\|_2^2\)$ This design ensures that each discrete token represents a coherent short-term motion primitive, establishing a geometrically consistent discrete manifold with reliable inverse decoding back to 3D joint space.
2. Kinematic Phase Masking: Automated Discovery of Skill-Critical Action Windows
In goal-directed athletic activities, execution proficiency is not uniformly distributed throughout the movement, but is concentrated around the operative moment that governs action successโsuch as the takeoff phase in a layup, the instant of ball strike in a penalty kick, or the point of maximal extension in a martial arts strike. To prevent arbitrary masking from corrupting uninformative context, ExpertEdit employs an interpretable scalar kinematic signal \(h(t)\) to locate the skill-critical peak index \(t^*\): $\(t^* = \arg\max_{t \in \{1,\dots,T\}} h(t)\)$ Specifically, basketball actions utilize vertical root velocity to capture takeoff; soccer penalty kicks employ lower-limb jerk magnitude to locate ball impact; karate reverse punches detect postural extremeness; and karate kicks track peak foot acceleration. Once \(t^*\) is identified, a masked span of length \(\ell = \max(2, \lfloor \alpha T \rfloor)\) with \(\alpha = 0.3\) is symmetrically centered around \(t^*\). Operating directly on raw motion dynamics, this mechanism bypasses manual frame-level annotation while reliably targeting the phase where expert differences are most pronounced.
3. Motion Infilling Training & Inference: Bidirectional Context-Conditioned Manifold Projection
After substituting the critical span with learnable [MASK] embeddings, ExpertEdit deploys a bidirectional Transformer, MotionInfiller, to reconstruct the masked expert token segment. Unlike unidirectional causal forecasting, bidirectional infilling conditions simultaneously on both past approach motion and subsequent follow-through context:
$\(\mathcal{L}_{\text{MLM}} = -\sum_{i=t^*-h}^{t^*+h} \log p_\theta(k_i^{\text{exp}} \mid \mathbf{k}_{\setminus [t^*-h:t^*+h]}^{\text{exp}})\)$
During inference on an unaligned novice sequence \(\mathbf{X}^{\text{nov}}\), the sequence is first mapped to tokens via \(E_\phi\), masked at its kinematic peak \(t^*\), and passed to MotionInfiller to predict expert-quality replacement tokens in a single forward pass. Decoder \(D_\phi\) maps the completed token sequence back into continuous joint rotations \(\hat{\mathbf{p}}_t\). By seamlessly stitching these refined rotations into the unmasked frames while preserving the original translation \(\mathbf{r}_t\) and root orientation \(\mathbf{o}_t\), ExpertEdit achieves a fine-grained balance between the performer's personal movement context and expert-level limb execution.
Loss & Training¶
ExpertEdit employs a technique-specific training paradigm. For each distinct athletic technique \(\tau\) (e.g., Mikan layup, penalty kick, front kick), separate Pose Tokenizer and MotionInfiller models are trained independently. This specialization enables each model to acquire a focused, sharp expert motion prior without the mode collapse or cross-skill interference common in heterogeneous multi-task motion datasets. Training relies exclusively on unpaired expert demonstration clips from Ego-Exo4D ("Late Expert" annotations) and Kyokushin Karate (1st-3rd dan black belts and 1st-3rd kyu). Optimization minimizes \(\mathcal{L}_{\text{VQ}}\) and \(\mathcal{L}_{\text{MLM}}\) without requiring paired supervision or reinforcement learning rewards.
Key Experimental Results¶
Main Results¶
To rigorously evaluate skill refinement quality, the authors established the first standardized benchmark spanning two multi-sport datasets across eight techniques (three basketball drills, one soccer kick, and four karate actions). Evaluation is conducted using two core metrics: - Pose Improvement (\(P\) โ): Measures relative reduction in Procrustes-Aligned Mean Per-Joint Position Error (PA-MPJPE) between edited motions and temporally aligned expert references, compared against the original novice motion error. - FID Improvement (\(F\) โ): Evaluates alignment with the expert motion distribution using Frรฉchet distance in the latent feature space of a technique-level skill classifier, reported as relative percentage improvement over the novice baseline.
The quantitative results across basketball, soccer, and karate techniques are presented in the tables below.
| Sport Domain | Technique | Metric | ExpertEdit (Ours) | SimMotionEdit (Supervised) | TMED (Supervised) | FLAME (Inference-time) |
|---|---|---|---|---|---|---|
| Basketball | Mikan Layup | \(P\) (%) โ \(F\) (%) โ |
6.18 5.72 |
2.26 1.28 |
1.87 1.00 |
2.35 1.45 |
| Basketball | Reverse Layup | \(P\) (%) โ \(F\) (%) โ |
5.87 12.08 |
2.31 4.88 |
2.20 4.30 |
2.09 6.13 |
| Basketball | Jumpshot | \(P\) (%) โ \(F\) (%) โ |
5.34 7.66 |
3.95 1.24 |
2.58 2.03 |
3.13 1.54 |
| Soccer | Penalty Kick | \(P\) (%) โ \(F\) (%) โ |
6.03 9.14 |
3.20 2.38 |
2.57 1.95 |
2.90 2.22 |
| Karate | Reverse Punch | \(P\) (%) โ \(F\) (%) โ |
2.07 1.36 |
2.20 1.30 |
0.92 1.08 |
- - |
| Karate | Spinning Back Kick | \(P\) (%) โ \(F\) (%) โ |
1.79 4.23 |
1.43 3.30 |
0.86 2.25 |
- - |
| Karate | Front Kick | \(P\) (%) โ \(F\) (%) โ |
1.88 9.73 |
1.45 5.05 |
0.38 2.90 |
- - |
| Karate | High Roundhouse Kick | \(P\) (%) โ \(F\) (%) โ |
2.96 6.18 |
3.30 5.88 |
2.25 3.37 |
- - |
Note: Baselines SimMotionEdit and TMED received privileged training on 16k novice-expert pseudo-pairs, whereas ExpertEdit trained strictly on unpaired expert videos. FLAME was not evaluated on Karate because its pretrained checkpoint is tied to SMPL parameterization whereas Karate utilizes custom 39-joint mocap data.
Ablation Study & Human Evaluation¶
The authors conducted a blinded A/B human preference study on rendered motion videos. Evaluators with playing experience in basketball and soccer answered: "Which edited motion would be more useful to improve your form if you were the performer in the video?"
| Evaluation Domain | Preferred ExpertEdit | Preferred SimMotionEdit (Strongest Baseline) | Neither / Indistinguishable | ExpertEdit Win Rate (Non-Neutral Decisions) |
|---|---|---|---|---|
| Basketball (78 trials) | 42.3% | 26.9% | 30.8% | 61.1% |
| Soccer (34 trials) | 47.1% | 23.5% | 29.4% | 66.7% |
| Overall (112 trials) | 43.8% | 25.9% | 30.4% | 62.8% |
Key Findings¶
- Superiority Over Privileged Baselines: Despite receiving zero paired supervision or instructional text, ExpertEdit achieves 2x to 4x higher gains in both pose error reduction (\(P\)) and expert distribution alignment (\(F\)) across ball sports compared to fully supervised diffusion baselines. In reverse layups and penalty kicks, FID improvements reach 12.08% and 9.14%, respectively.
- Distribution Realism Over Point Regression: On Karate reverse punches and high roundhouse kicks, while SimMotionEdit achieves marginally lower per-joint Euclidean error (\(P\)), ExpertEdit achieves strictly higher FID improvements (\(F\)). This demonstrates that regression baselines often yield blurry intermediate poses, whereas discrete manifold infilling respects dynamic biomechanical constraints.
- Biomechanical Quality of Edits: Qualitative visualizations confirm that ExpertEdit reliably introduces subtle expert adjustments, such as elevating the shooting-side knee during layup takeoff, repositioning the shooting palm directly beneath the ball on jumpshots, and completing full leg extension and follow-through on karate kicks and soccer strikes.
Highlights & Insights¶
- Localized Skill Projection vs. Global Re-synthesis: Rather than treating motion editing as an open-ended generative task or full-body style transfer, the framework exploits the concentrated nature of motor skill differences, modifying only the operative phases while preserving personal trajectory context.
- Pair-Free and Prompt-Free Practicality: While curating paired novice-expert trials with descriptive text corrections is prohibitively expensive, unpaired expert video clips are abundant in digital sports broadcasts and online tutorials. Demonstrating that MLM infilling suffices to transfer skill representations unlocks scalable real-world application.
- Standardized Multi-Sport Benchmark: By introducing DTW temporal warping, sagittal reflection alignment, and human expert validation, the work contributes a reproducible, rigorous evaluation pipeline for fine-grained athletic motion editing.
Limitations & Future Work¶
- Handcrafted Kinematic Statistics: Peak frame localization currently relies on domain heuristics (e.g., vertical velocity for basketball, foot acceleration for kicks). Integrating multimodal large language models (MLLMs) to automatically deduce key biomechanical indicators from exercise manuals represents a promising next step.
- Absence of Scene and Object Interactions: Operating purely in 3D skeletal space abstracts away physical interactions with balls, hoops, or playing surfaces. If a novice initiates a shot too far from the basket, local joint refinement alone cannot correct spatial discrepancies without explicit contact physics.
- Inflexible Temporal Pacing: Because frame duration is strictly frozen to match the source video, the model cannot emulate the explosive acceleration or elongated hang-time typical of elite athletes, pointing toward the need for adaptive temporal warping.
Related Work & Insights¶
- vs. MotionFix / TMED [5]: MotionFix relies on large-scale paired triplets with natural language instructions. Under general prompts like "make it smoother," diffusion baselines fail to deduce specific joint corrections. ExpertEdit foregoes text conditioning entirely, relying on the intrinsic expert manifold.
- vs. SimMotionEdit [32]: SimMotionEdit predicts a 1D temporal mask guided by text prompts, but struggles with temporal localization when prompts lack fine-grained timestamps. ExpertEdit anchors edits directly onto physical kinematic extrema.
- vs. ExpertAF [3]: ExpertAF requires synchronized novice-expert video pairs annotated with audio commentary. ExpertEdit bypasses paired supervision during training and inference, dramatically reducing annotation overhead.
Rating¶
- Novelty: โญโญโญโญโญ [Pioneers prompt-free, un-paired skill-driven 3D motion editing via kinematic contextual infilling]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive validation across 8 techniques, 3 sports, quantitative geometric/distributional metrics, and athlete human preference studies]
- Writing Quality: โญโญโญโญโญ [Clean formulation, rigorous terminology, and insightful biomechanical explanations]
- Value: โญโญโญโญโญ [Highly actionable for AI-assisted sports coaching, athletic rehabilitation, and robot demonstration synthesis]