Skip to content

Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Human Understanding / Diffusion Model
Keywords: motion style transfer, human motion diffusion, continuous style control, endpoint supervision, extrapolation evaluation

TL;DR

To address the lack of fine-grained continuous intensity control in human motion diffusion, this paper proposes Motion Style Slider, an endpoint-supervised framework that models style direction in a frozen latent embedding space and controls generation via a scalar slider \(\alpha\), leveraging diffusion denoising alongside linearity and monotonicity latent regularizers to achieve smooth, monotonic, and extrapolatable style scaling without requiring intermediate ground-truth motions.

Background & Motivation

Generating expressive and style-controllable 3D human motion is critical for game production, character animation, and interactive virtual reality. While recent human motion diffusion models have substantially advanced motion realism and physical plausibility, existing motion style transfer systems predominantly treat style as a discrete category (such as happy, sad, or angry) or perform transfer toward a fixed target reference instance. Consequently, fine-grained control over stylistic intensity remains fundamentally underexplored. Under the absence of explicit intermediate supervision, simple latent interpolations often fail to ensure monotonic progression of stylistic traits, let alone support reliable extrapolation beyond observed endpoints.

In practical game and animation studios, character movement is directed through iterative, relative adjustments such as "make the walk slightly less gloomy," "push the angry hop further," or "make it twice as energetic." Because artistic style perception is subjective and context-dependent, absolute intensity metrics (e.g., a globally standardized "2× style") cannot be universally established. Instead, the critical industrial requirement is predictable relative controllability: as a control slider increases, the stylistic nuance must intensify in a strictly monotonic order, transition smoothly across frames, and exhibit moderate extrapolation capability. However, motion capture datasets typically only provide paired endpoints containing a neutral motion and a fully stylized performance; continuous intermediate-intensity motion sequences do not exist in practice.

This paper tackles this real-world production bottleneck by reframing continuous style control around an endpoint-supervised relative coordinate system. The core idea is to model stylistic variations as directed vectors within a learned motion semantic embedding space, condition the diffusion process with a continuous scalar slider \(\alpha\), and impose latent linearity and monotonicity regularizers alongside the denoising objective, thereby synthesizing smooth, monotonic, and realistically extrapolatable motion trajectories from endpoint pairs alone.

Method

Overall Architecture

The objective of Motion Style Slider is to take a neutral content motion \(m_c\) and a stylized endpoint motion \(m_s\), and synthesize an action sequence \(\hat{m}_\alpha\) governed by a continuous intensity parameter \(\alpha \ge 0\). At \(\alpha = 0\), the model reproduces the neutral content motion \(m_c\); at \(\alpha = 1\), it reconstructs the target style endpoint \(m_s\); and for \(\alpha \in (0, 1)\) or extrapolation ranges \(\alpha > 1\), it generates motions along a smooth and monotonic style trajectory. Built upon a pretrained Motion Diffusion Model (MDM) backbone, the system extracts content tokens and directional style tokens using frozen encoders, which are subsequently injected into the denoiser via lightweight trainable adaptor layers.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Endpoint Motion Pair<br/>Content mc and Style ms"] --> B["Style Direction Construction<br/>Frozen TMR Encoding and Vector Estimation"]
    B --> C["Slider Condition Generation<br/>Continuous Scalar Scaling and Projection Token"]
    C --> D["Endpoint-Weighted Diffusion Denoising<br/>Content Preservation and MDM Adaptor Injection"]
    D --> E["Latent Trajectory Regularization<br/>Linearity and Monotonicity Penalties"]
    E --> F["Continuous Styled Motion Output<br/>Interpolated and Extrapolated Sequences"]

During training, the diffusion backbone denoises perturbed motion sequences conditioned on content tokens and scalar-modulated style tokens. Crucially, in addition to endpoint reconstruction supervision, intermediate and extrapolated points are regularized by projection operators in the latent feature space, ensuring smooth and monotonic intensity progression across arbitrary slider values.

Key Designs

1. Style Direction Construction: Establishing a Calibrated Style Axis in Normalized Latent Space

Interpolating directly in the joint coordinate space inevitably induces joint disarticulation, foot sliding, and physical unnaturalness. To circumvent this, the method extracts style vectors within a highly structured semantic motion embedding space. A frozen contrastive text-to-motion retrieval model (TMR) serves as the style encoder \(E_s(\cdot)\), mapping any motion sequence onto a unit hypersphere: \(\|z\|_2 = 1\). For any given training pair, the embeddings of the content and stylized endpoints are computed as \(z_c = E_s(m_c)\) and \(z_s = E_s(m_s)\). To balance instance-specific motion dynamics against global category consistency, three direction estimation modes are introduced:

\[d_{\text{instance}} = z_s - z_c, \qquad d_{\text{prototype}} = \mu_s - \mu_c\]

In instance mode, the direction is defined as the exact pair difference \(d_s = d_{\text{instance}}\); in prototype mode, the global cluster centroid difference between style and neutral sets is utilized \(d_s = d_{\text{prototype}}\); and hybrid mode blends both via weight \(\beta \in [0, 1]\): \(d_s = (1-\beta)d_{\text{prototype}} + \beta d_{\text{instance}}\). Normalizing the embedding space ensures that stylistic vectors reflect pure semantic orientation rather than magnitude shifts, enabling the scalar \(\alpha\) to serve as a well-behaved relative physical coordinate.

2. Slider Condition Generation: Endpoint-Aware Sampling and Dynamic Feature Injection

To support arbitrary continuous intensities while retaining structural action fidelity, the model calculates the modulated style embedding along the estimated direction: \(z_\alpha = z_c + \alpha d_s\). A linear projection module maps this vector into a style conditioning token \(c_{\text{style}}\). Simultaneously, the content motion \(m_c\) is processed by a frozen content encoder \(E_c(\cdot)\) through masked temporal average pooling (omitting the CLS token) to yield a content embedding \(h_c\), which is mapped via a learnable projection module \(P_c\) into a content token \(c_{\text{content}}\). Both tokens are injected into every layer of the pretrained MDM denoiser via trainable adaptor layers.

During training, sampling \(\alpha\) purely from a uniform distribution causes the model to underfit the exact boundary conditions. The framework addresses this via an endpoint-aware sampling distribution: sampling \(\alpha = 0\) with probability \(p_0\), \(\alpha = 1\) with probability \(p_1\), and sampling intermediate or extrapolated values from a uniform distribution \(\mathcal{U}(0, \alpha_{\max})\) otherwise. This skewed distribution heavily anchors the base reconstruction capabilities while simultaneously exposing the adaptor to continuous slider shifts.

3. Latent Trajectory Regularization: Linearity and Monotonicity Enforcement Without Midpoint Ground-Truth

Because ground-truth motion capture data for intermediate states \(\alpha \in (0, 1)\) and over-reactions \(\alpha > 1\) are completely absent, the diffusion denoising objective alone cannot prevent stylistic drift or backward collapse between endpoints. To guide intermediate generation along the desired style trajectory, the framework leverages the endpoint difference vector \(\Delta = z_s - z_c\) to define a scalar projection function:

\[g(m) = \left\langle E_s(m) - z_c, \frac{\Delta}{\|\Delta\|_2} \right\rangle\]

This projection quantifies the progress of a generated sequence along the target style axis. Using this metric, two unsupervised regularizers are formulated:

\[\mathcal{L}_{\text{lin}} = \left( g(\hat{m}_\alpha) - \alpha g(m_s) \right)^2\]
\[\mathcal{L}_{\text{mono}} = \max\left(0, \gamma - \left( g(\hat{m}_{\alpha_2}) - g(\hat{m}_{\alpha_1}) \right)\right) \quad (\text{where } \alpha_2 > \alpha_1)\]

where \(\gamma \ge 0\) is a separation margin. The linearity loss \(\mathcal{L}_{\text{lin}}\) penalizes deviation from proportional progression along the style axis, flattening undesirable manifold curvature; the monotonicity loss \(\mathcal{L}_{\text{mono}}\) suppresses reversal anomalies where a higher slider value inadvertently yields weaker stylistic expression. Together, they constrain generated motions to traverse a smooth, monotonic manifold without requiring midpoint ground-truth.

Loss & Training

The total optimization objective integrates diffusion denoising with both latent geometric regularizers:

\[\mathcal{L} = \mathcal{L}_{\text{diff}} + \lambda_{\text{lin}} \mathcal{L}_{\text{lin}} + \lambda_{\text{mono}} \mathcal{L}_{\text{mono}}\]

Here, \(\mathcal{L}_{\text{diff}}\) assigns full weight (1.0) to endpoint samples (\(\alpha \in \{0, 1\}\)) and downweights intermediate samples by a factor \(w_{\text{mid}}\). Optimization is performed using AdamW starting from pretrained MDM weights. To preserve general human motion kinematics, the original diffusion backbone remains frozen, and only the lightweight Transformer adaptors and condition projection layers are trained. At test time, given a pair \((m_c, m_s)\) and any user-specified slider value \(\alpha\) (e.g., 0.5, 1.2, 2.0), the model executes a single diffusion sampling trajectory to synthesize a high-fidelity stylized animation aligned with the original duration mask.

Key Experimental Results

Main Results

Evaluation is conducted across three public motion style benchmarks (PerMo, Bandai-Namco, Xia) and a newly captured game-industry Over-Reaction motion capture test set. The intensity evaluation set covers \(\alpha \in \{0, 0.5, 1.0, 1.5, 2.0\}\). Standard metrics include Fréchet Motion Distance (FMD), Content Recognition Accuracy (CRA), Style Recognition Accuracy (SRA), Monotonicity Violation rate (MonoViol), Linearity Error (LinErr), and Extrapolation Error (ExtraErr) against real held-out over-reaction captures.

Table 1 reports the macro-averaged results across all benchmark datasets:

Method FMD ↓ CRA ↑ SRA ↑ MonoViol ↓ LinErr ↓ ExtraErr ↓
DeepMotionEditing 0.588 0.290 0.060 0.25 0.326 1.196
MCM-LDM 0.381 0.808 0.265 0.11 0.358 1.174
Ours (Motion Style Slider) 0.457 0.606 0.295 0.05 0.280 1.109

Table 2 details the per-dataset performance across the three primary benchmarks (formatted as Ours / MCM-LDM):

Dataset FMD ↓ CRA ↑ SRA ↑ MonoViol ↓ LinErr ↓
PerMo 0.531 / 0.474 0.700 / 0.880 0.140 / 0.080 0.03 / 0.22 0.421 / 0.464
Bandai-Namco 0.143 / 0.237 0.714 / 1.000 0.714 / 0.429 0.01 / 0.06 0.259 / 0.368
Xia 0.279 / 0.346 0.108 / 0.351 0.405 / 0.216 0.07 / 0.05 0.261 / 0.407

Ablation Study

On the comprehensive PerMo benchmark (33 styles, 10 content actions), ablations isolate the contribution of direction estimation strategies and individual regularizers:

Variant \(\mathcal{L}_{\text{lin}}\) \(\mathcal{L}_{\text{mono}}\) Direction Mode FMD ↓ CRA ↑ SRA ↑ MonoViol ↓ LinErr ↓
Ours (Full Model) Yes Yes instance 0.531 0.70 0.14 0.03 0.421
Dir=proto Yes Yes proto 0.584 0.60 0.06 0.04 0.513
Dir=hybrid (\(\beta=0.3\)) Yes Yes hybrid 0.588 0.70 0.12 0.03 0.509
w/o \(\mathcal{L}_{\text{lin}}\) No Yes instance 0.605 0.54 0.06 0.03 0.422
w/o \(\mathcal{L}_{\text{mono}}\) Yes No instance 0.545 0.56 0.04 0.04 0.434

Furthermore, a blind user study with 11 professional animation evaluators across 6 matched case triplets demonstrated that our method achieves a distinguishability score of 4.92 (vs. 2.36 for MCM-LDM), steadily increasing stylization strength ratings across low, medium, and high slider values (3.92 / 4.61 / 5.18 vs. 3.56 / 3.59 / 3.46), and a 50% monotonic pass rate (vs. 0% for MCM-LDM).

Key Findings

  • Instance direction mode achieves the strongest SRA (0.14) and lowest LinErr (0.421) on PerMo, whereas prototype direction averages out actor-specific dynamics and collapses SRA to 0.06, indicating that stylistic traits in motion are inherently intertwined with individual kinematic execution.
  • Disabling linearity regularization \(\mathcal{L}_{\text{lin}}\) degrades FMD from 0.531 to 0.605 and halves SRA, showing that uncontrolled manifold curvature causes abrupt perceptual jumps; discarding monotonicity regularization \(\mathcal{L}_{\text{mono}}\) worsens intensity violations and harms content retention.
  • In out-of-range extrapolation (\(\alpha = 2.0\)), our method achieves an ExtraErr of 1.109 against true over-reaction captures, outperforming MCM-LDM (1.174) and DeepMotionEditing (1.196), confirming that the learned vector direction captures a physically and stylistically meaningful tangent rather than an unconstrained interpolation artifact.

Highlights & Insights

  • Formulates style intensity as a relative control coordinate rather than an absolute physical unit, perfectly aligning generation capabilities with real-world artistic iterative revision workflows.
  • Employs dual latent geometric regularizers—style projection linearity and monotonicity—to enforce smooth, predictable slider responses without requiring unobtainable intermediate motion ground-truth.
  • Introduces an over-reaction motion capture evaluation benchmark featuring professional actors performing exaggerated expressions, establishing a rigorous standard for assessing motion extrapolation beyond standard training boundaries.

Limitations & Future Work

  • The fidelity of the constructed style direction relies heavily on clean kinematic synchronization and semantic consistency within the endpoint motion pair; noisy or misaligned endpoints can skew the control axis.
  • Extreme extrapolation (\(\alpha > 2.0\)) can occasionally induce foot sliding or joint jitter due to the lack of explicit biomechanical contact and physics constraints.
  • While cross-content stylization is demonstrated qualitatively, the quantitative benchmark is restricted to paired inputs; expanding the slider framework to text prompts and unpaired reference motions remains an exciting frontier.
  • vs DeepMotionEditing (Aberman et al., 2020): Relies on autoencoder disentanglement for discrete transfer and lacks a continuous scalar conditioning interface, resulting in severe motion artifacts (FMD 0.588) and low style accuracy (0.060).
  • vs MCM-LDM (Song et al., 2024): Offers strong single-point generation quality and content preservation, but lacks explicit direction anchoring and trajectory regularization, frequently suffering from non-monotonic reversals (MonoViol 0.11, 0% user monotonic pass rate). Motion Style Slider provides substantially superior intensity distinguishability and monotonic predictability.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Reframing motion style transfer as continuous relative control via endpoint supervision and latent trajectory geometry is insightful and directly applicable.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive validation across three benchmark datasets, a newly captured extrapolation suite, ablation studies, and rigorous user studies.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear industrial motivation, solid problem formulation, and elegant mathematical exposition.
  • Value: ⭐⭐⭐⭐☆ Highly impactful for digital character animation, game asset authoring, and interactive graphics pipelines.