Skip to content

SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation

Conference: ECCV 2026
Paper: ECCV Official Page
Code: https://mpi-lab.github.io/SparseCtrl-HOI
Area: Video Generation
Keywords: Human-Object Interaction, Video Generation, Sparse Temporal Control, Rotary Position Embedding, Multimodal Large Language Model

TL;DR

Addressing the excessive annotation cost and motion rigidity of dense pose control in Human-Object Interaction (HOI) video generation, this paper proposes SparseCtrl-HOI, a DiT-based sparse temporal control framework that anchors interaction keyframes at specified timestamps via Time-Controlled Rotary Positional Embedding (TiRoPE), injects MLLM motion priors via a Q-Former, and decouples training into two stages to synthesize physically plausible, high-fidelity live-streaming e-commerce videos.

Background & Motivation

Human-Object Interaction (HOI) video generation aims to synthesize realistic dynamic videos of humans manipulating diverse objects, holding profound commercial promise and scientific value for virtual reality, digital entertainment, and AI-driven live-streaming e-commerce. In live commerce, virtual avatars are required to present continuous, fine-grained product demonstrations, demanding that generative models simultaneously preserve the streamer's facial identity and commodity textures while faithfully modeling intricate spatial-temporal hand-object coordination, deformation, and physical contact.

Existing HOI video generation pipelines predominantly follow a dense guidance paradigm: one family relies on frame-wise 2D/3D body poses, hand meshes, or relative coordinate maps to strictly steer the synthesis process; another family explores sparse spatial cues (such as sparse keypoints or trajectories) but still depends on decoupled guidance or motion densification to interpolate them into dense, frame-level constraints. Conditioning on full-frame dense poses overrides the generative foundation model's inherent physical priors, degrading diffusion transformers into mere texture renderers that produce repetitive, unnatural motions. Furthermore, any tracking errors or jitter in external pose estimators directly propagate into the synthesized video, while procuring dense annotations for custom e-commerce products remains prohibitively labor-intensive.

To circumvent these fundamental bottlenecks, a more scalable and natural approach is to shift to sparse temporal control—supplying key interaction states only at discrete timestamps and allowing the model to autonomously hallucinate smooth, physically coherent transitions. The core idea is to concatenate sparse interaction keyframes along the temporal axis with noisy latents, temporally anchor these states via Time-Controlled Rotary Positional Embedding (TiRoPE), extract high-level kinematic priors using a multimodal large language model (MLLM) to steer intermediate dynamics through a Q-Former, and resolve feature entanglement via a decoupled two-stage training strategy.

Method

Overall Architecture

SparseCtrl-HOI is built upon the continuous-time flow matching paradigm with a Wan2.1-DiT backbone. It takes three multimodal inputs: a reference image \(I_{\text{ref}}\) specifying identity, a speech audio sequence for lip synchronization, and a sparse set of interaction keyframes \(I_1^{\text{ho}}, \dots, I_K^{\text{ho}}\) situated at designated timestamps \(t_1, \dots, t_K\). Encoded by a causal 3D VAE, the reference and keyframe latents are injected asymmetrically into the DiT to propagate appearance and interaction geometries through self-attention. TiRoPE establishes direct temporal resonance at designated timestamps, while an MLLM-driven motion prior injection module guides the intermediate frames with physically grounded dynamics.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multimodal Inputs<br/>Reference Image + Audio + Sparse HOI Keyframes"] --> B["Asymmetric Multimodal Condition Injection<br/>Channel Broadcast Reference + Temporal Append Keyframes"]
    B --> C["Time-Controlled Rotary Positional Embedding (TiRoPE)<br/>Keyframe Assigned Target Timestamp Temporal Index"]
    D["Qwen2.5-VL Dynamics Reasoning<br/>Extract High-Level Transition Semantics from Keyframes"] --> E["Motion Prior Injection Module<br/>MLP Projection + 4-Module Q-Former into 64 Tokens"]
    C --> F["Decoupled Two-Stage Training Strategy<br/>Stage 1 Appearance LoRA + Stage 2 Motion Cross-Attention"]
    E --> F
    F --> G["High-Fidelity HOI Video Generation<br/>Identity Preserved + Keyframes Aligned + Physically Plausible"]

Key Designs

1. Asymmetric Multimodal Condition Injection: Channel-Wise Identity Broadcasting and Temporal Keyframe Concatenation

Concatenating sparse keyframe latents directly along the channel dimension with video latents forces non-keyframe steps to be zero-padded, resulting in extreme feature sparsity, slow training convergence, and corrupted pre-trained priors. SparseCtrl-HOI introduces an asymmetric injection scheme: the single reference image latent \(\mathbf{z}_{\text{ref}} \in \mathbb{R}^{1 \times \frac{H}{8} \times \frac{W}{8} \times C}\) is broadcast temporally and concatenated along channels with the noisy video latent \(\mathbf{z}_{\text{v}}^{\text{noise}} \in \mathbb{R}^{\frac{T}{4} \times \frac{H}{8} \times \frac{W}{8} \times C}\), providing consistent global identity conditioning. Concurrently, the \(K\) sparse interaction keyframe latents \(\mathbf{z}^{\text{ho}} \in \mathbb{R}^{K \times \frac{H}{8} \times \frac{W}{8} \times C}\) are channel-padded with random Gaussian noise \(\boldsymbol{\epsilon}_{\text{ho}}\) and concatenated along the temporal axis:

\[\mathbf{z}_{\text{in}} = \text{Concat}_T\left(\text{Concat}_C(\mathbf{z}_{\text{v}}^{\text{noise}}, \mathbf{z}_{\text{ref}}), \; \text{Concat}_C(\boldsymbol{\epsilon}_{\text{ho}}, \mathbf{z}^{\text{ho}})\right)\]

This allows the DiT's inherent spatio-temporal self-attention layers to dynamically query hand poses and object textures from keyframes. To prevent background interference and avoid trivial copy-paste shortcuts, keyframes undergo morphological dilation and Gaussian boundary smoothing on hand-object masks, coupled with HSV color jittering and scaling on hand regions to force the network to prioritize kinematic structure over local skin tone.

2. Time-Controlled Rotary Positional Embedding (TiRoPE): Explicit Frequency Alignment Across Timestamps

Concatenating keyframes along the temporal dimension provides appearance tokens, but standard 3D RoPE assigns them arbitrary sequential positional indices at the end of the sequence, leaving the model blind to when each interaction pose should materialize. TiRoPE breaks the rigid tie between tensor physical memory indices and temporal positional coordinates: while noisy video tokens receive standard continuous spatio-temporal RoPE, the \(i\)-th interaction keyframe latent \(\mathbf{z}_i^{\text{ho}}\) is explicitly assigned the exact temporal index of its target timestamp \(t_i\).

Because rotary position embeddings compute inner products based on relative coordinate differences, assigning temporal index \(t_i\) drives the relative temporal phase offset between the keyframe and its target generated frame to zero. This triggers strong attention affinity within the self-attention mechanism, seamlessly transferring the prescribed hand-object pose and contact relationship at the designated moment while leaving spatial positional indices unaltered to preserve 2D spatial layouts.

3. Motion Prior Injection Module: Cross-Keyframe Kinematic Reasoning and Q-Former Compression

While TiRoPE secures precise alignment at sparse anchor frames, the extensive intervals between them lack explicit intermediate constraints, often causing unrealistic object morphing, penetration, or levitation during large viewpoint transformations. SparseCtrl-HOI introduces Qwen2.5-VL to perform high-level physical reasoning over the ordered sequence of keyframes, prompted to infer hand trajectory, object orientation, and hand-object contact zone migration. The raw semantic representations from the final MLLM layer \(\mathbf{R}_{\text{mllm}}\) capture high-level dynamics.

To condense these lengthy representations and bridge the semantic gap to the generative diffusion space, a lightweight Q-Former with 4 cascaded modules (each containing self-attention, cross-attention, and an FFN) is deployed. Following an initial linear MLP projection, 64 learnable query vectors \(\mathbf{Q}\) distill the motion priors via cross-attention:

\[\mathbf{C}_{\text{mllm}} = \text{Q-Former}(\mathbf{Q}, \text{MLP}(\mathbf{R}_{\text{mllm}}))\]

These compact motion guidance tokens \(\mathbf{C}_{\text{mllm}}\) are then fed into newly inserted Motion Cross-Attention layers within each DiT Transformer block, effectively guiding the intermediate frames to hallucinate physically coherent, smooth transitions.

4. Decoupled Two-Stage Training Strategy: Segregating Appearance Transfer from Motion Synthesis

MLLMs excel at abstract kinematic reasoning but frequently suffer from visual hallucinations regarding fine-grained micro-textures; conversely, DiTs excel at high-fidelity visual rendering. End-to-end joint training causes severe feature entanglement, where the network over-relies on MLLM representations for texture synthesis, transferring visual hallucinations into foreground product artifacts. SparseCtrl-HOI implements a decoupled training protocol:

  • Stage 1: Appearance Feature Transfer: The VAE is frozen, patch embeddings and audio packages are unfrozen, and LoRA parameters are applied to DiT self-attention and text cross-attention layers. Trained for 15,000 steps with flow matching loss, the model masters appearance extraction from reference images and precise keyframe pose reproduction guided by TiRoPE.
  • Stage 2: Motion Transition Inference: All Stage 1 LoRA weights are frozen to safeguard baseline synthesis and visual fidelity. Only the newly added Q-Former and Motion Cross-Attention layers are optimized for 10,000 steps, dedicating this stage entirely to mapping \(\mathbf{C}_{\text{mllm}}\) into continuous motion trajectories without compromising surface texture clarity.

Loss & Training

The architecture adopts a continuous-time flow matching objective. For data latent \(\mathbf{z}_0\) and Gaussian noise \(\boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\), the linear interpolation is \(\mathbf{z}_t = (1 - t)\mathbf{z}_0 + t\boldsymbol{\epsilon}\), and the DiT velocity field \(v_\theta\) minimizes:

\[\mathcal{L}_{\text{flow}}(\theta) = \mathbb{E}_{t \sim \mathcal{U}(0,1), \, \mathbf{z}_0 \sim p_{\text{data}}, \, \boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0},\mathbf{I})} \left\| v_\theta(\mathbf{z}_t, t, c) - (\boldsymbol{\epsilon} - \mathbf{z}_0) \right\|_2^2\]

Training is executed on the 4,800 clips of the SparseHOI-5K training split using AdamW with a constant learning rate of \(5 \times 10^{-6}\), a global batch size of 8, and \(K=4\) sparse interaction keyframes randomly sampled per video.

Key Experimental Results

Main Results

Evaluations are conducted on the newly curated live commerce benchmark SparseHOI-5K (50 identity-disjoint test subjects) and the public AnchorCrafter test set. Evaluated baselines include AnchorCrafter (dense pose/mesh guidance), OmniAvatar and its keyframe-augmented baseline OmniAvatar-BL, alongside general video generators VACE and Phantom. Metrics cover distribution fidelity (FID, FVD), temporal consistency (VBench Motion Smoothness MS, Temporal Flickering TF, Aesthetic Quality AQ), audio alignment (Sync-C), pixel-level warping smoothness via optical flow (MS-RAFT), and MLLM-assessed physical plausibility (HOI-VLM, 1-5 scale).

Dataset Method FID ↓ FVD ↓ MS ↑ TF ↑ AQ ↑ Sync-C ↑ MS-RAFT ↓ HOI-VLM ↑
AnchorCrafter AnchorCrafter [45] 91.9 726.4 0.99 0.98 0.46 - 0.43 2.90
VACE [19] 249.0 3809.2 0.98 0.97 0.418 - 0.60 2.90
Phantom [28] 254.3 2544.3 0.96 0.98 0.43 - 0.53 2.50
SparseCtrl-HOI (Ours) 91.4 725.8 0.99 0.99 0.48 - 0.27 3.00
SparseHOI-5K AnchorCrafter [45] 258.4 2859.8 0.99 0.98 0.43 - 0.79 1.18
OmniAvatar [9] 131.2 1207.0 0.99 0.99 0.43 7.04 0.47 2.76
OmniAvatar-BL 129.3 1110.9 0.98 0.98 0.43 6.99 0.49 2.78
VACE [19] 212.9 1613.3 0.98 0.98 0.43 - 0.65 2.86
Phantom [28] 254.2 2701.7 0.99 0.98 0.44 - 0.48 1.56
SparseCtrl-HOI (Ours) 89.1 629.6 0.99 0.99 0.43 6.98 0.45 2.92

Ablation Study

Ablations on SparseHOI-5K validate key architectural components, condition injection configurations, and inference keyframe density \(K\), tracking timestamp structural similarity (Ti-SSIM) within a \(\pm 10\) frame window.

Group Variant / Config FID ↓ FVD ↓ Ti-SSIM ↑ MS-RAFT ↓ Note
Component w/o Data Aug. 95.14 709.14 0.65 0.52 Lacks blur/jittering; produces visible boundary cut-and-paste artifacts
w/o TiRoPE 99.28 697.84 0.52 0.64 Loses explicit timestamp anchoring; Ti-SSIM drops by 0.14
w/o M.P.I. 90.86 651.18 0.66 0.53 Missing semantic kinematic guidance; optical flow warping error worsens
Full model 89.17 629.61 0.66 0.45 Best overall trade-off and visual fidelity
Injection Scheme Reference along temporal axis 111.81 822.67 0.60 0.66 Weakens global identity conditioning across all frames
Keyframes along channel axis 202.54 1292.19 0.50 0.79 Heavy zero-padding and implicit propagation cause severe degradation
Full model 89.17 629.61 0.66 0.45 Optimal separation of global appearance and sparse temporal states
Inference \(K\) \(K = 2\) 115.64 795.32 0.67 0.45 Loose boundary control; higher synthesis variance
\(K = 3\) 102.25 664.77 0.67 0.44 Monotonic improvement as anchor density increases
\(K = 4\) (default) 89.17 629.61 0.66 0.45 Balanced annotation overhead and generative quality
\(K = 6\) 82.20 544.05 0.66 0.48 Tighter trajectory control enhances fidelity
\(K = 8\) 81.34 515.84 0.67 0.46 Lowest FVD (515.84) while maintaining consistent timestamp alignment

Key Findings

  • TiRoPE is essential for sparse temporal anchoring: Disabling TiRoPE causes Ti-SSIM to collapse from 0.66 to 0.52 (a >21% degradation), proving that standard RoPE fails to bind keyframe features to target moments, causing severe temporal drift.
  • Motion prior injection prevents intermediate physical distortion: Omitting MLLM motion priors raises MS-RAFT optical flow error from 0.45 to 0.53. In qualitative analysis, objects undergo unnatural morphing and hand-object penetration under large rotations without semantic priors.
  • Decoupled training eliminates multimodal hallucination artifacts: Joint single-stage training entangles MLLM feature extraction with DiT rendering, resulting in conspicuous grain and surface noise on manipulated commodities; the two-stage protocol successfully shields texture synthesis.
  • Scalable temporal elasticity at test time: Although trained solely on \(K=4\) keyframes, the model seamlessly generalizes from \(K=2\) to \(K=8\) during inference, retaining stable Ti-SSIM values (0.66–0.67) while consistently improving FVD, demonstrating true arbitrary-timestamp controllability.

Highlights & Insights

  • Asymmetric injection with TiRoPE frequency anchoring: Combining channel-wise broadcast for static identity with temporal concatenation for sparse interaction states, coupled with custom RoPE index rewriting, establishes a zero-overhead, highly effective temporal control mechanism.
  • Bridging MLLM reasoning and diffusion generation via Q-Former: Rather than tasking MLLMs with direct pixel synthesis, using Qwen2.5-VL to conceptualize dynamic trajectory shifts and distilling it into 64 guidance tokens provides an exemplary blueprint for multimodal kinematic control.
  • Curating the SparseHOI-5K live commerce benchmark: Supplying 4,850 high-resolution (512×512, 25 FPS) clips across 34 presenters and 1,000+ commodities with complete multi-modal masks and inpainting baselines addresses a major void in commercial HOI research.

Limitations & Future Work

  • Fine-grained bimanual dexterous manipulation: The current framework is optimized for one-handed presentation or structured two-handed holding. Highly intricate topological interactions (e.g., finger interweaving, unscrewing small bottle caps) can still trigger subtle occlusion artifacts.
  • Unseen back-face object geometry: Keyframes only supply visible perspectives. If an object undergoes a complete 360-degree rotation revealing a complex unseen back pattern, the model relies on generative hallucination rather than deterministic 3D reconstruction.
  • Future directions: Integrating parametric 3D hand models (e.g., MANO) as auxiliary geometric priors or multi-view 3D commodity meshes could further elevate high-precision interactive avatar generation.
  • vs AnchorCrafter [45]: AnchorCrafter relies on dense per-frame 2D/3D skeletons and depth maps, suffering severe hand distortion (FVD 2859.8 on SparseHOI-5K) under estimation noise. SparseCtrl-HOI requires only 4 sparse keyframes, achieving superior fidelity (FVD 629.6) with dramatically lower annotation costs.
  • vs OmniAvatar [9]: OmniAvatar is limited to portrait speech and subtle torso gesturing driven by a single static image; SparseCtrl-HOI extends the flow-matching DiT framework into controllable human-object physical interactions while preserving comparable lip synchronization (Sync-C 6.98 vs 7.04).
  • vs HOMA [18] & VHOI [50]: While HOMA and VHOI accept sparse spatial trajectories, they convert them into dense frame-wise paths via heuristic densification, leaving them vulnerable to interpolation jitter. SparseCtrl-HOI leverages semantic MLLM priors directly inside DiT latent space for organic non-linear kinematics.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering sparse temporal control for HOI video generation with elegant TiRoPE and MLLM motion prior injection.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous validation across two benchmarks, 10+ standard and custom metrics, in-depth ablation studies, and a newly built 4.8k dataset.
  • Writing Quality: ⭐⭐⭐⭐⭐ Exceptionally clear narrative, structured method descriptions, and thoroughly transparent technical formulations.
  • Value: ⭐⭐⭐⭐⭐ Highly relevant to AI live-streaming e-commerce and interactive digital humans; public code and dataset provide immense practical utility.