Optimizing Mesh Animation from Video via Shape Flow Guidance¶
Conference: ECCV 2026
Paper: CVF Open Access
Area: 3D Vision
Keywords: Shape Flow Guidance, Video-driven Mesh Animation, Skeletal Animation Decoupling, 3D Supervision, Latent ODE
TL;DR¶
Addressing the severe shape collapses and motion artifacts caused by conventional 2D rendering-based reconstruction losses in video-driven mesh animation, this paper proposes Shape Flow Guidance (SFG) to extract temporally coherent 3D shape sequences via training-free intervention on a pretrained mesh generator's sampling trajectory, paired with a decoupled skeletal animation model for high-fidelity 4D mesh optimization.
Background & Motivation¶
Extracting mesh animations from monocular video holds immense practical value in video games, virtual reality, and embodied simulation. To maintain seamless compatibility with standard computer graphics pipelines (such as skeletal rigging, skinning, and texture mapping), the synthesized dynamic meshes must strictly preserve topological consistency—namely, sharing identical sets of vertices and triangular faces across all frames. A conventional strategy optimizes per-vertex displacements starting from a canonical static mesh, driven by differentiable rendering losses that match rendered frames with the reference video. However, explicit polygonal meshes exhibit inherent topological rigidity, and discontinuous rasterization leads to sparse gradients, rendering gradient-based optimization ill-suited for capturing large-scale, complex non-rigid deformations.
Prior endeavors have explored innovations across representations, auxiliary supervision, and operational paradigms. Some introduce hybrid Gaussian-mesh representations, yet volumetric Gaussian splats compromise manifold surface quality during violent motions. Others incorporate 2D optical flow or keypoint tracking priors via reprojection losses, yet still suffer from severe depth ambiguities along the projection axis. Still another line of work generates per-frame 3D meshes independently and attempts cross-frame non-rigid registration, which remains brittle and heavily prone to collapse under rapid non-rigid movements. The root cause of these failures lies in the fundamental limitations of 2D projection supervision: occluded and unseen regions receive zero supervisory signal and drift arbitrarily, while in visible areas, the ambiguous inverse mapping from 2D pixel residuals to 3D displacements easily leads to catastrophic geometric stretching and unnatural artifacts.
This paper tackles the challenge from a fresh perspective: rather than treating 2D video as an indirect projection penalty, it leverages pretrained 3D mesh generative models (e.g., Rectified Flow models) to elevate 2D video dynamics directly into explicit, continuous 3D spatio-temporal geometric guidance. Core idea: intervene in the sampling Ordinary Differential Equation (ODE) of a pretrained mesh generator in a training-free manner to distill a temporally coherent 3D Shape Flow Guidance (SFG), and decouple the animation model into global root transformation and local skeletal deformation, allowing explicit 3D geometry to drive complex local deformations while confining 2D rendering losses to simple global motion.
Method¶
Overall Architecture¶
Given a static canonical mesh \(M_0\) and a monocular reference video \(\{I^l\}_{l=0}^L\) captured from a fixed viewpoint (where \(M_0\) can be generated from the first frame if unavailable), the objective is to optimize a topology-consistent sequence of per-frame vertex displacements \(\{D_l\}_{l=1}^L\), producing dynamic meshes \(M_l=(V_0+D_l, F_0, T^0)\). The framework operates in two distinct stages: the first stage is Shape Flow Guidance Elicitation, where synchronized latent paths anchored to the initial mesh are integrated with an advection-diffusion partial differential equation to yield a smooth, temporally consistent sequence of 3D shapes. The second stage is Shape Flow Guidance Utilization, which leverages an automatic rigging module to establish a skeletal kinematic tree, decouples global root trajectory from local joint transformations, aligns the root-space deformation to the 3D shape flow via bidirectional point-to-mesh distance, refines surface details via a post-skinning correction network, and performs global differentiable rendering alignment.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Monocular Reference Video & Static Mesh"] --> B["Synchronized Latent Path Construction<br/>Anchor sampling trajectories to the initial mesh"]
B --> C["Advection-Diffusion PDE Intervention<br/>2D spatio-temporal flow field suppresses jitter"]
C --> D["Temporally Coherent 3D Shape Flow (SFG)"]
A --> E["Automatic Rigging & Kinematic Decoupling<br/>Factor out root joint from local articulation"]
D & E --> F["Shape Flow Guidance in Root Space<br/>Consistent barycentric anchor sampling"]
F --> G["Post-Skinning Correction Network<br/>Pose-conditioned MLP fits non-rigid residuals"]
G --> H["Differentiable Rendering & Regularization<br/>Global alignment and geometric smoothness"]
H --> I["Topology-Consistent 4D Mesh Animation"]
Key Designs¶
1. Synchronized Latent Path Construction: Anchoring geometric origin to ensure local deformation inheritance
Independently generating meshes for each video frame yields severe inter-frame jitter, arbitrary bounding-box orientations, and unpredictable variations in occluded regions. To enforce structural consistency from the outset, the framework anchors all subsequent generation paths to the latent representation of the initial frame. Given initial mesh latent \(x_1^0\) and target frame latents \(x_1^l\), and assuming shared initial noise across frames (\(x_0^l = x_0^0\)), synchronized paths are defined as \(s_t^l = x_1^0 + x_t^l - x_t^0\). The velocity field driving each path is directly derived as the difference between original velocity fields in the rectified-flow model: $\(\frac{ds_t^l}{dt} = v(x_t^l, t, I^l) - v(x_t^0, t, I^0)\)$ This formulation guarantees that all intermediate shape trajectories originate from the identical canonical mesh at sampling step \(t=0\), strictly inheriting its global orientation, base proportions, and unseen backside geometry, evolving only along directions conditioned on video appearance variations.
2. Advection-Diffusion PDE Intervention: Modeling spatio-temporal dynamics to eliminate temporal jitter
While the synchronized 1D paths enforce consistency relative to the canonical frame, individual trajectories remain decoupled across the frame index \(l\), leading to subtle frame-to-frame geometric flickering during sampling. To bridge this gap, the method extends the 1D flow into a unified 2D continuous flow field spanning both sampling time \(t\) and frame index \(l\), governed by an advection-diffusion partial differential equation (PDE): $\(\frac{\partial s_t^l}{\partial t} + \lambda_1 \mathrm{A}(s_t^l) + \lambda_2 \mathrm{D}(s_t^l) = v(x_t^l, t, I^l) - v(x_t^0, t, I^0)\)$ Here, the advection operator \(\mathrm{A}(s_t^l) \approx s_t^l - s_t^{l-1}\) models directional temporal momentum transfer, pulling the current latent state toward its immediate predecessor. The diffusion operator \(\mathrm{D}(s_t^l) \approx s_t^{l+1} - 2s_t^l + s_t^{l-1}\) implements a discrete temporal Laplacian, providing isotropic smoothing across consecutive frames. Discretizing this PDE yields a coupled system of ODEs that natively produces temporally continuous, flicker-free 3D shape flows without any model retraining.
3. Kinematic Decoupling & Root-Space Shape Flow Alignment: Fully isolating local deformation from global trajectory
The elicited 3D shape flow resides in a normalized canonical bounding box and captures exclusively local deformations, stripping away camera perspectives and global translations. Forcing raw 3D mesh vertices to match the unaligned shape flow would destroy spatial trajectory. To resolve this mismatch, the authors employ an autoregressive automatic rigging module (UniRig) on the static mesh and reformulate Linear Blend Skinning (LBS) by explicitly factoring out the global transformation matrix of the root joint, \(T_0 = [R_0 \mid t_0]\): $\(\mathbf{v}' = T_0 \sum_{j=0}^{J-1} w_j T'_j B_j^{-1} \mathbf{v}\)$ where \(T'_j\) represents the joint transformation relative to the root joint coordinate frame (root space). Complex local articulations are optimized purely within root space, supervised directly by bidirectional distance to the shape flow, whereas the low-dimensional root transformation is delegated to 2D rendering supervision. To avoid noisy gradients from random surface resampling, fixed barycentric coordinate anchors are pre-sampled on the static mesh triangles and tracked dynamically across optimization steps, guaranteeing stable surface convergence.
4. Post-Skinning Correction Network: Overcoming linear skinning rigidity and absorbing rigging inaccuracies
Standard linear blend skinning is intrinsically limited by rigid bone transformations, making it incapable of expressing fine-grained non-rigid surface dynamics such as muscle bulging, clothing wrinkles, and organic breathing. Furthermore, automatic rigging models occasionally predict imperfect kinematic hierarchies. To bridge this expressive gap, an MLP-based Post-Skinning Correction (PSC) network is integrated within root space. Conditioning on the articulated pose, the network predicts per-vertex residual offset vectors added directly to the LBS-deformed vertices prior to applying global root transformation. This architecture establishes a dual-tier deformation model: LBS enforces large-scale structural plausibility, while the PSC network smoothly captures high-frequency geometric nuances and absorbs minor skeleton binding errors.
Loss & Training¶
The overall optimization objective for dynamic mesh parameters is formulated as a multi-term objective balancing 3D shape guidance, 2D rendering fidelity, and spatio-temporal regularizations: $\(\mathcal{L} = \mathcal{L}_{\text{shape}} + \mathcal{L}_{\text{recon}} + \lambda_{\text{reg}} (\mathcal{L}_{\text{joint}} + \mathcal{L}_{\text{arap}} + \mathcal{L}_{\text{nc}})\)$ 1. 3D Shape Loss \(\mathcal{L}_{\text{shape}}\): Measures the symmetric bidirectional distance between consistently sampled deformed mesh points \(P\) in root space and target shape flow triangular faces \(F\), summing the point-to-mesh and mesh-to-point squared distances. 2. 2D Reconstruction Loss \(\mathcal{L}_{\text{recon}}\): Employs differentiable rasterization to render the globally transformed mesh into RGB color images and silhouette masks, penalizing discrepancies against the reference video to supervise the root joint transform \(T_0\). 3. Geometric Regularizations: \(\mathcal{L}_{\text{joint}}\) penalizes abrupt acceleration in inter-frame joint rotations to ensure temporal smoothness; \(\mathcal{L}_{\text{arap}}\) applies the As-Rigid-As-Possible energy to suppress non-uniform local triangle distortion; \(\mathcal{L}_{\text{nc}}\) enforces normal vector consistency across adjacent faces to eliminate surface creases and folds. With hyperparameters \(\lambda_1=\lambda_2=0.1\) and \(\lambda_{\text{reg}}=10\), the full pipeline converges in only 200 epochs on a single NVIDIA A100 GPU.
Key Experimental Results¶
Main Results¶
The evaluation benchmark comprises 40 demanding video sequences (20 sourced from Consistent4D, STAG4D, and V2M4, plus 20 challenging animal and human motion clips curated from Sketchfab). Competing baselines include DreamMesh4D (Gaussian-mesh hybrid optimization via SDS), Puppeteer (projection-based 2D keypoint tracking optimization), and V2M4 (per-frame generation with non-rigid frame-to-frame registration). Evaluation metrics include perceptual similarity LPIPS, semantic alignment CLIP, temporal video quality FVD, 3D geometric fidelity Uni3D (comparing 8,192 surface points against reference views), average runtime for 32 frames, and a 20-subject user preference study.
| Method | LPIPS ↓ | CLIP ↑ | FVD ↓ | Uni3D ↑ | Runtime (32 frames) ↓ | Preference ↑ |
|---|---|---|---|---|---|---|
| DreamMesh4D [NeurIPS 2024] | 0.1213 | 0.8643 | 931.50 | 0.2761 | ~60 min | 0.0% |
| Puppeteer [NeurIPS 2025] | 0.1143 | 0.8964 | 805.31 | 0.3009 | ~10 min | 6.8% |
| V2M4 [ICCV 2025] | 0.1025 | 0.8791 | 763.89 | 0.2963 | ~20 min | 8.5% |
| Ours (SFG, Trellis) | 0.0791 | 0.9442 | 544.16 | 0.3486 | ~5 min | 84.8% |
Ablation Study¶
The ablation study validates the critical roles of Shape Flow Guidance (\(\mathcal{L}_{\text{shape}}\)), the Advection-Diffusion temporal intervention (ADT), and the Post-Skinning Correction network (PSC).
| Config | LPIPS ↓ | CLIP ↑ | FVD ↓ | Uni3D ↑ | Note |
|---|---|---|---|---|---|
| Full Model | 0.0791 | 0.9442 | 544.16 | 0.3486 | Full proposed pipeline |
| w/o SFG | 0.1447 | 0.8232 | 1057.73 | 0.2907 | Severe collapse without 3D geometric supervision |
| w/o ADT | 0.0954 | 0.9215 | 654.82 | 0.3289 | Degradation due to inter-frame trajectory jitter |
| w/o PSC | 0.1015 | 0.9106 | 689.27 | 0.3205 | Limited by rigid LBS, missing fine surface dynamics |
Cross-generator generalization experiments further confirm that replacing the Trellis backbone with Hunyuan3D 2.1 (HY3D) yields equally superior results (LPIPS 0.0774, CLIP 0.9465, FVD 524.13, Uni3D 0.3531, 5 min runtime), verifying that the PDE-guided shape flow framework is fully model-agnostic.
Key Findings¶
- Explicit 3D guidance is essential to avoid geometric collapse: Removing \(\mathcal{L}_{\text{shape}}\) causes FVD to skyrocket from 544.16 to 1057.73 and Uni3D to plunge to 0.2907. This proves that 2D rendering losses alone cannot constrain 3D degrees of freedom in non-rigid settings.
- Fast convergence unlocks significant speedup: Because 3D shape flow provides direct spatial gradients without depth ambiguity, the optimization converges in only 200 epochs. With full-frame batch parallelization, the pipeline completes a 32-frame sequence in 5 minutes, representing a \(2\times\) to \(12\times\) speedup over existing optimization baselines.
- Post-skinning correction affords robust error tolerance: Pure linear blend skinning (w/o PSC) incurs a 0.028 drop in Uni3D score when confronted with complex animal morphology or imperfect kinematic skeletons. The residual MLP effectively compensates for articulation inaccuracies.
Highlights & Insights¶
- Training-free generative sampling intervention: Modulating sampling trajectories via an advection-diffusion PDE turns an off-the-shelf single-frame 3D generator into a temporally smooth 4D shape engine with zero model fine-tuning.
- Hierarchical decomposition of motion spaces: By decoupling motion into localized root-space deformation (guided by 3D geometry) and global trajectory (guided by 2D rendering), the framework circumvents conflicting gradient objectives between camera-frame translation and local body flexing.
- Production-ready asset output: Rather than outputting unstructured Gaussian clouds or neural radiance fields, the method produces rigged, textured, topology-preserving polygonal meshes ready for immediate integration into standard DCC tools and game engines.
Limitations & Future Work¶
- Strict assumption of topological constancy: The framework cannot model topology-altering phenomena such as surface tearing, fracture, or volumetric fluid splashing.
- Constrained to static camera viewpoints: The system assumes monocular video captured from a fixed camera, leaving the disentanglement of moving-camera perspectives and dynamic background SLAM for future exploration.
- Potential self-intersection under extreme deformation: While bidirectional point-to-mesh loss provides robust shape envelope matching, rapid foldings may occasionally induce subtle mesh self-intersections without explicit collision penalty terms.
Related Work & Insights¶
- vs DreamMesh4D: DreamMesh4D optimizes a hybrid Gaussian-mesh representation via slow Score Distillation Sampling (taking ~60 min) and produces noisy surfaces; the proposed method leverages direct 3D shape flow supervision on explicit meshes, running in 5 minutes with smooth surface topology.
- vs Puppeteer: Puppeteer projects 3D vertices to 2D to track video keypoints, which suffers from severe depth-axis ambiguity; the proposed method lifts 2D video to a 3D shape flow, guiding deformations directly in Euclidean 3D space and preventing axial distortion.
- vs V2M4: V2M4 generates independent frames and performs multi-step sequential non-rigid registration that easily breaks during rapid motions; the proposed method maintains unified temporal coherence via PDE intervention and decoupled skeletal rigging, providing far superior robustness and parallel efficiency.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Elegantly adapts continuous advection-diffusion PDEs to intervene in rectified flow sampling for training-free 3D shape flow distillation.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 40 complex video benchmarks, extensive baselines, rigorous ablation, user studies, and cross-generator verification.
- Writing Quality: ⭐⭐⭐⭐⭐ Well-structured narrative, lucid motivation, precise mathematical formulations, and polished illustrations.
- Value: ⭐⭐⭐⭐⭐ Bridges the critical gap between monocular video animation and production-ready rigged 3D assets with remarkable speed and fidelity.