Skip to content

JacobianAvatar: Temporally Consistent Semi-rigid Avatar Reconstruction from a Monocular Video

Conference: ECCV 2026
Paper: ECCV Official Page
Area: Human Understanding
Keywords: avatar reconstruction, neural Jacobian fields, semi-rigid deformation, Poisson solver, 3D Gaussian splatting

TL;DR

JacobianAvatar introduces hierarchical neural Jacobian fields integrated with a screened Poisson solver, signed distance-based normal regularization, and deformation-guided residual flow loss to reconstruct temporally consistent, high-fidelity animatable avatars with detailed clothing dynamics from a single monocular video.

Background & Motivation

Reconstructing photorealistic and animatable human avatars from a casual monocular video is a long-standing goal in computer vision and graphics. Modern approaches commonly adopt a semi-rigid deformation paradigm that pairs articulated skeletal motion via Linear Blend Skinning (LBS) with learned non-rigid local correctives to capture fine clothing wrinkles and pose-dependent geometry. However, most existing methods regress per-vertex offsets conditioned solely on instantaneous body poses. This setup inherently ignores history-dependent factors such as cloth inertia, contact forces, and motion velocity, while relying excessively on frame-independent photometric reconstruction losses. Consequently, these models suffer from temporal flickering, unnatural stretching, and "texture-copying artifacts" where 2D surface patterns are erroneously baked into the underlying 3D geometry.

The fundamental tension stems from the ill-posed nature of monocular capture: severe self-occlusions (such as armpits, inner thighs, and the back) and sparse single-view observations yield ambiguous surface gradients. Implicit representations like NeRF and Neural SDFs struggle with slow rendering speeds and geometric artifacts when extracting surface meshes via Marching Cubes, whereas explicit 3D Gaussian Splatting (3DGS) representations attached to canonical meshes are constrained by the rigidity of underlying LBS templates and lack intrinsic area- and orientation-preserving priors for unseen regions. Directly predicting free-form deformations across faces causes standard Poisson surface reconstruction to diverge in unobserved areas, breaking the global anatomical structure.

The angle of attack in JacobianAvatar is to parameterize semi-rigid clothing deformations as local deformation gradients using neural Jacobian fields on an explicit canonical mesh. Core idea: predict pose-dependent deformation gradients via coarse-to-fine hierarchical neural Jacobian fields, anchor unobserved surfaces using a screened Poisson solver and an SDF-based normal regularizer, and enforce cross-frame sub-pixel temporal consistency with a deformation-guided residual flow loss.

Method

Overall Architecture

The JacobianAvatar pipeline combines explicit canonical mesh optimization with hierarchical deformation modeling and 3DGS rendering. It begins by initializing a canonical mesh from the Momentum Human Rigs (MHR) template. A preprocessing step jointly refines canonical vertex positions and per-frame body poses via differentiable mesh rendering guided by normal maps and masks predicted by a pretrained vision foundation model (Sapiens), yielding a refined base mesh \(M^c\) and its twice-subdivided counterpart \(M^f\). Two multilayer perceptrons (MLPs) then predict face Jacobian matrices in a coarse-to-fine hierarchy: the coarse network captures large-scale posture deformations, while the fine network refines subtle wrinkles. The resulting Jacobian fields deform the mesh into pose space through a screened Poisson solver via differentiable Cholesky decomposition. Finally, 3D Gaussians are anchored to the deformed mesh faces, with their scale, local offsets, and residual dynamics trained under multi-view photometric and residual flow constraints for real-time, photo-quality animation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Monocular Video Frames + Initial MHR Canonical Mesh"] --> B["Preprocessing Stage<br/>Differentiable Mesh Optimization with Sapiens Priors"]
    B --> C["Hierarchical Neural Jacobian Fields<br/>Coarse-to-Fine MLPs Predicting Per-Face Gradients"]
    C --> D["Screened Poisson Solver<br/>Cholesky Decomposition for Continuous Vertex Positions"]
    D --> E["Geometric Constraints & Temporal Regularization<br/>SDF Normal Alignment + Residual Flow Loss"]
    E --> F["Mesh-Anchored 3DGS Rendering<br/>Face-Level Offsets & Photometric Output Optimization"]

Key Designs

1. Hierarchical Neural Jacobian Fields: Decoupling Global Dynamics from High-Frequency Wrinkles Directly regressing 3D vertex displacements often leads to self-intersections and mesh topological collapse under abrupt motions. Instead, JacobianAvatar parameterizes local deformations through deformation gradientsโ€”represented as a \(3 \times 3\) Jacobian matrix \(J_i\) for each triangle face. To capture both macroscopic body flexions and delicate garment folds, the model employs a two-tier hierarchical architecture. On the coarse mesh \(M^c\), a spatial tri-plane feature extractor \(\mathcal{E}^c\) and an MLP \(\psi^c\) predict the coarse Jacobian field \(J_i^c = \psi^c(\mathbf{z}_i^c, \theta)\) conditioned on pose \(\theta\). On the subdivided mesh \(M^f\), the fine MLP \(\psi^f\) ingests the concatenation of the coarse spatial feature \(\mathbf{z}_i^c\) and fine feature \(\mathbf{z}_i^f\), predicting high-frequency correctives \(J_i^f\). This coarse-to-fine schedule ensures stable convergence of the global body envelope before learning fine-grained dynamic wrinkles.

2. Screened Poisson Solver: Suppressing Geometric Divergence in Unobserved Regions Given a predicted Jacobian field \(\mathbf{J}\), reconstructing globally continuous vertex coordinates \(\Phi\) entails solving a Poisson equation. However, the standard Poisson formulation only minimizes gradient differences; in monocular sequences where body areas such as the back or armpits remain unseen, unconstrained faces undergo extreme stretching and topological distortions. To eliminate this ambiguity, JacobianAvatar incorporates a screening term that penalizes deviations from the resting canonical mesh \(\mathcal{V}\): $$ (L + \lambda_{\Phi} I)\Phi = \mathcal{A} \nabla \mathbf{J} + \lambda_{\Phi} \mathcal{V} $$ where \(L\) denotes the cotangent Laplacian matrix, \(I\) is the identity matrix, and \(\mathcal{A}\) is the face mass matrix. Setting \(\lambda_{\Phi}\) to 5 for the coarse mesh and 1 for the fine mesh provides an effective regularization anchor, solved efficiently through differentiable Cholesky decomposition, thereby preventing unseen vertices from drifting wildly.

3. Signed Distance-based Jacobian Regularization: Stabilizing Global Hull with Volumetric Geometry Priors Severe occlusions in monocular videos can cause face normals to flip or collapse inward during gradient descent. To safeguard structural integrity, JacobianAvatar constructs a 3D discrete Signed Distance Function (SDF) voxel grid from the initial coarse mesh \(M^c\). The spatial gradients of this static SDF grid provide reference surface normals for deformed vertices and face centroids. The model minimizes the cosine discrepancy between the predicted mesh normals and the spatial SDF gradients \(\mathbf{n}_{\text{SDF}}\): $$ \mathcal{L}{\text{SDF}} = \lambda}} \sum_{a_i \in \mathcal{F}} \left( 1 - \hat{\mathbf{n}i^a \cdot \mathbf{n}}}(x_i) \right) + \lambda_{\text{vertex}} \sum_{v_j \in \mathcal{V}} \left( 1 - \hat{\mathbf{n}j^v \cdot \mathbf{n}(v_j) \right) $$ By enforcing normal alignment with neighboring volumetric iso-surfaces, the regularizer ensures that unobserved regions maintain realistic anatomical curvatures rather than collapsing into noisy spikes.}

4. Deformation-Guided Residual Flow Loss: Enforcing Cross-Frame Sub-pixel Motion Consistency Frame-by-frame photometric losses cannot prevent high-frequency temporal flickering or texture sliding during motion. JacobianAvatar directly links 3D vertex kinematics to 2D optical flow fields. For consecutive frames at \(t\) and \(t-1\), the 3D vertices deformed by NJF and LBS are projected onto the image plane to construct a 2D projected displacement map \(\mathbf{W}^t\). Rather than comparing \(\mathbf{W}^t\) directly against noisy external optical flow, the method feeds \(\mathbf{W}^t\) as an initial motion hypothesis into a pretrained flow network \(\mathcal{F}\) (such as WAFT): $$ \Delta\mathbf{W}^t = \mathcal{F}(\mathbf{I}^t, \mathbf{I}^{t-1}, \mathbf{W}^t), \quad \mathcal{L}_{\text{residual}} = |\Delta\mathbf{W}^t|_1 $$ When the projected 3D deformation matches the true image motion, the predicted residual flow \(\Delta\mathbf{W}^t\) evaluates to zero. This formulation penalizes non-physical temporal jitter in self-occluded areas like the armpits and ensures sub-pixel motion coherence across the entire sequence.

Loss & Training

The network is optimized in sequential stages. During mesh optimization, the objective aggregates image reconstruction terms (\(\mathcal{L}_{\text{L1}}, \mathcal{L}_{\text{lpips}}, \mathcal{L}_{\text{ssim}}\)), Sapiens-guided normal loss \(\mathcal{L}_{\text{normal}}\), mask loss \(\mathcal{L}_{\text{mask}}\), the proposed \(\mathcal{L}_{\text{residual}}\), \(\mathcal{L}_{\text{SDF}}\), and Laplacian surface regularizers. In the subsequent 3DGS stage, 3D Gaussians are anchored to each face and optimized under scale, offset, and residual flow penalties. Training is conducted on a single NVIDIA RTX 6000 Ada GPU and takes approximately 10 hours per sequence.

Key Experimental Results

Main Results

JacobianAvatar was evaluated on three real-world monocular benchmarksโ€”MonoPerfCap, DNA-Rendering, and NeuManโ€”as well as the synthetic SynWild dataset with ground-truth meshes. Performance was benchmarked against leading implicit SDF methods (Vid2Avatar, LSAvatar, FacAvatar) and explicit 3DGS baselines (ExAvatar, GoMAvatar). Metrics include rendering quality (PSNR, SSIM, LPIPS \(\times 100\)) and geometric reconstruction precision (Chamfer Distance \(CD \times 1000\), Normal Error \(NE\), and \(F_1\) scores at 1 cm and 2 cm thresholds).

Dataset Metric Ours Prev. SOTA Gain / Advantage
MonoPerfCap PSNR โ†‘ 31.83 30.47 (ExAvatar) +1.36 dB
MonoPerfCap SSIM โ†‘ 0.978 0.980 (ExAvatar) -0.002
MonoPerfCap LPIPS (ร—100) โ†“ 1.67 1.95 (GoMAvatar) -0.28
DNA-Rendering PSNR โ†‘ 29.90 29.46 (Vid2Avatar) +0.44 dB
DNA-Rendering LPIPS (ร—100) โ†“ 2.19 2.37 (LSAvatar) -0.18
SynWild CD (ร—1000) โ†“ 2.46 2.55 (LSAvatar) -0.09
SynWild Normal Error (NE) โ†“ 0.091 0.091 (Vid2Avatar) Tied for best
SynWild F1@1cm โ†‘ 0.397 0.346 (Vid2Avatar) +0.051 (+14.7%)
SynWild F1@2cm โ†‘ 0.681 0.647 (Vid2Avatar) +0.034 (+5.3%)

Ablation Study

On the 00000_random scene of the SynWild dataset, the authors ablated each critical component to measure its isolated geometric impact:

Config CD (ร—1000) โ†“ NE โ†“ Note
Full model 3.01 0.102 Best overall geometric and normal fidelity
w/o \(\mathcal{L}_{\text{SDF}}\) 3.51 0.113 Without SDF regularizer, unobserved back regions distort heavily; CD degrades by 16.6%
w/o \(\mathcal{L}_{\text{residual}}\) 3.01 0.106 Without residual flow, self-occluded armpits suffer motion noise; normal error increases
w/o Screened Poisson solver 3.28 0.106 Standard Poisson causes unconstrained surface stretching in unobserved regions
w/o Coarse-to-fine Strategy 2.96 0.107 Lacking coarse guidance leads to high-frequency surface ripple noise (higher NE)

Key Findings

  • SDF regularization provides essential hull stability: Omitting \(\mathcal{L}_{\text{SDF}}\) worsens Chamfer Distance from 3.01 to 3.51, showing that geometric priors from volume gradients are vital for anchoring single-view unobserved boundaries.
  • Residual flow loss eliminates localized occlusion artifacts: While \(\mathcal{L}_{\text{residual}}\) has minimal impact on global Chamfer Distance (affecting only a small fraction of vertices), it significantly reduces Normal Error (from 0.106 to 0.102) by smoothing motion jitter around the armpits and thighs.
  • Explicit mesh formulation prevents texture copying: Unlike implicit SDF representations that embed 2D clothing textures into reconstructed surface normals, JacobianAvatar preserves clean, crease-accurate geometry without baking photometric details into the mesh.

Highlights & Insights

  • Adapting Neural Jacobian Fields to Monocular Captures: While NJF was originally developed for fully supervised 3D mesh alignment, this work demonstrates that coupling NJF with a screened Poisson formulation enables stable deformation estimation under weak, single-view 2D supervision.
  • Residual Flow Optimization as a Sub-pixel Regularizer: Instead of supervising deformations directly with noisy 2D flow vectors, measuring the residual flow \(\Delta\mathbf{W}\) through a pretrained optical flow architecture creates a closed-loop refinement mechanism that penalizes temporal discrepancies at sub-pixel accuracy.
  • Hybrid Mesh-3DGS Synergy: Anchoring 3D Gaussians onto a dynamically deformed explicit mesh combines the structural integrity and topology preservation of mesh mechanics with the photorealism and real-time efficiency of 3DGS.

Limitations & Future Work

  • Handling Extremely Loose Apparel: The canonical deformation pipeline still relies on skinning weights from an underlying nude human template (MHR / SMPL-X). Garments with independent physics, such as flowing dresses or open coats, remain difficult to model purely through body-driven Jacobian fields.
  • Absence of Expressive Facial & Hand Rigging: The current implementation concentrates on body and garment deformations, omitting detailed facial blendshapes or fine-grained hand articulated models.
  • Future Directions: Integrating physical cloth simulation priors directly into the unsupervised Jacobian training and coupling full-body rigging with expressive facial models will further enhance fidelity.
  • vs. Vid2Avatar / LSAvatar / FacAvatar: These implicit NeRF/SDF approaches suffer from slow rendering speeds, Marching Cubes extraction artifacts, and texture-copying entanglements. JacobianAvatarโ€™s explicit formulation yields smooth, noise-free geometry with a 14.7% higher F1@1cm score on SynWild.
  • vs. ExAvatar / GoMAvatar: Mesh-anchored 3DGS baselines rely heavily on rigid template bindings or basic displacement offsets, struggling with intricate clothing dynamics. JacobianAvatarโ€™s screened Poisson deformation and residual flow yield superior novel-view rendering quality (31.83 dB vs. 30.47 dB PSNR on MonoPerfCap).

Rating

  • Novelty: โญโญโญโญโ˜† [Novel application of screened Neural Jacobian Fields and deformation-guided residual flow to monocular avatar reconstruction]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive evaluations across 4 benchmarks evaluating both 2D rendering fidelity and 3D surface geometry]
  • Writing Quality: โญโญโญโญโญ [Clear motivation, solid mathematical exposition, and disciplined ablation analyses]
  • Value: โญโญโญโญโ˜† [Provides an effective blueprint for combining explicit mesh mechanics with 3D Gaussian Splatting for real-time digital humans]