RegHead: Non-Humanoid Head Blendshapes via Feed-Forward Registration¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://snap-research.github.io/RegHead/
Area: 3D Vision
Keywords: Non-Humanoid Avatars, Blendshapes, Feed-Forward Registration, Stochastic Anchors, Real-Time Retargeting
TL;DR¶
To overcome the absence of vertex correspondence and paired expression annotations for non-humanoid 3D heads, RegHead curates a dataset of approximately 20k identities under a unified semantic expression vocabulary, introducing a topology-agnostic dense stochastic anchor representation and a coarse-to-fine feed-forward registration network that predicts corresponded blendshapes without ground-truth deformation supervision and enables real-time retargeting from human tracking.
Background & Motivation¶
Creating animatable 3D non-humanoid head avatars—such as stylized creatures, animals, and fictional characters—is increasingly critical for AR telepresence, virtual avatars, and gaming. Conventional linear blendshape models provide an interpretable, low-dimensional interface that naturally supports real-time facial animation and cross-identity retargeting via a fixed expression vocabulary. However, existing automated blendshape pipelines are predominantly designed around human-centric parametric priors (such as 3DMMs and FLAME). When applied to non-humanoid heads with diverse snout geometries, stylized eyes, and atypical jaw articulations, parametric anatomical assumptions fail completely, forcing practitioners to rely on labor-intensive manual rigging or slow per-instance non-rigid registration optimization.
Building semantic blendshapes for diverse non-humanoid heads in a scalable feed-forward manner faces three core bottlenecks. First, large-scale paired expression observations with consistent semantic labels across identities are virtually nonexistent in the non-humanoid domain, as generic image or video generative models cannot reliably reproduce a target expression without identity drift. Second, modern image-conditioned 3D generative pipelines can produce expression-specific meshes, but these raw shapes lack vertex correspondence across expressions; recovering correspondence via per-sequence test-time optimization remains prohibitively slow, taking minutes to hours per identity. Third, non-humanoid facial motions are characterized by highly localized, species-specific deformations (e.g., subtle eyelid movement and specialized mouth stretching) that coarse global deformation networks or sparse skeletal rigs fail to preserve.
The angle of attack in RegHead is to liberate the pipeline from predefined skeletons and human templates, expanding an artist-curated vocabulary into a large-scale synthetic dataset and amortizing per-instance non-rigid registration into a fast feed-forward network. The core idea is to curate a 20k non-humanoid dataset with a shared expression vocabulary, introduce a topology-agnostic dense stochastic anchor motion representation, and train a coarse-to-fine feed-forward registration network with differentiable Gaussian rendering supervision to predict corresponded blendshapes in a single forward pass.
Method¶
Overall Architecture¶
The input to RegHead consists of a neutral reference mesh \(\tilde{\mathcal{M}}_0\) and a set of unregistered target expression meshes \(\{\tilde{\mathcal{M}}_t\}_{t=1}^T\) for an identity under a predefined expression vocabulary \(\mathcal{E} = \{e_t\}_{t=1}^T\). The output is a set of corresponded semantic blendshapes \(\{B_t\}_{t=0}^T\) sharing identical mesh topology and vertex ordering with the neutral shape.
The feed-forward pipeline operates in three coupled stages: First, input meshes are voxelized to extract multi-modal voxel tokens enriched with unprojected multi-view semantic segmentation cues. Second, dense stochastic anchors are sampled over the neutral surface to serve as unconstrained local motion controllers, and a global matcher aggregates scene-level context via cross-attention to predict coarse anchor warps. Third, a structured multi-radius stencil gathers local neighborhood tokens around the coarsely warped anchors, allowing a local matcher to compute fine residual transformations before deformations are propagated to dense surface query points via Linear Blend Skinning (LBS), trained end-to-end under differentiable point-based Gaussian rendering.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Neutral & Target Expression Meshes<br/>Voxelization & Multi-View Feature Unprojection"] --> B["Dense Stochastic Anchor Sampling<br/>Template-Free Surface Control Nodes & LBS Weights"]
B --> C["Global Coarse Matching Stage<br/>Long-Range Cross-Attention for Large Support Motions"]
C --> D["Structured Multi-Radius Stencil Gathering<br/>Validity-Aware Neighbor Feature Extraction"]
D --> E["Local Refinement & Anchor Graph Fusion<br/>Local Cross-Attention for Residual Similarity Transforms"]
E --> F["Output: Corresponded Semantic Blendshapes<br/>Linear Blendshape Animation & Real-Time Face Retargeting"]
Key Designs¶
1. Dense Stochastic Anchor Motion Representation: Template-Free Localized Deformation To handle arbitrary non-humanoid head anatomies without hand-crafted skeletons or predefined vertex templates, RegHead defines motion over a set of \(K=3000\) stochastic anchors sampled directly on the neutral surface. Crucially, anchor positions are resampled randomly at every training iteration rather than optimized as static parameters. This stochastic resampling enforces anchor-layout invariance, encouraging the network to generalize across arbitrary control point distributions. Deformation from anchors to dense surface query points \(Q=\{q_j\}_{j=1}^N\) is achieved via Linear Blend Skinning using precomputed normalized inverse Mahalanobis distances over the \(K'=10\) nearest anchors. This sparse-to-dense formulation maintains linear computational complexity with respect to the anchor count while providing expressive local support.
2. Coarse-to-Fine Feed-Forward Matching: Bridging Global Articulation and Micro-Deformations Predicting non-rigid correspondence directly from local voxel neighborhoods fails when large-scale motions (e.g., wide jaw openings) exceed the local search radius, causing queries to map to incorrect anatomical regions. Conversely, pure global self-attention lacks spatial precision for subtle eyelid movements. RegHead addresses this via a two-stage coarse-to-fine architecture. In the first stage, the Global Matcher performs cross-attention between neutral anchor tokens and subsampled global target voxel tokens, predicting an initial coarse transform \(\mathcal{T}_k^{t,0}\) that brings each anchor into the correct convergence basin. In the second stage, a Local Matcher constructs a structured multi-radius stencil around each coarsely aligned anchor, sampling thin concentric shell bands in voxel space while masking out empty background voxels via a validity mask. Local cross-attention then predicts residual transformations \(\Delta \mathcal{T}_k^t\), yielding the final composite transform \(\mathcal{T}_k^t = \Delta \mathcal{T}_k^t \circ \mathcal{T}_k^{t,0}\).
3. Correspondence-Free Supervision via Differentiable Gaussian Splatting Because non-humanoid meshes across different synthetic expressions have no ground-truth vertex correspondences, direct 3D displacement supervision is impossible. RegHead overcomes this by treating both the deformed query point cloud \(\hat{\mathbf{Q}}^t\) and target surface samples \(\mathbf{P}^t\) as isotropic 3D Gaussian splats, rendering them under randomly sampled camera views \(\pi\). The network minimizes photometric color differences, perceptual loss (LPIPS), and semantic segmentation loss between rendered projections: $\(\mathcal{L}_{\mathrm{rend}} = \mathbb{E}_{\pi}\Big[ \|\mathcal{R}_{\mathrm{rgb}}(\hat{\mathbf{Q}}^t;\pi) - \mathcal{R}_{\mathrm{rgb}}(\mathbf{P}^t;\pi)\|_1 + \lambda_{\mathrm{lpips}}\mathcal{L}_{\mathrm{lpips}} + \lambda_{\mathrm{seg}}\|\mathcal{R}_{\mathrm{seg}}(\hat{\mathbf{Q}}^t;\pi) - \mathcal{R}_{\mathrm{seg}}(\mathbf{P}^t;\pi)\|_1 \Big]\)$ Combined with an auxiliary Chamfer distance loss on coarsely deformed anchors \(\mathcal{L}_{\mathrm{coarse}} = \mathrm{CD}(\mathbf{A}^t, \mathbf{P}_{\mathrm{sub}}^t)\), this point-based differentiable rendering avoids mesh rasterization artifacts and guides continuous surface deformation entirely without pairwise ground-truth correspondences.
4. Real-Time Expression Retargeting via Lightweight Translation Network Once topological correspondence is established across the semantic blendshape basis, real-time animation is achieved by mapping human performance tracking signals to non-humanoid blendshape weights. An off-the-shelf facial tracker extracts frame-wise 3DMM coefficients \(\mathbf{w}^{\mathrm{hum}}\) and head pose from a monocular webcam video. A lightweight offline MLP \(g_\psi\) translates these human expression coefficients into non-humanoid blendshape weights \(\mathbf{w}^t = g_\psi(\mathbf{w}^{\mathrm{hum}})\). Because \(g_\psi\) requires only a single forward pass without per-frame numerical optimization, the resulting avatar reproduces both tracked rigid head poses and fine-grained facial expressions at interactive real-time frame rates.
Loss & Training¶
The overall training objective combines the differentiable rendering loss \(\mathcal{L}_{\mathrm{rend}}\) with the coarse geometric alignment loss \(\mathcal{L}_{\mathrm{coarse}}\): $\(\mathcal{L} = \mathcal{L}_{\mathrm{rend}} + \lambda_{\mathrm{coarse}} \mathcal{L}_{\mathrm{coarse}}\)$ The model is trained end-to-end using AdamW (\(\beta_1=0.9, \beta_2=0.999\), weight decay \(10^{-4}\)) with a learning rate of \(1\times 10^{-4}\) on 32 NVIDIA A100 (80GB) GPUs using mixed precision for 2000 epochs, reaching convergence around epoch 1500.
Key Experimental Results¶
Main Results¶
The method is evaluated on a test split of 200 unseen non-humanoid identities with 7 expressions each. RegHead is benchmarked against ActionMesh (feed-forward general mesh animation baseline), V2M4 (optimization-based 4D registration), and T2Bs (generative video-to-blendshape alignment). Metrics include image-space rendering quality (PSNR, SSIM, LPIPS), multi-view visible first-hit surface geometry (Chamfer distance CD, depth error D-Err, normal consistency N-Cons, silhouette IoU Sil. IoU), and per-identity execution time.
| Method | PSNR ↑ | SSIM ↑ | LPIPS ↓ | CD ↓ | D-Err. ↓ | N-Cons. ↑ | Sil. IoU ↑ | Runtime (Time) ↓ |
|---|---|---|---|---|---|---|---|---|
| ActionMesh (feed-forward) | 22.2731 | 0.8989 | 0.0654 | 0.00513 | 0.01032 | 0.9103 | 0.9739 | \(\sim\)5 s |
| V2M4 (per-instance opt.) | 22.7849 | 0.9056 | 0.0660 | 0.00959 | 0.01827 | 0.9039 | 0.9567 | \(\sim\)100 s |
| T2Bs (generative opt.) | 23.4349 | 0.9064 | 0.0568 | 0.00387 | 0.00711 | 0.9160 | 0.9851 | \(\sim\)1800 s |
| RegHead (Ours) | 23.8421 | 0.9078 | 0.0551 | 0.00358 | 0.00613 | 0.9192 | 0.9865 | \(\sim\)5 s |
Ablation Study¶
Ablations examine design choices in the stochastic anchor representation (Table a) and the feed-forward registration pipeline (Table b):
| Component Group | Configuration | PSNR ↑ | SSIM ↑ | LPIPS ↓ | Findings & Analysis |
|---|---|---|---|---|---|
| Motion Representation | Full model (Ours, K=3000 stochastic) | 23.842 | 0.9078 | 0.0551 | Optimal trade-off between coverage and local fidelity |
| Fixed anchors | 23.323 | 0.9044 | 0.0610 | PSNR drops by 0.52 dB, showing that stochastic resampling prevents layout overfitting | |
| Sparse anchors (300 anchors) | 22.226 | 0.8987 | 0.0629 | Severe drop across all metrics due to insufficient local control capacity | |
| w/o segmentation channel (w/o seg.) | 23.391 | 0.9056 | 0.0575 | Loss of unprojected eye semantic prior harms localized boundary precision | |
| Registration Architecture | Full model (Ours, Global + Local) | 23.842 | 0.9078 | 0.0551 | Coarse-to-fine hierarchy achieves superior accuracy |
| w/o local matching (w/o local) | 22.138 | 0.8963 | 0.0635 | Drastic performance degradation (PSNR drops 1.70 dB); local matching is critical for fine expressions | |
| w/o global matching (w/o global) | 23.660 | 0.9062 | 0.0577 | Noticeable drop; global initialization prevents local matching from getting trapped in local minima | |
| kNN neighborhood (kNN neighbors) | 23.055 | 0.9044 | 0.0574 | Standard kNN clusters unevenly; structured multi-radius stencil ensures balanced spatial context |
Key Findings¶
- Local matching is the primary driver of fine expression accuracy: Removing the local matching module causes the largest performance drop across all ablations (PSNR plummets by 1.704 dB), demonstrating that global cross-attention alone is insufficient to resolve localized non-rigid facial nuances.
- Stochastic resampling prevents layout overfitting: Switching from dynamic stochastic anchors to fixed anchor coordinates incurs a 0.52 dB penalty in PSNR, confirming that random sampling forces the network to learn smooth, layout-invariant deformation fields.
- Superior speed-accuracy Pareto frontier: RegHead achieves higher rendering fidelity and visible surface accuracy than optimization-based baselines (V2M4 and T2Bs) while operating at \(\sim\)5 seconds per identity—over 360\(\times\) faster than T2Bs (\(\sim\)1800 s).
Highlights & Insights¶
- Template-free parameterization for non-humanoid morphology: By departing from fixed skeletons and FLAME/3DMM topology, dense stochastic anchors establish a universal deformation interface adaptable to arbitrary animal and fantasy facial geometries.
- Structured multi-radius stencil sampling: Gathering local voxel tokens via concentric spherical shells with validity masking proves far more robust than Euclidean kNN, preventing token clustering and background noise in empty voxel regions.
- Unsupervised registration via differentiable point rendering: Demonstrating that high-quality non-rigid geometric correspondence can be learned purely through multi-view 2D rendering losses provides a versatile blueprint for non-rigid 3D registration tasks lacking ground-truth point pairings.
Limitations & Future Work¶
- Anatomical extrapolation for extreme long-tail species: For morphologies deviating drastically from typical vertebrate anatomy (e.g., multi-eyed creatures or insect mandibles), semantic expression definitions become ambiguous, occasionally yielding unnatural deformation transfers.
- Sensitivity to upstream 3D generation artifacts: Because RegHead acts as a feed-forward registration model over raw reconstructed expression meshes, topological holes, severe surface noise, or identity drift present in upstream generations can propagate into the predicted blendshapes.
- Future directions: Integrating contact-aware topological tearing to handle open-to-closed mouth transitions cleanly, and exploring joint self-supervised skeletal binding with dense surface skinning.
Related Work & Insights¶
- vs ActionMesh [31]: ActionMesh targets generic object motion and whole-body dynamics, lacking explicit local feature matching; consequently, it fails to capture delicate facial expressions such as subtle eyelid closures. RegHead achieves significantly higher geometric precision and expression fidelity in comparable feed-forward runtimes (\(\sim\)5 s).
- vs V2M4 [4] & T2Bs [22]: V2M4 and T2Bs rely on expensive per-instance optimization loops (taking from several minutes to half an hour per identity) and frequently get trapped in local geometric minima. RegHead amortizes non-rigid registration into a single forward pass, slashing processing time to 5 seconds while outperforming both in surface depth accuracy and silhouette IoU.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates the first feed-forward blendshape generation framework for non-humanoid avatars, introducing effective dense stochastic anchors and structured stencil matching.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 200 identities combining image rendering metrics, visible surface geometry, thorough ablations, and real-time retargeting.
- Writing Quality: ⭐⭐⭐⭐⭐ Highly organized and technically rigorous, clearly dissecting core bottlenecks and architectural decisions.
- Value: ⭐⭐⭐⭐⭐ Provides an efficient, practical pipeline for avatar creation in gaming, VR/AR, and social media, significantly lowering technical barriers for non-humanoid character animation.