SplatCtrlA: Generalizable Single Image to Fully Controllable 3D Avatar¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Area: Human Understanding
Keywords: 3D Gaussian Splatting, Single-Image Human Reconstruction, Controllable Avatar, SMPL-X, Feed-forward Generation
TL;DR¶
Addressing texture degradation, structural distortion, and uncoordinated facial/hand dynamics in single-image drivable avatar creation, SplatCtrlA presents an end-to-end feed-forward framework built on SMPL-X UV Gaussian parameterization, powered by a 28.5k 3D-consistent synthetic Gaussian avatar dataset, a decoupled input-aware decoder, and a Gaussian-blending facial enhancement module to achieve second-level reconstruction of photorealistic, fully controllable 3D avatars.
Background & Motivation¶
Reconstructing drivable, photorealistic 3D human avatars from a single unconstrained photograph is a foundational capability across immersive virtual reality, telepresence, gaming, and embodied human-robot interaction. However, monocular 3D avatar generation represents a fundamentally ill-posed inverse problem. The single input viewpoint contains inherently incomplete appearance cues, making non-rigid deformation modeling, self-occluded texture inpainting, and nuanced articulation of facial micro-expressions and fine hand gestures exceptionally difficult to generalize across unseen subjects. Recent 2D diffusion-based video animation methods synthesize convincing dynamic motions, but they suffer from severe identity drift, temporal flickering over extended sequences, and heavy computational latency during iterative multi-step denoising, lacking underlying rigid 3D geometric consistency.
Semi-explicit 3D representations such as 3D Gaussian Splatting (3DGS) provide an ideal foundation due to real-time differential rasterization and strict multi-view coherence. Nevertheless, mainstream single-image 3D reconstruction pipelines remain confined to static geometries without animatable rigs. While personalized optimization methods incorporating diffusion priors can recover articulated avatars, they require hours of per-identity fitting, rendering them unsuitable for interactive deployment. Recently, feed-forward large reconstruction models (LRMs) have been adapted for human avatars (e.g., IDOL, LHM, GUAVA). Yet, these approaches face dual bottlenecks: first, existing synthetic training datasets produced via 2D video diffusion models lack multi-view 3D rigidity, introducing inter-frame geometric jitter and severe finger artifacts that degrade network training; second, conventional image encoders inevitably discard high-frequency textures, and standard uncalibrated mesh back-projections suffer from acute misalignments around loose clothing and silhouettes while neglecting expressive facial and hand micro-structures.
This paper's angle of attack is to tightly bind parametric SMPL-X priors with 3DGS in UV parameter space, eliminate synthetic data inconsistency at its source by compiling an explicit 3DGS dataset, and design an input-aware decoding mechanism that directly injects input imagery into the regression pipeline. Core idea: SplatCtrlA constructs an SMPL-X UV Gaussian feed-forward framework trained on 28.5k geometrically consistent 3D avatars, deploying a decoupled, semantic-guided input-aware decoder alongside a zero-convolution Gaussian-blending facial module and hand-focused Laplacian regularization to achieve single-image, whole-body, fully controllable avatar generation in seconds.
Method¶
Overall Architecture¶
SplatCtrlA operates through an efficient feed-forward paradigm that maps a single full-body image \(I\) and a cropped, super-resolved head image \(I_{head}\) into a canonical 3D Gaussian field represented as UV attribute maps, which can subsequently be articulated into arbitrary target poses and expressions via Linear Blend Skinning (LBS).
The complete inference pipeline comprises four primary stages: First, a frozen Sapiens-1B vision foundation model extracts global body and local facial tokens, which are aligned into the SMPL-X UV feature space \(F_{uv}\) by a scaling UV Transformer. Second, the Input-aware Decoder executes a two-step decoupled regression: an initial UV Mesh Inverse projects tracked canonical mesh vertices to sample visible input colors and predict Gaussian geometric displacements, followed by a semantic-guided UV Gaussian Inverse that projects deformed 3D Gaussian positions to condition color and scale decoding. Third, a specialized high-resolution head avatar model is seamlessly merged via a StyleUNet-based Gaussian Blending Network equipped with zero convolution. Finally, end-to-end training is supervised by view-space photometric losses and explicit Gaussian geometric regularizations.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Single Full-body Image + Cropped Super-resolved Head"] --> B["UV Gaussian Parameterization & Dual-Branch Feature Encoding<br/>Sapiens token extraction + UV Transformer feature alignment"]
B --> C["Decoupled Input-aware Decoding<br/>UV Mesh Inverse for geometry + Semantic-guided Gaussian Inverse for color"]
C --> D["Gaussian Blending-based Facial Enhancement<br/>Dedicated head model + Zero-convolution StyleUNet integration"]
D --> E["Gaussian Geometric Regularization<br/>Position/scale bounding + 100x hand Laplacian smoothing"]
E --> F["Canonical UV Gaussian Avatar<br/>Arbitrary body pose, gesture, and expression animation via LBS"]
Key Designs¶
1. UV Gaussian Parameterization & Dual-Branch Feature Encoding: Unifying 3D Radiance Fields and 2D Neural Representations in UV Space
Directly regressing unstructured 3D point clouds or spatial voxels poses significant hurdles for deep convolutional or attention-based architectures. To combine the rendering speed of 3D Gaussian Splatting with the structural advantages of 2D generative networks, the framework defines 3D Gaussians over the 2D UV unwrapping space of the parametric SMPL-X mesh. Each Gaussian primitive is parameterized as \(G = \{\mu, \alpha, r, s, c\}\), denoting center position \(\mu\), opacity \(\alpha\), rotation quaternion \(r\), 3D scale \(s\), and RGB color \(c\). Rather than regressing unconstrained coordinates, the model exploits SMPL-X surface topology by learning local positional and rotational offsets relative to canonical mesh anchors \(\hat{\mu}\) and \(\hat{r}\), formulated as \(\mu = \hat{\mu} + \delta \mu\) and \(r = \hat{r} \cdot \delta r\).
To preserve subtle facial cues that are otherwise diluted in full-body image tokens, the encoding stage employs a dual-branch strategy. Full-body input \(I\) and facial crop \(I_{head}\) are concurrently processed through a frozen Sapiens-1B backbone to yield tokenized feature representations: $\(F = \mathcal{E}_{sapiens}(I) \oplus \mathcal{E}_{sapiens}(I_{head})\)$ The combined global token \(F\) is concatenated with learnable positional embeddings \(F_{pos}\) along the patch channel and passed through a \(D\)-layer UV Transformer. Self-attention layers naturally aggregate long-range context across the human body, facilitating accurate hallucination of occluded and posterior body regions before a CNN upsampling block produces the high-resolution feature map \(F_{uv} \in \mathbb{R}^{768 \times 768 \times C}\).
2. Decoupled Input-aware Decoding: Preserving High-Frequency Appearance via Mesh and Semantic-Guided Gaussian Inversion
Relying exclusively on deep latent features for color regression inevitably leads to over-smoothed results and loss of fine garment patterns. While inverse texture mapping can directly transfer pixels from the input image onto UV space, naive projection based solely on a parametric mesh fails due to loose clothing offsets, silhouette misalignments, and tracking errors. To overcome this limitation, the authors design a decoupled two-step decoding strategy.
The initial stage performs UV Mesh Inverse for geometry decoding: for each UV pixel corresponding to Gaussian \(g_k\), its anchor position \(\hat{\mu}_k\) on the tracked SMPL-X mesh is rasterized, projected onto input image \(I\), and gated by a visibility mask \(M_{uv}\) to extract the initial inverse RGB map \(\hat{I}_{uv}\). The geometry decoder \(D_{geo}\) takes \(F_{uv} \oplus \hat{I}_{uv}\) as input to predict the Gaussian deformation field \(\delta \mu\) and opacity \(\alpha\). In the subsequent stage, because the actual Gaussian centers \(\hat{\mu} + \delta \mu\) reside on the clothing/hair surface rather than the bare mesh, a calibrated UV Gaussian Inverse re-projects these deformed 3D points back to the image plane to extract color conditioning map \(I_{uv}\). To prevent silhouette edge bleed where background pixels contaminate body boundaries due to tracking noise, 2D semantic guidance \(S\) is enforced: each Gaussian carries a body-part semantic tag, and erroneous projections are dynamically snapped to the nearest pixel sharing identical semantics. The color decoder \(D_{color}\) then ingests \(F_{uv} \oplus I_{uv}\) to predict rotational residual \(\delta r\), scale \(s\), and color \(c\).
3. Gaussian Blending-based Facial Enhancement: Local Prior Specialization via Zero-Convolution StyleUNet
Full-body synthetic datasets naturally provide insufficient spatial resolution for the facial region, which occupies only a minor pixel fraction of full-body frames. Consequently, generic full-body avatars struggle to synthesize crisp teeth, eye gaze, and expressive dynamic wrinkles required for realistic talking avatars.
To resolve this limitation, SplatCtrlA incorporates a specialized head reconstruction model trained exclusively on 5,253 high-resolution real human head video sequences. The head model mirrors the whole-body architecture but reduces the UV Transformer depth to \(D/2\) and processes only head tokens. To eliminate boundary discontinuities and color seams between the body and head models across the neck region, the authors introduce a lightweight Gaussian Blending Network. Composed of a StyleUNet followed by a zero convolution layer, it takes the concatenated UV attribute maps from both models under a facial UV mask and outputs a seamlessly blended final UV attribute map. The zero-initialized convolution guarantees that the network predicts zero residual offsets at initialization, preserving the stable body avatar while progressively learning refined boundary blending without corrupting global structure.
4. Gaussian Geometric Regularization: Stabilizing Expressive Articulation via Scale Bounds and Hand-Weighted Laplacian Smoothing
Monocular depth ambiguity frequently leads unconstrained Gaussians into degenerate solutions, producing needle-like spiky ellipsoids, non-physical floating artifacts, or splintered digits under large driving deformations. SplatCtrlA introduces three geometric regularizers to stabilize the canonical Gaussian field.
First, a displacement threshold loss \(\mathcal{L}_{pos}\) enforces that Gaussian centers remain bounded within an \(\epsilon_{pos}\) neighborhood around the SMPL-X surface: $\(\mathcal{L}_{pos} = \frac{1}{N_{gs}} \sum_{i=1}^{N_{gs}} \max(\|\delta \mu_i\|_2 - \epsilon_{pos}, 0)\)$ Second, an anisotropic scale ratio constraint \(\mathcal{L}_{sca}\) penalizes extreme needle-like shapes: $\(\mathcal{L}_{sca} = \frac{1}{N_{gs}} \sum_{i=1}^{N_{gs}} \left( \frac{s_i^{max}}{s_i^{min}} - \epsilon_{sca} \right)\)$ Third, a mesh-topology-guided Laplacian smoothing term \(\mathcal{L}_{lap}\) enforces spatial continuity over position offsets \(\delta \mu\), scales \(s\), and colors \(c\) across neighboring Gaussians. Crucially, recognizing that hand regions are highly prone to structural collapse during animation, the authors scale the hand Laplacian loss weight by a factor of 100 (\(\times 100\)), ensuring strictly coherent finger articulations during dynamic re-posing.
Loss & Training¶
The framework is optimized end-to-end via differential rendering under target driving views and poses. The total objective integrates photometric supervision and geometric regularizations: $\(\mathcal{L}_{total} = \mathcal{L}_{photometric} + \mathcal{L}_{reg}\)$ The photometric term computes pixel-level L1 loss and perceptual LPIPS loss against ground-truth frames: $\(\mathcal{L}_{photometric} = \lambda_{rgb}\mathcal{L}_1(I_{gt}, I_{pred}) + \lambda_{per}\mathcal{L}_{lpips}(I_{gt}, I_{pred})\)$ The regularization objective unifies positional, scale, and Laplacian terms: $\(\mathcal{L}_{reg} = \lambda_{pos}\mathcal{L}_{pos} + \lambda_{sca}\mathcal{L}_{sca} + \lambda_{lap}\mathcal{L}_{lap}\)$
The model is trained on 8 NVIDIA A100 GPUs over 5 days with batch size 8. Optimization utilizes Adam with an initial learning rate of \(4 \times 10^{-4}\) and a 6,000-step warmup. The UV resolution is set to \(768 \times 768\), managing approximately 320,000 Gaussians. Regularization bounds are set to \(\epsilon_{pos} = 0.05\) and \(\epsilon_{sca} = 5\), with loss balancing weights \(\lambda_{rgb} = 30\), \(\lambda_{per} = 10\), \(\lambda_{pos} = 10\), \(\lambda_{sca} = 50\), and \(\lambda_{lap} = 10^4\). The head model and Gaussian blending network are subsequently trained for 15,000 iterations each, using a blending learning rate of \(4 \times 10^{-6}\).
To provide solid 3D geometric supervision, the authors constructed the 28.5k-identity Gaussian Avatar Dataset. High-diversity portrait images synthesized via Flux were animated using MimicMotion video diffusion to supply multi-pose dynamic frames; individual 3D Gaussian avatars were rapidly reconstructed via per-identity fitting, enabling multi-view, multi-pose rendering with strict 3D consistency. The training corpus is augmented by THuman2.1 (real 3D scans), HuGe100K (curated subset), and 5,253 real head video captures to maximize generalization.
Key Experimental Results¶
Main Results¶
Quantitative benchmarking is conducted across three public benchmarks: THuman2.0 test set (100 real 3D scan subjects), NeuMan test set (all 6 video sequences), and Seamless Interaction (100 randomly sampled high-fidelity motion sequences). Performance is benchmarked against leading feed-forward Gaussian avatar models: IDOL, GUAVA, and LHM.
| Dataset | Method | MSE(\(\times 10^{-3}\)) ↓ | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|---|---|---|---|---|
| THuman2.0 | IDOL | 10.841 | 19.995 | 0.912 | 0.0907 |
| GUAVA | 9.727 | 20.239 | 0.886 | 0.0765 | |
| LHM | 9.951 | 20.276 | 0.917 | 0.0858 | |
| Ours (SplatCtrlA) | 6.440 | 22.542 | 0.931 | 0.0653 | |
| NeuMan | IDOL | 7.262 | 22.072 | 0.961 | 0.0424 |
| GUAVA | 7.365 | 21.875 | 0.904 | 0.0666 | |
| LHM | 6.469 | 23.139 | 0.964 | 0.0332 | |
| Ours (SplatCtrlA) | 5.904 | 23.489 | 0.964 | 0.0329 | |
| Seamless | IDOL | 4.551 | 23.685 | 0.929 | 0.0650 |
| GUAVA | 4.483 | 24.125 | 0.929 | 0.0645 | |
| LHM | 3.776 | 24.383 | 0.935 | 0.0582 | |
| Ours (SplatCtrlA) | 3.595 | 24.858 | 0.936 | 0.0542 |
A comprehensive user study across 32 participants was also conducted to record the percentage of times each method was judged as the most photorealistic:
| Category | IDOL | GUAVA | LHM | Ours (SplatCtrlA) |
|---|---|---|---|---|
| Face Quality | 4.8% | 13.5% | 17.3% | 64.4% |
| Hand Structure | 3.9% | 20.4% | 7.8% | 67.9% |
| Clothed Body | 2.9% | 11.5% | 10.6% | 75.0% |
| Overall Quality | 1.9% | 9.6% | 10.6% | 77.9% |
Ablation Study¶
Ablation experiments evaluated on the Seamless benchmark isolate each proposed component. For fair evaluation of full-body regression, facial regions were masked out during quantitative ablation metrics:
| Config | MSE(\(\times 10^{-3}\)) ↓ | PSNR ↑ | SSIM ↑ | LPIPS ↓ | Note |
|---|---|---|---|---|---|
| Full model | 3.527 | 24.949 | 0.9365 | 0.0527 | Full architecture with 3DGS dataset |
| w/o GA data | 5.283 | 23.037 | 0.9326 | 0.0676 | 2D synthetic data only; PSNR drops 1.91 dB |
| w/o input-aware | 5.281 | 23.114 | 0.9328 | 0.0684 | Pure latent decoding; blurred clothing textures |
| w/o GauInverse | 5.846 | 22.733 | 0.9320 | 0.0669 | Coarse mesh inverse only; severe outer boundary artifacts |
| w/o semantic | 3.953 | 24.477 | 0.9328 | 0.0583 | Missing semantic correction; edge color bleeding |
Key Findings¶
- 3D-consistent synthetic supervision sets the performance ceiling: Removing the custom Gaussian Avatar dataset (w/o GA data) and relying solely on 2D video diffusion data induces a massive 1.91 dB drop in PSNR and degrades perceptual quality. 2D diffusion frames carry minute geometric inconsistencies between viewpoints, which forces the network to learn blurred, conflicting Gaussian positions that corrupt finger structures.
- Decoupled input awareness prevents projection misalignment: Removing input-aware conditioning (w/o input-aware) results in washed-out textures. Furthermore, performing only standard mesh inverse mapping without Gaussian-offset re-projection (w/o GauInverse) causes severe distortion along clothing silhouettes (MSE deteriorates to 5.846), confirming that Gaussians must guide their own sampling once deformed off the canonical mesh.
- Dedicated facial blending and hand regularization eliminate uncanny artifacts: Qualitative comparisons show that the dedicated head model and zero-convolution StyleUNet recover sharp oral cavity structures and expressive gaze. Meanwhile, the \(100\times\) Laplacian constraint on hand Gaussians ensures articulated fingers remain intact under extreme re-posing, securing a 67.9% user study preference compared to 7.8% for LHM.
Highlights & Insights¶
- Harmonious coupling of UV parameterization with 3DGS: Representing 3D Gaussian attributes directly within 2D SMPL-X UV maps establishes a natural bridge between discrete 3D spatial points and modern 2D vision Transformers, maintaining real-time rendering speed while unlocking scalable pre-trained 2D priors.
- Cascaded geometry-color inverse projection: By splitting projection into a mesh-guided geometry step and a subsequent Gaussian-guided, semantically corrected color step, the model bypasses tracking misalignments and captures complex garment textures extending beyond the bare parametric body.
- Zero-convolution local blending pattern: Merging a specialized local model (head) into a global full-body avatar via StyleUNet initialized with zero convolution provides a plug-and-play recipe that can be seamlessly extended to other detail-sensitive regions, such as footwear or accessories.
Limitations & Future Work¶
- Vulnerability to complex directional illumination: Because the training pipeline builds upon synthetic data with relatively uniform diffuse lighting, the model tends to bake directional highlights and cast shadows into the static Gaussian albedo, producing lighting discrepancies under novel rendering views.
- Difficulty with high-frequency protruding geometry: Due to the topological bias of the SMPL-X UV sheet, extremely loose flowing garments (e.g., long dresses) or voluminous hair strands that diverge significantly from the body surface remain challenging to reconstruct with sharp geometric boundaries.
- Future directions: Integrating intrinsic albedo-shading decomposition into the data generation pipeline to support dynamic relighting, and exploring sparse point-cloud diffusion hybrid priors to better accommodate unconstrained loose clothing.
Related Work & Insights¶
- vs IDOL [CVPR 2025]: While IDOL pioneers feed-forward Gaussian human creation from a single image, it trains on 2D synthetic video data (HuGe100K) lacking strict multi-view 3D rigidity and omits explicit geometric scale/offset constraints, leading to splintered fingers and blurry textures. SplatCtrlA's 3D-consistent avatar dataset and decoupled input-aware decoder achieve a 2.55 dB PSNR advantage on THuman2.0.
- vs LHM [ICCV 2025]: LHM trains on extensive real human videos for robust full-body poses, but lacks decoupled facial expression and hand articulation modeling, yielding rigid mask-like faces and collapsing fingers. SplatCtrlA integrates an independent head model with zero-convolution blending and hand Laplacian regularization, outperforming LHM in face user preference by 64.4% to 17.3%.
- vs GUAVA [ICCV 2025]: GUAVA applies inverse texture mapping but is restricted to frontal upper-body talking heads and is fragile against tracking misalignments. SplatCtrlA enables full 360-degree whole-body avatars with robust semantic-guided Gaussian projection.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ An exceptionally well-engineered feed-forward framework combining 3DGS synthetic datasets, decoupled input-aware inverse projection, and zero-convolution blending.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarking across THuman2.0, NeuMan, and Seamless Interaction, supported by extensive ablations and a 32-subject user study.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear structure, insightful motivation, and transparent presentation of experimental and architectural details.
- Value: ⭐⭐⭐⭐☆ Significantly advances practical single-image controllable 3D avatar generation, offering immediate value for virtual reality, game development, and telepresence.