Skip to content

TimeWalker: Personalized Neural Space for Lifelong Head Avatars

Conference: ECCV 2026
Paper: ECCV Official
Code: https://TimeWalker2026.github.io/
Area: 3D Vision
Keywords: Head Avatar, Lifelong Modeling, Neural Parametric Model, 2D Gaussian Splatting, Dynamic Mesh Reconstruction

TL;DR

Addressing the core limitation of conventional head avatar pipelines that only model momentary static captures or short video clips, TimeWalker introduces a personalized neural linear combination space and Dynamic 2D Gaussian Splatting (DNA-2DGS) to reconstruct and explicitly animate lifelong 3D head avatars with robust identity consistency across ages from unstructured in-the-wild photo collections.

Background & Motivation

Significant progress has been made in 3D human head avatar modeling, spanning classical 3DMMs to modern NeRFs, 3D Gaussian Splatting, and diverse animatable neural head representations. However, existing paradigms implicitly rely on a foundational assumption: an individual's cranial shape and geometry are static over time (Shape Invariance). Consequently, modern avatars operate as momentary replicas captured from single-session photogrammetry, calibrated multi-view camera rigs, or short monocular video sequences. Yet, according to theories of self-identity in sociology and developmental psychology, human identity is not defined by an isolated moment, but is constructed continuously through a lifelong narrative spanning various stages of life.

Extending 3D head avatar modeling to a lifelong scale introduces three fundamental challenges: First, longitudinal modeling invalidates the shape invariance assumption; musculoskeletal growth and aging reshape craniofacial geometry, while evolving hairstyles, facial textures, and skin wrinkles introduce drastic physical variations, making long-horizon identity consistency preservation extremely difficult. Second, capturing calibrated multi-view or temporal continuous data across decades is practically impossible; one's lifetime is documented through unstructured image collections with sparse, uncalibrated viewpoints, irregular camera distributions, and uneven lighting. Third, existing animation frameworks are tailored to static reference topologies, struggling to achieve explicit, multi-dimensional disentangled control (expression, shape, age period, viewpoint) while simultaneously producing high-fidelity surface meshes from low-quality in-the-wild observations.

To overcome severe data sparsity across life stages without relying on massive cross-identity pretraining, this work revives the additive combination principle of classical 3DMMs and advances it into a neural linear combination feature space. Core idea: model a person's lifelong identity as the additive combination of a shared canonical average head representation and a set of moment-specific attribute neural bases dynamically pruned via Dynamo, coupled with Dynamic 2D Gaussian Splatting (DNA-2DGS) that enforces inverse warping guidance and deferred dynamic meshing to achieve explicit multi-axis disentangled animation and artifact-free surface reconstruction.

Method

Overall Architecture

TimeWalker aims to learn an interpretable, scalable, and steerable personalized neural parametric space from an individual's unstructured cross-age photo collections. The system architecture adopts a two-tier deformation strategy: the first tier uses a neural linear combination space to deform canonical 2D Gaussian Surfels (initialized from a neutral FLAME template) to a moment-specific static avatar representing a particular lifestage; the second tier utilizes DNA-2DGS with FLAME expression coefficients to execute differential-geometry-guided inverse warping for motion animation, alongside a deferred warping strategy for extracting watertight dynamic surface mesh sequences.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Unstructured Cross-Age Photo Collection"] --> B["Canonical Average Representation<br/>FLAME-initialized 2DGS"]
    B --> C["1. Dynamic Neural Basis-Blending Module<br/>Multi-resolution Hashgrids with Active Pruning"]
    C --> D["2. Residual Global Compensation<br/>Lifestage Residual Vector & Deform/Color MLPs"]
    D --> E["Moment-Specific Static Avatar<br/>Additive Attribute Deformation"]
    E --> F["3. Dynamic Gaussian Inverse Warping<br/>Frenet Frame Differential Affine Transform"]
    E --> G["4. Deferred Warping Dynamic Meshing<br/>Poisson Meshing followed by Vertex Warping"]
    F --> H["Disentangled Multi-Scale Avatar Rendering"]
    G --> I["High-Fidelity Dynamic Mesh Sequences"]

Key Designs

1. Dynamic Neural Basis-Blending Module: Adaptive Compact Modeling of Temporal Variations Assigning an independent feature grid to each individual lifestage is parameter-redundant, prone to overfitting transient surface cues (lighting, accessories), and fails to capture deep invariant identity traits shared across decades. To address this, TimeWalker introduces the Dynamo module. The framework assumes that any lifestage can be linearly blended from \(N\) neural head bases: $\(f(\mathbf{x}_c) = \sum_{i=1}^N \omega_i H_i(\mathbf{x}_c)\)$ where each basis \(H_i\) is compactly parameterized via multi-resolution hashgrids and queried using trilinear interpolation. To eliminate redundant parameters and enforce learning shared structural variations, Dynamo incorporates an adaptive pruning mechanism: during training, blending weights \(\omega_i^k\) across all life stages \(k\) are inspected at intervals of \(q\) iterations. If a basis consistently satisfies \(\mathbb{I}(\omega_i^k < \kappa) = 1\) below a preset threshold \(\kappa\) across life stages, it is flagged as uninformative and deactivated. This pruning yields a minimal, highly expressive set of neural bases at convergence, effectively balancing representational capacity with cross-age generalization.

2. Residual Global Compensation and Additive Gaussian Attribute Deformation: Bridging Basis Gaps Linear basis blending alone can struggle to capture extreme local appearance variations and stage-specific lighting shifts. TimeWalker introduces a learnable global residual embedding \(\mathbf{l}_{\text{res}}\) for each life stage, which is concatenated with the blended feature \(f(\mathbf{x}_c)\) and fed into \(\text{MLP}_{\text{deform}}\). This network predicts attribute offsets for the 2D Gaussian surfels: position shift \(\delta \mathbf{x}\), rotation offset \(\delta \mathbf{r}\), scale offset \(\delta \mathbf{s}\), and opacity offset \(\delta \sigma\). An intermediate deformation feature is subsequently passed to \(\text{MLP}_{\text{color}}\) to derive spherical harmonic coefficient offsets \(\delta \mathbf{C}\). These learned offsets are additively combined with the canonical average Gaussian attributes: $\([\mathbf{x}_d, \mathbf{r}_d, \mathbf{s}_d, \sigma_d, \mathbf{C}_d] = [\mathbf{x}_c, \mathbf{r}_c, \mathbf{s}_c, \sigma_c, \mathbf{C}_c] + [\delta \mathbf{x}, \delta \mathbf{r}, \delta \mathbf{s}, \delta \sigma, \delta \mathbf{C}]\)$ Through self-supervised analysis-by-synthesis, this formulation isolates a neutral, moment-specific static head avatar for any target life stage without requiring explicit 3D supervision.

3. Dynamic Gaussian Inverse Warping: Differential Frame-Guided Expression Animation Once moment-specific Gaussian surfels are derived, using a generic implicit neural deformation field to animate expressions often fails on fine ocular/labial motions and leads to entanglement between identity and expression. Conversely, rigidly attaching surfels to a discrete mesh sacrifices geometric flexibility. DNA-2DGS addresses this by incorporating an explicit inverse warping operator rooted in differential geometry. Preprocessing fits a tracked mesh \(M_{\text{def}}\) and a canonical template \(M_{\text{canon}}\) with identical topology. For each Gaussian surfel \(\mathbf{x}_d\), a nearest-triangle query defines a local transformation matrix using Frenet frames: $\(F = L_{\text{def}} \cdot \Lambda^{-1} \cdot L_{\text{canon}}^{-1}\)$ where \(\Lambda\) accounts for local triangular surface scaling. Surfel positions are updated via \(\mathbf{x}'_d = F \cdot \mathbf{x}_d\). This explicit geometric guidance moves and rotates the 2D Gaussian disks faithfully with facial action units, maintaining strict geometric realism prior to 2D Gaussian rasterization.

4. Deferred Warping Dynamic Meshing: Bypassing Per-Frame Poisson Reconstruction Bottlenecks While 2D Gaussian Surfels provide strong surface-normal consistency for static reconstruction, applying Screened Poisson Surface Reconstruction per frame on animated dynamic scenes is computationally prohibitive and prone to topological tearing around rapid speech regions (e.g., inner mouth and lip boundaries). TimeWalker introduces a deferred warping strategy (Defer-Warping): rather than meshing after motion deformation, Poisson surface reconstruction is applied strictly to the static, moment-specific representation of each life stage. This yields an appearance-specific, clean base mesh free from motion artifacts. Dynamic animation is subsequently applied by directly warping the vertices of the reconstructed static mesh using the tracked deformation field. This guarantees temporally coherent mesh topology and achieves efficient dynamic mesh sequence extraction.

Loss & Training

The framework is optimized end-to-end using a joint objective covering photometric reconstruction, geometric smoothness/normal consistency, and deformation regularization: $\(\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{image}} + \mathcal{L}_{\text{geometry}} + \mathcal{L}_{\text{regulation}}\)$ Before full optimization, a warm-up phase freezes the Neural Head Basis updates and optimizes only the canonical 2D Gaussian Surfels to converge onto an identity-wide average head shape. Adaptive densification and pruning during this warm-up establish a stable geometric anchor for downstream basis blending.

Key Experimental Results

Main Results

Experiments are conducted on the newly introduced TimeWalker-1.0 benchmark (40 celebrities, 2.7 million frames). Quantitative evaluations on 5 representative identities across 8–13 life stages follow two protocols: Protocol-1 (1 vs 1), where a single unified model fits all life stages of an identity; Protocol-2 (1 vs N), where baselines train dedicated models per life stage while TimeWalker maintains a single model.

Protocol Method Representation PSNR ↑ SSIM ↑ LPIPS ↓ ID Score ↑
Protocol-1 (1 vs 1) INSTA NeRF 20.68 0.697 0.299 51.59
Protocol-1 (1 vs 1) INSTA++ (PAV re-impl.) NeRF 26.39 0.879 0.139 85.37
Protocol-1 (1 vs 1) FlashAvatar 3DGS 22.14 0.771 0.267 66.97
Protocol-1 (1 vs 1) TimeWalker (Ours) 2DGS 27.28 0.949 0.071 93.71
Protocol-2 (1 vs N) Gaussian Surfels 2DGS 26.98 0.950 0.141 82.73
Protocol-2 (1 vs N) GS++ (Dynamic) 2DGS 27.61 0.948 0.134 76.98
Protocol-2 (1 vs N) INSTA NeRF 25.47 0.860 0.170 87.96
Protocol-2 (1 vs N) FlashAvatar 3DGS 24.90 0.848 0.165 84.17
Protocol-2 (1 vs N) TimeWalker (Ours, Single Model) 2DGS 27.28 0.949 0.071 93.71

Ablation Study

Table 1: Ablation on Dynamo Hashgrid Count and Adaptive Pruning | Config | PSNR ↑ | SSIM ↑ | LPIPS ↓ | Note | |---|---|---|---|---| | w/o Dynamo (Coordinate only) | 21.69 | 0.767 | 0.197 | Severe collapse without basis blending | | 1 Hashgrid | 24.84 | 0.890 | 0.119 | Insufficient capacity in single global grid | | All Hashgrid (Matches stages) | 26.86 | 0.938 | 0.078 | Slight overfitting from independent memorization | | 20 Hashgrids (Fixed excessive) | 26.94 | 0.935 | 0.077 | Parameter redundancy | | Ours (Dynamo Adaptive) | 27.20 | 0.941 | 0.077 | Dynamic pruning balances capacity & generalization |

Table 2: Multi-Lifestage Joint Training vs Single-Stage Independent Training | Training Data Config | PSNR ↑ | SSIM ↑ | LPIPS ↓ | Note | |---|---|---|---|---| | One stage data | 23.69 | 0.959 | 0.072 | Lacks cross-age structural priors; lower PSNR | | Full stages (Ours) | 27.22 | 0.955 | 0.070 | Cross-stage knowledge sharing yields +3.53 dB PSNR gain |

Key Findings

  • In the rigorous 1-vs-1 setup where one model represents an entire lifetime, traditional approaches like INSTA and FlashAvatar suffer severe blurring and identity corruption. TimeWalker achieves an ArcFace ID score of 93.71 (vs 85.37 for INSTA++ and 66.97 for FlashAvatar) alongside an ultra-low LPIPS of 0.071, verifying the robustness of its disentangled representation.
  • Under Protocol-2 (1 vs N), where baselines enjoy individual model overfitting per life stage, TimeWalker's single unified model outperforms them in perceptual fidelity (LPIPS 0.071 vs 0.134–0.170) and ID consistency (93.71 vs 76.98–87.96). This indicates that the shared personalized space acts as a strong regularizer that compensates for unobserved poses and lighting in sparse stages.
  • In geometric surface evaluations, COLMAP and vanilla static Gaussian Surfels fail completely under sparse monocular views, producing fractured surfaces with heavy floaters. TimeWalker extracts smooth, anatomically accurate facial meshes via its deferred warping strategy.

Highlights & Insights

  • Reinvigorating classical 3DMM linear additive principles within modern neural representations: modeling lifelong facial transformation as "canonical average + dynamically blended neural bases + differential motion warping" provides clean mathematical tractability and natural linear latent interpolation for smooth "time traveling."
  • Adaptive basis pruning via Dynamo: pruning inactive hashgrids based on lifetime-wide blend weights prevents memorization of superficial accessories or transient illumination, ensuring that surviving bases capture core invariant morphological variations.
  • Deferred dynamic meshing paradigm: recognizing that Poisson surface reconstruction is brittle in the presence of dynamic motion artifacts, isolating meshing to the static neutral stage and transferring motion onto mesh vertices cleanly resolves the trade-off between reconstruction fidelity and runtime efficiency.

Limitations & Future Work

  • Reliance on FLAME priors: complex internal oral cavity structures (such as tongue dynamics) and extreme non-linear facial expressions absent from the FLAME topology can cause localized stretching or visual artifacts.
  • In-the-wild longitudinal data noise: historical archival photographs suffer from severe color cast, variable resolution, motion blur, and inaccurate boundary alpha matting, which can degrade appearance consistency under extreme profile angles.
  • Inability to extrapolate unobserved life periods: the model performs smooth interpolation across observed stages within its personalized latent space, but cannot hallucinate unrecorded epochs (e.g., childhood or advanced senescence) without external generative priors. Combining personalized neural spaces with large-scale 3D foundation models is a promising future trajectory.
  • vs INSTA / PAV: INSTA animates NeRF avatars via mesh-guided warping fields for short video clips but degrades on long-horizon unstructured collections. PAV introduces appearance embeddings but relies on a single implicit deformation field, lacking explicit multi-axis disentanglement between expression, shape, and age, while incurring substantial volume rendering costs.
  • vs FlashAvatar / GaussianAvatars: Both bind 3D Gaussians to dynamic facial meshes for fast rendering, but their reliance on volumetric ellipsoids and momentary single-sequence setups prevents them from generalizing across life stages or extracting watertight, non-noisy surface meshes.
  • vs GAGAvatar / One-shot Generative Avatars: While one-shot talking head models offer broad generalizability from large pretraining, their identity consistency across ages is markedly lower (ArcFace ID Score ~63.39 vs TimeWalker's 93.71), and they lack true 3D multi-view consistency for high-fidelity asset preservation.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates the lifelong 3D head avatar task and introduces a principled neural linear combination space with adaptive basis pruning and deferred dynamic meshing.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Introduces the TimeWalker-1.0 benchmark (40 subjects, 2.7M frames), accompanied by rigorous 1-vs-1 and 1-vs-N protocols, cross-stage reenactments, and mesh ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Well-grounded motivation from psychological identity theory, mathematically coherent formulations, and clear architectural diagrams.
  • Value: ⭐⭐⭐⭐⭐ Serves as a vital benchmark for digital aging in VFX, memory asset preservation, and longitudinal 3D human modeling.