title: >- [Paper Note] D-Rex : Diffusion Rendering for Relightable Expressive Avatars description: >- [ECCV 2026][Human Understanding][Relightable Avatars] Decouples relighting from avatar modeling as an image-space post-process via video diffusion LoRA fine-tuning, achieving expressive and view-consistent full-body rendering. tags: - ECCV 2026 - Human Understanding - Relightable Avatars - Video Diffusion - Expressive Avatars date: 2026-09-19 content_hash: c530f3ddd877dfd3
D-Rex : Diffusion Rendering for Relightable Expressive Avatars¶
Conference: ECCV 2026
Paper: ECCV Official
Project: https://vcai.mpi-inf.mpg.de/projects/DRex/
Area: Human Understanding
Keywords: relightable avatars, full-body dynamic avatars, video diffusion model, facial expression control, image-space rendering
TL;DR¶
Addressing the trade-off between relighting and expressive animation caused by explicit 3D intrinsic decomposition in full-body avatars, D-Rex decouples relighting as an image-space post-process, driving albedo-like renderings from an off-the-shelf white-light avatar and generating photorealistic, view- and temporally consistent relit avatars under arbitrary HDR environments via LoRA fine-tuning of a video diffusion model.
Background & Motivation¶
Building animatable, relightable full-body human avatars capable of photorealistic synthesis under arbitrary illumination driven by body pose and facial expressions is a foundational pillar for virtual production, telepresence, and interactive gaming. In recent years, non-relightable avatar frameworks utilizing neural rendering and 3D Gaussian splatting have made dramatic strides in handling complex cloth dynamics and high-fidelity facial performance. However, once relighting capability is demanded, existing pipelines almost universally revert to explicit 3D intrinsic decomposition—estimating albedo, roughness, normals, and ambient occlusion—coupled with analytic BRDF physically based rendering (PBR). Because inverting pose-dependent intrinsics from multi-view video is severely ill-posed, geometry registration errors degrade appearance modeling, leaving existing relightable avatars virtually incapable of delivering expressive, nuanced facial animations.
The core tension lies in the tight coupling of complex light transport effects (subsurface scattering, Fresnel reflectance, specular highlights, and self-shadowing) with explicit 3D geometric optimization. Constraining both physical reflectance and non-rigid geometric deformation simultaneously introduces severe optimization bottlenecks, forcing prior methods to sacrifice facial expressiveness and overall system modularity. In parallel, diffusion-based generative relighting models demonstrate superior synthesis of intricate illumination effects in image space, yet they typically lack explicit control over full-body articulated pose and facial expression, and struggle to guarantee multi-view and long-term temporal coherence when applied directly.
This paper's angle of attack is to abandon explicit 3D material inversion altogether, recasting full-body relighting as an independent image-space post-processing translation. Core idea: decouple relighting from avatar modeling by driving high-fidelity albedo-like renderings with an independent white-light avatar, while adapting a pre-trained video diffusion rendering model via lightweight LoRA on single-image frame pairs to synthesize view- and temporally coherent relit performances under arbitrary HDR environment maps.
Method¶
Overall Architecture¶
The D-Rex framework operates in two cleanly decoupled stages: front-end white-light expressive avatar synthesis and back-end video diffusion relighting. Given target skeletal pose \(\theta\), FLAME facial expression \(\psi\), virtual camera viewpoint \(v\), and an HDR environment map \(l\), an expressive avatar model (EVA) trained exclusively on light stage flat-lit frames first generates an albedo-like, shadow-free driving image \(I_{\text{WL}}\). Next, \(I_{\text{WL}}\) along with the view-projected, tone-mapped HDR map are fed into a LoRA-adapted video diffusion model (Cosmos DiffusionRenderer), which synthesizes complex light transport directly in latent space, yielding photorealistic, view- and temporally consistent relit output frames \(I_{\text{RL}}\).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Driving Inputs<br/>Pose θ + FLAME Expression ψ + View v + HDR Map l"] --> B["White-Light Expressive Avatar Synthesis<br/>Disentangled Deformable Mesh & UV Gaussians (EVA)"]
B --> C["Illumination Projection & Latent Encoding<br/>Reinhard Tone-Mapping & VAE Encoding"]
C --> D["Video Diffusion Relighting Synthesis<br/>LoRA-adapted Cosmos DiffusionRenderer"]
D --> E["Long-Sequence Linear Blending<br/>32-frame Window Smoothing Inter-chunk Drift"]
E --> F["Photorealistic Relit Avatar Video Output"]
Key Designs¶
1. White-Light Expressive Avatar Synthesis: Decoupling Geometry and Expression Appearance To resolve the fragility of analytic reflectance modeling under fine facial articulation and cloth wrinkles, D-Rex utilizes Expressive Virtual Avatars (EVA) as the underlying driving model. EVA decouples coarse dynamic body geometry from fine appearance, employing U-Nets to predict 3D Gaussian attributes in UV space anchored to a deformable template. To prevent background Gaussian floaters arising from slightly overestimated light stage segmentation masks, a background-compositing objective is introduced after 40,000 optimization steps: $$ \hat{I}{\text{Pred}} = I) $$ where }} \cdot I_{\alpha} + I_{\text{BG}} \cdot (1 - I_{\alpha\(I_{\text{Pred}}\) and \(I_{\alpha}\) are the predicted RGB and opacity maps, and \(I_{\text{BG}}\) is the ground-truth background under the dilated mask. This mechanism eliminates floating peripheral artifacts and delivers clean, crisp albedo-like renderings that faithfully reflect arbitrary FLAME expressions and full-body motions without being burdened by lighting calculations.
2. Video Diffusion Relighting Synthesis: Adapting Generative Priors via Single-Image Frame Pairs To map the flat-lit render \(I_{\text{WL}}\) to target HDR lighting \(l\), D-Rex adapts Cosmos DiffusionRenderer (DRCosmos), a large-scale video diffusion foundation model. Rather than retraining or using expensive OLAT linear basis compositing, D-Rex discards the original inverse G-buffer estimation module and retains only the forward rendering DiT. \(I_{\text{WL}}\) is treated as an approximation of albedo and inserted into the base color channel while remaining G-buffers (roughness, metalness, normals) are set to zero. Lightweight LoRA parameters \(\Delta\theta\) are injected into the forward DiT attention blocks, optimized via denoising score matching: $$ \mathcal{L}(\Delta\theta) = \mathbb{E}{\tau, z_0, \epsilon}\left[ | \epsilon - f, \tau) |_2^2 \right] $$ where }(z_{\tau}, g, c_{\text{env}\(g\) denotes the VAE latent encoding of \(I_{\text{WL}}\), and \(c_{\text{env}}\) is the view-projected HDR environment map tone-mapped using Reinhard operator. Crucially, the authors discover that fine-tuning solely on static, single-image flat-lit to relit frame pairs is sufficient to unlock the underlying video diffusion model's temporal prior, enabling coherent view rotation and lighting transitions across 57-frame sequences without demanding multi-frame temporal losses during training.
3. Long-Sequence Linear Blending: Suppressing Inter-Inference Temporal Drift While diffusion sampling exhibits exceptional view and illumination consistency within an individual 57-frame chunk, independent inference chunks can exhibit slight illumination or motion variance due to subtle human subject drift during capture (the sub-second temporal gap between interleaved flat-lit and relit frames). To achieve seamless long-form video synthesis, D-Rex implements a 32-frame overlapping linear blending window across consecutive chunks. This smoothly suppresses boundary discontinuities and flicker without requiring expensive auto-regressive recurrence or secondary temporal discriminator networks.
Loss & Training¶
The framework is trained in two modular stages: 1. EVA Avatar Training: Trained on 4 NVIDIA A40 GPUs for coarse deformable geometry (~4 days) followed by 1 NVIDIA A40 GPU for Gaussian appearance (~3 days), operating at \(540 \times 1024\) resolution with high-resolution localized crops. 2. Diffusion Relighting Fine-Tuning: LoRA fine-tuning of Cosmos DiffusionRenderer is executed on a single NVIDIA H100 GPU with a batch size of 5 for approximately 2 days (~30k iterations). Frames are downsampled to \(371 \times 704\) and zero-padded to match the native aspect ratio and latent structure of DRCosmos.
Key Experimental Results¶
Main Results¶
The evaluation is conducted across 4 captured human subjects displaying dynamic clothing folds and expressive performances. The test protocol withholds ~10% unseen lighting conditions, unseen motion sequences (frame indices \(\ge 12250\)), and 4 complete full-body camera views. Baselines comprise: MeshAvatar adapted for calibrated PBR [MA/PBR], a hybrid combining flat-lit MeshAvatar with diffusion relighting [MA/DR], and zero-shot relighting using IC-Light on EVA driving frames [EVA/IC].
| Method | Novel View/Motion PSNR ↑ | Novel View/Motion LPIPS ↓ | Novel View/Motion SSIM ↑ | Novel View/Motion/Light PSNR ↑ | Novel View/Motion/Light LPIPS ↓ | Novel View/Motion/Light SSIM ↑ |
|---|---|---|---|---|---|---|
| MA/PBR [7] | 22.61 | 0.095 | 0.899 | 22.38 | 0.095 | 0.899 |
| MA/DR [7,25] | 26.07 | 0.080 | 0.935 | 25.73 | 0.080 | 0.933 |
| EVA/IC [55] | 10.26 | 0.108 | 0.853 | 10.41 | 0.106 | 0.855 |
| D-Rex (Ours) | 26.62 | 0.077 | 0.937 | 26.30 | 0.077 | 0.936 |
Ablation Study¶
Ablation studies rigorously evaluate the diffusion generative prior, background preprocessing choices, and the trade-off between relighting and perceptual enhancement:
| Strategy / Configuration | PSNR ↑ | LPIPS ↓ | SSIM ↑ | Compute / Preprocess Budget | Key Behavioral Analysis |
|---|---|---|---|---|---|
| LoRA-adapted DR (Ours) | 27.62 | 0.079 | 0.945 | ~48h (1 GPU) | Matches full fine-tuning with minimal compute overhead |
| (a) Zero-shot w/o fine-tuning | 19.18 | 0.120 | 0.889 | 0h | Severe domain gap; cannot synthesize realistic avatar relighting |
| (b) Full model fine-tuning | 27.88 | 0.079 | 0.947 | ~192h (2 GPUs) | Comparable quality to LoRA but requires 4× GPU resources |
| (c) Training w/o prior (from scratch) | 21.27 | 0.096 | 0.910 | ~192h (2 GPUs) | Completely fails to converge within allocated training budget |
| Unmasked frame pairs (w/o BG mask, Ours) | 26.30 | 0.077 | 0.936 | Preprocess: 0s | Preserves clean black background, highly scalable |
| Masked frame pairs (w/ BG mask) | 26.82 | 0.073 | 0.939 | Preprocess: 16 days | Marginally closer color fidelity but prohibitively slow |
Key Findings¶
- Generative Diffusion Bypasses PBR Bottlenecks: MA/PBR struggles with material ambiguities, yielding dull skin tones and geometric artifacts in shadow zones (PSNR 22.38 dB). Replacing its rendering component with post-process diffusion (MA/DR) triggers a 3.35 dB leap, establishing that image-space diffusion outperforms explicit inverse rendering in capturing Fresnel glints and complex indirect illumination.
- Extreme Efficiency of LoRA Adaptation: LoRA fine-tuning on 1 GPU (~48h) effectively unlocks the pre-trained Cosmos prior, nearly matching full parameter fine-tuning on 2 GPUs (~192h), whereas training from scratch fails entirely (PSNR 21.27 vs 27.62 dB).
- Relighting vs. Enhancement Dilemma: Training the diffusion model directly on EVA-to-relit pairs (Relight + Enhance) sharpens blurred clothing and shoe details (LPIPS improves to 0.072), but hallucinations emerge in unconstrained regions such as eye gaze. Training on real flat-lit to relit pairs (Relight only) remains the more robust and generalizable configuration.
Highlights & Insights¶
- Paradigm Shift in Relightable Avatars: Instead of forcing the inverse optimization of physical BRDF parameters alongside non-rigid avatars, D-Rex delegates animation to white-light avatars and lighting to a generative video diffusion prior. This clean decoupling circumvents ill-posed inverse rendering and delivers expressive animation without compromise.
- Temporal and View Coherence from Paired Stills: A standout finding is that fine-tuning only on static paired images (flat-lit to relit) allows the model to preserve 3D view rotation consistency and temporal stability inherited from pre-trained video diffusion weights.
- Drop-in Upstream Compatibility: Because relighting is strictly formulated as a post-process, D-Rex is fully agnostic to the underlying avatar architecture; any future advance in white-light animatable humans can be seamlessly adopted.
Limitations & Future Work¶
- Lack of Real-Time Capability: The video diffusion model takes approximately 15 sampling steps, resulting in an inference frame rate of ~0.5 FPS. This limits deployment in real-time interactive VR/AR or telepresence. Future iterations could leverage flow-matching distillation or single-step consistency models to bridge this latency gap.
- Inter-Chunk Illumination Variance: While linear blending smooths transitions across 57-frame video chunks, minute variations in lighting intensity caused by capture motion drift can persist. Conditioning the model auto-regressively on initial context frames is a promising direction.
- Distant Occlusion Shadows: Shadowing effects originating from distant unobserved occluders are not present in the single-person dome capture, preventing the network from synthesizing large-scale external cast shadows.
Related Work & Insights¶
- vs MeshAvatar [7] & Relightable Codec Avatars [48]: PBR-based methods demand highly calibrated multi-view tracking and fragile material optimization, making full facial expression control virtually unattainable; D-Rex sidesteps intrinsic decomposition completely, achieving superior visual photorealism and expressive facial animation.
- vs DiffRelight [13] & IC-Light [55]: DiffRelight requires dense OLAT basis captures and is confined to facial portraits; IC-Light lacks person-specific 3D geometric awareness and produces severe overexposure and view inconsistency. D-Rex achieves full-body, free-viewpoint HDR relighting for dynamic subjects.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Decoupling relighting into an image-space generative post-process fundamentally circumvents the long-standing trade-off between relighting and expressive animation in digital avatars.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous multi-subject light stage captures, exhaustive comparisons against PBR and generative baselines, and comprehensive ablations on priors, compute, and data masking.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear problem formulation, detailed architectural descriptions, informative visuals, and honest discussions regarding limitations.
- Value: ⭐⭐⭐⭐⭐ Establishes a highly practical and high-fidelity blueprint for full-body avatar relighting in virtual production, visual effects, and gaming.