LumiTokens: 3D Relighting via Token-Space Lighting Transformation¶
Conference: ECCV 2026
Paper: ECCV Official
Project Page: https://neu-vi.github.io/LumiTokens
Area: 3D Vision
Keywords: 3D Relighting, Latent Scene Tokens, Progressive Relighting, Plücker Rays, Neural Rendering
TL;DR¶
LumiTokens reformulates 3D relighting as a direct closed transformation within an unstructured latent token space, utilizing self-attention between scene tokens and Plücker light-ray tokens to achieve view-consistent relit novel view synthesis and seamless progressive multi-light composition without explicit 3D geometry or rendering equations.
Background & Motivation¶
Multi-view 3D reconstruction techniques, such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), have dramatically advanced digital asset creation across film production, gaming, and augmented reality. Nevertheless, deploying these reconstructed 3D assets into new virtual environments requires re-rendering them under arbitrary target illumination. Classical inverse rendering pipelines tackle this by decomposing scenes into explicit intrinsic material attributes—such as surface normals, diffuse albedo, and surface roughness—and evaluating a physically based rendering equation under novel light configurations. Although inherently view-consistent, these methods demand dense multi-view captures and expensive per-scene numerical optimization. Furthermore, their physical expressiveness is fundamentally constrained by simplified BRDF assumptions, struggling with non-Lambertian phenomena such as subsurface scattering and high-order inter-reflections.
Generative and diffusion-based paradigms circumvent explicit material factorization by exploiting massive generative priors directly in view space. However, single-image relighting baselines inherently lack multi-view consistency. Even recent video and multi-view diffusion models that introduce cross-view attention remain view-space generative processes: the edited illumination is not retained as a persistent scene-level representation, forcing users to execute a full, costly generative pass for every lighting change. Recent hybrid feed-forward reconstruction architectures (e.g., RelitLRM and NeAR) eliminate per-scene optimization by predicting explicit 3DGS from sparse views, yet they remain bottlenecked by the geometric expressiveness of splatting and still require re-running their generative appearance branches for every new illumination condition.
Across all existing paradigms, relighting is treated either as explicit physical decomposition or as image/view-space synthesis, leaving lighting edits disconnected from a persistent, incrementally editable scene state. This paper takes a fundamentally different angle inspired by recent large view synthesis models (e.g., LVSM and SRT), which compress multi-view observations into unstructured, compact 1D latent tokens without fixed physical semantics. Core idea: formulate 3D relighting as a direct closed-loop transformation within the latent scene token space by parameterizing all light sources as unified Plücker ray tokens and performing bidirectional self-attention, thereby enabling native 3D lighting interaction and progressive multi-source composition without explicit 3D structures.
Method¶
Overall Architecture¶
The LumiTokens architecture comprises three modular components: a scene Encoder \(E\), a Scene Token Editor \(T\), and a multi-scale Decoder \(D\) equipped with a Dense Prediction Transformer (DPT) readout head. First, the encoder aggregates sparse posed input images along with their Plücker ray embeddings into a compact sequence of unstructured 1D scene tokens \(\mathbf{z} \in \mathbb{R}^{K \times d}\). Second, the Scene Token Editor takes the current scene tokens and target illumination—discretized into unified Plücker light-ray tokens \(\boldsymbol{\ell}\)—and modulates scene appearance via stacked self-attention layers to yield relit scene tokens \(\mathbf{z}'\). Finally, the decoder renders novel views under target camera rays directly from \(\mathbf{z}'\). Crucially, because the editor's output remains strictly inside the latent space \(\mathcal{Z}\), multi-step lighting edits can be applied recurrently in token space (\(\mathbf{z}^{(s-1)} \to \mathbf{z}^{(s)}\)) without intermediate image decoding.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multi-view Inputs + Plücker Rays<br/>Concatenated with learnable empty tokens z0"] --> B["Scene Encoder<br/>Aggregates spatial and photometric cues into scene tokens z"]
C["Target Lighting Parameterization<br/>Env maps, point lights, area lights unified into Plücker rays"] --> D["Scene Token Editor<br/>Global self-attention over joint scene and light tokens"]
B --> D
D -->|Update latent manifold| E["Relit Scene Tokens z'"]
E -->|Progressive relighting loop z s-1 to z s| D
E --> F["Multi-Scale DPT Decoder<br/>Renders sharp novel views conditioned on novel camera rays"]
Key Designs¶
1. Unified Plücker Light-Ray Tokenization: Bridging Abstract Latent Tokens with Native 3D Interaction
Unstructured scene tokens contain no explicit coordinate grids, mesh vertices, or surface normals, making native 3D spatial user manipulation non-trivial. LumiTokens bridges this gap by discretizing all illumination configurations into a set of \(L\) light rays sampled from the emitter towards the scene bounding volume, denoted as \(\boldsymbol{\ell} = \{\boldsymbol{\ell}_1, \dots, \boldsymbol{\ell}_L\}\). Each token is parameterized by its spatial geometry, color, and power: \(\boldsymbol{\ell}_i = (\mathbf{r}^\ell_i, \mathbf{c}_i, e_i)\), where \(\mathbf{c}_i \in \mathbb{R}^3\) represents spectral radiance, \(e_i\) is radiant power, and \(\mathbf{r}^\ell_i = (\mathbf{d}_i, \mathbf{m}_i)\) is the Plücker ray coordinate with unit direction \(\mathbf{d}_i\) and moment \(\mathbf{m}_i = \mathbf{o}_i \times \mathbf{d}_i\). For distant environment maps, the moment vector is nullified (\(\mathbf{m}_i = \mathbf{0}\)) to eliminate parallax; for local point lights, ray origins converge at the 3D source position such that direction and moment jointly encode exact 3D coordinates; for area lights, ray origins are distributed across the emitter surface to encode spatial extent and emission angles. This unified formulation maps arbitrary 3D lighting edits into identical token structures compatible with standard attention mechanisms.
2. Bidirectional Self-Attention Scene Token Editor: Modeling Implicit Visibility and Radiative Transfer
Injecting lighting information via standard unidirectional cross-attention would only allow light to modulate scene features without permitting scene geometry to query occlusion and shadow relationships. LumiTokens concatenates the \(K\) scene tokens \(\mathbf{z} = \{\mathbf{z}_1, \dots, \mathbf{z}_K\}\) and \(L\) light tokens \(\boldsymbol{\ell} = \{\boldsymbol{\ell}_1, \dots, \boldsymbol{\ell}_L\}\) into a single joint sequence, processing them through stacked self-attention Transformer blocks: $\(\mathbf{z}'_1, \dots, \mathbf{z}'_K, \boldsymbol{\ell}'_1, \dots, \boldsymbol{\ell}'_L = \text{Transformer}(\mathbf{z}_1, \dots, \mathbf{z}_K, \boldsymbol{\ell}_1, \dots, \boldsymbol{\ell}_L)\)$ Self-attention allows light tokens to attend across scene tokens to discern scene geometry, surface normals, and cast shadow boundaries, while enabling scene tokens to absorb incoming directional irradiance and synthesize specular highlights. After processing, the updated light tokens \(\boldsymbol{\ell}'\) are discarded, leaving only the relit scene representation \(\mathbf{z}'\).
3. Closed Latent-Space Progressive Relighting: Eliminating Compounding Drift of Pixel Re-Encoding
In realistic virtual production workflows, artists iteratively construct complex illumination by adding key, fill, and rim lights sequentially. Existing view-synthesis or diffusion models attempting multi-step lighting must decode full-resolution images at each step and re-encode them back into latent features, causing severe cumulative error drift and degradation of high-frequency details. LumiTokens constrains the transformation \(T\) to be approximately closed over the latent manifold \(\mathcal{Z}\), satisfying \(\mathbf{z}^{(s)} = T(\mathbf{z}^{(s-1)}, \boldsymbol{\ell}^{(s)}) \in \mathcal{Z}\). The entire interactive lighting sequence is performed strictly in token space; the decoder functions exclusively as an on-demand, read-only preview generator, completely avoiding decode-and-re-encode bottlenecks.
4. Multi-Scale DPT Decoder Head: Restoring High-Frequency Textures and Sharp Boundaries
Vanilla LVSM employs an independent per-token MLP readout head, which lacks inter-patch contextual integration and often produces checkerboard artifacts and blurry specular highlights. LumiTokens replaces this with a Dense Prediction Transformer (DPT) readout module attached to the decoder. The DPT head reshapes transformer output tokens into multi-scale feature pyramids and progressively fuses shallow spatial features with deep semantic representations via cascaded convolutional blocks, recovering crisp reflectance details, specular highlights, and clean shadow boundaries.
Loss & Training¶
During training, the encoder \(E\) and editor \(T\) are optimized end-to-end while keeping the pre-trained novel-view decoder \(D\) frozen to anchor generative spatial priors. The total objective is formulated as: $\(\mathcal{L} = \mathcal{L}_{\text{render}} + \alpha \mathcal{L}_{\text{inv}}\)$
- Rendering Loss \(\mathcal{L}_{\text{render}}\): To maintain latent stability across multi-step editing chains, training samples random sequences of \(S\) lighting conditions \(\boldsymbol{\ell}^{(1)}, \dots, \boldsymbol{\ell}^{(S)}\), recursively applies the editor in token space, and supervises decoded images against ground-truth relit images \(I_{\text{gt}}\) at each step via pixel-wise \(\ell_2\) and perceptual LPIPS loss: $\(\mathcal{L}_{\text{render}} = \|I_{\text{gt}} - \hat{I}\|_2^2 + \lambda \mathcal{L}_{\text{LPIPS}}(I_{\text{gt}}, \hat{I})\)$
- Lighting Invariance Loss \(\mathcal{L}_{\text{inv}}\): Multi-view input captures inevitably contain source lighting effects. To enforce separation between intrinsic scene identity and transient illumination, the model encodes paired multi-view captures of the same scene under two distinct illuminations \(L_A\) and \(L_B\), minimizing their latent Euclidean distance: $\(\mathcal{L}_{\text{inv}} = \|E(\{I_i^A, \mathbf{r}_i\}) - E(\{I_i^B, \mathbf{r}_i\})\|_2^2\)$ This regularizer compels the encoder to strip source-illumination artifacts (cast shadows and highlights) from the scene tokens \(\mathbf{z}\), delegating all light transport synthesis to the editor \(T\).
Key Experimental Results¶
Main Results¶
The authors evaluate performance across two rigorous tasks: multi-view relighting (relighting the original viewpoints to isolate lighting fidelity) and novel-view relighting (joint evaluation of relighting and novel view synthesis). Evaluation encompasses both synthetic objects (Objaverse and Polyhaven subsets) and real-world captures (Objects-with-Lighting and Stanford-ORB).
Table 1: Multi-view relighting accuracy (4 input views, evaluated on input poses)
(ILR: image-level scale alignment; SLR: scene-level global scale alignment across all views)
| Method | ILR PSNR↑ | ILR SSIM↑ | ILR LPIPS↓ | SLR PSNR↑ | SLR SSIM↑ | SLR LPIPS↓ |
|---|---|---|---|---|---|---|
| LightSwitch (ICCV 2025) | 21.22 | 0.868 | 0.105 | 21.19 | 0.868 | 0.131 |
| Neural Gaffer (NeurIPS 2024) | 28.40 | 0.947 | 0.030 | 28.23 | 0.955 | 0.037 |
| DiffusionRenderer (CVPR 2025) | 26.39 | 0.951 | 0.038 | 26.24 | 0.923 | 0.086 |
| LumiTokens (Ours) | 30.48 | 0.954 | 0.045 | 30.36 | 0.917 | 0.086 |
Table 2: Novel-view relighting accuracy and input view requirements
| Method | #Input | Synthetic PSNR↑ | Synthetic SSIM↑ | Synthetic LPIPS↓ | Objects-w-Light PSNR↑ | Objects-w-Light SSIM↑ | Stanford-ORB PSNR↑ | Stanford-ORB SSIM↑ |
|---|---|---|---|---|---|---|---|---|
| LightSwitch | 16 | 21.61 | 0.86 | 0.13 | 25.43 | 0.84 | 32.02 | 0.98 |
| NVDiffrecMC | 50 | 22.95 | 0.86 | 0.10 | 20.24 | 0.73 | 31.60 | 0.97 |
| TensoIR | 50 | 26.17 | 0.92 | 0.07 | 26.12 | 0.77 | 25.27 | 0.94 |
| LumiTokens (Ours) | 8 | 27.76 | 0.91 | 0.06 | 26.76 | 0.92 | 31.01 | 0.97 |
Ablation Study¶
Table 3: Ablation study on encoder fine-tuning and lighting invariance regularization
| Config | Encoder Status | Regularization Scheme | PSNR↑ | SSIM↑ | LPIPS↓ | Note |
|---|---|---|---|---|---|---|
| Baseline (LVSM setup) | Frozen | None | 25.06 | 0.910 | 0.073 | Fixed view-synthesis tokens cannot support relighting |
| Encoder Fine-tuned | Unfrozen | None | 26.12 | 0.927 | 0.046 | Unfreezing encoder yields +1.06 dB PSNR gain |
| Explicit Supervision | Unfrozen | Auxiliary Albedo Prediction Head | 27.09 | 0.934 | 0.057 | Predicts albedo directly, requires ground-truth labels |
| Implicit Invariance (Ours) | Unfrozen | Invariance Loss \(\mathcal{L}_{\text{inv}}\) | 27.05 | 0.952 | 0.045 | Highest structural and perceptual fidelity, zero GT labels |
Key Findings¶
- Cross-View Consistency Robustness: LumiTokens achieves state-of-the-art multi-view PSNR (30.48 dB), outperforming Neural Gaffer by over 2.0 dB. Crucially, when switching from image-level rescaling (ILR) to scene-level rescaling (SLR), LumiTokens drops by a negligible 0.12 dB with identical perceptual quality, whereas DiffusionRenderer severely degrades from 0.038 to 0.086 in LPIPS under SLR, proving that LumiTokens' shared token space enforces robust 3D coherence.
- Sparse-View Efficiency: On novel-view relighting, LumiTokens requires only 8 sparse input viewpoints to achieve 27.76 dB PSNR on synthetic benchmarks and 0.92 SSIM on Objects-with-Lighting. In contrast, classical inverse rendering baselines (NVDiffrecMC and TensoIR) require 50 views and extensive per-scene optimization while trailing in overall perceptual scores.
- Superiority of Token-Space Chaining: In 10-step sequential progressive lighting experiments, token-space editing preserves high structural fidelity across all steps. In contrast, pixel-space decode-and-re-encode baselines experience sharp degradation immediately after the first step due to lossy reconstruction loops.
Highlights & Insights¶
- Plücker Ray Tokens as a Universal 3D Light Interface: Transforming arbitrary light sources into Plücker ray tokens elegant solves the spatial grounding dilemma for geometry-free latent representations, unlocking intuitive 3D lighting manipulation.
- Progressive Relighting via Latent Manifold Invariance: Training the transformer editor on recurrent multi-step chains ensures that edited tokens remain inside the valid latent scene manifold, enabling arbitrary lighting additions without pixel-space degradation.
- Self-Supervised Lighting Disentanglement: The paired invariance loss \(\mathcal{L}_{\text{inv}}\) forces the encoder to discard transient source illumination without requiring costly ground-truth albedo or normal annotations.
Limitations & Future Work¶
- High-Frequency Specular Highlights on Extreme Reflectors: For intricate materials with intense mirror-like reflections or complex refraction, fixed-length 1D tokens may experience minor high-frequency detail compression.
- Camera Pose Sensitivity: The framework relies on accurate camera Plücker ray encodings; severe SfM pose jitter or uncalibrated inputs can perturb latent geometry alignment.
- Future Directions: Exploring pose-free self-supervised token backbones, and incorporating flow-matching continuous trajectory dynamics into the token editor for brush-level local lighting control.
Related Work & Insights¶
- vs. TensoIR / Relightable 3D Gaussians (Inverse Rendering): Classical methods perform per-scene gradient optimization under rigid BRDF rendering integrals; LumiTokens offers an instantaneous feed-forward pipeline free of explicit rendering equations and material parameter limits.
- vs. Neural Gaffer / LightSwitch (Diffusion Relighting): Diffusion methods perform 2D pixel-space denoising and lack a reusable 3D scene state; LumiTokens maintains a persistent, editable 3D token representation that can be updated incrementally with minimal compute.
- vs. RelitLRM / NeAR (Hybrid Models): Hybrid models require lifting features into explicit 3DGS geometry and re-running generative appearance passes; LumiTokens executes entirely in token space, supporting progressive multi-light composition natively.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ (Pioneering 3D relighting and progressive editing entirely within an unstructured latent token space)
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ (Thorough validation across multi-view, novel-view, progressive chaining, and real-world datasets)
- Writing Quality: ⭐⭐⭐⭐⭐ (Clear logical progression, insightful motivation, and rigorous architectural rationale)
- Value: ⭐⭐⭐⭐⭐ (Establishes a highly efficient, composable paradigm for 3D neural rendering and interactive virtual production)