Consistent Feature Transport for Image Relighting¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/Dixin-Lab/CFT
Area: Image Generation
Keywords: Image Relighting, Feature Transport, Rectified Flow, Portrait Generation, Image Editing
TL;DR¶
Reformulating image relighting as an illumination feature transport problem, this paper proposes Consistent Feature Transport (CFT) under rectified flow, using the parallelogram law and cross-instance illumination trajectory supervision to isolate lighting changes from identity and geometry, establishing new state-of-the-art results on a newly constructed large-scale complex portrait relighting benchmark.
Background & Motivation¶
Image relighting aims to alter the illumination conditions of a given image—including lighting direction, intensity, color temperature, and shadow interactions—while strictly preserving non-lighting content such as facial identity, scene geometry, surface materials, and background context. Driven by the rapid progress of diffusion models and flow matching algorithms, image relighting has unlocked immense practical value across portrait enhancement, virtual cinematography, and visual effects production. However, under complex lighting scenarios, simultaneously achieving fine-grained shadow manipulation and pristine structural fidelity remains an open challenge.
Existing diffusion-based relighting paradigms fall into three main categories, each suffering from notable deficiencies: control-based methods inject illumination signals (such as environment maps or shading priors) as auxiliary conditions to guide denoising, but depend heavily on the completeness and precision of external controls, often struggling with intricate multi-source interreflections; decomposition-based methods attempt to factorize images into intrinsic components like albedo and shading before editing, yet intrinsic decomposition is ill-posed under unconstrained lighting and prone to error propagation, producing color casts and shading artifacts; direct image-to-image translation and trajectory-guided editing (e.g., bridge matching or inversion-free editing) rely on source anchoring to keep structure intact, but treat relighting as a generic instance-level modification without an explicit mechanism to learn illumination-specific feature displacement, limiting editing expressiveness. Furthermore, available portrait datasets are predominantly acquired under simplistic, controlled laboratory setups, lacking complex lighting effects like Tyndall beams, neon hues, and Venetian blind shadows.
The key insight of this paper is that relighting should not be approached as a generic image-to-image translation, but as a directed transport of illumination features in latent space. When directly supervising the source-to-target trajectory using the identical image pair, generative models inherently overfit instance-specific geometry and scene differences. Only by decoupling the illumination displacement from identity cues can we accomplish pure lighting edits. Core idea: reformulate image relighting as an illumination feature transport problem, jointly optimizing noise-to-image generation and parallelogram-derived source-to-target transport under rectified flow, and supervising the direct velocity field using cross-instance image pairs sharing identical lighting transformations (CFT) to isolate illumination variations from underlying content.
Method¶
Overall Architecture¶
The framework is built upon rectified flow, jointly modeling the distributions of noise prior, source images, and target relit images. Given a source image \(x_{\text{src}}\), a target relit image \(x_{\text{tgt}}\), and a target lighting text prompt \(c_{\text{tgt}}\), a velocity field model \(v_\theta\) learns both the generative path from Gaussian prior to the target latent representation and an auxiliary reconstruction path that recovers the source image under a neutral lighting condition \(c_{\text{src}}\). Crucially, leveraging the linearity of rectified flow, the framework analytically derives a direct transport velocity \(v_t^{\text{direct}}\) between source and target latents via the parallelogram law. It then supervises this direct path using an alternative image pair \((x'_{\text{src}}, x'_{\text{tgt}})\) that undergoes the identical lighting transformation on entirely different subjects and scenes, compelling the velocity field to capture pure illumination feature transport.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Data<br/>Source image and target lighting prompt"] --> B["Dual Rectified Flow Generation<br/>Noise-to-target generation and source reconstruction"]
B --> C["Parallelogram Direct Transport<br/>Linear flow trajectory derivation"]
C --> D["Cross-Instance Consistent Feature Transport<br/>Shared-lighting supervision for content decoupling"]
D --> E["Large-Scale Portrait Relighting Dataset<br/>Three-model ensemble quality filtering"]
E --> F["Output Relit Image<br/>Physical plausibility and structural fidelity"]
Key Designs¶
1. Dual Rectified Flow Generation: Anchoring Target Synthesis and Source Geometry Reconstruction Standard conditional rectified flow models only the single generation trajectory from a Gaussian prior \(\pi_0\) to the target latent representation, governed by the velocity matching loss: $$ \mathcal{L}1 = \mathbb{E} - z_0) \right|^2 $$ However, relying solely on target generation causes the model to drift away from high-frequency source cues, degrading facial identity and fine textures. To remedy this, the authors introduce a source reconstruction pathway starting from the same prior sample }},\, t} \left| v_\theta(z_t^{\text{tgt}}, t, x_{\text{src}}, c_{\text{tgt}}) - (z_{\text{tgt}\(z_0\) under a neutral lighting condition \(c_{\text{src}}\): $$ \mathcal{L}2 = \mathbb{E} - z_0) \right|^2 $$ This reconstruction objective leaves lighting untouched while firmly grounding facial geometry and background semantics within the velocity backbone, serving as a steadfast structural anchor for subsequent illumination edits.}},\, t} \left| v_\theta(z_t^{\text{src}}, t, x_{\text{src}}, c_{\text{src}}) - (z_{\text{src}
2. Parallelogram Direct Transport: Direct Editing Path Enabled by Linear Flow Geometry Owing to the linear ODE formulation of rectified flow, latent trajectories follow straight interpolation paths across time \(t \in [0, 1]\). Invoking the geometric parallelogram law, intermediate states along the direct path between source latent \(z_{\text{src}}\) and target latent \(z_{\text{tgt}}\) can be algebraically approximated as: $$ z_t^{\text{direct}} = z_{\text{src}} + z_t^{\text{tgt}} - z_t^{\text{src}} $$ Differentiating with respect to time, the instantaneous tangent velocity along this direct editing route is simply the difference between the target generation velocity and the source reconstruction velocity: $$ v_t^{\text{direct}} = v_\theta(z_t^{\text{tgt}}, t, x_{\text{src}}, c_{\text{tgt}}) - v_\theta(z_t^{\text{src}}, t, x_{\text{src}}, c_{\text{src}}) $$ This closed-form formulation circumvents computationally burdensome diffusion inversion iterations, furnishing an explicit vector field that directly pushes the source image into the target lighting domain.
3. Cross-Instance Consistent Feature Transport: Eliminating Sample-Specific Content Correlations If one naively supervises \(v_t^{\text{direct}}\) with the original image pair \((x_{\text{src}}, x_{\text{tgt}})\), the network inevitably learns to memorize instance-specific geometry and non-lighting disparities, producing spurious entanglements between lighting and identity. The defining innovation of CFT is decoupling this supervisory signal: the pipeline pairs the editing trajectory with an auxiliary image pair \((x'_{\text{src}}, x'_{\text{tgt}})\) that experiences the exact same lighting transformation (e.g., from overhead noon sunlight to neon street lighting) on an entirely different individual and scene, supervising the direct velocity field with \((z'_{\text{tgt}} - z'_{\text{src}})\): $$ \mathcal{L}3 = \mathbb{E}{z'{\text{src}},\, z' - (z'}},\, t} \left| v_t^{\text{direct}{\text{tgt}} - z') \right|^2 $$ Because the input image and the supervisory target share zero identity or background overlap, the model is compelled to isolate and transport only lighting-specific feature representations, eliminating content leakage at its root.}
4. Large-Scale Complex Portrait Relighting Dataset: Multi-Source Synthesis and Ensemble Filtering To overcome the lighting simplicity of existing benchmarks, the authors gather neutral, high-quality portrait base images from open-source collections, design comprehensive lighting descriptions across directions, standard setups (Rembrandt, Tyndall beam, Blue Hour), and photometric properties (color temperature, intensity, shadow softness), and synthesize candidate target images with state-of-the-art generators. An ensemble-based filtering pipeline comprising QwenVL, EditScore, and GPT-4o audits each candidate across three criteria: lighting quality, content consistency, and physical plausibility. A final manual review yields 34,695 pristine image pairs (34,249 training pairs and 446 test pairs) balanced across indoor (48.1%) and outdoor (51.9%) scenes with complex lighting effects.
Loss & Training¶
The joint optimization loss function of the complete CFT framework is formulated as: $$ \mathcal{L} = \mathcal{L}_1 + \mathcal{L}_2 + \alpha \mathcal{L}_3 $$ where hyperparameter \(\alpha\) governs the regularization weight of consistent feature transport, with \(\alpha = 0.1\) empirically proven optimal. The models are trained across 8 NVIDIA H20 GPUs using Adam with a learning rate of \(1 \times 10^{-4}\) and a batch size of 4 for 4,000 steps until full convergence.
Key Experimental Results¶
Main Results¶
Quantitative evaluations are conducted on the test split of the constructed portrait relighting benchmark across pixel-level fidelity (SSIM, PSNR), perceptual similarity (LPIPS), distribution realism (FID), and lighting-specific evaluation (Estimated Irradiance error IE, computed as MAE between predicted and ground-truth irradiance maps).
| Method | Paradigm Category | SSIM ↑ | PSNR ↑ | LPIPS ↓ | FID ↓ | IE ↓ |
|---|---|---|---|---|---|---|
| IC-light | Control-based generation | 0.7436 | 15.3314 | 0.2209 | 47.3196 | 15.3575 |
| LBM | Image-to-image translation | 0.7908 | 15.5954 | 0.2267 | 40.8663 | 16.9276 |
| FlowEdit | Trajectory inversion editing | 0.7133 | 13.6775 | 0.3298 | 78.1236 | 23.2227 |
| Intrinsic Edit | Intrinsic decomposition | 0.7148 | 15.4219 | 0.2646 | 57.3994 | 14.2292 |
| Latent Intrinsic | Latent intrinsic editing | 0.8425 | 18.1809 | 0.2379 | 35.0617 | 16.2387 |
| Nano Banana | Pretrained without tuning | 0.7339 | 14.9196 | 0.2507 | 46.5012 | 15.9686 |
| Flux-Kontext (original) | Pretrained base model | 0.6386 | 14.1543 | 0.2700 | 53.3793 | 18.2834 |
| Qwen-Image-Edit (w.o. CFT) | Control-based fine-tuning | 0.8820 | 20.1170 | 0.1140 | 30.2147 | 11.5505 |
| Qwen-Image-Edit (w. CFT) | Ours | 0.8894 | 20.9176 | 0.1132 | 25.5529 | 10.8670 |
| Flux-Kontext (w.o. CFT) | Control-based fine-tuning | 0.9153 | 22.9192 | 0.0985 | 20.3817 | 7.8155 |
| Flux-Kontext (w. CFT) | Ours | 0.9202 | 23.5105 | 0.0946 | 18.7338 | 7.4231 |
Ablation Study¶
Ablations on Flux-Kontext examine individual loss components and alternative supervision formulations for the direct transport branch \(\mathcal{L}_3\).
| Config | Note | SSIM ↑ | PSNR ↑ | LPIPS ↓ | FID ↓ |
|---|---|---|---|---|---|
| \(\mathcal{L}_1\) | Standard unidirectional rectified flow baseline | 0.9153 | 22.9192 | 0.0985 | 20.3817 |
| \(\mathcal{L}_1 + \mathcal{L}_2\) | Adding source reconstruction without direct transport | 0.9141 | 22.8299 | 0.1018 | 21.6054 |
| \(\mathcal{L}_1 + \mathcal{L}_3\) | Adding transport loss without source reconstruction | 0.9161 | 23.3778 | 0.0972 | 19.3191 |
| Original pairs as GT for \(\mathcal{L}_3\) | Supervising direct path with \((x_{\text{src}}, x_{\text{tgt}})\) | 0.9141 | 22.8031 | 0.1003 | 20.2961 |
| Random pairs as GT for \(\mathcal{L}_3\) | Supervising direct path with random pairs \((x^{\text{rand}}_{\text{src}}, x^{\text{rand}}_{\text{tgt}})\) | 0.9135 | 22.7958 | 0.1035 | 19.5538 |
| CFT (full model) | \(\mathcal{L}_1 + \mathcal{L}_2 + 0.1 \mathcal{L}_3\) (shared-lighting cross-instance) | 0.9202 | 23.5105 | 0.0946 | 18.7338 |
Key Findings¶
- Cross-instance supervision is essential to disentangle lighting: When the direct trajectory \(\mathcal{L}_3\) is supervised by the original image pair itself, PSNR falls from 22.9192 to 22.8031, underperforming even the vanilla \(\mathcal{L}_1\) baseline. This confirms that self-pair supervision forces the flow field to overfit identity and background shifts. In contrast, cross-instance supervision with matching lighting elevates PSNR to 23.5105 while maintaining the lowest trajectory velocity variance across all diffusion timesteps (Figure 5).
- Synergistic cooperation between \(\mathcal{L}_2\) and \(\mathcal{L}_3\): While adding \(\mathcal{L}_2\) alone brings no gain, combining \(\mathcal{L}_2\) with \(\mathcal{L}_3\) proves vital: source reconstruction locks in high-frequency geometry, preventing feature transport from distorting structural fidelity.
- Critical balance in transport weight \(\alpha\): Testing \(\alpha \in \{0.05, 0.1, 0.3, 0.6\}\) reveals that \(\alpha = 0.1\) hits the sweet spot. Smaller values collapse toward unconstrained conditional generation, while larger values (0.3, 0.6) aggressively warp the generative dynamics, degrading PSNR and perceptual realism.
- Superior zero-shot real-world transferability: Evaluated directly on the Multi-Illumination Dataset without any fine-tuning, CFT-enhanced Flux-Kontext outperforms the standard fine-tuned baseline, boosting SSIM from 0.6924 to 0.7325, lifting PSNR by 0.63 dB, and dropping Estimated Irradiance error from 28.96 to 27.34.
Highlights & Insights¶
- Attribute editing as vector flow transport: By exploiting the linear geometry of rectified flow, CFT transforms complex relighting from brittle conditional generation into an analytical vector transport problem, avoiding slow inversion steps.
- Cross-instance supervision trick: Guiding the velocity field using alternative image pairs undergoing the identical physical transformation elegantly decouples attribute shifts from content nuances at the data supervision level.
- Extensible across general image editing tasks: Applied to style transfer on OmniStyle (achieving PSNR 14.1224 vs baseline 13.8040), CFT demonstrates that consistent feature transport serves as a generalizable paradigm for disentangled visual attribute manipulation.
Limitations & Future Work¶
- Trade-off between deterministic transport and generative diversity: In one-to-many editing tasks like artistic style transfer, enforcing strict transport consistency slightly narrows stylistic diversity, leading to a minor uptick in FID (e.g., from 82.23 to 85.64 on OmniStyle).
- Physical limitations under extreme self-occlusions: While filtered rigorously by three vision-language models, purely 2D diffusion priors still lack the rigid physical precision of ray-tracing engines when rendering intricate fine-scale shadows (such as sub-surface scattering through hair or self-shadowing fingers).
- Future directions: Integrating explicit 3D surface normal priors into the transport loss and devising content-adaptive transport weights to handle dynamically shadowed video portraits.
Related Work & Insights¶
- vs IC-light: IC-light requires external environment map conditioning and labor-intensive scene capture priors; CFT requires no auxiliary maps, learning illumination transport directly in latent space with higher physical realism.
- vs Latent Intrinsic: Latent Intrinsic explicitly separates albedo and shading in latent space, where decomposition inaccuracies propagate and cause color fringing; CFT bypasses explicit intrinsic decomposition via end-to-end flow transport.
- vs LBM (Latent Bridge Matching): LBM performs instance-level image-to-image translation without explicit illumination constraints, often yielding muted lighting edits on complex scenes; CFT directly regulates the illumination velocity field.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Pioneering the integration of dual rectified flow transport with cross-instance trajectory supervision for image relighting.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous evaluation across multiple base architectures, real-world zero-shot transfer, user studies, comprehensive ablations, and style transfer extension.
- Writing Quality: ⭐⭐⭐⭐⭐ Lucid mathematical formulations with clear physical intuition and transparent empirical reporting.
- Value: ⭐⭐⭐⭐⭐ Sets a new standard for controllable diffusion editing and provides a valuable large-scale benchmark for portrait relighting.