Through Van Gogh’s Eyes: Global Style Transfer with Diffusion Model¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Image Generation
Keywords: Global Style Transfer, Diffusion Models, h-space Guidance, Content Alignment Guidance, Artistic Image Synthesis
TL;DR¶
This paper introduces Global Style Transfer (GST), a many-to-one artistic synthesis paradigm that learns a distribution-level style representation from an artist's full artwork collection via residual h-space modulation under a fixed prompt, complemented by training-free CLIP perceptual guidance for content alignment and style-driven deformation.
Background & Motivation¶
Artistic image synthesis aims to recreate the unique expressive visual identity and aesthetic perspective of a target master painter. For years, the field has been largely divided between two main paradigms: conventional neural style transfer (NST) and text-to-image (T2I) diffusion models conditioned on artist names (e.g., "in the style of Van Gogh"). However, traditional style transfer relies on a single or few reference artworks in a one-to-one setup, which inevitably overfits to the localized low-level color palette and brushstrokes of a single exemplar while failing to capture the broader stylistic distribution formed across an artist's extensive oeuvre. On the other hand, name-conditioned T2I diffusion models provide synthesis flexibility but suffer from severe text-dependent artistic bias, frequently collapsing into stereotypical visual motifs from a handful of iconic masterpieces (such as swirling patterns from The Starry Night), thereby neglecting the true breadth of an artist's style.
The core tension lies in the fact that capturing an artist's authentic visual identity requires aggregating high-order distributions across their full artistic corpus, yet existing mechanisms are either constrained by pixel-level statistics of individual instances or distorted by cross-modal linguistic priors. Escaping this dilemma requires bypassing both single-image reference reliance and text-mediated conditioning, establishing a mechanism that learns directly from collective visual statistics while permitting controlled geometric deformation over content scenes.
This paper addresses this challenge by reformulating artistic synthesis as a many-to-one Global Style Transfer (GST) paradigm. Core idea: train a lightweight style extraction function in the semantic h-space of the diffusion U-Net bottleneck under a constant, zero-variance prompt to learn text-independent visual style offsets from multiple artworks, coupled with training-free Content Alignment Guidance based on Tweedie's approximation and deep CLIP perceptual features to preserve semantic structure while enabling artist-specific geometric deformation.
Method¶
Overall Architecture¶
Global Style Transfer takes a single content image \(I_c\) and a collection of \(K\) artworks \(I_s = \{i_s^1, i_s^2, \dots, i_s^K\}\) by a target artist as input. The entire pipeline comprises two coordinated components: offline Global Style Guidance (GSG) and inference-time Content Alignment Guidance (CAG). During offline training, GSG pairs all artworks with a unified generic prompt ("A painting") and optimizes a lightweight Style Extraction Function (SEF) \(f_t\) via noise reconstruction to capture a residual global style offset \(\Delta \mathbf{h}_t\) in the U-Net bottleneck \(h\)-space. At inference time, the content image \(I_c\) is mapped into a noisy latent \(z_T\) via DDIM Inversion. During the subsequent denoising trajectory, \(\Delta \mathbf{h}_t\) is injected into the bottleneck via an asymmetric reverse process to drive global stylization, while CAG computes perceptual gradients between Tweedie-approximated clean images and noisy representations using deep CLIP features to steer content fidelity and structural deformation.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Artwork Collection Is and Content Image Ic"] --> B["Stage 1: Global Style Guidance (GSG)<br/>Train lightweight SEF ft in h-space under fixed prompt"]
B --> C["Stage 2: DDIM Inversion & Asymmetric Sampling<br/>Invert content to zT, inject global residual offset Δht"]
C --> D["Stage 3: Content Alignment Guidance (CAG)<br/>Tweedie approximation and deep CLIP gradient latent update"]
D --> E["Output: Artist-level Stylized Image x0"]
Key Designs¶
1. Global Style Guidance: Learning Text-Independent Style Distributions in h-Space To eliminate linguistic bias from text prompts and avoid overfitting to specific exemplars, the framework models style semantics in the intermediate activation space of the U-Net bottleneck (\(h\)-space). Empirical analysis reveals that even under progressive Gaussian noise corruption across diffusion timesteps \(t\), features in \(h\)-space remain robustly separable between different artists in t-SNE projections, confirming that \(h\)-space preserves stable artist-level semantic style. The authors define a lightweight Style Extraction Function (SEF) \(f_t(h_t; \theta)\) parameterized as a single-hidden-layer MLP with zero-initialization to ensure an unbiased starting state. During training, all artworks in \(I_s\) are paired with an identical generic prompt \(y = \text{“A painting”}\), driving prompt conditioning variance to zero and forcing \(f_t\) to learn purely from visual statistics via noise reconstruction: $\(\mathcal{L}_{\text{SEF}} := \mathbb{E}_{z \sim \mathcal{E}(x), \epsilon_t, t} \left\lVert \epsilon_\theta(z_t, t, \tau_\phi(y) \mid \Delta \mathbf{h}_t) - \epsilon_t \right\rVert_2^2\)$ where \(\Delta \mathbf{h}_t = f_t(h_t; \theta)\). During reverse diffusion, an asymmetric reverse process applies this modulation exclusively to the predicted clean latent component \(P_t(\epsilon_\theta(z_t \mid \Delta \mathbf{h}_t))\) while keeping the directional trajectory component \(D_t(\epsilon_\theta(z_t))\) unchanged, achieving faithful global style transfer without degrading generative reconstruction quality.
2. Content Alignment Guidance: Deep Perceptual Gradient Steering for Structure and Deformation Conventional style transfer either over-constrains low-level edges or destroys global composition during stylization. Genuine artistic creation preserves high-level topological semantics while allowing distinctive brushwork and geometric distortion. Inspired by conditional score-based diffusion formulations and Tweedie's formula, the authors devise a training-free guidance mechanism. At each reverse diffusion step \(t\), Tweedie's formula approximates the clean latent \(\tilde{z}_0\) from noisy state \(z_t\), yielding a decoded coarse image \(\tilde{x}_0 \approx \mathcal{D}(\tilde{z}_0 \mid z_t)\). The perceptual distance is computed against the current decoded image \(x_t = \mathcal{D}(z_t)\) using feature activations from layer \(l\) of a pre-trained CLIP image encoder: $\(\ell(z_t) = \left\lVert \mathcal{E}_{\text{CLIP}}^l(\tilde{x}_0) - \mathcal{E}_{\text{CLIP}}^l(x_t) \right\rVert_2\)$ The gradient of this perceptual loss with respect to latent \(z_t\) acts as a conditioning score term to update the latent: \(\tilde{z}_t = z_t - s \nabla_{z_t} \ell(z_t)\), where \(s\) controls guidance strength. By selecting a deep layer (Layer 11) of CLIP, guidance focuses strictly on abstract semantic composition rather than low-level color or texture edges, granting the diffusion process the necessary flexibility for expressive, artist-specific geometric deformations.
Loss & Training¶
The lightweight SEF \(f_t\) is trained with the Adam optimizer while keeping the base diffusion backbone frozen. Artworks are randomly sampled from the artist's collection, perturbed through the forward diffusion process, and fitted using the residual noise reconstruction loss under the fixed prompt. Training epoch analysis demonstrates that while 10 to 20 epochs yield generic stylization, 200 epochs successfully capture nuanced brush techniques, stroke dynamics, and palette distributions specific to individual masters. During sampling, standard Classifier-Free Guidance (CFG) combined with DDIM inversion produces high-fidelity results without fine-tuning on the test content images.
Key Experimental Results¶
Main Results¶
Evaluation was conducted on WikiArt (containing over 80,000 artworks across 1,000 artists) and VanGogh2Photo (real photographic landscapes paired with artistic counterparts). Quantitative metrics include ArtFID (style and content fidelity), FID (overall generative quality), CFSD (content preservation), CLIP-Div (stylistic diversity relative to real artwork collections, reported as Ours/Real), and 1-Precision (1-Prec., measuring avoidance of memorization). Baselines include vanilla Stable Diffusion (S.D), Textual Inversion, Custom Diffusion, and LoRA-based fine-tuning.
| Artist | Method | FID (↓) | ArtFID (↓) | CFSD (↓) | CLIP-Div (↑) (Ours/Real) | 1-Prec (↑) |
|---|---|---|---|---|---|---|
| Van Gogh | S.D | 11.97 | 23.78 | 0.7428 | 0.178 | 0.257 |
| Textual Inversion | 7.67 | 13.68 | 0.1674 | 0.220 | 0.550 | |
| Custom Diffusion | 9.49 | 14.71 | 0.1208 | 0.289 | 0.933 | |
| LoRA based Fine-tuning | 11.01 | 21.38 | 0.1886 | 0.261 | 0.913 | |
| Ours (GST) | 9.46 | 19.25 | 0.2896 | 0.297 / 0.335 | 0.988 | |
| Chagall | S.D | 16.26 | 32.44 | 0.3117 | 0.224 | 0.844 |
| Textual Inversion | 11.72 | 21.15 | 0.1683 | 0.238 | 0.877 | |
| Custom Diffusion | 18.06 | 26.70 | 0.1108 | 0.289 | 0.962 | |
| LoRA based Fine-tuning | 17.30 | 32.59 | 0.1674 | 0.266 | 0.928 | |
| Ours (GST) | 13.25 | 26.39 | 0.2626 | 0.311 / 0.293 | 0.977 | |
| Renoir | S.D | 16.72 | 33.70 | 0.1584 | 0.225 | 0.987 |
| Textual Inversion | 12.14 | 21.99 | 0.1100 | 0.238 | 0.895 | |
| Custom Diffusion | 15.15 | 22.77 | 0.1160 | 0.294 | 0.987 | |
| LoRA based Fine-tuning | 15.36 | 29.23 | 0.1749 | 0.279 | 0.988 | |
| Ours (GST) | 12.27 | 24.70 | 0.1566 | 0.318 / 0.313 | 0.999 |
Ablation Study¶
The ablation study analyzes the role of Content Alignment Guidance (CAG) in balancing content preservation and stylization, along with the impact of CLIP feature layer depth.
| Config / Variant | FID (↓) | ArtFID (↓) | Note |
|---|---|---|---|
| Ours w/o CAG | 11.05 | 22.45 | Removing CAG causes severe structural distortion; ArtFID worsens (+3.20) |
| Ours (Full Model, Layer 11) | 9.46 | 19.25 | Full model: deep semantic guidance balances artistic deformation with content structure |
| CLIP Layer Depth Ablation (Qualitative) | |||
| - Low/Mid Layers (Layer 1~9) | - | - | Over-constrains low-level colors and textures, introducing artificial building-like artifacts |
| - High Layer (Layer 11) | - | - | Encodes high-level scene composition, ensuring robust structure and natural stylization |
Key Findings¶
- Stylistic Diversity Matches Real Art Distribution: Across all tested masters, GST achieves CLIP-Diversity scores remarkably consistent with real artwork collections (e.g., Van Gogh 0.297 vs. Real 0.335, Chagall 0.311 vs. Real 0.293, Renoir 0.318 vs. Real 0.313), far exceeding baseline models that overfit to narrow visual modes.
- Superior Avoidance of Memorization: GST achieves near-perfect 1-Precision scores (0.999 for Renoir, 0.988 for Van Gogh), demonstrating that the synthesized outputs are genuine novel generations sampling from the learned style manifold rather than memorized replications of specific paintings.
- Crucial Decoupling via Deep Semantic Guidance: Omitting CAG sharply degrades ArtFID from 19.25 to 22.45. Furthermore, constraining guidance to deeper semantic layers (Layer 11) decouples high-level composition from localized painterly deformation, whereas lower layers impose rigid texture constraints that conflict with artistic stroke synthesis.
Highlights & Insights¶
- Many-to-One Paradigm Shift: Replaces the conventional assumption that style equals a single exemplar with a distribution-level formulation over an artist's full portfolio, resolving instance-level overfitting in style transfer.
- Zero-Variance Text-Independent h-Space Guidance: Restricting training prompts to a constant generic token (
"A painting") removes textual variance, allowing the residual offset in the diffusion U-Net bottleneck to capture purely visual artistic statistics. - Training-Free Perceptual Latent Steering: Connecting Tweedie-approximated clean projections with deep CLIP perceptual gradients allows zero-shot structural preservation with expressive painterly deformations without test-time fine-tuning.
Limitations & Future Work¶
- Dependency on Artwork Corpus Size: Establishing a representative global distribution requires dozens to hundreds of cataloged paintings, which restricts applicability to artists with very few surviving works.
- Multi-Period Style Evolution: Many painters underwent distinct artistic periods (e.g., Picasso's Blue Period vs. Cubism); aggregating all works into a single distribution blends these periods, suggesting the need for temporal or cluster-based subspace modeling.
- Hyperparameter Sensitivity: Balancing style guidance weight \(w\) and content guidance scale \(s\) requires manual adjustment, occasionally leading to boundary bleeding on highly complex photographic inputs.
Related Work & Insights¶
- vs. Conventional Neural Style Transfer (e.g., StyleInjection, StyTR2, CSGO): Traditional methods apply one-to-one transfers; averaging multiple stylized instances produces uncoordinated, blurred artifacts. GST directly models the global manifold for coherent, multi-attribute synthesis.
- vs. Personalization Methods (Textual Inversion, Custom Diffusion, DreamBooth): Personalization approaches often suffer from prompt entanglement or mode collapse toward iconic compositions. GST relies entirely on visual statistics, achieving diversity metrics virtually identical to authentic artwork collections.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Introduces the Many-to-One Global Style Transfer paradigm, resolving both single-image overfitting and text-induced bias.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across fidelity, diversity, and memorization metrics with rigorous ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical formulations, intuitive visualizations, and well-structured algorithms.
- Value: ⭐⭐⭐⭐⭐ Provides a foundational methodology for computational art history, digital cultural preservation, and controllable diffusion synthesis.