Skip to content

InstantRetouch: Personalized Image Retouching without Test-time Fine-tuning

Conference: ECCV 2026
Paper: ECCV Official Link
Area: Image Generation
Keywords: Personalized Image Retouching, Asymmetric Auto-Encoder, Retrieval-Augmented Retouching, Tuning-Free Style Transfer, Photorealistic Style Transfer

TL;DR

RefRetouch (titled InstantRetouch) achieves generic personalized image retouching without test-time fine-tuning by coupling a LoRA-adapted Siamese SigLIPv2 encoder with a color-space conditional MLP decoder to disentangle content from retouching style, and introduces retrieval-augmented retouching (RAR) to dynamically aggregate relevant reference styles based on photometric content similarity while generalizing out of the box to photorealistic style transfer.

Background & Motivation

Image retouching enhances photographs to elevate visual aesthetics and convey specific artistic narratives, but because aesthetic preferences are inherently subjective, retouching tastes vary substantially across individuals. Conventional learning-based retouching models (such as 3D LUTs and deep bilateral learning networks) fit static transformations learned from curated expert datasets. When adapting to a new user's taste, they typically demand collecting dozens of paired images and retraining the underlying network, severely curtailing real-world usability. While several attempts have emerged to personalize retouching from a handful of reference pairs, prior methodologies built on style classification (e.g., StarEnhancer), contrastive metric learning (e.g., PieNet), or generic residual auto-encoders (e.g., MSM) produce latent representations that heavily entangle retouching style with scene semantics, leading to poor generalization across unseen camera viewpoints and compositional distributions.

The core tension lies in the content-dependent nature of real-world user retouching: users do not enforce an identical adjustment across all photographs, but instead dynamically calibrate brightness, contrast, and color balance according to specific scene conditions (e.g., underexposed indoor scenes versus brightly lit outdoor landscapes). Existing personalization methods routinely compress multiple reference samples into a single global vector via naive average pooling, completely stripping out the fine-grained relationship between query image content and individual reference styles. Furthermore, recent in-context visual generation approaches based on large diffusion transformers suffer from severe computational overhead and struggle to scale beyond one or two references due to limited effective context windows.

This paper's angle is that true content-style disentanglement requires an architectural asymmetry that deprives the decoder of spatial memory, forcing it to act strictly as a per-pixel color mapper while the encoder abstracts semantic cues. Core idea: construct an asymmetric auto-encoder pairing a Siamese LoRA-adapted SigLIPv2 encoder with a lightweight per-pixel color-space conditional MLP decoder, and introduce Retrieval-Augmented Retouching (RAR) to retrieve and aggregate top-\(K\) reference style latents based on photometric similarity, enabling entirely tuning-free, content-aware personalized retouching.

Method

Overall Architecture

RefRetouch operates through two core phases: style extraction and adaptive application. During style extraction, an asymmetric auto-encoder processes a user's paired reference example through a Siamese vision encoder to produce a low-dimensional, content-disentangled style latent vector. During adaptive application, the Retrieval-Augmented Retouching (RAR) module evaluates the photometric cosine similarity between the unseen query image and the reference inputs, retrieves the top-\(K\) most contextually relevant reference pairs, aggregates their style latents via Softmax weighting, and feeds the composite latent into a lightweight conditional MLP decoder operating in color space to synthesize the retouched query.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["User Reference Pairs (x_i, y_i)"] --> B["Asymmetric Auto-Encoder<br/>Siamese SigLIPv2 extracts style latent z_i"]
    C["Query Input Image x_q"] --> D["Retrieval-Augmented Retouching RAR<br/>Extract tone embeddings & compute s_i"]
    B --> E["Top-K Selection & Softmax Aggregation<br/>Aggregate into query style latent z_q"]
    D --> E
    E --> F["Color-Space Conditional MLP Decoder<br/>Inject z_q into per-pixel mapping"]
    C --> F
    F --> G["Personalized Retouched Output y_hat_q"]

Key Designs

1. Asymmetric Auto-Encoder: Structural and Operational Asymmetry for Content-Style Disentanglement

Traditional auto-encoders commonly employ mirror-symmetric U-Net designs with skip connections that predict residual maps. While expressive, spatial convolutional decoders easily memorize spatial layout and high-level scene composition, inadvertently leaking semantic content into the latent style space. RefRetouch establishes structural and domain asymmetry: the encoder \(E_\theta\) employs a Siamese vision backbone initialized from SigLIPv2, where trainable LoRA adapters (rank 16) are inserted into all attention linear projections to process the original reference image \(x \in \mathcal{X}\) and retouched target \(y \in \mathcal{Y}\). Pooled representations from both branches are concatenated and projected to a 2048-dimensional style latent \(z = E_\theta(x, y)\). In contrast, the decoder \(D_\phi\) is deliberately designed as a lightweight conditional MLP operating strictly in per-pixel color space (\(\mathbb{R}^3 \to \mathbb{R}^3\)). Because the decoder is stripped of spatial receptive fields, it cannot reconstruct local textures from parametric memory and is forced to interpret \(z\) purely as global color and tone transformation rules.

2. Layer-wise Additive Latent Injection: Stable and Lightweight Tone Modulation

To effectively steer the per-pixel mapping without excessive compute, the conditional MLP decoder comprises an input layer mapping raw RGB pixels to a 128-dimensional hidden representation, followed by three conditional MLP blocks with hidden dimensions \([256, 512, 3]\). All intermediate blocks apply ReLU activations, while a terminal Sigmoid squashes the output into the normalized dynamic range \([0, 1]\). Rather than relying on computationally heavy cross-attention mechanisms or adaptive layer normalization (AdaLN) which can introduce instability in dense point-wise transformations, the 2048-dimensional style latent \(z\) is projected via a tiny linear layer with ReLU to match each block's hidden dimension and added directly to the layer features. Empirical ablation confirms that this additive conditioning delivers optimal reconstruction fidelity while eliminating spatial overfitting.

3. Retrieval-Augmented Retouching (RAR): Photometric Retrieval and Context-Adaptive Aggregation

When presented with an extensive user portfolio \(\mathcal{P} = \{(x_i, y_i)\}_{i=1}^M\) spanning varied lighting environments and color schemes, global style averaging smears fine-grained adjustments across disparate scenes. RAR overcomes this by leveraging the LoRA-adapted SigLIPv2 backbone to extract content embeddings \(c_q\) for query image \(x_q\) and \(\{c_i\}\) for each reference input \(x_i\). Adaptation shifts the backbone representation away from abstract semantics toward exposure, color distribution, and tonal characteristics. Given cosine similarity \(s_i = \frac{c_q \cdot c_i}{\|c_q\| \|c_i\|}\), RAR retrieves the top-\(K\) nearest neighbors \(\mathcal{N}_K\) (default \(K=3\)) and computes a normalized Softmax weighting modulated by temperature \(\tau=0.1\):

\[w_i = \frac{\exp(s_i / \tau)}{\sum_{j \in \mathcal{N}_K} \exp(s_j / \tau)}, \quad z_q = \sum_{i \in \mathcal{N}_K} w_i z_i\]

The resulting composite style latent \(z_q\) encapsulates the retouching transformation most appropriate for the query image's specific tonal distribution, yielding the final output \(\hat{y}_q = D_\phi(x_q, z_q)\).

4. Zero-Shot Photorealistic Style Transfer: Synthetic Paired Formulation for Unpaired Inputs

RefRetouch generalizes seamlessly to unpaired photorealistic style transfer without task-specific retraining. Given an arbitrary unretouched content image \(x_q\) and an unpaired artistic style photograph \(y_{style}\), the method simply recycles \(x_q\) as a pseudo-input, feeding \((x_q, y_{style})\) into the Siamese encoder to extract \(z = E_\theta(x_q, y_{style})\). Because the asymmetric auto-encoder maps only relative chromatic and tonal shifts, it completely bypasses semantic discrepancies between \(x_q\) and \(y_{style}\). Applying \(z\) to \(x_q\) through the conditional MLP decoder transfers the color palette and atmosphere of \(y_{style}\) while preserving 100% of the content image's photorealistic structure.

Loss & Training

The framework is pre-trained end-to-end on a curated dataset of 95,000 paired images transformed by 800 authentic Adobe Lightroom presets. The optimization objective is a pixel-level \(\ell_1\) reconstruction loss over the training distribution \(\mathcal{D}\):

\[\mathcal{L}_{\text{recon}} = \mathbb{E}_{(x, y) \sim \mathcal{D}} \| D_\phi(x, E_\theta(x, y)) - y \|_1\]

The network is optimized using Adam for 180,000 iterations with a batch size of 8 on random \(384 \times 384\) crops, enforcing precise color matching while preserving training efficiency.

Key Experimental Results

Main Results

RefRetouch is evaluated across three rigorous benchmarks: single-style consistency (VCIRB), multi-style consistent groups (PPR10K-Groups), and diverse, inconsistent historical user edits (MIT-FiveK and PPR10K). The table below details the quantitative performance on VCIRB and PPR10K-Groups:

Dataset Method PSNR (1 Ref)↑ PSNR (Max Ref)↑ SSIM (1 Ref)↑ SSIM (Max Ref)↑ LPIPS (1 Ref)↓ LPIPS (Max Ref)↓
VCIRB StarEnhancer 21.00 21.27 (4 Ref) 0.864 0.868 (4 Ref) 0.123 0.120 (4 Ref)
VCIRB MSM 17.69 18.53 (4 Ref) 0.822 0.836 (4 Ref) 0.164 0.150 (4 Ref)
VCIRB MSM† (Retrained) 24.35 25.28 (4 Ref) 0.900 0.905 (4 Ref) 0.097 0.093 (4 Ref)
VCIRB VisualCloze 19.81 12.05 (4 Ref) 0.765 0.378 (4 Ref) 0.173 0.591 (4 Ref)
VCIRB VisualCloze† (Retrained) 23.91 24.70 (4 Ref) 0.830 0.825 (4 Ref) 0.077 0.079 (4 Ref)
VCIRB RefRetouch (Ours) 29.13 29.82 (4 Ref) 0.963 0.965 (4 Ref) 0.048 0.046 (4 Ref)
PPR10K-Groups StarEnhancer 19.16 19.36 (6 Ref) 0.875 0.866 (6 Ref) 0.102 0.119 (6 Ref)
PPR10K-Groups MSM 18.86 18.59 (6 Ref) 0.877 0.878 (6 Ref) 0.154 0.155 (6 Ref)
PPR10K-Groups MSM† (Retrained) 19.27 18.16 (6 Ref) 0.876 0.866 (6 Ref) 0.149 0.170 (6 Ref)
PPR10K-Groups VisualCloze 15.04 7.74 (6 Ref) 0.453 0.246 (6 Ref) 0.220 0.844 (6 Ref)
PPR10K-Groups VisualCloze† (Retrained) 21.12 20.98 (6 Ref) 0.832 0.812 (6 Ref) 0.097 0.112 (6 Ref)
PPR10K-Groups RefRetouch (Ours) 22.51 25.27 (6 Ref) 0.932 0.961 (6 Ref) 0.065 0.043 (6 Ref)

Note: † indicates baseline models re-trained or fine-tuned on the author's collected Lightroom preset dataset.

On the multi-style, inconsistent benchmarks with 20, 50, and 100 historical references from MIT-FiveK and PPR10K, RefRetouch achieves outstanding performance: on MIT-FiveK (100 references), it scores 23.34 dB PSNR (0.920 SSIM), surpassing tuning-free MSM (22.05 dB) and competitive with methods requiring user-specific training such as RSFNet (22.66 dB) and AdaInt (22.49 dB). On PPR10K (100 references), it attains 22.47 dB PSNR and 0.942 SSIM, consistently beating all tuning-free alternatives.

Ablation Study

The table below validates the architectural components of the asymmetric auto-encoder on a held-out validation set:

Exp № Encoder Backbone Decoder Architecture Conditioning Method PSNR (dB)↑ SSIM↑
1 (Full Model) SigLIPv2 + LoRA Color-Space Conditional MLP Additive Injection (Add) 32.40 0.977
2 ResNet-50 Color-Space Conditional MLP Additive Injection (Add) 30.00 0.971
3 SigLIPv2 (Frozen) Color-Space Conditional MLP Additive Injection (Add) 23.41 0.931
4 SigLIPv2 + LoRA Spatial-Domain U-Net Residual Addition 29.58 0.944
5 SigLIPv2 + LoRA Color-Space Conditional MLP Adaptive Layer Norm (AdaLN) 31.93 0.977
6 SigLIPv2 + LoRA Color-Space Conditional MLP Cross-Attention 22.66 0.880

In the ablation of the RAR retrieval mechanism on MIT-FiveK with 100 available references: setting \(k=1\) achieves 21.94 dB PSNR; \(k=3\) reaches 23.34 dB; \(k=5\) yields 23.82 dB; whereas naively averaging all 100 available references degrades performance down to 21.94 dB.

Key Findings

  • Asymmetric decoding enforces true disentanglement: Restricting the decoder to a per-pixel color-space MLP delivers a +2.82 dB PSNR advantage over a spatial U-Net (Exp 1 vs Exp 4). Depriving the decoder of spatial convolutional filters prevents semantic leakage, forcing the encoder to encapsulate purely photometric transformations in \(z\).
  • LoRA adaptation redirects visual focus: A frozen SigLIPv2 achieves only 23.41 dB PSNR, whereas adding LoRA adapters improves performance to 32.40 dB. Embedding retrieval visualizations confirm that fine-tuning pivots the model's feature space from semantic categorization to tone, exposure, and color temperature discrimination.
  • Selective retrieval prevents style interference: In multi-reference environments, naive global pooling blends mutually conflicting adjustments across diverse scenes. RAR's top-\(K\) aggregation yields monotonic performance gains as reference counts grow (improving from 22.51 dB at 1 reference to 25.27 dB at 6 references on PPR10K-Groups).

Highlights & Insights

  • Dual Structural-Spatial Asymmetry: Pairs a high-capacity vision foundation model with an ultra-lightweight, zero-receptive-field MLP decoder, resolving the persistent trade-off between style expressiveness and semantic overfitting.
  • Seamless RAG Transfer to Photometric Conditioning: Adapts retrieval-augmented generation to continuous visual tone modeling, allowing the system to handle complex user preferences across varied scenes without retraining.
  • Zero-Shot Multi-Task Generalization: Reusing the query image as a pseudo-input enables out-of-the-box photorealistic style transfer without architectural modifications, demonstrating the robustness of the learned style manifold.

Limitations & Future Work

  • Global-Only Color Transformation: Operating on isolated pixels without spatial coordinates means the model cannot execute localized adjustments such as facial skin smoothing or sky-mask exposure recovery.
  • Limited Control Over High-Contrast Gradients: In scenes with severe dynamic range challenges or complex backlit shadows, 1D global style latents struggle to handle spatially non-uniform gradient curves.
  • Future Directions: The authors suggest extending the framework by integrating spatial decomposition or region-level semantic masks to facilitate localized, mask-guided personalized retouching.
  • vs StarEnhancer / PieNet: StarEnhancer relies on a discrete style classifier, failing on open-ended styles; PieNet employs coarse contrastive metric learning. RefRetouch utilizes continuous reconstruction-driven style latents with LoRA adaptation, capturing nuanced color adjustments.
  • vs MSM (Masked Style Modeling): MSM uses a spatial U-Net that entangles semantic layout with style, and trains on synthetically degraded pseudo-pairs. RefRetouch adopts an asymmetric MLP decoder and trains on real-world Lightroom presets, achieving substantially superior reconstruction and style fidelity.
  • vs VisualCloze & In-Context Generation: Large visual in-context diffusion models incur massive compute costs and degrade when scaling to multiple references. RefRetouch operates with millisecond latency, scales effectively to 100+ references, and preserves photorealistic detail without generative artifacts.

Rating

  • Novelty: ⭐⭐⭐⭐ [The dual-asymmetry auto-encoder and retrieval-augmented retouching offer an elegant, principled formulation]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Thorough benchmarking across four distinct operational setups with comprehensive ablations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, self-contained mathematical formulation, and well-structured empirical analysis]
  • Value: ⭐⭐⭐⭐⭐ [Delivers highly practical, real-time personalized retouching on edge devices without test-time retraining]