Skip to content

RefDiT: Local Attribute Guidance in Reference-Based Image Generation

Conference: ECCV 2026
Paper: ECCV 2026
Area: Image Generation
Keywords: Reference-based image generation, Diffusion Transformer, local attribute guidance, triplet decomposition, cross-attention control

TL;DR

Addressing the failure of prior reference-guided methods to selectively guide generation from local regions in multi-object reference images, RefDiT introduces structured attribute triplet decomposition and a triplet-conditioned LoRA training strategy, enabling high-fidelity local attribute transfer without inference-time architecture overhead.

Background & Motivation

Reference-guided image generation has progressed rapidly, with subject personalization (e.g., DreamBooth) and style transfer (e.g., StyleDrop, B-LoRA, ZipLoRA) becoming staples in creative workflows. However, existing mainstream pipelines operate under an overly restrictive assumption: they assume the reference image depicts a solitary focal subject or embodies a single homogeneous global style. By training lightweight Low-Rank Adaptation (LoRA) adapters or learning a single identifier token, these methods compress all visual attributes of the reference into a monolithic global conditioning vector. When confronted with complex, multi-object real-world scenes where distinct objects exhibit disparate attributes, this coarse global conditioning triggers severe attribute bleeding. For instance, when presented with a scene containing a mug, a cat, and a toy cube, prior methods cannot isolate the mug to guide the generation of a "cup and saucer," frequently transferring irrelevant background wall hues onto the target object.

Recent DiT-based image editing methods (such as Flux.1 Kontext and FlowEdit) exhibit impressive capabilities in editing local attributes within an image. Nevertheless, their core objective is iteratively modifying existing pixel regions rather than generating entirely new objects guided by reference features. Treating novel image synthesis as a cascade of in-place edit steps quickly causes the model to diverge: cumulative editing errors erode source reference context within a few steps, ultimately yielding severe visual distortion. The fundamental tension is clear: users often want to synthesize completely new visual scenes borrowing the color, material, or style of a specific localized element, yet existing architectures lack a mechanism to autonomously establish region-level spatial correspondence between text attribute identifiers and reference image regions without requiring laborious manual bounding boxes or segmentation masks.

To resolve this limitation, this work rethinks the granularity of conditioning signals in reference guidance, moving from holistic global conditioning to entity-relational local attribute decomposition. Core idea: extract unstructured multi-object scenes into spatial-supervision-free object triplets with five fine-grained attribute dimensions via an MLLM, and introduce weight-shared auxiliary triplet attention streams alongside a dual-branch denoising consistency objective in multi-modal Diffusion Transformers (DiTs) to autonomously bind text attribute identifiers to local reference visual regions.

Method

Overall Architecture

RefDiT consists of two coordinated components: the Structural Conditioning Module (SCM), which extracts object triplets tagged with fine-grained attributes from the reference image and constructs both global and local prompt conditions; and the Triplet-Conditioned LoRA Training module, which teaches the multi-modal joint attention layers of the DiT to ground attribute identifier tokens to specific local image regions.

At inference time, given a reference image \(I\), a target text prompt \(P\), and an optional user guidance control \(M_c\), SCM first identifies all entities and interactions in the reference. It then matches the target prompt (or honors explicit user selection) to retrieve the most semantically relevant triplet, appending its attribute identifier tokens to the inference prompt. The adapted DiT performs standard diffusion denoising directly on the prompt, synthesizing novel objects that faithfully inherit local attributes without any architectural modification at inference.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Reference Image + Target Prompt + Optional Control"] --> B["Structured Triplet Attribute Decomposition<br/>MLLM extracts attribute triplets & builds identifiers"]
    B --> C["Weight-Shared Triplet Auxiliary Attention Streams<br/>Image patch & triplet text multi-modal joint alignment"]
    C --> D["Dual-Branch Denoising Consistency Constraint<br/>Global reconstruction & local triplet consistency optimization"]
    D --> E["Generated Image<br/>Faithfully inherits local attributes aligned with prompt"]

Key Designs

1. Structured Triplet Attribute Decomposition: Escaping Global Identifier Entanglement

Prior personalization methods rely on a single identifier token per image or per object to capture all reference details, leading to acute ambiguity in multi-object scenes. To provide unambiguous fine-grained control without spatial masks, SCM leverages a multi-modal large language model (such as GPT-4o or InternVL-3.5) to parse reference image \(I\) into a structured triplet set \(S_T = \{T_c^n\}\), where the \(n\)-th triplet is formalized as \([O_i, R_{ij}, O_j]\) with objects \(O_i, O_j\) and interaction \(R_{ij}\). For each object \(O_i\), SCM extracts five canonical attribute dimensions: \(A_{type} = \{\text{color, material, size, shape, global appearance}\}\). The global appearance attribute captures intrinsic textural cues (e.g., glossy, neon, matte, grainy) that standard color and material tags overlook. Each attribute value is converted into a dedicated identifier <attr_{type_j}_{value_{ij}}>. During generation, SCM matches inference prompts to local triplets via semantic label similarity or explicit user control \(M_c\), injecting the corresponding attribute tokens into the inference prompt for targeted local guidance.

2. Weight-Shared Triplet Auxiliary Attention Streams: Progressively Establishing Region-Level Correspondence

Directly training DiT weights under global reconstruction loss causes attribute identifier tokens to attend uniformly across the entire image space rather than focusing on their source objects. To induce localized attention, RefDiT leverages the multi-modal joint attention mechanism in MM-DiT: at each layer \(l\), patch embeddings \(x^l\) and global text tokens \(G_c^l\) are projected into concatenated queries, keys, and values \(q_{xc}^l, k_{xc}^l, v_{xc}^l\) to compute softmax attention. RefDiT establishes an auxiliary attention stream for each matched triplet \(T_{c,i}\), initializing visual candidate features with raw reference patches (\(T_{x,i}^0 = x_r^0\)). Crucially, this auxiliary stream shares the exact same DiT and LoRA projection weights \(\{Q_x, K_x, V_x\}\) and \(\{Q_c, K_c, V_c\}\):

\[T_{x,i}^{l+1}, T_{c,i}^{l+1} = \text{softmax}\left(\mathbf{q}_{txc}^l {\mathbf{k}_{txc}^l}^\top\right) \mathbf{v}_{txc}^l\]

where \(\mathbf{q}_{txc}^l, \mathbf{k}_{txc}^l, \mathbf{v}_{txc}^l\) are concatenated projections from triplet visual candidates and triplet text tokens. Across successive layers, the localized semantics of the triplet text pull visual attention progressively toward the relevant spatial coordinates of the object. This forces the LoRA weights to bind specific attribute tokens to localized visual patches during training, while leaving the inference-time architecture completely untouched.

3. Dual-Branch Denoising Consistency Constraint: Balancing Macro Semantics and Micro Attribute Fidelity

To prevent the model from overfitting to isolated local patches at the expense of global image coherence and unmentioned background context, RefDiT optimizes a composite objective balancing global reconstruction against local triplet consistency. The global reconstruction branch \(L_r\) enforces accurate denoising of the full reference latent conditioned on global signal \(G_c\). Concurrently, the triplet region consistency branch \(L_{c,i}\) measures denoising error when the network \(\epsilon_\Theta\) is conditioned on specific local triplet signals \(T_{c,i}\). The complete training objective is defined as:

\[\mathcal{L}_{\text{training}} = \alpha L_r + (1 - \alpha) \sum_{i=1}^K L_{c,i}\]

where \(K\) is the number of matched triplets, and \(\alpha \in [0, 1]\) balances global reconstruction and local attribute alignment. Empirical evaluations confirm that setting \(\alpha = 0.5\) provides the ideal equilibrium: higher \(\alpha\) values dilute local attribute precision into vague global stylization, whereas lower \(\alpha\) values compromise scene naturalness and overall image coherence.

Loss & Training

The framework is optimized under the Rectified Flow matching objective. Experiments are conducted using Stable Diffusion 3.5 and Flux as base DiT architectures. Following diffusion block heuristics, LoRA adapters are injected into multi-modal attention layers 12–24 and 30–37, with base weights and text encoders frozen. The LoRA rank is set to 32 with scaling factor \(\alpha = 0.5\), trained for 750 steps with batch size 1 using Adam at a learning rate of \(5 \times 10^{-5}\).

Key Experimental Results

Main Results

The evaluation benchmark comprises 500 test images synthesized from 60 complex multi-object reference images. In addition to standard CLIP image similarity (CLIP-I) and prompt alignment (CLIP-T), the authors introduce Attr-Match (measuring five-attribute matching accuracy via LLaMA-3.2-11B across a constrained vocabulary) and Attr-SIM (open-vocabulary attribute caption similarity).

Model CLIP-I ↑ CLIP-T ↑ Attr-Match ↑ Attr-SIM ↑
DLoRA (DreamBooth-LoRA) 0.71 0.29 0.48 0.41
B-LoRA 0.70 0.29 0.42 0.44
K-LoRA 0.77 0.27 0.58 0.50
UnZipLoRA 0.73 0.30 0.52 0.46
ProSpect 0.68 0.21 0.28 0.33
SDEdit 0.75 0.22 0.33 0.35
FlowEdit 0.73 0.24 0.38 0.38
FluxKontext 0.79 0.28 0.60 0.52
Commercial Models
Gemini-Banana 0.78 0.29 0.74 0.65
GPT-5 0.82 0.31 0.84 0.70
Ours
RefDiT (SD 3.5 DiT) 0.80 0.32 0.88 0.70
RefDiT (Flux DiT) 0.83 0.31 0.86 0.74

Ablation Study

Ablations conducted on SD 3.5 systematically evaluate the contributions of the SCM module, Triplet Conditioning (Triplet Cond.), the loss weighting parameter \(\alpha\), and alternative MLLM extractors.

Configuration SCM Triplet Cond. Loss Weight \(\alpha\) CLIP-I ↑ CLIP-T ↑ Attr-SIM ↑ Attr-Match ↑
M1 (DreamBooth baseline) \(\times\) \(\times\) - 0.71 0.29 0.43 0.48
M2 (Global SCM only) \(\checkmark\) \(\times\) - 0.77 0.27 0.57 0.66
M3 (Emphasizing global loss) \(\checkmark\) \(\checkmark\) 0.7 0.79 0.32 0.68 0.81
M4 (Emphasizing local triplet loss) \(\checkmark\) \(\checkmark\) 0.3 0.80 0.30 0.66 0.82
M5 (Open-source MLLM: InternVL-3.5) \(\checkmark\) \(\checkmark\) 0.5 0.73 0.30 0.63 0.80
RefDiT (Full model, GPT-4o) \(\checkmark\) \(\checkmark\) 0.5 0.80 0.32 0.70 0.88

Key Findings

  • Substantial Local Attention Alignment Gains: Moving from M1 to M2 shows that structured identifiers alone lift Attr-SIM from 0.43 to 0.57. Adding triplet-conditioned auxiliary attention elevates Attr-SIM to 0.70. Spatial cross-attention mask validation against SAM ground truth demonstrates a +0.17 mIoU improvement over M2.
  • Criticality of Loss Balance: Setting \(\alpha = 0.5\) strikes the optimal balance between global coherence and local specificity. Setting \(\alpha = 0.7\) inhibits local attribute fidelity, while \(\alpha = 0.3\) slightly compromises global scene harmony (Attr-SIM drops to 0.66).
  • Extractor Robustness: Replacing proprietary GPT-4o with open-source InternVL-3.5 (M5) achieves 0.80 Attr-Match and 0.63 Attr-SIM, still substantially outperforming all open-source baselines and proving that SCM is robust across underlying foundation vision-language models.
  • Strong Human Preference: In a blind user study comprising 45 participants and 450 total evaluations, users preferred RefDiT over UnZipLoRA in 92% of pairwise trials, over DreamBooth in 88%, and over FluxKontext in 80%. When benchmarked against commercial state-of-the-art GPT-5, 28% preferred RefDiT, 48% rated them comparable, and only 24% preferred GPT-5.

Highlights & Insights

  • Training-Auxiliary, Inference-Free Formulation: The auxiliary triplet attention stream operates exclusively during LoRA optimization to guide spatial localization; inference runs standard DiT feedforward execution with zero memory or latency penalties.
  • Five-Dimensional Attribute Disentanglement: Decomposing holistic appearance into color, material, size, shape, and global appearance eliminates semantic ambiguity, preventing cross-object attribute leakage.
  • Transferable Cross-Modal Attention Calibration: Binding localized text tokens to visual patch streams through weight-shared multi-modal attention provides a generalizable paradigm for multi-subject personalization, localized texture control, and virtual try-on.

Limitations & Future Work

  • Dependence on MLLM Parsing Accuracy: In scenes with severe occlusion or rare niche objects, extraction errors in object detection or attribute labeling can propagate directly into the conditioning prompt.
  • Per-Image Optimization Overhead: Optimizing LoRA parameters for 750 steps per reference image introduces moderate computational overhead. Future research could explore distilling the triplet-grounding mechanism into a zero-shot, feedforward cross-attention adapter.
  • Complex Dynamic Relational Composition: Current conditioning models physical contact and interaction statically; synthesizing scenes with intricate geometric deformations or dynamic multi-object kinematics remains challenging.
  • vs B-LoRA / K-LoRA / UnZipLoRA: These methods decompose images into global content and style LoRA blocks, which conflates multiple objects into an entangled representation in multi-object scenes; RefDiT introduces object-level triplets and dedicated attribute tokens for selective local control.
  • vs Flux.1 Kontext / FlowEdit: DiT editing approaches rely on iterative latent modification of an existing image, which rapidly degrades source context when generating novel objects; RefDiT extracts decoupled attribute knowledge from references and injects it cleanly into text-to-image synthesis.
  • vs ProSpect / MATTE: Previous attribute control methods relied on U-Net textual inversion step-wise scheduling; RefDiT exploits the native dual-stream joint attention of modern DiT architectures and dual-branch denoising consistency, achieving significantly superior spatial precision.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering framework for local attribute-level reference guidance using structured triplets and auxiliary DiT attention streams.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across curated multi-object benchmarks, new fine-grained metrics (Attr-Match, Attr-SIM), extensive ablations, and rigorous user studies.
  • Writing Quality: ⭐⭐⭐⭐⭐ Exceptionally clear narrative, disciplined methodology formulation, and well-substantiated architectural motivations.
  • Value: ⭐⭐⭐⭐⭐ Offers critical insights and architectural solutions for fine-grained reference-conditioned synthesis in the Diffusion Transformer era.