Skip to content

RCEdit-500K: Reference Completion for Image-Conditioned Image Editing

Conference: ECCV 2026
Paper: ECCV 2026
Code: https://huggingface.co/datasets/carpedkm/RCEdit-500K
Area: Image Generation
Keywords: image-conditioned image editing, reference completion, dataset curation, weak-instruction augmentation, multimodal editing

TL;DR

Reformulates image-conditioned image editing (ICIE) data curation as a reference-completion problem by synthesizing type-specific reference images for existing high-quality text-guided editing triplets, building RCEdit-500K—the first unified open dataset spanning 477K quadruplets across six edit types.

Background & Motivation

Text-conditioned image editing (TCIE) has seen rapid progress, but natural language inherently struggles to communicate fine-grained or distribution-level visual attributes. Subtleties like lighting atmosphere, photographic color palettes, fine material textures, and artistic styles are difficult to specify through text prompts alone. This limitation drives the demand for image-conditioned image editing (ICIE), where visual manipulation is guided jointly by an input anchor image, a descriptive text instruction, and a visual reference image. However, while closed-source proprietary systems exhibit strong multi-reference editing capabilities, the open-source community lacks a unified, large-scale ICIE dataset. Existing public datasets only cover narrow subsets such as simple object insertion or material transfer at negligible scales of 10K to 40K samples, leaving open-source ICIE models severely lagging behind.

The primary obstacle lies in the standard data construction paradigm. Conventional forward synthesis attempts to generate all four elements—input image, reference image, instruction, and edited target—from scratch using chained generative models. In practice, this multi-stage synthesis is computationally prohibitive and compounds model biases, yielding perceptual artifacts and severe distributional degradation that undermine visual fidelity. Conversely, existing high-quality TCIE corpora (such as Pico-Banana-400K and GPT-IMAGE-EDIT-1.5M) already supply three of the four required components: the input image, the editing instruction, and the corresponding edited target. The only missing element is a compatible reference image.

Recognizing this fundamental insight, this paper shifts away from error-prone forward synthesis and reformulates ICIE data construction as a reference-completion problem. Core idea: complete existing curated TCIE triplets into ICIE quadruplets by synthesizing aligned reference images via edit-type-specific pipelines, integrating weak-instruction augmentation to enforce visual conditioning, and pruning noisy candidates with a strict five-dimensional VLM post-filter.

Method

Overall Architecture

In the ICIE setting, each sample is represented as a quadruplet \(\langle I_{\text{in}}, I_{\text{ref}}, T_{\text{ins}}, I_{\text{tgt}} \rangle\). The input image \(I_{\text{in}}\) serves as an anchor whose unedited regions and global spatial structure must remain intact; the reference image \(I_{\text{ref}}\) supplies visual conditioning; the instruction \(T_{\text{ins}}\) specifies the intended operation; and \(I_{\text{tgt}}\) is the edited target image. The reference-completion pipeline starts from curated TCIE triplets \(\langle I_{\text{in}}, T_{\text{orig}}, I_{\text{tgt}} \rangle\). First, a vision-language model (GPT-4o) classifies the edit into one of six categories, synthesizes specialized prompts, and adapts the original text instruction. Next, four specialized generative pathways produce aligned reference images. A weak-instruction augmentation step deliberately abstracts entity details to prevent textual bypass. Finally, a five-dimensional VLM filtering strategy evaluates candidate quadruplets along orthogonal axes to eliminate low-quality samples, yielding 477K high-quality training pairs.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Curated TCIE Triplets<br/>Input + Original Text + Target"] --> B["VLM Classification & Adaptation<br/>Categorize 6 Types / Generate Prompts / Rewrite Instruction"]
    B --> C["Four Type-Specific Synthesis Pipelines<br/>Object Segmentation / Spatial Cue / Inpainting / Style Transfer"]
    C --> D["Weak-Instruction Augmentation<br/>Abstract Entity Details / Enforce Visual Conditioning"]
    D --> E["Five-Dimensional VLM Post-Filtering<br/>Independent Scoring / Strict Axis-Wise Rejection"]
    E --> F["RCEdit-500K Unified Quadruplet Dataset<br/>477K Samples / 1024 Resolution / Concrete & Abstract"]

Key Designs

1. Four Type-Specific Reference Synthesis Pipelines: Decoupled Visual Sourcing by Edit Semantics Direct joint synthesis often disrupts the visual correspondence between reference cues and the anchor image. To preserve the semantic fidelity of the original edit, four distinct synthesis pipelines handle the six edit categories: Add, Replace, Remove, Background, Style, and Alter. For Add and Replace, Grounded-SAM-2 segments the edited target subject from \(I_{\text{tgt}}\), which is then either resynthesized on a neutral background using Flux-Klein-9B or composited onto a novel scene to provide an isolated reference object. For Remove, pure text descriptions are frequently ambiguous when scenes contain multiple instances of the same class; the pipeline localizes the target subject in \(I_{\text{in}}\) and supplies an explicit spatial cue—such as a tight bounding box overlay or a cropped exemplar patch—to cleanly resolve instance ambiguity. For Background replacement, a TCIE inpainting model removes foreground objects while retaining the background environment to yield a pristine scene layout. For abstract Style and Alter tasks, where open-source models struggle with pure concept disentanglement, text-to-image models first generate an intermediate visual subject, which is subsequently converted into the desired artistic style or surface texture using GPT-Image-1.5.

2. Weak-Instruction Augmentation: Suppressing Textual Shortcuts to Enforce Visual Grounding When an editing instruction provides exhaustive descriptive details about the reference entity (for instance, "add the fluffy bright-yellow puppy wearing a red collar from image 2 into image 1"), capable multimodal models frequently bypass the reference image entirely, falling back on their pre-trained text-to-image priors to execute the edit purely from text. To compel models to attend to visual features in \(I_{\text{ref}}\), the pipeline introduces a weak-instruction augmentation strategy. A VLM deliberately abstracts fine-grained attributes in the prompt (e.g., weakening the instruction to "add the animal from image 2 into image 1"). During training, mixing weak instructions with standard adapted instructions forces the model to extract necessary visual details directly from the reference image, establishing genuine cross-image visual conditioning rather than relying on textual shortcuts.

3. Five-Dimensional VLM Post-Filtering: Strict Multi-Axis Quality Assurance Multi-stage generation inherently accumulates subtle errors across segmentation, inpainting, and transfer stages. Prior data curation methods rely on holistic VLM ratings or CLIP similarity, which conflate distinct error modes and fail to verify four-way cross-component consistency. To guarantee rigorous quality, this work designs a five-dimensional VLM-based filter that independently scores: (i) Reference Compatibility with the original TCIE pair, (ii) Reference Reasonableness, (iii) Adapted Instruction Correctness, (iv) Original Pair Correctness, and (v) Ref-Target Similarity. A candidate quadruplet is retained only if every single axis surpasses its calibrated threshold, operating under an axis-wise veto policy. Cross-VLM evaluations with Claude Opus 4.6 and Gemini 3.1 Pro along with human verification achieve a Cohen's \(\kappa\) exceeding human inter-annotator agreement (0.44 vs. 0.38) and a high precision of 96.7%, pruning approximately 12.5% of defective samples.

Loss & Training

RCEdit-500K provides effective supervision across diverse model families: 1. Diffusion Models: Flux-Kontext and Qwen-Image-Edit-2509 are adapted using Low-Rank Adaptation (LoRA) with rank \(r=256\) and scaling factor \(\alpha=256\). Parameter updates are confined to the linear attention projection weights while maintaining the base flow-matching or diffusion objective. 2. Autoregressive Models: Janus-Pro undergoes full parameter fine-tuning. During ICIE inference, visual tokens produced by the understanding encoder and the generation encoder are explicitly concatenated, allowing the autoregressive transformer to model spatial anchor retention and reference feature transfer within a unified sequence.

Key Experimental Results

Main Results

Evaluation is conducted on the MultiBanana benchmark two-reference subset (excluding the makeup category). All models are evaluated along five axes rated 1–10 by GPT-4o: Instruction Alignment (IA), Reference Consistency (RC), Background-Subject Match (BSM), Physical Realism (PR), and Visual Quality (VQ).

Model Class Model IA↑ RC↑ BSM↑ PR↑ VQ↑ Avg↑
Closed-source Nano Banana 4.59 4.00 5.13 5.72 5.44 4.67
Closed-source GPT-Image-1.5 5.87 4.94 5.40 5.89 5.85 5.51
Open-source Baseline Flux-Kontext 2.94 2.80 2.39 2.94 2.81 2.82
Open-source Baseline Omnigen2 3.62 3.28 4.50 5.17 5.05 3.93
Open-source Baseline Qwen-Image-Edit-2509 3.60 4.06 4.83 5.61 5.39 4.31
Open-source Baseline Dreamomni2 4.45 3.61 5.07 5.83 5.80 4.54
Open-source Baseline Qwen-Image-Edit-2511 4.15 4.18 4.95 5.68 5.59 4.58
Open-source Baseline Flux-Klein-4B 4.69 4.01 5.05 5.62 5.57 4.71
Open-source Baseline Flux-Klein-9B 4.78 4.09 5.01 5.61 5.50 4.75
Open-source Baseline Janus-Pro (vanilla) 1.02 1.01 1.02 1.04 1.11 1.03
Fine-tuned Adaptation Janus-Pro + RCEdit-500K 4.65 3.81 4.52 5.04 4.91 4.43
Fine-tuned Adaptation Flux-Kontext + RCEdit-500K 3.52 3.42 4.76 5.50 5.27 4.04
Fine-tuned Adaptation Qwen-Image-Edit-2509 + RCEdit-500K 4.08 4.03 5.11 5.86 5.72 4.56

Ablation Study

Ablation studies analyze the impact of post-filtering, weak-instruction augmentation, and training dataset quality on MultiBanana using Flux-Kontext LoRA.

Table 1: Ablation on Pipeline Components (MultiBanana 2-reference)

Configuration IA↑ RC↑ BSM↑ PR↑ VQ↑ Avg↑ Note
Full Model 3.52 3.42 4.76 5.50 5.27 4.04 Complete pipeline with five-axis filter and weak instructions
w/o post-filtering 3.60 3.30 3.62 4.20 4.04 3.62 Severe degradation in BSM, PR, and VQ due to noisy outliers
w/o weak instruction 3.67 3.35 3.84 4.50 4.27 3.74 RC drops; model relies on textual priors rather than reference cues

Table 2: Comparison with Alternative ICIE Datasets (MultiBanana 2-reference, Flux-Kontext LoRA)

Dataset Size IA↑ RC↑ BSM↑ PR↑ VQ↑ Avg↑
Baseline (no LoRA) 0 2.94 2.80 2.39 2.94 2.81 2.82
ImgEdit 60K 3.05 2.83 2.58 3.10 2.92 2.92
AnyEdit (w/o structure) 40K 3.02 2.65 2.10 2.57 2.44 2.68
AnyEdit 40K 2.17 1.87 4.55 5.15 4.93 2.97
RCEdit-60K (matched scale) 60K 3.62 3.28 3.78 4.45 4.16 3.68
RCEdit-500K (full) 477K 3.52 3.42 4.76 5.50 5.27 4.04

Key Findings

  • Post-filtering safeguards physical plausibility: Omitting the five-dimensional filter causes the overall score to drop from 4.04 to 3.62, with Physical Realism (PR) plunging from 5.50 to 4.20 and BSM dropping from 4.76 to 3.62. Unfiltered pipeline noise directly harms contextual harmony.
  • Weak instructions compel visual conditioning: Without weak instructions, IA shows an artificial increase (3.67 vs. 3.52) while Reference Consistency (RC) declines (3.35 vs. 3.42). Qualitative inspections confirm that without weak instructions, generated objects diverge from the reference image and mimic generic textual descriptions.
  • Superiority over legacy ICIE data at matched scale: Under a controlled 60K volume comparison, RCEdit-60K (3.68) decisively outperforms ImgEdit (2.92) and AnyEdit (2.97) across all five axes, confirming that reference completion preserves superior real-image distributions.
  • Autoregressive models unlock editing capabilities: Janus-Pro jumps from an unusable baseline score of 1.03 to 4.43 upon full fine-tuning with RCEdit-500K, outperforming specialized diffusion editors like Qwen-Image-Edit-2509 (4.31) and demonstrating that data curation is the core driver for ICIE performance.

Highlights & Insights

  • Inverse completion beats forward generation: Rather than synthesizing all four quadruplet items and compounding artifacts, completing only the missing reference image preserves the rich diversity and authentic quality of curated TCIE datasets at low computational overhead.
  • Countering textual bias via weak instructions: Intentionally stripping specific descriptive tokens from prompts eliminates shortcut pathways, forcing multi-modal models to attend directly to visual features in the reference image.
  • Orthogonal five-dimensional audit: Applying an axis-wise threshold over five distinct evaluation axes achieves a high 96.7% precision, ensuring dataset cleanliness without conflating orthogonal failure modes.

Limitations & Future Work

  • Exclusion of dynamic motion editing: Due to the expressive limitations of single static reference images, temporal motion deformations are intentionally excluded from the taxonomy.
  • Dependency on proprietary APIs for abstract concepts: Open-source models currently struggle to cleanly disentangle abstract artistic styles and micro-textures; the pipeline still relies on GPT-Image-1.5 for the Style and Alter categories.
  • Complex multi-instance spatial boundaries: Although bounding box and cropped overlays effectively guide object removal, refining tight contours in heavily occluded scenes warrants further investigation.
  • vs. ImgEdit [36] & AnyEdit [37]: Prior collections only feature small visual-editing subsets (10K–40K samples) restricted to concrete object addition or simple material transfer; RCEdit-500K scales up to 477K quadruplets while covering six diverse categories spanning concrete entities and abstract attributes.
  • vs. DreamOmni2 [34] & USO [32]: Multi-image personalization synthesizes unconstrained combinations of subjects, whereas ICIE strictly enforces that the input image serves as a spatial anchor whose non-edited context must be rigorously preserved.
  • Methodological Transfer: The weak-instruction strategy represents a generalizable paradigm for any multimodal conditioning pipeline where dominant language priors threaten to overpower visual, spatial, or audio conditioning signals.

Rating

  • Novelty: ⭐⭐⭐⭐☆ [Reframes ICIE data construction as reference completion and introduces weak-instruction augmentation to prevent modality collapse]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluated across both diffusion and autoregressive architectures on the MultiBanana benchmark with rigorous ablations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, structured taxonomy, detailed statistical distributions, and comprehensive experimental analysis]
  • Value: ⭐⭐⭐⭐⭐ [Fills a critical gap by releasing 477K unified quadruplets and a low-cost curation pipeline for the open-source community]