Unlocking Complex Image Editing via Natively Interleaved Visual Textual CoT with Deep Confidence Reasoning¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/zhentao-zou/MURE
Area: Multimodal VLM
Keywords: image editing, natively interleaved visual-textual CoT, unified multimodal model, deep confidence reasoning, spatial reasoning
TL;DR¶
MURE proposes a unified multimodal framework that natively interleaves textual reasoning with intermediate visual rationales (precise masks, new object concepts) in an autoregressive stream, reinforced by Multimodal Deep Confidence (MMDC) tree pruning to achieve state-of-the-art complex image editing.
Background & Motivation¶
Instruction-guided image editing has witnessed surging popularity, eliminating the burdensome requirement of manual inpainting masks in conventional workflows. However, dominant contemporary diffusion-based architectures largely rely on direct cross-attention conditioning, endeavoring to translate abstract linguistic instructions into edited pixels within a single forward pass. When confronting real-world scenarios characterized by intricate object intersections, fine-grained spatial configurations, and physical interactionsโsuch as removing an entity alongside its reflection in a mirrorโthese ungrounded direct mappings frequently suffer from semantic drift, positional inaccuracy, and severe physical inconsistencies.
To inject deliberative cognitive capacity into editing, recent explorations have attempted to borrow Chain-of-Thought (CoT) prompting from large language models. Nonetheless, pure textual CoT can only parse instructions at a high semantic level; discrete text tokens are fundamentally incapable of describing irregular pixel contours and dense spatial layouts. Subsequent efforts augmenting CoT with bounding-box coordinates (e.g., GoT) mitigate placement ambiguity to some degree, but still fail when representing complex non-convex geometry or overlapping shapes. Concurrently, modular agent-based pipelines that dispatch external tools (like SAM for segmentation followed by off-the-shelf inpainting) suffer from rigid execution flows, absence of KV cache reuse, and catastrophic error accumulation across independent stages.
The key insight of this work is that human visual editing inherently mirrors an interleaved mental simulationโcontinually alternating between verbal planning and vivid visual drafting. Hence, the reasoning trajectory should be formulated as a natively multimodal autoregressive sequence within a unified latent space. Core idea: construct a unified monolithic model MURE that generates natively interleaved textual and visual rationales, coupled with a Multimodal Deep Confidence (MMDC) tree search that prunes low-quality intermediate visual candidates via reward model scoring to eliminate cascading errors.
Method¶
Overall Architecture¶
MURE is instantiated on top of a unified multimodal autoregressive backbone (initialized from BAGEL), taking an input image and a natural language edit instruction to autoregressively decode an interleaved sequence of textual reasoning tokens and intermediate visual tokens, culminating in the final edited image. The process sequentially executes target localization (predicting a textual rationale and a dense spatial mask), conceptualization of new content (generating descriptive text and a standalone object rendering), and context-aware fusion. At each visual reasoning juncture, the MMDC mechanism explores candidate branches in parallel and prunes sub-optimal paths before subsequent reasoning proceeds.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Image + Edit Instruction"] --> B["Natively Interleaved CoT Modeling<br/>Autoregressive Text & Visual Tokens"]
B --> C["Multimodal Deep Confidence MMDC<br/>Tree Search & Reward Model Pruning"]
C --> D["Two-Stage CoT Rationales<br/>Mask Prediction โ New Object Concept"]
D --> E["Contextual Fusion & Generation<br/>High-Fidelity Final Edited Image"]
Key Designs¶
1. Natively Interleaved Visual-Textual CoT and Monolithic Autoregressive Modeling: Bridging Symbolic Reasoning and Dense Visual Rationales
To overcome the expressive barrier of text-only CoT and the rigid error propagation of multi-tool agent systems, MURE formulates the editing reasoning trajectory as a continuous autoregressive stream:
$$
{s^{(1)}, v^{(1)}, s^{(2)}, v^{(2)}, \ldots, v^{(k-1)}, s^{(k)}, O} \sim \mathbf{f}_\theta(\cdot \mid I, P)
$$
where \(s^{(i)}\) denotes textual reasoning segments, \(v^{(i)}\) denotes visual rationales (such as positional binary masks or intermediate new object renderings), and \(O\) represents the final edited canvas. The transitions between text tokens and visual latent representations are demarcated by specialized delimiter tokens โจvisual startโฉ and โจvisual endโฉ. By executing inside a monolithic multimodal model, all preceding textual thoughts and visual hypotheses share a unified latent space and reuse the global KV cache. This unified design preserves end-to-end differentiability, minimizes memory footprint, and empowers the model to perform intrinsic self-correction conditioned on prior intermediate visual outcomes.
2. Deep Confidence Tree Pruning (MMDC): Suppressing Intermediate Visual Degradation and Cascading Failure Recognizing that diffusion/flow sampling stochasticity and generative hallucinations can introduce degraded intermediate visual artifacts that irreversibly derail downstream synthesis, MMDC discards unguided parallel sampling or simple majority voting. Instead, at each intermediate visual generation step \(k\), MMDC branches out \(N\) parallel candidate visual samples \(\{v^{(k,1)}, \ldots, v^{(k,N)}\}\). An auxiliary visual-language reward evaluator \(R_\theta\) (parameterized by Qwen2.5-VL-7B) assesses each candidate conditioned on the cumulative multimodal history, computing a deep confidence score \(S_{k,i}\): $$ S_{k,i} = R_\theta(I, P, s^{(1)}, v^{(1)}, \ldots, s^{(k)}, v^{(k,i)}) $$ A greedy selection policy prunes all inferior branches, retaining solely the optimal candidate \(i^* = \arg\max_{i \in \{1,\ldots,N\}} S_{k,i}\) to anchor all subsequent autoregressive text and visual tokens. This phased tree search enforces a rigorous quality checkpoint at each visual transition node, guaranteeing high physical fidelity along the entire generation trajectory.
3. Multimodal CoT Data Construction Pipeline and Logic Regularization: Inducing Deep Visuo-Physical Priors Due to the complete absence of paired multimodal interleaved editing traces in existing corpora, the authors introduce the CoT-Edit-14K benchmark dataset spanning 10 representative editing subtasks. The pipeline utilizes LLM classification to tailor task-specific CoT layouts, applies adaptive region-weighting for mask synthesis, leverages in-context prompts to render new object concepts, and employs Qwen2.5-VL to curate backward-aligned verbal rationales. Training on this dataset yields a profound "logic regularizer" effect: mastering tightly bound visual-physical intermediate states forces the network to internalize deep physical and spatial laws, thereby conferring significant generalizability improvements even onto conventional text-only editing routines.
Loss & Training¶
The optimization of MURE is jointly driven by causal language modeling and rectified flow velocity matching. For text tokens, standard cross-entropy loss \(\mathcal{L}_{\mathrm{CE}}^{\text{text}}\) is computed over positions \(T\) corresponding to textual rationales. For visual token generation, following the Rectified Flow framework, clean latents \(z_0^{(i)}\) and standard Gaussian noise latents \(z_1^{(i)}\) define linear trajectories \(z_t^{(i)} = t z_0^{(i)} + (1-t) z_1^{(i)}\), where model \(f_\theta\) is trained via mean squared error \(\mathcal{L}_{\mathrm{MSE}}^{\text{image}}\) to regress the velocity field across all visual rationales and the final image. The composite objective is formulated as: $$ \mathcal{L}{\mathrm{total}} = \lambda}} \mathcal{L{\mathrm{CE}}^{\text{text}} + \mathcal{L} $$ where }}^{\text{image}\(\lambda_{\mathrm{CE}}\) serves as a balancing hyperparameter governing the optimization dynamics between discrete linguistic reasoning and continuous visual velocity prediction.
Key Experimental Results¶
Main Results¶
MURE was comprehensively evaluated across three established image editing benchmarksโMagicBrush, Emu Edit, and SmartEditโagainst prominent UNet, DiT, autoregressive, and unified multimodal baselines:
| Methods | CoT Paradigm | MagicBrush L1 โ | MagicBrush CLIP-I โ | MagicBrush DINO โ | Emu CLIP-I โ | Emu CLIP-Out โ | Emu DINO โ |
|---|---|---|---|---|---|---|---|
| InstructP2P (CVPR23) | None | 0.114 | 0.851 | 0.744 | 0.856 | 0.292 | 0.773 |
| MagicBrush (NeurIPS23) | None | 0.074 | 0.908 | 0.847 | 0.877 | 0.298 | 0.807 |
| UltraEdit (NeurIPS24) | None | 0.066 | 0.904 | 0.852 | 0.880 | 0.304 | 0.847 |
| FluxEdit (HF 2025) | None | 0.114 | 0.779 | 0.663 | 0.852 | 0.282 | 0.760 |
| ICEdit (arXiv25) | None | 0.060 | 0.928 | 0.853 | 0.907 | 0.305 | 0.866 |
| Bagel (arXiv25) | Text-only | 0.067 | 0.923 | 0.856 | 0.869 | 0.308 | 0.824 |
| MURE (Ours) | Text & Images | 0.049 | 0.943 | 0.877 | 0.920 | 0.301 | 0.897 |
On the challenging SmartEdit benchmark evaluating spatial understanding and world-model simulation, MURE consistently outclasses all prior specialized models:
| Methods | Spatial Reasoning PSNR โ | Spatial Reasoning SSIM โ | Spatial Reasoning LPIPS โ | Spatial Reasoning CLIP Score โ | World Model PSNR โ | World Model SSIM โ | World Model LPIPS โ | World Model CLIP Score โ |
|---|---|---|---|---|---|---|---|---|
| InstructP2P | 21.576 | 0.721 | 0.089 | 22.762 | 24.234 | 0.707 | 0.083 | 19.413 |
| SmartEdit-7B | 22.049 | 0.731 | 0.087 | 23.611 | 25.258 | 0.742 | 0.055 | 20.950 |
| SmartEdit-13B | 23.596 | 0.751 | 0.068 | 23.536 | 25.757 | 0.747 | 0.051 | 20.777 |
| Bagel | 23.823 | 0.892 | 0.083 | 23.842 | 28.076 | 0.839 | 0.060 | 20.767 |
| MURE (Ours) | 25.611 | 0.897 | 0.065 | 23.947 | 28.694 | 0.883 | 0.062 | 21.298 |
Ablation Study¶
The ablation experiments quantify the cumulative impact of natively interleaved CoT (ICT) and the MMDC search width (\(N\)) on the MagicBrush and Emu benchmarks:
| Configuration | Search Width \(N\) | MagicBrush L1 โ | MagicBrush CLIP-I โ | MagicBrush DINO โ | Emu CLIP-I โ | Emu DINO โ | Note |
|---|---|---|---|---|---|---|---|
| Text-only Baseline (Bagel) | - | 0.067 | 0.923 | 0.856 | 0.869 | 0.824 | Pure verbal CoT, weak spatial constraints |
| + ICT (Interleaved CoT) | 1 | 0.058 | 0.936 | 0.856 | 0.880 | 0.836 | Integrates intermediate mask and object visual cues |
| + MMDC Greedy Pruning | 3 | 0.052 | 0.941 | 0.872 | 0.913 | 0.887 | Expands candidate search width to 3 |
| + MMDC Deep Search | 5 | 0.049 | 0.943 | 0.877 | 0.920 | 0.897 | Full model with robust intermediate verification |
On challenging subsets requiring rigorous spatial reasoning (selected via GPT-4o), MURE displays dramatic performance gains: on MagicBrush challenging cases, L1 drops from 0.081 to 0.052 (a 35.8% error reduction), while PSNR on SmartEdit spatial reasoning improves by 1.79 dB.
Key Findings¶
- Crucial Role of Interleaved Visual Cues for Preservation: Generating intermediate localization masks reduces L1 error by over 13.4%, successfully isolating the edit zone and completely halting unprompted hallucinations across background elements.
- Monotonic Scaling with Search Width: Scaling test-time search width from \(N=1\) to \(N=5\) monotonically improves all fidelity and alignment metrics, corroborating the presence of test-time compute scaling laws in generative visual reasoning.
- Logic Regularization Spillover: Even on tasks executed solely through textual reasoning (e.g., Object Action Change, Obj.Act.Chg.), MURE secures a +17.14% relative enhancement over Bagel, demonstrating that interleaved multimodal pretraining fundamentally enriches underlying visuo-physical cognition.
Highlights & Insights¶
- Natively Interleaved Visual Tokens as Cognitive Scaffolds: Transcending purely linguistic chains, MURE natively embeds pixel-level rationales (masks, partial visual concepts) directly into the autoregressive thought process, effectively uniting symbolic abstraction with dense geometric precision.
- Intermediate Quality Gating Over Answer-Level Ensembling: Recognizing that generative stochasticity compounds exponentially in long reasoning chains, MMDC introduces targeted local branch pruning via a lightweight reward model, delivering remarkable trajectory reliability with minimal test-time compute.
- Pioneering Multimodal CoT Benchmark CoT-Edit-14K: Curates and releases 14,000 vetted multimodal editing chains across 10 diverse operation categories, establishing a crucial foundation for subsequent research in unified generative autoregressive models.
Limitations & Future Work¶
- Substantial Test-Time Latency and Compute Overhead: Autoregressively decoding multiple textual blocks alongside multiple intermediate image latents, combined with candidate evaluation under MMDC, incurs notably higher inference latency and VRAM consumption compared to single-step diffusion models.
- Limited Applicability to Global Artistic Edits: Interleaved spatial decomposition offers minimal utility for global style transfer or holistic color grading, where explicit geometric masks and component conceptualization are largely redundant.
- Reward Model Spatial Bias: Empirical audits indicate that Qwen2.5-VL exhibits a subtle semantic-over-spatial bias, occasionally prioritizing semantic keyword presence over sub-pixel boundary adherence. Training dedicated spatially grounded reward evaluators remains an important open direction.
Related Work & Insights¶
- vs ICEdit / FluxEdit: Standard diffusion and DiT inpainting models rely on one-shot latent conditional denoising, frequently failing when tasks involve compound multi-object constraints; MURE explicitly decomposes the task into sequential, self-verifying sub-problems.
- vs GoT / MINI-CoT: GoT utilizes bounding-box coordinates to express spatial layout, which cannot delineate non-rectangular or disjoint regions; MURE directly generates dense visual masks and explicit object renderings.
- vs Multi-Agent Tool-Chaining (e.g., TIE): Orchestrating independent models (LLM + SAM + Diffusion) lacks a unified latent context and forfeits KV cache efficiency, resulting in irrecoverable cascade errors; MURE achieves seamless end-to-end coherence within a monolithic architecture.
Rating¶
- Novelty: โญโญโญโญโญ Pioneering introduction of natively interleaved visual-textual CoT and unified autoregressive modeling for complex image editing.
- Experimental Thoroughness: โญโญโญโญโญ Evaluated across three competitive public benchmarks with targeted spatial splits, scaling analysis, and reward alignment validations.
- Writing Quality: โญโญโญโญโญ Rigorous methodology, crisp motivation, well-balanced mathematical formulations, and clear architectural diagrams.
- Value: โญโญโญโญโญ Provides an exemplary blueprint and public dataset for unlocking test-time visual reasoning in unified generative models.