Editing Everything Everywhere All at Once¶
Conference: ECCV 2026
arXiv: arXiv:2606.31278
Paper: ECCV Official
Code: https://github.com/Blowing-Up-Groundhogs/mice
Area: Segmentation
Keywords: Multi-Instance Editing, Multimodal Diffusion Transformers, Flow Matching, Attention Bias, Instance-Aware Smoothing
TL;DR¶
To tackle severe attribute leakage and harsh boundary stitching in multi-instruction concurrent image editing, MICE introduces a training-free, architecture-agnostic continuous attention biasing scheme that achieves strict attribute binding and natural visual harmonization in a single forward pass.
Background & Motivation¶
Modifying multiple distinct instances or regions in a single image concurrently has emerged as a high-value frontier for text-guided image manipulation. While multi-turn sequential editing pipelines offer an intuitive baseline, sequential modifications accumulate significant visual degradation as the iteration count increases (typically beyond six turns). Even worse, when target instances overlap or occlude one another, sequential workflows become highly path-dependent and sensitive to execution order. Concurrent editing performs all transformations within a single forward pass, providing indispensable efficiency and determinism. Nevertheless, existing concurrent models face a fundamental bottleneck: when numerous textual instructions simultaneously target different visual regions, global cross-attention mechanisms cause cross-talk among conditioning signals, triggering catastrophic attribute leakage and feature entanglement across instances.
Existing multi-instance editing strategies generally fall into three paradigms, all suffering from major operational limitations. Multi-branch approaches (e.g., MultiDiffusion, LoMOE, LayerEdit) decompose the task into independent parallel denoising paths and perform latent-space gradient descent, which incurs severe \(O(N)\) memory and computational explosion in dense scenes. Global latent regularization techniques (e.g., MDE-Edit) aggregate multiple loss terms, but back-propagation gradient competition causes large objects to overwhelm smaller targets. Attention-based isolation techniques (e.g., IDAttn) enforce spatial disentanglement via binary hard masks (0 and \(-\infty\)), which create noticeable boundary seams, break natural illumination transitions, and heavily rely on fragile, layer-dependent manual heuristics across network depths.
This paper's angle of attack is that rather than replicating generation paths or engineering ad-hoc layer schedules, one can exploit the additive attention bias of joint attention operators to establish a continuous negative-decay space. Core idea: propose MICE, a training-free framework that maps spatial distances into continuous negative log-probability attention biases via instance-aware Gaussian smoothing, achieving strict intra-instance semantic binding, gradual boundary harmonization, and absolute cross-instance suppression within a single forward pass.
Method¶
Overall Architecture¶
MICE builds upon Multimodal Diffusion Transformer (MMDiT) architectures trained with conditional rectified flow matching for text-conditioned image editing. Given a reference image \(I_{ref}\), a set of target instance segmentation masks \(M = \{m_n\}_{n=1}^N\), and corresponding editing instructions \(T = \{t_n\}_{n=1}^N\), visual latents (noisy latents \(Z_{latent}\) and reference context \(Z_{context}\)) and text prompt tokens \(Z_{prompt}\) are concatenated into a global joint sequence \(Z = Z_{prompt} \parallel Z_{latent} \parallel Z_{context}\), which is processed through joint self-attention across all transformer blocks.
While standard MMDiT layers set the additive attention bias \(\mathbf{B}\) to 0, MICE substitutes a unified, smoothly-disentangled bias matrix \(\mathbf{B} \in \mathbb{R}^{|Z| \times |Z| \le 0}\) across all layers. Derived from downsampled instance masks and processed via instance-aware Gaussian smoothing followed by log-probability mapping, this bias regulates token interactions continuously without requiring per-layer adjustments.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input Reference Image + Instance Masks + Editing Prompts"] --> Down["Mask Downsampling to VAE Latent Grid"]
Down --> Smooth["Instance-Aware Gaussian Smoothing<br/>Independent Gaussian convolution + zeroing at neighboring entity boundaries"]
Smooth --> LogB["Log-Probability Continuous Attention Bias<br/>Mapped to negative log bias to govern spatial attention falloff"]
LogB --> Rules["Multimodal Joint Token Interaction Rules<br/>Partitioning instance prompts/background latents/context tokens"]
Rules --> MMDiT["MMDiT Joint Self-Attention Computation<br/>Identical bias applied uniformly across all transformer layers"]
MMDiT --> Out["Generated Concurrently Edited Image with Natural Harmonization"]
Key Designs¶
1. Instance-Aware Gaussian Smoothing: Eliminating hard seams while preventing neighbor attribute intrusion Using strict binary segmentation masks or rectangular bounding boxes yields pronounced drawbacks: binary masks completely sever communication between edited instances and background context, resulting in harsh, cut-and-paste boundaries; conversely, applying standard unconstrained Gaussian filtering across all masks causes overlapping or touching instances to blur into each other, creating fresh inter-instance semantic bleeding. MICE addresses this with an instance-aware smoothing mechanism. For the binary mask \(m_n \in \{0, 1\}^{H \times W}\) of instance \(n\), downsampled to the VAE latent resolution, a 2D Gaussian convolution with kernel size \(k\) yields an initial smoothed representation \(s_n\). Crucially, for every latent location, if that position belongs to the original non-smoothed mask of any other instance \(n' \neq n\), its smoothed value in \(s_n\) is strictly overridden to 0. This guarantees unconstrained smooth falloff into neutral background regions while maintaining an impermeable geometric barrier between neighboring edited objects.
2. Log-Probability Continuous Attention Bias: Architecture-agnostic unified single-matrix regulation Prior attention-masking methods alternate between binary masks and unrestricted attention at heuristically chosen layers, limiting architectural portability. MICE discards binary constraints, mapping the smoothed mask values \(s_n \in [0, 1]\) directly into the negative log-probability space to enforce exponential distance decay. For any visual token \(Z_j\) (latent or context) outside the strict bounds of instance \(n\), its attention bias relative to an instance token \(Z_i\) is formulated as:
where \(\tau\) is a temperature hyperparameter governing falloff steepness, and \(\epsilon = 10^{-30}\) ensures numerical stability. When token \(j\) lies inside the instance interior, \(s_n[j]=1\), yielding a bias of 0 (unrestricted self-attention); when \(j\) is distant, \(s_n[j] \to 0\), forcing the bias toward \(-\infty\) (complete cross-instance isolation); in boundary buffer zones, \(s_n[j] \in (0, 1)\) supplies a calibrated negative bias allowing essential context and illumination blending. Because this continuous formulation inherently balances isolation with harmonization, the exact same bias matrix is deployed across all transformer layers without architectural tuning.
3. Multimodal Joint Token Interaction Rules: Binding instance prompts while preserving context integrity Partitioning tokens by modality into prompts \(Z_{prompt}\) (instance prompts \(Z_{prompt, n}\) and optional global prompt \(Z_{prompt, g}\)), generated latents \(Z_{latent}\) (unedited background \(Z_{latent, u}\) and instance regions \(Z_{latent, n}\)), and reference context \(Z_{context}\) (background \(Z_{context, u}\) and source instance regions \(Z_{context, n}\)), MICE establishes structured interaction rules: - Instruction Disentanglement: Distinct instance prompts \(Z_{prompt, n}\) are mutually barred with \(-\infty\) bias, eliminating prompt-level text leakage. - Background Consistency: Unedited background latents \(Z_{latent, u}\) interact with zero bias across global prompts and reference context, preserving global layout and visual fidelity. - Directional Multimodal Biasing: Cross-attention between generated latents and instance prompts is modulated by \(\mathcal{B}_n(j)\), focusing semantic synthesis on the intended target zone and its transition boundary. - Context Protection: Reference background tokens \(Z_{context, u}\) are forbidden from attending to instance prompts (bias set to \(-\infty\)), shielding unedited regions from unintended prompt corruption.
Loss & Training¶
MICE operates as an entirely training-free, inference-time framework. It introduces zero learned parameters and avoids latent-space backpropagation during sampling. During rectified flow matching inference, the static attention bias matrix \(\mathbf{B}\) is computed from user masks and passed directly into the standard joint attention operator:
On state-of-the-art flow matching models like FLUX.2 [klein] 4B/9B, synthesis converges in only 4 sampling steps, while 28 steps are employed for larger models such as FLUX.1 Kontext-dev and FLUX.2-Dev.
Key Experimental Results¶
Main Results¶
Evaluation is conducted on LoMOE-Bench (64 images, averaging 3 edits per image) and the newly curated, dense MICE-Bench (260 images, averaging 8.5 edits per image, with up to 40 edits). Metrics comprise Target CLIP Score (Tgt C, global instruction alignment), Localized CLIP Score (Loc C, region-specific edit faithfulness), Mean Absolute Error on Background (MAEB, background preservation, lower is better), and Attempt Rate (AR%, percentage of attempted edits exceeding pixel-change thresholds).
The table below reports quantitative comparisons across methods on both benchmarks:
| Model / Strategy | Backbone | LoMOE Tgt C โ | LoMOE Loc C โ | LoMOE MAEB โ | LoMOE AR [%] โ | MICE-Bench Tgt C โ | MICE-Bench Loc C โ | MICE-Bench MAEB โ | MICE-Bench AR [%] โ |
|---|---|---|---|---|---|---|---|---|---|
| LoMOE | SD 2.0 | 26.00 | 29.40 | 6.66 | 92.19 | 24.99 | 25.46 | 7.76 | 99.07 |
| LayerEdit | SDXL | 25.61 | 29.07 | 7.43 | 97.92 | 23.77 | 24.38 | 9.32 | 99.51 |
| Qwen-Image-Edit | Proprietary | 26.34 | 28.85 | 7.13 | 88.02 | 25.26 | 25.33 | 26.47 | 98.83 |
| FLUX.2 [klein] 9B (End-to-End) | FLUX.2 9B | 26.57 | 29.45 | 8.63 | 99.48 | 25.66 | 26.34 | 27.12 | 100.00 |
| Gemini 3 Pro Image (Commercial API) | Proprietary | 26.82 | 29.67 | 7.24 | 97.37 | 25.64 | 27.09 | 13.08 | 98.97 |
| IDAttn | FLUX.2 9B | 25.67 | 29.26 | 6.41 | 98.44 | 24.68 | 25.54 | 5.06 | 73.01 |
| MICE (Ours) | FLUX.2 9B | 26.49 | 30.30 | 8.88 | 98.44 | 25.24 | 27.79 | 13.89 | 98.97 |
| MICE + SAM3 (Practical Setup) | FLUX.2 9B | 26.22 | 30.41 | 8.94 | 96.88 | 25.26 | 27.73 | 13.36 | 98.54 |
Across diverse backbones, integrating MICE consistently boosts localized editing accuracy (Loc C): on FLUX.1 Kontext from 28.26/25.43 to 29.13/25.54; on FLUX.2 [klein] 4B from 28.68/25.97 to 29.57/26.85; and on FLUX.2-Dev 32B from 29.00/26.25 to 30.28/26.91.
Ablation Study¶
On the FLUX.2 [klein] 4B backbone, the authors systematically ablate spatial editing area formulations (Table 1):
| Config | LoMOE Tgt C โ | LoMOE Loc C โ | LoMOE MAEB โ | LoMOE AR [%] โ | MICE-Bench Tgt C โ | MICE-Bench Loc C โ | MICE-Bench MAEB โ | MICE-Bench AR [%] โ | Note |
|---|---|---|---|---|---|---|---|---|---|
| Bounding Boxes | 25.94 | 29.55 | 9.38 | 97.92 | 25.00 | 26.57 | 10.66 | 98.78 | Over-inclusive boxes degrade background fidelity |
| Original Seg. M. | 25.78 | 29.59 | 7.60 | 96.35 | 24.95 | 26.63 | 8.15 | 95.36 | Rigid boundaries lack smooth contextual blending |
| Gaussian Seg. M. | 25.77 | 29.55 | 7.08 | 96.35 | 25.00 | 26.19 | 7.63 | 87.30 | Unchecked Gaussian bleeding corrupts dense edits |
| Our Seg. M. | 25.77 | 29.57 | 7.07 | 96.35 | 25.15 | 26.85 | 7.84 | 92.48 | Mutual boundary reset delivers optimal isolation & harmony |
In an LLM-as-judge double-blind tournament powered by Gemini 3.1 Pro across 900 pairwise comparisons (Table 5), MICE achieves top ELO ratings on both datasets (1342.53 on LoMOE and 1568.46 on MICE-Bench), substantially outperforming IDAttn (1412.73) and baseline FLUX.2 (1172.16).
Key Findings¶
- Robustness in Dense Regimes: Under high-density edits (averaging 8.5 modifications on MICE-Bench), vanilla end-to-end models suffer drastic background degradation (MAEB surging to 26-27), whereas MICE preserves unedited scene geometry while achieving superior localized instruction adherence over proprietary leaders like Gemini 3 Pro.
- Memory & Runtime Scalability: Multi-branch pipelines (MultiDiffusion, LoMOE) exhaust 40GB A100 VRAM as edit counts surpass 15 due to linear memory scaling; in contrast, MICE maintains near-constant memory overhead and runtime, gracefully scaling to 40 concurrent edits.
- Viability with Automated Segmenters: Paired with SAM3-generated zero-shot concept masks, MICE yields virtually identical localized fidelity to ground-truth masks (Loc C 27.73 vs 27.79), confirming plug-and-play efficacy for interactive user-in-the-loop applications.
Highlights & Insights¶
- Instance-Aware Geometric Suppression: Zeroing smoothed Gaussian weights at the true boundary of adjacent instances elegantly reconciles continuous spatial blending with strict semantic disentanglement.
- Architecture-Agnostic Attention Biasing: Formulating soft falloff in the negative log-probability space allows a single unified bias matrix to be plugged into any MMDiT layer without manual depth-tuning.
- Orthogonal Coexistence of Global & Local Edits: By keeping unbiased communication channels for global prompts and background latents, the model seamlessly blends dense local object replacements with full-frame artistic restyling and global relighting.
Limitations & Future Work¶
- Heavy 3D Occlusion and Enclosure: When multiple targets strongly occlude or visually nest inside one another in 3D space, 2D boundary resets can compress the available latent area of occluded entities, leading to boundary artifacts.
- Extreme Miniature Targets: Highly tiny objects correspond to negligible token counts after VAE downsampling; Gaussian diffusion may overly dilute their activation weights, slightly lowering the attempt rate.
- Future Directions: Integrating explicit depth-aware 3D bounding geometry, and introducing content-adaptive scheduling for temperature \(\tau\) and smoothing radius \(k\).
Related Work & Insights¶
- vs LoMOE / MultiDiffusion: Multi-branch diffusion and latent optimization incur prohibitive \(O(N)\) memory expansion and collapse under dense edits; MICE operates end-to-end via attention biasing, ensuring constant memory footprint and high throughput.
- vs IDAttn: IDAttn relies on binary attention masks and architecture-specific layer partitioning heuristics that cause noticeable seams; MICE utilizes continuous Gaussian log-biasing applied uniformly across all layers, delivering superior harmonization and model portability.
- vs Qwen-Image-Edit / Unconstrained End-to-End Models: Purely prompt-guided models lack localized geometric conditioning, leading to attribute confusion and background degradation; MICE provides explicit spatial attention routing to anchor edits precisely.
Rating¶
- Novelty: โญโญโญโญโญ Formulates instance-aware Gaussian smoothing mapped into continuous log-probability attention biases to resolve the trade-off between isolation and harmonization.
- Experimental Thoroughness: โญโญโญโญโญ Introduces the 260-sample high-density MICE-Bench and validates across multiple MMDiT backbones, automated metrics, and LLM-as-judge tournaments.
- Writing Quality: โญโญโญโญโญ Rigorous mathematical formulation, crisp narrative motivation, and comprehensive empirical analyses.
- Value: โญโญโญโญโญ Training-free, lightweight, and capable of executing up to 40 concurrent edits, offering immediate practical utility for downstream generative graphics.