Skip to content

SR-Edit: Region-Aware Image Editing via Self-Refinement

Conference: ECCV 2026
Paper: ECCV 2026 Official
Area: Image Generation
Keywords: image editing, diffusion model, flow matching, Doob's h-transform, self-refinement

TL;DR

To tackle unintended non-edit drift and artifact-prone heuristic feature interventions in mask-free image editing, SR-Edit extracts precise edit masks from model predictions via pixel-space self-feedback and enforces non-edit preservation via Doob's h-transform guidance aligned with native sampling dynamics.

Background & Motivation

Instruction-driven image editing based on diffusion models and flow matching backbones has achieved substantial progress. In practical editing applications, users routinely demand that the model accurately apply targeted semantic modifications while strictly preserving background environments, non-target objects, fine-grained geometric textures, and subject identities. However, standard generative pipelines initialize sampling from global noise distributions or unconstrained latent trajectories, where the denoising dynamics inherently operate over the entire image plane. Consequently, even when the intended semantic modification is successfully realized, non-target regions frequently suffer from severe unintended drift, blurred high-frequency textures, and identity degradation.

To mitigate this drift, an intuitive paradigm is to explicitly partition the image into editable and non-editable regions, enforcing strict consistency constraints across the latter. Nevertheless, acquiring high-quality manual annotation masks is cost-prohibitive in interactive workflows. Existing mask-free alternatives encounter two critical bottlenecks: first, region estimation derived from cross-attention maps or latent-space feature discrepancies lacks sub-pixel precision, yielding coarse, fuzzy boundaries that fluctuate unstably across sampling steps; second, preservation mechanisms in methods such as SpotEdit and Follow-Your-Shape rely heavily on heuristic feature replacement, key-value (KV) cache swapping, or latent interpolation. These ad-hoc interventions distort the model's native sampling trajectory, converting the fidelity-preserving mechanisms themselves into a primary source of boundary tearing, edge jumps, and residual blur.

An empirical investigation of generative inference behavior reveals a distinctive pattern: well-trained editing models concentrate large, spatially coherent variations within true edit regions, whereas non-edit regions exhibit low-amplitude, scattered, and weakly structured residual fluctuations. Core idea: use the model's own decoded prediction as a stable pixel-space self-feedback signal to isolate precise edit masks via lightweight morphological and connectivity operations, and derive an additive masked residual guidance term from Doob's h-transform theory that seamlessly preserves native sampling dynamics across an iterative self-refinement loop.

Method

Overall Architecture

At inference time, SR-Edit organizes the generative process into a short initial stabilization stage followed by a cascade of SR-Edit blocks. Within each refinement block, a lightweight few-step probing trajectory generates an intermediate provisional edit \(I_{\mathrm{ref}}\). On the decoded pixel grid, a per-pixel difference map against the source image \(I_{\mathrm{in}}\) is constructed, followed by Otsu thresholding, binary morphological filtering, and connected-component area pruning to extract a sharp binary change mask \(M\). Conditioning on the non-edit complement mask \(\bar{M}\), non-edit preservation is formulated as a vanishing-noise observation problem under Doob's h-transform theory, yielding an additive drift correction for diffusion SDEs and a velocity correction for flow-matching ODEs that steers intermediate states toward the source image without disrupting native transport structures.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image I_in + Textual Edit Instruction"] --> B["Initial Stabilization & Few-Step Probing<br/>Few-step inference produces provisional prediction I_ref"]
    B --> C["Self-Feedback Pixel-Space Region Separation<br/>Difference map calculation + Otsu adaptive thresholding"]
    C --> D["Morphological Refinement & Connectivity Filtering<br/>Opening/closing cleans noise + prunes small components"]
    D --> E["h-Transform Dynamics-Preserving Correction<br/>Doob observation model derives masked residual guidance"]
    E -->|Next self-refinement iteration| B
    E -->|Complete multi-iteration refinement| F["Final High-Fidelity Edited Image Output"]

Key Designs

1. Self-Feedback Pixel-Space Region Separation: Leveraging model predictions for robust difference mapping

Rather than relying on latent representations or cross-attention maps whose spatial resolution is compromised by patch tokenization and nonlinear decoding distortions, SR-Edit computes per-pixel absolute color deviations directly in pixel space across all \(C\) channels: $\(d(p) = \frac{1}{C}\sum_{c=1}^C \left|I_{\mathrm{in}}(p, c) - I_{\mathrm{ref}}(p, c)\right|\)$ To eliminate heuristic threshold tuning across diverse editing magnitudes, the method utilizes Otsu's between-class variance maximization criterion: $\(t^\star = \arg\max_t \omega_0(t)\omega_1(t)\left(\mu_0(t) - \mu_1(t)\right)^2\)$ This adaptively bi-partitions the difference histogram into high-energy foreground modifications and low-energy background perturbations, producing an initial binary mask \(M_0(p) = \mathbb{I}[d(p) \ge t^\star]\).

2. Morphological Refinement & Connectivity Filtering: Eliminating scattered artifacts and enforcing spatial coherence

Initial Otsu thresholding often leaves isolated background speckles, tiny interior voids, and disconnected boundary segments. SR-Edit enforces geometric regularity through classical mathematical morphology. First, a binary morphological opening with a small circular structuring element of radius \(s_{\mathrm{open}}\) eliminates isolated speckles; subsequently, a closing operation with radius \(s_{\mathrm{close}}\) bridges narrow crevices and seals interior holes, generating an intermediate coherent mask \(M_1\). Finally, connected-component analysis labels all maximal contiguous foreground clusters, discarding components whose pixel area falls below a minimum size threshold \(A_{\mathrm{min}}\). This pure pixel-space post-processing guarantees crisp boundary alignment and eradicates spatial jitter.

3. Doob's h-Transform Dynamics-Preserving Correction: Rigorous distribution-derived non-edit preservation

Addressing the severe edge discontinuities and manifold corruption induced by heuristic KV injection, SR-Edit establishes preservation through stochastic conditioning theory. The non-edit region (complement mask \(\bar{M}\)) preservation constraint \(\bar{M}x_0 = \bar{M}x_{\mathrm{src}} \triangleq y\) is formulated via a vanishing-noise observation likelihood \(Y = \bar{M}x_0 + \xi\), where \(\xi \sim \mathcal{N}(0, \sigma_{\mathrm{obs}}^2 I)\) as \(\sigma_{\mathrm{obs}} \to 0\). By Doob's h-transform, the conditional score factorizes additively into the base score and the log-gradient of the Doob function \(h(t, x_t) \triangleq \mathbb{E}[p(y \mid x_0) \mid x_t, \mathcal{C}]\): $\(\nabla_{x_t} \log p_t(x_t \mid y, \mathcal{C}) = \nabla_{x_t} \log p_t(x_t \mid \mathcal{C}) + \nabla_{x_t} \log h(t, x_t)\)$ Evaluating \(h(t, x_t)\) via a plug-in approximation with the model's clean image prediction \(\hat{x}_0(x_t)\) and absorbing the local Jacobian sensitivity into a scheduled scalar weight \(\lambda(t)\) yields the practical masked residual update: $\(\nabla_{x_t} \log h(t, x_t) \approx -\lambda(t) \bar{M} \left(\hat{x}_0(x_t) - x_{\mathrm{src}}\right)\)$ For continuous-time diffusion SDEs, this enters purely as an additive drift vector without modifying the diffusion coefficient \(g(t)\). For flow-matching ODEs \(\frac{dx_t}{dt} = u_\theta(x_t, t)\), minimizing the quadratic preservation energy \(E_t(x_t) = \frac{1}{2\sigma_{\mathrm{obs}}^2} \|\bar{M}(\hat{x}_0(x_t) - x_{\mathrm{src}})\|_2^2\) produces an identical additive velocity correction: $\(\frac{dx_t}{dt} = u_\theta(x_t, t) - \gamma(t)\lambda(t) \bar{M} \left(\hat{x}_0(x_t) - x_{\mathrm{src}}\right)\)$ This pulls non-edit regions toward source fidelity while preserving the intrinsic probability flow transport structure.

4. Probing-Based Multi-Iteration Self-Refinement: Progressive convergence of region boundaries and trajectories

Establishing a static mask early in the sampling process risks over-constraining the model when semantic layout remains unresolved. SR-Edit structures the inference into multiple (default \(N=2\)) self-refinement iterations starting at 30% of the total denoising steps. Each iteration executes a fast few-step probing rollout (stride=3) to estimate an updated, crisper mask, followed by h-transform guidance steps. Successive iterations operate on cleaner intermediate representations, sharpening region localization and non-edit preservation in lockstep with the evolving synthesis.

Key Experimental Results

Main Results

Evaluated on ImgEdit-Bench across 497 regional editing instances spanning replace, adjust, background, remove, add, and action tasks, SR-Edit consistently boosts non-edit preservation and overall quality across both diffusion backbones (InstructPix2Pix, AnyEdit) and flow-matching backbones (Qwen-Image-Edit 2511, Step1X-Edit v1p2).

Backbone & Method CLIP ↑ SSIM ↑ PSNR ↑ DISTS ↓ LPIPS ↓ Judge ↑
InstructPix2Pix 26.68 0.67 16.58 0.20 0.46 2.69
+ SR-Edit (Ours) 25.07 0.68 16.94 0.18 0.43 2.75
AnyEdit 25.15 0.70 18.95 0.17 0.41 2.82
+ SR-Edit (Ours) 25.21 0.85 19.19 0.13 0.35 2.99
Qwen-Image-Edit 2511 26.16 0.61 14.92 0.23 0.46 3.84
+ Follow-Your-Shape 26.03 0.61 14.96 0.23 0.47 3.59
+ SpotEdit 26.11 0.70 16.55 0.20 0.32 3.91
+ SR-Edit (Ours) 26.80 0.67 17.40 0.18 0.31 3.94
Step1X-Edit v1p2 25.89 0.68 15.96 0.20 0.39 4.00
+ Follow-Your-Shape 25.84 0.67 15.93 0.20 0.39 4.03
+ SpotEdit 25.91 0.75 16.77 0.17 0.31 4.08
+ SR-Edit (Ours) 26.09 0.77 16.48 0.14 0.27 4.01

Ablation Study

Ablation experiments conducted on Qwen-Image-Edit 2511 examine region estimation strategies, iteration budgets, and individual post-processing modules.

Configuration / Variant CLIP ↑ SSIM ↑ PSNR ↑ DISTS ↓ LPIPS ↓ Judge ↑
SR-Edit (Full Model, 2 Iterations) 26.80 0.67 17.40 0.18 0.31 3.94
Reconstruction-based Region Estimation 26.95 0.64 16.90 0.21 0.35 3.88
Single-Iteration Refinement (Single Iter) 25.90 0.60 16.10 0.24 0.41 3.60

Evaluating the necessity of each stage within the mask post-processing pipeline confirms that every component provides indispensable, complementary benefits:

Post-Processing Component Ablation Precision (P) ↑ Recall (R) ↑ IoU ↑ F1 ↑
w/o Otsu Adaptive Separation 0.10 1.00 0.10 0.16
w/o Opening Filter 0.69 1.00 0.69 0.76
w/o Closing Filter 1.00 0.91 0.91 0.95
w/o Connected-Component (CC) Filtering 0.76 1.00 0.76 0.84

Key Findings

  • Unintended drift is effectively suppressed without sacrificing semantics: On AnyEdit, SR-Edit boosts SSIM from 0.70 to 0.85 and slashes LPIPS from 0.41 to 0.35. On Step1X-Edit, LPIPS drops substantially from 0.39 to 0.27, demonstrating that theoretically sound residual guidance preserves complex high-frequency background textures far better than heuristic KV swapping.
  • Iterative refinement prevents early over-constraining: Fixing the mask after a single early estimation step (Single Iter) causes the Judge score to drop steeply to 3.60 and CLIP to fall to 25.90. In contrast, multi-iteration refinement progressively focuses the mask boundaries as semantic features crystallize, maintaining creative flexibility in edit zones.
  • Few-step probing avoids reconstruction instability: Single-step reconstruction suffers severe variance and sudden drops in IoU/F1 at intermediate sampling steps (e.g., around step 32), whereas multi-step probing rollouts (stride=2 or 3) provide smooth, noise-resilient feedback.

Highlights & Insights

  • Principled guidance grounded in distribution transforms: By substituting ad-hoc KV caching and attention overwriting with Doob's h-transform conditioned drift and velocity adjustments, SR-Edit establishes an elegant mathematical foundation for training-free, localized diffusion manipulation.
  • Pixel-space self-feedback sidesteps latent decoder distortion: Recognizing that patch-level latent features undergo highly nonlinear transformations during decoding, the method performs thresholding and morphological filtering directly in pixel space, achieving sub-pixel precision with minimal computational overhead.
  • Generalizable sampling trajectory correction paradigm: The probe-and-correct closed loop can be readily extended beyond 2D image editing to other continuous generative trajectories, such as video inpainting, audio editing, and conditional scientific diffusion modeling.

Limitations & Future Work

  • Computational overhead from probing rollouts: Performing few-step forward probing within each SR-Edit block introduces additional forward evaluation steps, moderately increasing total inference latency compared to direct single-pass generation.
  • Complex multi-target and structural deformation boundaries: In scenarios involving extreme pose modifications or fragmented foreground objects across wide areas, Otsu thresholding and connectivity thresholds risk either merging subtle background regions or filtering legitimate small modifications.
  • Future directions: Investigating adaptive probing step schedules to lower inference FLOPs, and porting the Doob's h-transform velocity formulation to spatiotemporal flow matching models for long-context localized video editing.
  • vs Follow-Your-Shape: Follow-Your-Shape relies on trajectory divergence maps and heuristic KV caching schedules, which can introduce residual blur in the background; SR-Edit cleanses difference boundaries via morphological self-feedback and applies principled h-transform guidance, eliminating blur and boundary tearing.
  • vs SpotEdit: SpotEdit uses single-step reconstruction and intermediate KV interpolation, which suffers from numerical instability at specific sampling steps and causes visible edge jumps; SR-Edit utilizes multi-step probing rollouts and vector field corrections that perfectly preserve native sampling continuity.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Rigorous derivation of Doob's h-transform preservation guidance for both diffusion SDEs and flow-matching ODEs without training.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Thorough validation across 4 top-tier backbones, full ImgEdit-Bench evaluation, and meticulous component ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Elegant mathematical progression connecting physical intuition, empirical observation, and formal stochastic control.
  • Value: ⭐⭐⭐⭐⭐ A clean, plug-and-play solution that significantly elevates the fidelity and visual cleanliness of modern open-source image editing backbones.