OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/Zhouqm-Git/osor
Area: Image Generation
Keywords: object removal, image inpainting, efficient diffusion, shadow and reflection removal, mask robustness
TL;DR¶
OSOR introduces a one-step diffusion inpainting framework for object removal that combines an occupancy-guided multi-scale discriminator for boundary supervision with a lightweight alpha head for adaptive latent compositing, achieving complete shadow/reflection elimination and mask robustness while outperforming multi-step baselines at 4x to 30x faster inference.
Background & Motivation¶
In real-world photographic post-processing, object removal is a quintessential image editing task that requires eliminating not only the target object itself, but also restoring the underlying uncorrupted background and erasing non-local environmental interactions (such as cast shadows, specular surface reflections, and contact ambient occlusions). While traditional GAN-based inpainting models run quickly via a single forward pass, their limited representational capacity yields blurry textures, structural distortion, and unconvincing completions in complex scenes. Modern multi-step diffusion inpainting models (such as SDXL-Inpainting and FLUX Fill) substantially boost generative realism; however, their iterative denoising procedures demand dozens of network evaluations, incurring severe latency (often several to dozens of seconds per image) that prevents interactive on-device deployment.
Compounding this computational bottleneck are two critical failure modes in interactive editing: lack of effect-awareness and poor mask-robustness. In practice, user-drawn masks are rarely pixel-perfect: casual users frequently provide conservative masks that enclose only the salient foreground object while omitting surrounding shadows and reflections, leaving conspicuous ghosting artifacts in the output, or conversely draw excessively loose masks that overwrite valid background. Naively applying adversarial step distillation from global image generation to inpainting fails because single-step models lack the multi-step corrective iterations needed to smooth out seam artifacts. Furthermore, public instruction-guided editing datasets are riddled with semantic misalignments and global scene drifts, failing to provide reliable effect-aware supervision.
To resolve these tensions, this paper reformulates object removal as a single-step latent restoration task initialized from an intermediate noised latent, systematically addressing discriminator supervision granularity, mask-adaptive soft compositing, and automated effect data curation. Core idea: stabilize single-step diffusion inpainting via an occupancy-guided multi-scale patch discriminator to eliminate boundary seams, equip the backbone with a lightweight alpha head to infer true affected extents for adaptive latent compositing, and build the 280K-pair effect-aware CORNE dataset via a semantic-anchored verification pipeline (SAVP) to achieve fast, effect-aware, and mask-robust removal in a single step.
Method¶
Overall Architecture¶
OSOR is constructed upon pretrained diffusion inpainting backbones (e.g., SDXL-Inpainting and FLUX Fill) and operates in the latent space of a pretrained VAE. Given an input shot image \(x\) and a user-provided initial mask \(m\), the model first encodes the latent representation \(\bar{z} = \mathcal{E}(x)\), followed by adding intermediate forward perturbation noise at timestep \(t\) (empirically optimized at \(t=400\)) to yield \(z_t\). The diffusion backbone is conditioned on the concatenated context tuple \(c = \langle \bar{z}, m, e \rangle\) (where \(e\) denotes the text embedding of the static prompt "Remove the instance of object"), executing a single-step forward pass to estimate the clean restored background latent \(\hat{z}_0\).
To accomplish boundary-consistent and robust removal, OSOR adopts a two-phase training curriculum: 1. Phase I (Boundary-consistent One-step Removal): Trained under ground-truth effect-aware masks \(m_{\text{gt}}\) with hard latent blending to freeze unedited regions. An occupancy-guided multi-scale discriminator provides fractional patch-level supervision across 4 OpenCLIP ConvNeXt feature levels to eradicate boundary seam artifacts; 2. Phase II (Alpha-aware Robust Removal with Adaptive Blending): A lightweight alpha projection head is appended to the terminal backbone layer. Given conservative or degraded conditioning masks \(m_{\text{in}}\), the model predicts an opacity map \(\hat{\alpha}\) spanning the entire affected zone (including omitted shadows/reflections) and performs soft latent compositing, shielding end-users from tedious shadow tracing.
The end-to-end framework and training flow are illustrated below:
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Image x and Initial Mask m"] --> B["SAVP Verification Pipeline<br/>Curating Effect-Aware CORNE Dataset"]
B --> C["One-Step Latent Restoration<br/>Intermediate Timestep Forward Noising"]
C --> D["Occupancy-Guided Discriminator<br/>Phase I: Multi-Scale Area-Pooled Targets"]
D --> E["Lightweight Alpha Head<br/>Phase II: Incomplete Mask Adaptation & Soft Blending"]
E --> F["Clean Background Output Without Residual Effects"]
Key Designs¶
1. SAVP Semantic-Anchored Verification Pipeline: Mining High-Purity Effect Pairs from Noisy Instruction Triplets
High-quality paired supervision is essential for learning non-local physical interactions like shadows and reflections. SAVP processes noisy single-edit triplets from instruction corpora (e.g., NHR-Edit). It first builds a multi-feature difference heatmap spanning log-luminance, chromaticity, and gradient magnitude, extracting connected-component candidate bounding boxes \(b_{\text{diff}}\). Concurrently, GroundingDINO predicts open-vocabulary semantic boxes \(b_{\text{sem}}\) from the instruction text. Triplets with high fragmentation or excessive scattered noise are rejected immediately. Remaining boxes in \(b_{\text{diff}}\) are sorted by area and matched against \(b_{\text{sem}}\) under IoU and scale-ratio constraints to filter out global scene collapse. Next, SAM2 produces a tight object-core mask \(m_{\text{obj}}\) prompted by the validated boxes \(b_{\text{val}}\), which is fused with the difference region into an effect-aware target mask \(m_{\text{fuse}} = m_{\text{obj}} \cup m_{\text{diff}}^{\text{val}}\). By computing the effect residual \(m_{\text{eff}} = m_{\text{fuse}} \setminus m_{\text{obj}}\) and selecting effect-heavy samples based on \(\|m_{\text{eff}}\|_1 / \|m_{\text{fuse}}\|_1\), SAVP constructs the CORNE dataset comprising 280K high-fidelity paired images.
2. Occupancy-Guided Discriminator: Area-Pooled Fractional Supervision for Seamless One-Step Restoration
Directly applying adversarial distillation to image inpainting causes severe boundary seams because discriminator logit grids lose spatial precision near mask borders. Nearest-neighbor downsampling produces rigid binary targets that make partially covered boundary patches over-confident, whereas Gaussian smoothing relies on arbitrary bandwidth parameters unrelated to physical geometry. OSOR introduces occupancy-guided patch targets (OG-Patch Targets) by applying 2D area pooling to the target mask at each of the 4 discriminator scales, yielding exact fractional occupancies \(\tilde{w}_k \in [0, 1]\). Built on a frozen OpenCLIP ConvNeXt trunk and lightweight patch heads, the multi-scale discriminator minimizes: $$ \mathcal{L}D(w) = \sum_k \mathbb{E}\left[ -\log D\xi^k(x^{\text{bg}}) \right] + \sum_k \mathbb{E}\left[ -(1-\tilde{w}k) \odot D\xi^k(\hat{x}) - \tilde{w}k \odot \left(1 - D\xi^k(\hat{x})\right) \right] + \lambda_{\text{r1}}\mathcal{R}{\text{r1}} $$ The generator adversarial objective incorporates mask-area normalization to balance gradients across varied object sizes: $$ \mathcal{L}}}(w) = \sum_k \mathbb{E}\left[ \frac{\sum_p \tilde{wk(p) D\xi^k(\hat{x})_p}{\sum_p \tilde{w}_k(p) + \varepsilon} \right] $$ This fractional formulation provides smooth, continuous gradients along boundary transitions, completely resolving seam artifacts.
3. Lightweight Alpha Head: Inferring Complete Effect Extents Under Imperfect User Masks
When a user provides a tight mask omitting cast shadows, standard hard latent compositing (\(z_{\text{out}} = m_z \odot \hat{z}_0 + (1 - m_z) \odot \bar{z}\)) strictly retains the unmasked shadows, making clean removal impossible. In Phase II, OSOR augments the terminal projection of the backbone (the final convolution in SDXL or linear projection in FLUX) to output an extra logit channel \(\ell_\theta\), producing a soft alpha map \(\hat{\alpha} = \sigma(\ell_\theta)\) with virtually zero added computation. The output latent is formed via soft compositing: $$ z_{\text{out}} = \hat{\alpha} \odot \hat{z}_0 + (1 - \hat{\alpha}) \odot \bar{z} $$ During training, conditioning masks \(m_{\text{in}}\) are sampled from conservative variants (tight object masks \(m_{\text{obj}}\), eroded masks, random hole droppings, and spatial shifts), while supervision is enforced against the full effect-aware target \(m_{\text{gt}}\) via combined BCE and Dice objectives: \(\mathcal{L}_{\alpha} = \lambda_{\text{bce}}\operatorname{BCE}(\ell_\theta, m_{\text{gt}}^z) + \lambda_{\text{dice}}\operatorname{Dice}(\hat{\alpha}, m_{\text{gt}}^z)\). The diffusion backbone's deep semantic priors are thus harnessed to automatically extrapolate and eradicate omitted shadows and reflections.
Loss & Training¶
OSOR utilizes parameter-efficient fine-tuning: the vast majority of the pretrained diffusion weights remain frozen, and only lightweight LoRA adapters along with the terminal output projection layers are updated. - Phase I: The guidance map is set to \(w = m_{\text{gt}}\). Generator weights \(\theta\) and discriminator weights \(\xi\) alternate min-max optimization under normalized L1 reconstruction, LPIPS perceptual loss, and occupancy adversarial loss: $$ \mathcal{L}G(m)|}}) = \lambda_{\text{rec}} \frac{|m_{\text{gt}} \odot (\hat{x} - x^{\text{bg}1}{|m|}1 + \varepsilon} + \lambda}} \text{LPIPS}(\hat{x}, x^{\text{bg}}) + \lambda_{\text{adv}} \mathcal{L{\text{adv}}(m) $$ - }Phase II: Initialized from Phase I checkpoints, the model receives incomplete masks \(m_{\text{in}}\) while evaluating reconstruction and adversarial losses strictly on \(m_{\text{gt}}\) (preventing the trivial shortcut of shrinking \(\hat{\alpha}\)). The full minimax formulation optimizes: $$ \min_\theta \max_\xi \mathcal{L}G(m}}) + \mathcal{L\alpha - \mathcal{L}_D(m) $$}
Key Experimental Results¶
Main Results¶
Evaluation spans paired-background benchmarks CORNE-Val and RORD-Val, alongside AnimeEraseBench, TextEraseBench, OmniPaint-Bench, and RemovalBench. Metrics comprise full-reference fidelity (PSNR, SSIM, LPIPS), distribution realism (FID, CMMD), Contextual Fracture Distance (CFD) for seam quality, and per-image latency on an NVIDIA A100.
The table below presents core comparative results on CORNE-Val, RORD-Val, and AnimeEraseBench under both Object Mask and Object-Effect Mask settings:
| Dataset | Method | Latency (s)โ | Object Mask FIDโ | Object Mask PSNRโ | Object Mask SSIMโ | Object Mask CFDโ | Object-Effect PSNRโ |
|---|---|---|---|---|---|---|---|
| CORNE-Val | OmniPaint | 26.12 | 18.6077 | 28.436 | 0.9166 | 0.1950 | 27.950 |
| ObjectClear | 6.44 | 22.0457 | 29.207 | 0.9191 | 0.2040 | 29.012 | |
| AttentiveEraser | 8.69 | 25.4818 | 25.650 | 0.8636 | 0.2244 | 25.602 | |
| FLUX-Fill (Multi-step) | 25.27 | 59.5678 | 23.124 | 0.9130 | 0.2766 | 23.629 | |
| OSOR (SDXL) | 0.42 | 29.3846 | 27.262 | 0.8371 | 0.1608 | 27.422 | |
| OSOR (FLUX) | 0.80 | 12.3241 | 32.031 | 0.9373 | 0.1538 | 32.189 | |
| RORD-Val | OmniPaint | 16.51 | 27.5658 | 23.293 | 0.7899 | 0.3707 | 23.078 |
| ObjectClear | 8.69 | 25.4667 | 25.804 | 0.8500 | 0.3248 | 25.633 | |
| AttentiveEraser | 8.68 | 36.5740 | 22.798 | 0.7023 | 0.3608 | 23.080 | |
| OSOR (SDXL) | 0.42 | 35.5989 | 23.767 | 0.6827 | 0.2644 | 23.826 | |
| OSOR (FLUX) | 0.62 | 27.3621 | 25.599 | 0.8175 | 0.2465 | 25.613 | |
| AnimeErase | OmniPaint | 24.93 | 27.4034 | 25.542 | 0.8573 | 0.4369 | 25.351 |
| ObjectClear | 10.18 | 36.6865 | 25.429 | 0.8409 | 0.3872 | 25.530 | |
| OSOR (SDXL) | 0.48 | 55.7614 | 23.090 | 0.6857 | 0.2820 | 23.130 | |
| OSOR (FLUX) | 0.89 | 26.8352 | 27.780 | 0.8859 | 0.2703 | 27.871 |
Performance across the other three object-only benchmarks reinforces OSOR's advantages: on TextEraseBench, OSOR(FLUX) achieves 31.997 dB PSNR and 19.6169 FID (outperforming ObjectClear's 29.689 dB and 29.1359 FID) in 0.83s (12.2x speedup); on OmniPaint-Bench, OSOR(FLUX) records 49.1927 FID and 24.936 dB PSNR in 1.06s versus OmniPaint's 34.13s (32.2x speedup).
Ablation Study¶
Ablations on discriminator target construction on RORD-Val under object-effect masks establish the necessity of fractional occupancy guidance:
| Patch Target Type | PSNRโ | SSIMโ | FIDโ | LPIPSโ | CFDโ | Note |
|---|---|---|---|---|---|---|
| Hard Mask (Nearest-neighbor) | 22.463 | 0.6739 | 56.1508 | 0.2007 | 0.3399 | Binary downsampling causes boundary overconfidence |
| Gaussian Soft | 21.992 | 0.6708 | 64.1290 | 0.2103 | 0.3430 | Fixed filter bandwidth diverges from true patch area |
| Occupancy (Ours) | 22.696 | 0.6741 | 50.0639 | 0.1855 | 0.3013 | Area pooling yields exact fractional targets, eliminating seams |
Sweeping forward noising timestep \(t \in \{200, 400, 600, 800\}\) confirms that \(t=400\) provides the best trade-off (FID 50.0639, PSNR 22.696 dB), avoiding the under-conditioning of \(t=200\) (FID 60.1520) and the over-distortion of \(t=800\) (FID 51.8961).
Ablating the alpha head and Phase II training on RORD-Val (SDXL backbone) demonstrates significant resilience to conservative inputs:
| Conditioning Mask | Training Phase & Blending | PSNRโ | SSIMโ | FIDโ | LPIPSโ | CFDโ |
|---|---|---|---|---|---|---|
| Object-only \(m_{\text{obj}}\) | Phase I (Hard Blending) | 21.965 | 0.6726 | 59.8017 | 0.2062 | 0.3133 |
| Object-only \(m_{\text{obj}}\) | Phase II (Alpha Soft Blending) | 23.767 | 0.6827 | 35.5989 | 0.1795 | 0.2644 |
| Effect-aware \(m_{\text{gt}}\) | Phase I (Hard Blending) | 22.696 | 0.6741 | 50.0639 | 0.1855 | 0.3013 |
| Effect-aware \(m_{\text{gt}}\) | Phase II (Alpha Soft Blending) | 23.826 | 0.6823 | 35.6523 | 0.1797 | 0.2624 |
Key Findings¶
- Alpha head is transformative for imperfect masks: Under conservative object-only masks, Phase I hard blending leaves residual shadows unaddressed (FID 59.8017). Phase II alpha compositing cuts FID by 24.2 points to 35.5989 and raises PSNR by 1.80 dB, confirming the backbone's ability to self-extrapolate the true editing scope.
- Occupancy guidance is essential for seam-free restoration: Area-pooled fractional supervision lowers the Contextual Fracture Distance from 0.3399 to 0.3013, preventing boundary discontinuity artifacts common in single-step generation.
- Superior efficiency without quality trade-offs: OSOR(FLUX) outperforms the 28-step FLUX Fill across all metrics while slashing runtime from 25.27s to 0.80s, confirming the viability of one-step latent restoration.
Highlights & Insights¶
- Area-pooled patch targets for inpainting boundary supervision: While standard generative discriminators evaluate whole images, local editing models suffer primarily at boundary interfaces. Employing 2D area pooling to calculate continuous spatial coverage fractions effectively stabilizes single-step adversarial learning and is easily extensible to other localized editing tasks.
- Zero-overhead alpha head via backbone projection reuse: Instead of deploying an auxiliary segmentation model, OSOR expands the output channels of the terminal projection layer, repurposing the backbone's intrinsic semantic representations to estimate soft transparency maps with negligible parameter overhead.
- Autonomous high-purity data filtering loop (SAVP): The combination of multi-feature physical difference mapping, open-vocabulary semantic alignment, SAM2 boundary completion, and effect residual thresholding provides a scalable blueprint for purifying noisy web-scale editing corpora.
Limitations & Future Work¶
- Extreme large-scale occlusions and complex global relighting: When removing objects that cast complex shadows across non-planar reflective surfaces or occupy over half the frame, single-step inference can occasionally struggle with global geometric consistency.
- Alpha boundary ambiguity on rough specular materials: On highly reflective water or rough mirror surfaces, the predicted alpha map can occasionally appear slightly diffuse or overly dilated.
- Future directions: Integrating physical illumination reasoning from multimodal LLMs into the alpha head, or exploring few-step (2-4 step) deterministic refinement loops to further boost geometric fidelity while preserving sub-second interactive speed.
Related Work & Insights¶
- vs OmniPaint / ObjectClear / OmniEraser: Prior state-of-the-art models handle effect-aware removal through iterative multi-step diffusion (20-50 steps) or self-attention redirection, incurring latency between 6 and 34 seconds. OSOR is the first framework to achieve effect-awareness and mask-robustness within a single step, attaining sub-second speed (<1s) while beating these baselines in perceptual quality.
- vs Adversarial Diffusion Distillation (ADD / DMD): Standard distillation frameworks focus on unconditional or text-to-image distribution matching and overlook spatial continuity along masked borders. OSOR's latent restoration formulation and occupancy-guided discriminator bridge this gap for local mask-conditioned editing.
Rating¶
- Novelty: โญโญโญโญโญ Formulates the first effect-aware, mask-robust one-step diffusion inpainting framework with an occupancy-guided discriminator and an integrated alpha head.
- Experimental Thoroughness: โญโญโญโญโญ Evaluated across 6 benchmark datasets and 7 leading baselines, supported by detailed multi-metric comparisons and solid ablations.
- Writing Quality: โญโญโญโญโญ Rigorous mathematical modeling, clean logical progression, and direct targeting of real-world deployment challenges.
- Value: โญโญโญโญโญ Exceptional practical utility for interactive on-device image editing, edge processing, and latency-critical computational photography.