Skip to content

Less Supervision, Better Generalization: Weakly Supervised Fake Region Localization in Diffusion-Edited Images

Conference: NeurIPS2026 (task-list assignment; the version read is an arXiv preprint)
arXiv: 2609.36882v1
Area: AI Safety
Keywords: weakly supervised localization, image forensics, reconstruction prior, multiple instance learning, cross-generator generalization

TL;DR

ReGFLoW trains a dual-encoder, VAE-reconstruction-guided regional scoring network with image-level real/fake labels, improving average cross-generator F1 in mixed partial-edit/full-synthesis evaluation at a substantial in-domain localization cost; it does not beat the strongest mask-supervised baseline's OOD average on partial edits alone.

Background & Motivation

Image-level real/fake classification indicates whether an image is suspicious but does not locate the changes. Localizing diffusion edits also differs from semantic segmentation: texture within an object, backgrounds, and illumination can change without artifacts following object boundaries. Densely supervised methods such as MaskCLIP and TruFor can learn clear editing regions on controlled data, but their training masks typically describe the intended edit rather than the complete traces left by generation.

The paper identifies two constraints arising from this distinction. First, instruction-based or hosted editing services often expose inputs, instructions, and outputs without reliable pixel-level artifact labels. Second, latent encoding, decoding, and blending can affect regions outside the mask; labeling all those locations as real can suppress weak, transferable evidence. Appendix A further distinguishes conditioning masks, interface selections, and forensic supervision masks, explaining why before–after differences combine geometric misalignment, exposure changes, and texture harmonization rather than providing reliable labels directly.

Removing masks does not automatically solve localization. Image-level labels require only enough evidence to identify a fake image, potentially producing a few activated locations or widespread false positives. Expanding responses through semantic similarity can also mistake an entire person or object for an edited region. Core idea: use the reconstruction behavior of a frozen diffusion-model VAE as spatial guidance, inject it into both features and scores, and learn regional evidence with artifact-centric multiple instance learning without forcing artifacts to follow editing-mask boundaries.

Method

Overall Architecture

ReGFLoW takes an image under examination and its reconstruction-residual prior, producing a regional localization map and a separate image-level real/fake prediction. Dual-encoder Representation combines frozen CLIP global semantics with trainable MAE local features; Reconstruction-Guided Scoring injects frozen SD1 VAE residuals into local features and regional logits to obtain a low-resolution score map.

During training, Artifact-Centric Multiple Instance Learning aggregates score-map evidence and uses image-level labels to order the evidence of real and fake samples; the image classification branch separately combines global representations with high-response local regions. At test time, Image-Wise Adaptive Calibration adjusts the threshold of the upsampled score map before producing a binary mask. It does not enter the training losses and is not equivalent to statistically validated posterior probability calibration.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400, 'subGraphTitleMargin': {'top': 8, 'bottom': 16}}}%%
flowchart TD
    I["Image under examination"] --> A["Dual-encoder Representation"]
    subgraph B["Reconstruction-Guided Scoring"]
        R["Frozen SD1 VAE<br/>reconstruction residual and prior encoding"] --> F["Gated feature correction<br/>and score bias"]
    end
    I --> R
    A --> F
    F --> S["Regional score map"]
    S -->|training| C["Artifact-Centric Multiple Instance Learning"]
    Y["Image-level real/fake labels"] -.-> C
    A --> H["Image classification branch"]
    S -->|high-response local representations| H
    Y -.-> H
    S -->|testing| D["Image-Wise Adaptive Calibration"]
    D --> O["Localization mask"]

Key Designs

1. Dual-encoder Representation: retain global semantics and local artifact cues

High-level semantics alone can anchor attention to objects, whereas artifacts may appear as weak textures or local inconsistencies. The model uses frozen CLIP ViT-L/14 for global visual features and an ImageNet-pretrained MAE ViT-B/32 to initialize a trainable local artifact encoder, exchanging information through intermediate cross-attention. The CLIP encoder is frozen, not the entire localization network: the local encoder and task-specific heads still learn from real/fake supervision.

The local branch receives 512×512 inputs. The main text describes a 32×32 local feature grid pooled to 16×16 with 768 channels, giving multiple instance learning a more compact instance set and easing propagation of sparse image-level gradients. The backbone name in the appendix cannot replace the complete side-adapter network's dimensional description; this note does not derive a different grid size from that name alone.

The image classification branch selects CLS features from the global encoder, aggregates them with single-layer attention, and concatenates the result with local evidence. Local evidence comes from the highest-response 10% of locations: hard selection is followed by masked softmax attention based on standardized scores and a two-layer MLP. This hard top-10% selection belongs to the classification branch and must not be confused with the soft aggregation used for localization MIL below.

2. Reconstruction-Guided Scoring: let the prior correct features and scores separately

The prior comes from one encoding–decoding pass through a frozen SD1 VAE, not reverse diffusion, denoising sampling, or diffusion inversion. The rationale is that content generated through latent diffusion may lie closer to the autoencoder's learned manifold and exhibit smaller or more regular reconstruction residuals. Residual magnitude is not a direct real/fake label, however: the network must learn its relationship to local artifacts rather than simply classifying high-residual regions as fake.

\[ \mathbf{r}=\left|I-\mathcal{D}\!\left(\mathcal{E}(I)\right)\right|. \]

Main-text Equation (1) computes channel-wise absolute differences. A three-layer convolutional prior encoder converts residuals into spatial features and pools them to the same 16×16 grid with 32 prior channels. Appendix C instead describes loading an offline-provided 128×128 single-channel residual during training. The source does not fully specify channel reduction and preprocessing, so these descriptions cannot be treated as an unambiguously identical implementation detail.

Local and prior features pass through separate convolutional projectors and are concatenated to predict a correction offset, whose magnitude is controlled by a learnable scalar gate. A base scoring head predicts logits from corrected features. A separate bias branch reads both corrected and prior features to add reconstruction evidence directly to the logits.

\[ \widetilde{\mathbf{F}}=\mathbf{F}^{\mathrm{loc}}+\tanh(\gamma)\Delta(\mathbf{F}^{\mathrm{loc}},\mathbf{F}^{\mathrm{rec}}),\qquad \mathbf{S}=\mathbf{S}_{0}+\tanh(\beta)b(\widetilde{\mathbf{F}},\mathbf{F}^{\mathrm{rec}}). \]

Both scalar gates start at 0.05, so feature correction begins near identity but is not strictly zero. Strict zero initialization applies to the bias predictor's final 1×1 convolution: its initial output is zero, making initial scores equal to the base logits. This distinction prevents weak prior injection from being misrepresented as every reconstruction-related branch starting completely disabled.

3. Artifact-Centric Multiple Instance Learning: replace pixel labels with real/fake evidence ranking

Each 16×16 logit map forms a bag of 256 instances. A fake image needs generated evidence in some regions, whereas a real image should not contain strong artifact responses. Simple averaging dilutes small edits, while a fixed hard selection of highest-scoring regions does not accommodate area variation. Main-text Equation (5) uses softmax weights on image-wise standardized logits to aggregate the original logits:

\[ \hat{s}_{i}=\sum_{n=1}^{N}\alpha_{i,n}s_{i,n},\qquad \alpha_{i,n}=\frac{\exp(\kappa z_{i,n})}{\sum_{m=1}^{N}\exp(\kappa z_{i,m})},\qquad z_{i,n}=\frac{s_{i,n}-\mu_i}{\sigma_i+\epsilon}. \]

The mean and standard deviation are computed over one image's regional logits, and the sharpness parameter controls concentration on high-response instances. This behaves like soft top-k but does not perform a literal hard top-K cutoff: all locations remain in the differentiable aggregation. The ablation paragraph's phrase “adaptive top-k” should be understood through this formula rather than used to rewrite the algorithm.

All fake–real image pairs within a batch are then constrained to have fake aggregate scores exceed real aggregate scores by a margin. This supervises evidence ordering, not the assertion that every region of a fake image is fake, and it supplies no manually annotated local pseudo-mask.

\[ \mathcal{L}_{\mathrm{acmil}}=\frac{1}{|\mathcal{F}||\mathcal{R}|}\sum_{i\in\mathcal{F},\,j\in\mathcal{R}}\operatorname{softplus}\!\left(m-\hat{s}_{i}+\hat{s}_{j}\right). \]

Ranking can still identify only a few discriminative regions and cannot independently guarantee complete spatial coverage. The model therefore suppresses high responses on real images, encourages spatial smoothness on fake images, and combines classification with auxiliary regularization described in the appendix. Weak supervision means neither an absence of supervision nor an absence of pretrained-model priors.

4. Image-Wise Adaptive Calibration: convert relative regional scores into a sample-adaptive threshold

Weakly supervised ranking principally constrains relative evidence, leaving absolute score scales variable across images. At test time, logits are bilinearly upsampled to input resolution, and the method computes their image-wise mean and standard deviation together with the whole-image mean of raw sigmoid probabilities. Although called a “positive response ratio” in the source, the latter is a continuous probability mean, not the proportion of pixels above 0.5.

\[ r_i=\frac{1}{HW}\sum_{u,v}\sigma(U_{i,u,v}),\qquad d_i=\operatorname{clip}\!\left(\frac{r_i-c}{\max(c,\epsilon)},-1,1\right),\qquad k_i=\operatorname{clip}(k_0-gd_i,k_{\min},k_{\max}),\qquad \tau_i=\mu_i+k_i\sigma_i. \]

The method subtracts this threshold from logits, divides by the standard deviation, multiplies by a scale factor, and applies sigmoid before binarizing at 0.5. Default values are center 0.08, base coefficient 0.9, adjustment strength 0.50, coefficient bounds 0.40–1.20, scale factor 2.0, and stability constant 10^-6. A higher mean raw response generally reduces the threshold offset. This is distribution normalization and heuristic threshold adjustment, not a guarantee of reliable probabilistic calibration.

A Worked Example

Consider an image with a locally altered texture. CLIP supplies the global scene representation, the local encoder supplies spatial features, and the VAE produces an aligned reconstruction residual. The residual does not directly become an editing mask: it influences both feature correction and score bias.

Once the 16×16 score map is formed, soft MIL aggregation during training emphasizes high-response regions and compares their aggregate evidence with that of real images. The classification branch separately selects the top-response 10% of local locations for joint classification with global features. Test-time localization calibrates the complete upsampled logit map instead of performing the fake–real pairwise ranking used in training.

If an edit is smaller than the effective spatial scale of the regional grid, its evidence may occupy an instance containing both real and modified content. Interpolation back to pixel resolution cannot recover precise boundaries absent from that representation, explaining the small-edit limitation acknowledged in Appendix F.

Loss & Training

Main-text Equation (7) lists MIL ranking, image-level classification cross-entropy, real-image hard-negative suppression, and edge-aware smoothness. Real-image suppression takes the mean of the highest 10% of regional logits and applies softplus, targeting sparse high-score false positives. Fake-image smoothness penalizes probability differences between two-dimensional neighbors, reducing the penalty at stronger grayscale image edges; its default edge-weight parameter is 10.0.

Appendix B's Equation (9) adds fake-image sparsity, patch-level contrastive regularization, and optional peak separation. The sparsity term lightly penalizes mean fake-image sigmoid responses to discourage an all-positive solution. The contrastive term selects high-response fake embeddings and hard real embeddings, applying supervised contrastive learning with classes inherited from image labels; its default selection fraction and temperature are both 0.10. Peak separation has zero default weight and is not an active default training contribution.

These appendix terms must not be silently presented as identical to the main-text objective. Appendix B gives smoothness weight 0.03 and sparsity weight 3×10^-3. Appendix C gives smoothness weight 0.05, uses alternative names such as rank and contrast, and does not restate sparsity in that list. Both give real-suppression weight 0.05 and contrastive weight 0.02. These configuration descriptions are inconsistent and require implementation details or author clarification.

Appendix C reports 10 training epochs with AdamW, learning rate 5×10^-5, weight decay 0.05, no warmup, a 3.0 learning-rate multiplier for new heads, and random seed 42. It also states “6 GPUs, batch 8 per GPU, effective batch 64,” whereas direct multiplication gives 48. No gradient accumulation configuration explains the difference, so this note preserves the conflict instead of correcting it on the authors' behalf.

Pixel masks are not used in training and are reserved for evaluation. Reconstruction residuals are loaded with training samples, and the prior is also needed at test time. A frozen VAE adds no trainable VAE parameters, but this does not imply zero preprocessing or inference cost.

Key Experimental Results

Main Results

Experiments follow the OpenSDID protocol, training on SD1.5 and evaluating it in-domain while treating the other generators as cross-domain. P denotes partial edits only; P+F jointly trains and evaluates on partial edits and fully synthetic images. The table below extracts main-text Tables 1 and 2: every score is pixel-level F1, and OOD Avg averages the four cross-domain generators, excluding SD1.5.

Setting Method SD1.5 SD2.1 SDXL SD3 Flux.1 Avg. OOD Avg.
P MaskCLIP 75.8 63.2 35.2 49.8 18.4 48.5 41.7
P ReGFLoW 49.3 46.3 34.9 43.9 29.2 40.7 38.6
P+F MaskCLIP 95.1 73.2 20.8 17.7 4.8 42.3 29.1
P+F ReGFLoW 35.6 33.4 30.1 32.2 29.4 32.2 31.3

Under P, ReGFLoW scores higher on Flux.1 but lower than MaskCLIP on both OOD Avg and in-domain evaluation. Under P+F, the OOD Avg improvement from 29.1 to 31.3 does not mean it wins on every generator: it still trails substantially on SD2.1, and the in-domain change from 95.1 to 35.6 is a particularly large cost. The claimed benefit should be restricted to transfer under some cross-domain conditions rather than universal superiority over dense supervision.

For image-level detection in main-text Table 3, ReGFLoW's OOD F1/ACC is 76.22/80.93, compared with MaskCLIP's 69.28/76.33 and NPR's 73.71/75.74. Detection metrics cannot substitute for regional localization metrics. On Flux.1, for example, ReGFLoW's detection F1 is 56.78, below NPR's 67.62.

Ablation Study

The table below comes from main-text Table 4. Neither its caption nor the adjacent ablation discussion explicitly identifies the data split and generator scope for this table. The proximity of 49.38 to the main table's in-domain score does not justify calling it an OOD average. The final row simultaneously lacks both reconstruction modules and changes the objective, so it is not a single-variable MIL-removal experiment.

Config Score bias Feature fusion Artifact-centric MIL Pixel F1
ReGFLoW Yes Yes Yes 49.38
w/o Bias Predictor No Yes Yes 46.06
w/o Reconstruction Prior No No Yes 38.08
Score-Pooling BCE No No No 32.34

The full configuration outperforms the configuration retaining feature fusion without score bias, with a larger gap when the reconstruction prior is removed entirely. This supports dual-space injection, but does not show that reducing label density alone causes all gains, since architecture, spatial priors, and objectives also differ.

Main-text Table 5 separately adds one source domain while holding the total number of training images fixed, giving average unseen-domain gains of +4.40, +4.58, +3.52, and +3.82. Its SD1.5-only baseline differs numerically from Table 1 and should be interpreted as an independent domain-diversity experiment rather than merged with that protocol. Individual target domains need not improve: adding SDXL gives -0.08 on SD2.

Key Findings

  • Reconstruction priors provide learnable spatial cues beyond masks, but are neither pixel-level ground truth nor a rule that larger residuals imply fake content.
  • The clearest transfer gains concern the distant Flux.1 generator and the P+F OOD average; in-domain accuracy and the P OOD average expose substantial trade-offs.
  • Fixed-size multi-source experiments support domain diversity, but do not independently identify the causal contributions of weak labels, architecture, and priors.
  • Appendix E discusses failures involving tiny edits, compression artifacts, and high-frequency real textures. The text cache includes captions rather than interactive source figures, so qualitative descriptions do not replace image-by-image visual verification.

Highlights & Insights

  • The most transferable idea is turning reconstruction behavior from an image-level detection cue into a learnable local prior. Feature correction and logit bias intervene in representation and decision spaces without thresholding residuals into pseudo-masks.
  • Distinguishing intended editing regions from the spatial support of generated traces matters beyond annotation cost. Forensic targets need not follow semantic boundaries, making object-similarity expansion potentially misleading.
  • Evidence aggregation, real-image false positives, and final thresholding deserve separate treatment under weak supervision. Soft MIL, hard-negative suppression, and image-wise calibration offer this division of labor, but each still requires validation on the target distribution.

Limitations & Future Work

  • The authors acknowledge extra reconstruction cost and imprecise small-edit localization. The 16×16 score grid constrains boundary detail, and bilinear upsampling cannot remove that information bottleneck.
  • The paper critiques incomplete editing masks yet still uses them for pixel F1/IoU evaluation. Detecting genuine generated traces outside masks may count as false positives: low scores alone cannot disprove such traces, but the authors' explanation also cannot automatically excuse every out-of-mask response. More reliable artifact annotations or separate-target evaluation are needed.
  • OpenSDID results do not guarantee that the SD1 VAE prior transfers to fundamentally different generation pipelines, and no sufficient cost evidence supports a zero-extra-latency claim.
  • Single-channel residual preprocessing, batch arithmetic, and loss configuration remain unresolved. Main-text/appendix inconsistencies limit certainty about direct reproduction.
  • Test thresholds use fixed heuristic parameters, and sparsity regularization plus relative-distribution thresholding may affect fully synthetic image coverage. Future evaluations should report threshold sensitivity, real-image false positives, performance by edit size, and supervision-strength controls within the same architecture.
  • vs MaskCLIP / OpenSDI: Both combine CLIP and MAE for joint detection and localization, but ReGFLoW learns spatial evidence with image labels, reconstruction priors, and MIL. The comparison demonstrates cross-domain benefits alongside in-domain losses, not that this paper invented the basic dual-encoder design.
  • vs DIRE / AEROBLADE: These works establish reconstruction error as a generated-image detection cue. ReGFLoW guides regional features and scores with frozen VAE encoding–decoding residuals and should not be described as executing a DIRE-style reverse-diffusion pipeline.
  • vs semantic weakly supervised localization / WSCL: The former often propagates through object similarity, while the latter learns generic manipulation cues through multi-source and inter-patch consistency. ReGFLoW targets diffusion traces with reconstruction behavior as a dense prior, while still relying on pretrained visual models and image labels.
  • vs earlier diffusion weakly supervised localization research: A cited WACV study already investigates controlled diffusion-face localization. The paper's “first” claim should be restricted to its emphasized general diffusion-generated/edited image setting, not broadened into a claim that no earlier diffusion weakly supervised localization exists.

Rating

  • Novelty: 4/5 — Combines reconstruction priors, dual-space injection, and artifact-centric MIL for general diffusion-edit localization, but weakly supervised localization and dual encoders have precedents.
  • Experimental Thoroughness: 3/5 — Covers cross-generator transfer, mixed fake-image types, and domain diversity, but lacks sufficient supervision-strength controls, explicit ablation scope, and runtime analysis.
  • Writing Quality: 3/5 — The main mechanism is clear, while appendix loss weights, effective batch size, and prior format introduce reproducibility ambiguities.
  • Value: 4/5 — Informative for image forensics without dense masks, but deployment requires validation of in-domain accuracy, tiny edits, and false-positive risks.