Skip to content

PhysEdit: Physically Consistent Image Editing via Causal Enforcement

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/HiDream-ai/PhysEdit/
Area: Image Generation
Keywords: Image Editing, Causal Generation, Physical Consistency, Reinforcement Learning, Diffusion Transformer

TL;DR

Addressing the issue where conventional single-step image editing neglects physical laws—causing floating objects and missing shadows or reflections—PhysEdit reformulates editing into an autoregressive causal evolutionary process, combining single-step teacher consistency regularization with fine-grained physics-aware reinforcement learning to achieve physically consistent edits.

Background & Motivation

Instruction-guided image editing has witnessed rapid progress driven by diffusion transformer architectures and large-scale training paradigms, enabling intricate visual manipulations governed by open-vocabulary text prompts. However, prevailing editing models primarily optimize for semantic alignment and visual fidelity, frequently failing to respect underlying physical laws and causal relationships. Typical failure modes include removing the bottom support of stacked objects while the upper items remain floating in mid-air in violation of gravity, or removing a subject in front of a mirror while leaving its reflection or cast shadow intact. Such physical implausibility severely compromises the realism and practical utility of synthesized outputs.

This fundamental shortcoming stems from two primary architectural and data bottlenecks in existing paradigms: first, standard training datasets are confined to discrete, static input-output image pairs, offering sparse supervision that prevents models from internalizing continuous physical dynamics and mechanical/optical transitions; second, prevailing architectures perform a direct, non-causal single-step mapping, where bidirectional attention obliterates the arrow of time and completely bypasses the real-world chain of physical causality. While recent works attempt to simulate dynamics using video generation models, video synthesis inevitably introduces camera jitters and viewpoint drifts that degrade pixel-level consistency across unedited background regions.

To resolve the tension between continuous physical dynamics and spatial background preservation, this paper factorizes editing along the temporal causal evolution while maintaining spatial diffusion fidelity. Core idea: reformulate image editing as a continuous causal diffusion process over time, distill dynamic trajectories from video generative models into the PhysEdit-50K dataset, and integrate single-step teacher consistency regularization alongside MLLM-guided fine-grained physics reinforcement learning (GRPO) to enforce physical consistency.

Method

Overall Architecture

PhysEdit reformulates the conventional single-step direct mapping from a source image \(I\) and instruction \(c\) into a trajectory of \(N\) intermediate visual states \(x_{1:N}\). At each temporal step \(i\), spatial denoising operates within a continuous latent space via Rectified Flow, while the multi-step sequence strictly adheres to an autoregressive causal probability factorization conditioned on past history. To balance continuous physical causality with pixel-level fidelity in unaltered regions, the framework introduces a frozen single-step editing teacher model for consistency regularization. During post-training, a vision-language model decomposes complex instructions into fine-grained physical constraints, steering the policy through reinforcement learning toward strict adherence to physical laws.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Source Image and Instruction"] --> B["Dataset Construction and Fine-Grained Verification<br/>Synthesize trajectories via video priors and filter via MLLM rules"]
    B --> C["Causal Diffusion Transformer Architecture<br/>Autoregressive temporal factorization + Spatial flow matching"]
    C --> D["Self-Forcing Training Mechanism<br/>On-the-fly autoregressive rollout with KV caching"]
    C --> E["Single-Step Teacher Consistency Regularization<br/>Anchor unedited regions against video viewpoint drift"]
    D --> F["Physics-Aware Reinforcement Learning<br/>MLLM multi-constraint reward and GRPO policy optimization"]
    E --> F
    F --> G["Physically Consistent Final Edited Image"]

Key Designs

1. Dataset Construction and Fine-Grained Verification: Distilling Continuous Physical Trajectories from Video Priors

Existing physics-aware editing benchmarks provide only static pairs, preventing models from learning intermediate state transitions. The authors leverage the rich physical priors embedded in advanced video generative models (Wan2.2) to synthesize dynamic transformation videos \(V = \mathcal{F}_{vid}(I, c)\) conditioned on source image \(I\) and editing prompt \(c\), where the final frame represents the edited target \(I_e\). To eliminate semantic drift and physical hallucinations common in video models, an instance-adaptive fine-grained verification mechanism is introduced: an MLLM (Qwen2.5-VL-72B) decomposes instruction \(c\) into \(K\) granular physical constraints \(\{p_1, p_2, \dots, p_K\}\) (e.g., "stem refracts and bends upon water entry", "glass retains transparency and specular highlights", "small meniscus appears along the boundary"). Only video samples where the final frame satisfies all \(K\) binary checks are preserved, from which \(N\) uniformly sampled frames form the PhysEdit-50K dataset containing 51,460 qualified trajectories across 8 physical interaction categories.

2. Causal Diffusion Transformer Architecture: Decoupling Temporal Autoregression from Spatial Flow Matching

Single-step discrete mappings circumvent causal chains, whereas full-sequence bidirectional denoising allows future states to influence the past, violating temporal causality. PhysEdit adopts a Causal Diffusion Transformer (cDiT) that factorizes the joint editing trajectory distribution strictly along the time dimension:

\[p_\theta(x_{1:N} | c) = \prod_{i=1}^N p_\theta(x_i | x_{<i}, c)\]

For each discrete state \(i\), spatial generation follows the Rectified Flow framework, interpolating linearly between clean latent \(x_i\) and Gaussian noise \(\epsilon \sim \mathcal{N}(0, I)\) via \(z_{i,t} = (1-t) x_i + t \epsilon\) (\(t \in [0, 1]\)), while the DiT predicts the velocity field \(v_\theta\). During attention computation, attention operates causally over the historical context, ensuring physical effects propagate forward in time.

3. Self-Forcing Training Mechanism: Eliminating Exposure Bias with Efficient Rollout

Standard teacher-forcing injects ground-truth preceding frames during training, causing severe exposure bias as errors accumulate during inference. PhysEdit employs a self-forcing training strategy where the current state is conditioned directly on the model's own on-the-fly generated history \(\hat{x}_{<i}^\theta\). To maintain computational feasibility, a few-step Ordinary Differential Equation (ODE) solver is employed for rapid generation of preceding latents, combined with Key-Value (KV) caching to incrementally store intermediate Transformer activations, eliminating redundant computation across the history trajectory without sacrificing precision.

4. Single-Step Teacher Consistency Regularization: Countering Video Viewpoint Shifts to Anchor Backgrounds

Visual trajectories generated by video models often contain subtle global viewpoint shifts or camera jitter, which can cause background blurring or spatial distortion if trained solely on a flow matching objective. PhysEdit incorporates a frozen high-fidelity single-step editing model (such as OmniGen2 or Qwen-Image-Edit) as a teacher \(\phi_{teacher}\). When the cDiT predicts the final clean latent \(\hat{x}_N^\theta\), it is re-corrupted to \(z_t(\hat{x}_N^\theta)\) and passed to the teacher along with \(x_0\) and \(c\). The teacher predicts reference velocity \(v_{\text{teacher}}\) to derive a refined target:

\[\hat{x}_N^{\phi_{teacher}} = z_t(\hat{x}_N^\theta) - (1-t) v_{\text{teacher}}\]

The model is optimized via an explicit consistency loss:

\[\mathcal{L}_{\text{consist}}(\theta) = \mathbb{E}_{t, \epsilon} \left[ \| \text{sg}(\hat{x}_N^{\phi_{teacher}}) - \hat{x}_N^\theta \|_2^2 \right]\]

where \(\text{sg}(\cdot)\) denotes the stop-gradient operator. This loss penalizes spatial drift, aligning unedited regions with the high-fidelity manifold established by the single-step teacher.

5. Physics-Aware Reinforcement Learning: GRPO Policy Optimization via Fine-Grained Physical Rewards

To enhance generalization across complex, open-world physical causalities, PhysEdit introduces a second-stage reinforcement learning phase (PARL) utilizing an MLLM \(\mathcal{M}\) (Qwen2.5-VL-7B) as a physics expert. Rather than assigning a holistic visual quality score, the MLLM evaluates the edited image \(I^e\) against the \(K\) fine-grained physical constraints \(\{p^k\}_{k=1}^K\) extracted from prompt \(c\), defining a dense scalar reward as the constraint satisfaction rate:

\[R(I^e, c) = \frac{1}{K} \sum_{k=1}^K \mathbb{I}(\mathcal{M}(I^e, p^k) == \text{True})\]

Using Group Relative Policy Optimization (GRPO) with a KL divergence penalty, \(G\) candidate images \(\{I_i^e\}_{i=1}^G\) are sampled per prompt, and group-normalized relative advantages \(\hat{A}_i = (r_i - \text{mean}(r)) / \text{std}(r)\) guide the policy update, effectively aligning the generator with challenging physical phenomena such as structural collapse, optical refraction, and light-matter interactions.

Loss & Training

The training pipeline consists of two stages: 1. Supervised Fine-Tuning (SFT): Trained on PhysEdit-50K for 2 epochs with batch size 8 and constant learning rate 1e-5. A condition dropout rate of 0.1 is applied to prompts and input images to enable classifier-free guidance (CFG). The objective combines flow matching with consistency regularization: \(\mathcal{L} = \mathcal{L}_{\text{FM}} + \lambda \mathcal{L}_{\text{consist}}\). 2. Reinforcement Learning (PARL): Optimized via GRPO for 1 epoch with group size \(G=4\) under a frozen reference policy penalty, using Qwen2.5-VL-7B to score fine-grained physical constraint satisfaction.

Key Experimental Results

Main Results

Quantitative evaluations are conducted on the physics-oriented benchmark PICABench-Superficial (900 samples across 8 sub-dimensions: light propagation LP, light source effects LSE, reflection, refraction, deformation, causality, global state transition GST, and local state transition LST) and the general editing benchmark ImgEdit-Bench (734 samples). Accuracy (Acc) is scored by GPT-5.1, and Consistency (Con) measures PSNR over unedited background regions.

Model Backbone PICABench Overall Acc (%) ↑ PICABench Overall Con (dB) ↑ ImgEdit-Bench Overall ↑
Instruct-Pix2Pix SD 37.09 20.15 1.88
MagicBrush SD 37.91 20.91 1.90
AnyEdit SDXL 40.49 22.01 2.45
Step1X-Edit Custom 44.59 25.11 3.06
BAGEL MLLM 43.87 26.30 3.20
OmniGen2 OmniGen2 45.36 24.12 3.44
FLUX.1 Kontext FLUX.1 48.08 24.57 3.52
ChronoEdit Video DiT 58.83 27.91 4.16
Qwen-Image-Edit Qwen 57.21 27.50 4.16
PhysEdit (Ours) OmniGen2 56.72 27.23 3.75
PhysEditQwen (Ours) Qwen 60.70 28.35 4.29

Across complex physical interaction sub-dimensions, OmniGen2-based PhysEdit achieves 62.50% in Refraction (a 35.7% relative gain over Uniworld-V1 at 46.05%), 64.81% in LSE, and 63.33% in Reflection. When scaling to the Qwen backbone, PhysEditQwen reaches state-of-the-art performance with 60.70% Overall Accuracy and 28.35 dB Consistency on PICABench, surpassing the base Qwen-Image-Edit model by +4.44% on Deformation and +9.71% on Refraction.

Ablation Study

Ablations on PICABench-Superficial evaluate the individual impact of the Causal DiT (cDiT), Consistency Loss (CL), and Physics-Aware Reinforcement Learning (PARL), as well as trajectory length \(N\) and data curation strategies.

Config Core Modification Overall Acc (%) ↑ Overall Con (dB) ↑ LSE Acc (%) Refraction Acc (%) Causality Acc (%)
Base OmniGen2 single-step baseline 45.36 24.12 48.52 42.22 40.10
Base + cDiT Multi-step causal sequence generation 52.23 26.45 63.96 57.52 45.12
Base + cDiT + CL Add single-step teacher consistency loss 53.40 26.76 64.33 62.00 46.12
Base + cDiT + CL + PARL Full PhysEdit with RL fine-tuning 56.72 27.23 64.81 62.50 49.54
Trajectory length \(N=1\) Static start/end pairs 48.94 25.67 56.36 52.50 42.65
Trajectory length \(N=2\) 2-step intermediate rollout 50.55 25.94 59.67 57.36 44.25
Trajectory length \(N=4\) 4-step rollout (default) 52.23 26.45 63.96 57.52 45.12
Trajectory length \(N=6\) 6-step rollout 52.41 26.58 63.60 57.67 45.73
Base + SFT (static pairs) Fine-tuned on start/end pairs of 50K 48.94 25.67 56.36 52.50 42.65

Key Findings

  • Causal trajectory modeling drives the primary physical capability leap: Transitioning from a discrete single-step base model to cDiT yields the largest accuracy jump (+6.87%, from 45.36% to 52.23%, with LSE surging +15.44%), indicating that continuous dynamics cannot be internalized via discrete static mapping.
  • Consistency loss resolves the video prior drift dilemma: Incorporating CL enhances accuracy by +1.17% and improves consistency by 0.31 dB (Refraction surges by +4.48%), proving that grounding predictions against a frozen single-step teacher effectively suppresses video viewpoint shifts.
  • Reinforcement learning refines causality and complex interactions: PARL brings an additional +3.32% in Overall Accuracy and boosts the strict Causality category from 46.12% to 49.54%, demonstrating that fine-grained constraint verification provides a more informative steering signal than holistic reward metrics.
  • Trajectory length saturation: Accuracy improves sharply from \(N=1\) to \(N=4\) (+3.29% Acc), but plateaus between \(N=4\) and \(N=6\) (+0.18% Acc), confirming that a 4-frame trajectory sufficiently captures essential physical state transitions without incurring unnecessary compute.

Highlights & Insights

  • From discrete mapping to continuous causal rollout: While prior works frame image editing as instantaneous translation, PhysEdit introduces autoregressive temporal factorization over continuous diffusion latents, respecting the thermodynamic and optical arrow of time.
  • Complementary synergy between video dynamics and image teachers: Rather than choosing between dynamic video models (which suffer from camera drift) and static image editors (which lack physical understanding), PhysEdit fuses both: extracting temporal physical priors from video models while anchoring spatial background geometry with a single-step image teacher.
  • Rule-decomposed reward formulation for physical alignment: Instead of relying on noisy holistic aesthetic scores, the framework parses editing instructions into verifiable physical sub-rules and optimizes satisfaction rates via GRPO, providing a stable paradigm for physical reasoning alignment.

Limitations & Future Work

  • Limitations acknowledged by authors: Multi-step autoregressive diffusion increases training and inference latency compared to single-step models; generation fidelity remains partially bounded by the dynamic synthesis capacity of the upstream video generative prior.
  • Potential improvements: The trajectory length \(N=4\) is currently fixed globally; adaptive length allocation based on task complexity (e.g., shorter trajectories for simple color swaps, deeper rollouts for complex fluid or structural interactions) could optimize inference cost.
  • Broader applications: The causal diffusion framework can be extended to 3D scene editing, physical world simulators, and embodied AI agents requiring physically grounded forward dynamics prediction.
  • vs InstructPix2Pix / OmniGen2 / FLUX.1 Kontext: Conventional instruction models rely on static paired data and bidirectional single-step generation, failing when edits require gravity-induced collapse or optical reflection adjustments; PhysEdit resolves this through causal temporal factorization.
  • vs PICABench / UniReal: While these works leverage video models to generate physics training pairs, they collapse them into static start-and-end pairs; PhysEdit demonstrates that static supervision achieves only 48.94% accuracy, whereas continuous trajectory modeling unlocks 56.72%.
  • vs ChronoEdit / F2F: These methods cast image editing directly as full video generation, incurring severe computational costs and accumulating temporal blurring; PhysEdit combines temporal causality with spatial flow matching and consistency distillation, maintaining sharp image fidelity.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates physically consistent editing through causal diffusion architectures, dense video trajectory distillation, and fine-grained rule-based RL.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous evaluation across 8 dimensions of PICABench, 9 tasks of ImgEdit-Bench, and RISEBench, complemented by ablations on components, trajectory length, and data formats.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clean, rigorous exposition with well-defined mathematical formulations and thorough analysis of physical failure modes.
  • Value: ⭐⭐⭐⭐⭐ Establishes a practical, physically grounded trajectory generation paradigm bridging generative vision and physical world modeling.