h-Flow: Flexible Flow-based Image Editing via Doob's h-Transform¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Image Generation
Keywords: Flow Matching, Image Editing, Doob's h-Transform, Orthogonal Velocity Decomposition, FLUX
TL;DR¶
Addressing the entanglement between source image fidelity and target text alignment in flow-based image editing, this paper bridges deterministic Rectified Flow to Doob's h-Transform via an equivalent SDE, derives closed-form reconstruction guidance and semantic editing signals, and achieves completely decoupled, training-free control via orthogonal velocity projection.
Background & Motivation¶
Large-scale text-to-image Rectified Flow (RF) models such as FLUX.1 have emerged as a dominant generative paradigm, yet text-based image editing under this framework remains fundamentally challenging due to the need to balance two competing objectives: faithful alignment with the target text prompt (target alignment) and strict preservation of unedited source regions and structure (source consistency). Classical inversion-based pipelines attempt to map the source image to a noise latent via a forward ODE and re-integrate backward under the target prompt; however, discretization error accumulation throughout forward inversion and subsequent synthesis frequently induces structural drift, forcing prior methods to resort to architecture-specific heuristic injections such as attention sharing. Conversely, recent inversion-free dual-trajectory methods attempt to bypass inversion by coupling two trajectories directly, but they exhibit severe sensitivity to random noise seeds and integration steps, lacking a unified theoretical probabilistic grounding.
The fundamental tension stems from the trajectory dynamics in velocity space: the velocity direction required to reconstruct the source image structure and the velocity direction required to steer generation toward the target text prompt are generally non-parallel and often act in direct conflict. Conventional formulations superimpose or linearly scale these vectors directly within the same high-dimensional space, causing any amplification of editing strength to severely disrupt background fidelity, while strong structural regularization in turn suffocates prompt compliance. Crucially, while Doob's h-Transform provides a classical potential-theoretic tool to rigorously condition Markov processes on desired terminal states, it requires a non-degenerate diffusion coefficient in an SDE. In deterministic Rectified Flow models, the diffusion term is identically zero (\(g \equiv 0\)), causing continuous-time h-transform guidance terms to vanish identically and rendering the classic framework inapplicable.
This work resolves this theoretical impasse by bridging deterministic probability flows with diffusion bridge theory: recognizing that infinitely many SDEs share identical marginal distributions with a given ODE, the authors construct an equivalent SDE to activate Doob's h-Transform, reformulating editing as conditional generation under dual terminal constraints. Core idea: activate Doob's h-Transform within Rectified Flow via a marginal-preserving equivalent SDE bridge, derive closed-form geometric reconstruction velocities and pure semantic editing directions, and eliminate structural interference via orthogonal velocity projection for training-free, decoupled editing.
Method¶
Overall Architecture¶
h-Flow formulates text-guided image editing as sampling from a conditional posterior distribution \(p(\cdot \mid x_0^{src}, c^{edit})\) governed by two simultaneous terminal events. To circumvent the vanishing guidance issue caused by the zero diffusion coefficient (\(g(t) \equiv 0\)) in deterministic RF, the framework first constructs an equivalent Itô SDE sharing identical marginal distributions with the base model, yielding a tractable bridge probability-flow ODE (PF-ODE). Next, treating source structural fidelity and target semantic alignment as conditionally independent terminal events given the intermediate state \(x_t\), the joint harmonic function factorizes into a product of experts: a source preservation expert \(h_{rec}\) and a target editing expert \(h_{edit}\). Finally, the raw semantic editing velocity is projected onto the orthogonal complement of the reconstruction velocity field, completely decoupling semantic modification from structural anchoring before backward integration.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Source Latent & Target Text"] --> B["Equivalent SDE Bridge Construction<br/>Match marginals to activate continuous h-Transform"]
B --> C["Dual-Objective h-Expert Functions<br/>Derive closed-form reconstruction & editing directions"]
C --> D["Orthogonal Velocity Decomposition<br/>Project out parallel components to eliminate interference"]
D --> E["Interference-Free Reverse Integration"]
Key Designs¶
1. Equivalent SDE Bridge Construction: Activating h-Transforms in Deterministic Flows
Classical Doob's h-Transform requires the diffusion coefficient to be strictly positive (\(g(t) > 0\)) to ensure a non-degenerate transition kernel; in contrast, Rectified Flow operates via a deterministic ODE \(\frac{dx_t}{dt} = v_\theta(x_t, t)\) where \(g(t) \equiv 0\), causing the drift correction \(\frac{g^2(t)}{2} \nabla_{x_t} \log h\) to vanish identically. To resolve this structural incompatibility, the paper leverages the fact that infinitely many SDEs share identical marginals \(\{p_t\}\) with the deterministic flow. An equivalent Itô SDE is constructed with diffusion coefficient \(g_{eq}(t) = \sqrt{\frac{2t}{1-t}}\), satisfying \(g_{eq}(0)=0\) (no noise at the clean data endpoint) and \(g_{eq}(t)>0\) for all \(t \in (0, 1]\). Substituting \(g_{eq}\) into the reverse-time bridge PF-ODE yields an exact coefficient \(\frac{g_{eq}^2(t)}{2} = \frac{t}{1-t}\), establishing the conditioned bridge flow ODE: $\(\frac{dx_t}{dt} = v_\theta(x_t, t) - \frac{t}{1-t} \nabla_{x_t} \log h(x_t, t)\)$ This formulation provides a mathematically rigorous foundation for imposing future terminal constraints onto deterministic flow matching trajectories without modifying the pretrained unconditional prior.
2. Dual-Objective h-Expert Functions: Closed-Form Reconstruction and Pure Semantic Direction
Given the continuous bridge ODE, the composite terminal event is factorized under conditional independence as \(h(x_t, t) = h_{rec}(x_t, t) \cdot h_{edit}(x_t, t)\), translating in log-gradient space into an additive superposition of forces: \(\nabla \log h = \nabla \log h_{rec} + \nabla \log h_{edit}\). For source fidelity \(h_{rec}\), the terminal event is set to a Dirac delta constraint \(\delta(x_0 - x_0^{src})\). Under the isotropic Gaussian transition kernel of RF, expanding via Bayes' rule and the Tweedie posterior mean \(\hat{x}_0 = x_t - t v_\theta^{src}\) causes all time-dependent scale factors to cancel out, yielding the closed-form reconstruction velocity \(v_{rec} = (x_t - x_0^{src}) / t\), which points straight toward \(x_0^{src}\). To prevent model approximation error in \(v_\theta^{src}\) from corrupting this error-free anchor, an additive anchor velocity is defined: $\(\tilde{v}_{rec} = v_\theta(x_t, t, c^{src}) + \lambda_{rec} v_{rec}\)$ For target semantic alignment \(h_{edit}\), the terminal condition corresponds to the likelihood \(p(c^{edit} \mid x_0)\). By substituting the source-conditioned score \(\nabla \log p(x_t \mid c^{src})\) for the unconditional baseline, shared concepts between source and target prompts cancel out, resulting in a clean semantic delta: $\(\Delta_{edit} = v_\theta(x_t, t, c^{edit}) - v_\theta(x_t, t, c^{src})\)$ This velocity-level semantic difference avoids the background destruction common in standard classifier-free guidance.
3. Orthogonal Velocity Decomposition: Decoupling Structural Fidelity and Editing Control
Although the additive superposition of log-gradients is mathematically valid, in high-dimensional velocity space the raw editing direction \(\Delta_{edit}\) typically retains a non-zero projection along the reconstruction trajectory \(\tilde{v}_{rec}\). This parallel component either redundantly reinforces or destructively counteracts structural fidelity, creating the primary source of tension between fidelity and editability. h-Flow introduces an orthogonal projection operator that strips away the parallel component and preserves only the orthogonal complement: $\(\Delta_{edit}^\perp = \Delta_{edit} - \frac{\langle \Delta_{edit}, \tilde{v}_{rec} \rangle}{\|\tilde{v}_{rec}\|^2} \tilde{v}_{rec}\)$ Because \(\langle \Delta_{edit}^\perp, \tilde{v}_{rec} \rangle = 0\) by construction, the reverse editing ODE integrates along mutually independent subspaces: $\(v = \tilde{v}_{rec} + \lambda_{edit} \Delta_{edit}^\perp\)$ This guarantees that scaling \(\lambda_{edit}\) enhances semantic transformation without distorting source structures, while adjusting \(\lambda_{rec}\) anchors background fidelity without suppressing target modifications. The computation incurs negligible overhead, requiring only a single vector inner product and subtraction per step.
Loss & Training¶
h-Flow is a training-free and gradient-free test-time sampling algorithm. During inference, each reverse step requires only two forward evaluations of the velocity network \(v_\theta\) (under \(c^{src}\) and \(c^{edit}\) respectively). The implementation is evaluated on the FLUX.1-dev backbone with \(N=28\) Euler integration steps on \([0, 1]\), using default hyperparameter weights \(\lambda_{rec} = 1.2\) and \(\lambda_{edit} = 1.5\). The backward editing ODE is completely decoupled from forward initialization, allowing plug-and-play pairing with any forward process, such as RF-Inversion, standard ODE inversion, or one-step noise injection (SDEdit).
Key Experimental Results¶
Main Results¶
On both the single-aspect PIE-Bench benchmark (700 images across human, animal, and object categories) and the multi-aspect PIE-Bench++ benchmark, h-Flow was evaluated against leading inversion-based and inversion-free methods using the identical FLUX.1-dev backbone.
| Dataset | Method | Average Rank↓ | PSNR↑ | LPIPS↓ | SSIM↑ | MSE↓ | Structure Distance↓ | CLIP-Entire↑ | CLIP-Edited↑ | HPSv2↑ | Aesthetic Score (AS)↑ |
|---|---|---|---|---|---|---|---|---|---|---|---|
| PIE-Bench | SDEdit | 6.06 | 18.01 | 0.240 | 0.650 | 0.021 | 0.063 | 24.67 | 22.54 | 27.08 | 5.05 |
| FlowEdit | 4.39 | 21.92 | 0.111 | 0.833 | 0.009 | 0.028 | 25.19 | 22.54 | 27.70 | 4.87 | |
| ODEInv | 5.44 | 13.44 | 0.349 | 0.620 | 0.071 | 0.130 | 26.56 | 23.83 | 28.73 | 4.92 | |
| FireFlow | 5.89 | 18.15 | 0.207 | 0.738 | 0.026 | 0.057 | 25.69 | 22.56 | 26.45 | 4.63 | |
| RF-Inversion | 5.00 | 20.59 | 0.186 | 0.706 | 0.013 | 0.042 | 25.26 | 22.34 | 28.21 | 4.98 | |
| UniEdit-Flow | 3.00 | 29.45 | 0.060 | 0.900 | 0.002 | 0.011 | 25.80 | 22.34 | 26.77 | 4.96 | |
| RF-Solver | 3.22 | 22.90 | 0.140 | 0.819 | 0.007 | 0.031 | 26.09 | 22.68 | 27.24 | 5.48 | |
| h-Flow (Ours) | 2.33 | 23.88 | 0.110 | 0.840 | 0.006 | 0.026 | 25.80 | 22.80 | 27.79 | 4.98 | |
| PIE-Bench++ | SDEdit | 7.11 | 19.92 | 0.230 | 0.680 | 0.014 | 0.090 | 22.50 | 21.80 | 27.08 | 4.00 |
| FlowEdit | 4.11 | 22.27 | 0.109 | 0.841 | 0.009 | 0.031 | 24.94 | 24.24 | 27.70 | 4.86 | |
| ODEInv | 5.44 | 17.21 | 0.240 | 0.720 | 0.028 | 0.080 | 25.50 | 24.92 | 28.20 | 4.73 | |
| FireFlow | 5.67 | 18.95 | 0.188 | 0.760 | 0.021 | 0.060 | 25.71 | 24.77 | 26.45 | 4.56 | |
| RF-Inversion | 4.44 | 21.13 | 0.180 | 0.730 | 0.011 | 0.040 | 25.23 | 24.49 | 27.93 | 5.18 | |
| UniEdit-Flow | 3.56 | 29.46 | 0.060 | 0.900 | 0.018 | 0.018 | 25.37 | 24.39 | 26.96 | 4.89 | |
| RF-Solver | 2.89 | 23.69 | 0.129 | 0.832 | 0.006 | 0.033 | 25.79 | 24.89 | 27.24 | 5.43 | |
| h-Flow (Ours) | 2.78 | 24.30 | 0.109 | 0.850 | 0.005 | 0.029 | 25.48 | 24.54 | 27.67 | 4.92 |
Ablation Study¶
1. Effect of Orthogonal Projection (PIE-Bench)
| Config | SSIM↑ | Structure Distance↓ | CLIP-Entire↑ | CLIP-Edited↑ | Note |
|---|---|---|---|---|---|
| w/o projection | 0.834 | 0.026 | 25.37 | 22.79 | Raw \(\Delta_{edit}\) leaks components along \(\tilde{v}_{rec}\), degrading text alignment |
| w/ projection (Full model) | 0.840 | 0.026 | 25.80 | 22.80 | Completely eliminates mutual interference, boosting CLIP-Entire (+0.43) and SSIM |
2. Parameter Decoupling Analysis (PIE-Bench)
| Parameter Sweep | Value | SSIM↑ | Structure Distance↓ | CLIP-Entire↑ | CLIP-Edited↑ | Observation |
|---|---|---|---|---|---|---|
| Reconstruction Strength \(\lambda_{rec}\) (fixed \(\lambda_{edit}=1.5\)) |
0.0 | 0.830 | 0.026 | 25.97 | 22.92 | Low structural anchoring |
| 1.2 (Default) | 0.840 | 0.026 | 25.80 | 22.80 | Optimal trade-off; editing metrics barely change (-0.12) | |
| 1.5 | 0.856 | 0.025 | 25.65 | 22.73 | Higher fidelity with near-zero suppression of editability | |
| Editing Strength \(\lambda_{edit}\) (fixed \(\lambda_{rec}=1.2\)) |
0.0 | 0.846 | 0.020 | 22.84 | 20.49 | Pure reconstruction, minimal text modification |
| 1.5 (Default) | 0.840 | 0.026 | 25.80 | 22.80 | Major leap in text alignment (+2.31 CLIP-Edited) with strong fidelity | |
| 3.0 | 0.804 | 0.034 | 26.26 | 23.49 | Aggressive editing begins encroaching on background structure |
3. Plug-and-Play Compatibility with Diverse Forward Processes
| Base Forward Method | Integration | SSIM↑ | Structure Distance↓ | CLIP-Entire↑ | CLIP-Edited↑ | Impact of h-Flow |
|---|---|---|---|---|---|---|
| RF-Inversion | Baseline | 0.70 | 0.070 | 25.26 | 22.34 | Balanced baseline, limited semantic editing capability |
| + h-Flow (Ours) | 0.73 | 0.082 | 26.10 | 23.60 | Substantial gain in CLIP-Edited (+1.26) while maintaining fidelity | |
| SDEdit | Baseline | 0.65 | 0.063 | 24.67 | 22.54 | Noise perturbation limits structural consistency |
| + h-Flow (Ours) | 0.65 | 0.063 | 25.28 | 23.10 | Boosts text alignment (+0.56) with zero structural degradation | |
| ODE-Inv | Baseline | 0.62 | 0.130 | 26.56 | 23.83 | High editability accompanied by severe structural collapse |
| + h-Flow (Ours) | 0.70 | 0.077 | 26.44 | 23.65 | Acts as structural stabilizer: SSIM +0.08, Distance drops -0.053 |
Key Findings¶
- Top Overall Pareto Balance: UniEdit-Flow achieves extreme fidelity (PSNR 29.45) at the cost of poor semantic flexibility (HPS 26.77, unable to perform non-rigid changes like "kissing parrots"); conversely, ODEInv suffers severe structural collapse (Distance 0.130). h-Flow achieves the best average rank across all 9 metrics on both PIE-Bench (2.33) and PIE-Bench++ (2.78).
- Verified Control Decoupling: Sweeping \(\lambda_{rec}\) across \([0.0, 1.5]\) causes only a negligible 0.19-point fluctuation in CLIP-Edited while steadily lifting SSIM. Conversely, scaling \(\lambda_{edit}\) drives a 3.0-point swing in editing alignment, proving that orthogonal decomposition isolates control into mutually independent subspaces.
- Universal Structural Stabilization: When paired with naive ODE inversion (ODE-Inv), which otherwise suffers catastrophic drift, h-Flow's closed-form reconstruction term acts as a trajectory stabilizer, dramatically reducing structural distance from 0.130 to 0.077.
Highlights & Insights¶
- Probabilistic Bridge via Equivalent SDE: Elegant bridging of deterministic Rectified Flow with stochastic potential theory by constructing an equivalent Itô SDE with identical marginals, unlocking Doob's h-Transform for continuous flow matching.
- Subspace Velocity Projection: Borrowing insights from projected gradient descent, the orthogonal velocity projection eliminates the parallel interference component between editing and reconstruction, establishing true control decoupling.
- Zero-Overhead Test-Time Utility: Purely training-free and gradient-free, requiring merely two forward passes of \(v_\theta\) and one vector inner product per step, while offering seamless compatibility across diverse inversion frameworks.
Limitations & Future Work¶
- Author-Acknowledged Limitations: The framework currently focuses on 2D image editing and has not yet been extended to temporal video flows or high-dimensional 3D representations.
- Observed Practical Limitations: Orthogonal projection operates at the global velocity field level; for delicate localized edits requiring precise object boundaries without explicit masks, subtle global velocity leakage can still occur.
- Future Directions: Incorporating spatial attention gating or frequency-domain decoupling into local orthogonal projections, and generalizing the bridge formulation to multimodal text-to-video architectures.
Related Work & Insights¶
- vs UniEdit-Flow / RF-Solver: While prior works rely on complex numerical solver tailoring or empirical cross-attention hijacking, h-Flow offers an architecture-agnostic probabilistic derivation rooted in stochastic bridge theory.
- vs FlowEdit / FlowAlign: Rather than coupling dual noise-to-image ODE trajectories that are sensitive to random seeds and step schedules, h-Flow anchors the backward process with closed-form reconstruction guidance and orthogonal velocity filtering.
- vs h-Edit: Unlike h-Edit, which operates exclusively in diffusion SDEs and approximates source preservation using a text proxy, h-Flow extends to Rectified Flow and provides exact closed-form guidance directly from the source image.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Elegant mathematical bridge connecting Doob's h-Transform to deterministic flow matching.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive validation on PIE-Bench and PIE-Bench++, featuring parameter scans, projection ablations, and cross-pipeline compatibility tests.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical derivation and lucid structural clarity throughout.
- Value: ⭐⭐⭐⭐⭐ Provides a foundational, training-free paradigm for decoupled controllable generation in flow models.