Skip to content

V-HOLD: Stabilizing Flow Trajectories to Rethink the Edit–Preservation Trade-off

Conference: ECCV 2026
Paper: ECCV Official
Full Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-5540.txt
Area: Image Generation
Keywords: text-guided image editing, flow matching, trajectory stability, inversion-free image editing, source preservation

TL;DR

Addressing the edit–preservation trade-off in generative flow-based editing, this paper identifies integration trajectory instability as a primary driver of semantic degradation, and proposes V-HOLD, a training-free framework that stabilizes editing dynamics via anchor-step candidate selection and cross-step directional holding, improving semantic preservation while reducing model evaluation costs.

Background & Motivation

Recent breakthroughs in diffusion models and rectified flow generative architectures have significantly advanced text-guided image editing, allowing users to modify image semantics while preserving the structural context of the original image. Representative inversion-based and inversion-free frameworks, such as SDEdit, RFEdit, and FlowEdit, exploit powerful pretrained generative priors to perform diverse structural, stylistic, and semantic transformations. However, a pervasive challenge remains: stronger edits designed to strictly align with target text prompts often severely degrade subject identity, fine-grained details, and non-target background structures. This dilemma, widely recognized as the edit–preservation trade-off, has long been regarded as an inherent limitation of generative editing systems.

Prior work has almost exclusively evaluated editing performance through final text–image alignment metrics (e.g., CLIP similarity) and reconstruction fidelity, largely neglecting intermediate trajectory dynamics during numerical integration. In flow- and diffusion-based sampling, the transport trajectory between source and target latent states fundamentally dictates how source characteristics and target conditioning are merged. In inversion-free flow editing, models estimate instantaneous editing vectors at each step by resampling random Gaussian perturbations. However, this independent per-step estimation introduces substantial stochastic estimation noise, causing the latent update directions to exhibit acute step-to-step directional deviations and magnitude volatility.

Empirical investigation reveals a critical insight: this trajectory instability correlates strongly with degraded source semantic preservation, yet displays little association with improvements in target prompt alignment. In other words, much of the background distortion and semantic degradation observed in prior methods does not stem from intrinsic prompt-image semantic incompatibility, but rather from numerical turbulence in integration dynamics. The core idea is to stabilize the latent editing path by selecting a target-oriented velocity candidate aligned with a reference direction at anchor steps and holding it across multiple integration steps (V-HOLD), eliminating stochastic noise accumulation and simultaneously reducing velocity-field evaluations (NFE) without training.

Method

Overall Architecture

In flow-based image editing, latent states are transported through a learned velocity field to transform a source image into a target image reflecting an edit prompt. In inversion-free flow editing (such as FlowEdit), the instantaneous update direction is determined by the difference between target-conditioned and source-conditioned velocity fields. However, estimating this velocity difference at every step with newly sampled noise perturbations introduces high-frequency random noise, leading to erratic trajectories that disrupt preserved regions.

V-HOLD (Velocity Hold) partitions the overall integration timeline into uniform intervals of length \(H\). The model evaluates candidate update directions only at designated anchor steps, compares them with a running reference direction via cosine similarity to select the most coherent candidate, and then holds this selected velocity constant across the subsequent \(H\) steps. This design achieves smooth trajectory integration, bounds update volatility, and skips redundant neural network forward evaluations. The pipeline comprises shared-perturbation latent formulation, anchor candidate sampling and reference alignment, and cross-step directional holding with latent transport.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input source image & edit prompts<br/>Construct shared-perturbation source/target latents"] --> B["Anchor Candidate Selection & Reference Alignment<br/>Match K candidates against reference direction r via cosine similarity"]
    B --> C["Directional Holding & Latent State Integration<br/>Fix selected velocity and integrate latent states over H steps without NFE"]
    C -->|Within current anchor interval a+j < a+H| C
    C -->|Advance to next anchor step a+H| B
    C -->|Reach final integration step N| D["Output high-fidelity edited image aligned with target prompt"]

Key Designs

1. Anchor Candidate Selection & Reference Alignment: Filtering stochastic estimation noise

In inversion-free rectified flow editing, the instantaneous velocity difference between target and source conditions is formulated as: $\(\Delta v_\theta(z_t, t; \epsilon) = v_\theta(z_t^{\text{tgt}}, t; c_{\text{tgt}}) - v_\theta(z_t^{\text{src}}, t; c_{\text{src}})\)$ where \(z_t^{\text{src}} = (1 - \sigma_t) x^{\text{src}} + \sigma_t \epsilon\), and \(z_t^{\text{tgt}} = z_t + z_t^{\text{src}} - x^{\text{src}}\), sharing Gaussian noise \(\epsilon \sim \mathcal{N}(0, I)\). Because a single stochastic sample has non-negligible variance, per-step drift estimation introduces an estimation error \(\xi_k\) that causes trajectory zig-zagging.

V-HOLD restricts direction re-estimation to anchor steps \(a \in \{0, H, 2H, \dots\}\). At each anchor step \(a\), the algorithm samples \(K\) independent perturbations \(\{\epsilon^{(i)}\}_{i=1}^K\) to produce \(K\) velocity candidates: $\(\Delta v_a^{(i)} = \Delta v_\theta(z_{t_a}, t_a; \epsilon^{(i)}), \quad i = 1, \dots, K\)$ The system maintains a global reference editing direction \(r\), initialized at step \(0\) by the first candidate \(r = \Delta v_0^{(1)}\). At each subsequent anchor step, the cosine similarity between each candidate and \(r\) is evaluated, selecting the candidate with the highest alignment: $\(i^\star = \arg\max_{i} \frac{\langle \Delta v_a^{(i)}, r \rangle}{\|\Delta v_a^{(i)}\| \, \|r\|}\)$ The selected velocity vector is assigned as the active editing direction \(d_a = \Delta v_a^{(i^\star)}\), after which the reference vector is updated via \(r \leftarrow d_a\). This mechanism utilizes macroscopic semantic continuity to eliminate erratic lateral components, ensuring the chosen direction steadfastly tracks the target prompt.

2. Directional Holding & Latent State Integration: Suppressing trajectory volatility with minimal compute

Once an optimal velocity vector \(d_a\) is identified at anchor step \(a\), V-HOLD bypasses model evaluation for the subsequent \(H-1\) steps. Across the holding window \(j \in \{0, 1, \dots, H-1\}\), the update direction remains strictly constant: $\(d_{a+j} = d_a\)$ The latent state transitions according to the noise schedule \(\{\sigma_k\}\) along this locked vector: $\(z_{t_{a+j+1}} = z_{t_{a+j}} + (\sigma_{a+j+1} - \sigma_{a+j}) \, d_a\)$ Analyzing error propagation shows that when recomputing directions at every step, the total accumulated path incorporates the sum of independent estimation errors \(\sum_k \xi_k\), scattering updates across background pixels and causing hallucinations. By reducing the direction estimation frequency to \(1/H\), V-HOLD suppresses the injection of stochastic noise, naturally confining velocity energy to the intended editing mask. Furthermore, freezing the velocity field for \(H-1\) intermediate steps eliminates backbone forward passes, yielding substantial reductions in number of function evaluations (NFE), FLOPs, and wall-clock latency.

3. Trajectory Stability & Volatility Metrics: Formulating the dynamical mechanics of editing

To quantitatively dissect flow trajectories, the paper introduces rigorous mathematical formulations for trajectory stability. Given the realized sequence of update vectors \(\{d_k\}_{k=0}^{N-1}\), the step-wise directional consistency is quantified via the cosine similarity between consecutive steps: $\(\text{cos}_{\text{traj}}(k) = \frac{\langle d_k, d_{k-1} \rangle}{\|d_k\| \, \|d_{k-1}\|}\)$ The instantaneous directional instability is defined as its complement, \(\text{Instability}(k) = 1 - \text{cos}_{\text{traj}}(k)\). In parallel, the relative change in update magnitude is captured by \(\Delta_{\text{mag}}(k) = \frac{|\|d_k\| - \|d_{k-1}\||}{\|d_{k-1}\| + \delta}\) with \(\delta = 10^{-8}\), yielding the overall magnitude volatility across the trajectory: $\(\text{Volatility} = \frac{1}{N-1} \sum_{k=1}^{N-1} \Delta_{\text{mag}}(k)\)$ Empirical measurements show that while standard inversion-free baselines exhibit instability scores exceeding 0.4–0.6, V-HOLD reduces instability by 70%–85% and substantially dampens magnitude volatility. This confirms that semantic preservation degrades primarily under high directional turbulence, and stabilizing the trajectory successfully decouples preservation from editing strength.

Key Experimental Results

Main Results

Evaluations were conducted on two primary benchmarks: PIE-Bench (700 real-image editing pairs) and FlowEdit-Data (300 pairs of \(1024 \times 1024\) high-resolution images). Experiments employ FLUX.1-dev and SD3-medium as representative rectified flow backbones. Baselines comprise inversion-based approaches (RF-Inversion, RFEdit, FireFlow, DDIB, iRFDS, SDEdit) and inversion-free methods (FlowEdit, FlowAlign). Evaluation metrics include perceptual preservation LPIPS (lower is better), semantic drift DINO distance (lower is better), and text-image alignment CLIP score (higher is better).

Backbone Method PIE-Bench LPIPS ↓ PIE-Bench DINO ↓ PIE-Bench CLIP ↑ FlowEdit-Data LPIPS ↓ FlowEdit-Data DINO ↓ FlowEdit-Data CLIP ↑
FLUX.1-dev RF-Inversion 0.4338 0.8585 0.2536 0.5004 0.9451 0.2931
FLUX.1-dev RFEdit 0.3776 0.8617 0.2531 0.4012 0.8911 0.2835
FLUX.1-dev FireFlow 0.4478 0.9778 0.2698 0.4545 0.9845 0.3023
FLUX.1-dev FlowEdit 0.2712 0.8004 0.2583 0.2805 0.8574 0.2898
FLUX.1-dev V-HOLD (Ours) 0.2874 0.7862 0.2623 0.3201 0.8373 0.2928
SD3-medium DDIB 0.6751 1.1331 0.2726 0.6852 1.1262 0.3013
SD3-medium SDEdit 0.4415 0.9392 0.2731 0.4895 0.9724 0.3009
SD3-medium iRFDS 0.3998 0.8872 0.2652 0.7680 1.1317 0.3031
SD3-medium FlowAlign 0.3521 0.8640 0.2482 0.3278 0.9389 0.2910
SD3-medium FlowEdit 0.3201 0.8873 0.2781 0.2243 0.8831 0.2985
SD3-medium V-HOLD (Ours) 0.2957 0.8461 0.2723 0.2378 0.7547 0.2975

Trajectory Stability and Computational Efficiency

Trajectory stability metrics (median and interquartile range Q1–Q3 over 1,000 samples) and computational efficiency metrics (NFE, FLOPs, and per-image editing runtime) demonstrate the quantitative advantages of V-HOLD.

Backbone Method Inversion Instability ↓ Update Volatility ↓ NFE ↓ FLOPs (hook) ↓ Edit time (s) ↓
SD3-medium FlowEdit X 0.636 (0.598–0.678) 0.293 (0.265–0.311) 33 2.52e+14 3.420
SD3-medium FlowAlign X 0.546 (0.511–0.587) 0.215 (0.190–0.238) 33 2.87e+14 8.404
SD3-medium V-HOLD X 0.090 (0.064–0.113) 0.123 (0.098–0.167) 27 2.35e+14 2.177
SD3-medium Relative vs. FlowEdit — -85.8% -58.0% -18.2% -6.7% -36.3%
FLUX.1-dev FlowEdit X 0.421 (0.380–0.466) 0.315 (0.294–0.344) 24 9.53e+14 8.875
FLUX.1-dev V-HOLD X 0.108 (0.094–0.124) 0.223 (0.202–0.248) 21 2.47e+14 6.943
FLUX.1-dev Relative vs. FlowEdit — -74.3% -29.3% -12.5% -74.0% -21.8%

Ablation Study

The ablation investigates the individual impact of candidate sampling (\(K\)) and holding length (\(H\)). Default full setting: \(K=3, H=4\); w/o candidate denotes \(K=0\) (averaged over \(H \in \{2,4,6\}\)); w/o hold denotes \(H=0\) (recomputing every step, averaged over \(K \in \{1,3,5\}\)).

Backbone Config PIE-Bench LPIPS ↓ PIE-Bench CLIP ↑ PIE-Bench DINO ↓ FlowEdit-Data LPIPS ↓ FlowEdit-Data CLIP ↑ FlowEdit-Data DINO ↓
FLUX.1-dev w/o candidate 0.348 0.257 0.992 0.333 0.291 0.856
FLUX.1-dev w/o hold 0.340 0.254 0.995 0.324 0.282 0.866
FLUX.1-dev V-HOLD (Full: K=3, H=4) 0.287 0.262 0.786 0.320 0.293 0.837
SD3-medium w/o candidate 0.326 0.275 1.049 0.205 0.287 0.789
SD3-medium w/o hold 0.300 0.276 1.047 0.201 0.289 0.797
SD3-medium V-HOLD (Full: K=3, H=4) 0.296 0.272 0.846 0.238 0.297 0.755

Key Findings

  • Consistent Superiority in Semantic Preservation: Across all backbones and benchmarks, V-HOLD achieves the lowest DINO distance (e.g., reaching 0.7547 on FlowEdit-Data with SD3-medium compared to 0.8831 for FlowEdit), validating that preserving source identity and structure is strongly enhanced while keeping target prompt alignment competitive.
  • Drastic Reduction in Trajectory Instability: Compared to FlowEdit, V-HOLD slashes directional instability by 85.8% on SD3-medium and 74.3% on FLUX.1-dev, successfully preventing structural drift from diffusing into background regions.
  • Substantial Computational Savings: By skipping redundant velocity field evaluations during holding steps, V-HOLD reduces FLOPs by up to 74% on FLUX.1-dev and lowers wall-clock editing time from 8.88s to 6.94s, and down to 2.18s on SD3-medium.
  • Mutual Necessity of Candidates and Holding: Disabling either candidate selection or directional holding leads to severe degradation in DINO scores (e.g., rising from 0.786/0.846 to 0.992–1.049 on PIE-Bench). Holding without candidate selection accumulates biased errors, whereas candidate selection without holding fails to suppress high-frequency oscillations. The configuration \(K=3, H=4\) achieves the sweet spot between stability and flexibility.

Highlights & Insights

  • Revisiting the Edit–Preservation Trade-off: Rather than accepting semantic degradation as an unavoidable theoretical cost of strong prompt guidance, the paper demonstrates empirically that intermediate numerical integration instability is the true driver of collateral background damage.
  • Zero-Cost Plug-and-Play Generalization: V-HOLD requires no fine-tuning, no inversion ODE solving, and no spatial attention map caching. It integrates directly into standard flow integration loops via a few lines of vector selection logic.
  • Spontaneous Spatial Update Localization: Heatmap analysis of cumulative update magnitudes proves that stabilizing trajectories naturally focuses update energy onto the target object region, achieving clean, localized edits without explicit spatial segmentation masks.

Limitations & Future Work

  • Static Holding Intervals: The current scheme applies a uniform holding length \(H=4\) and candidate size \(K=3\) across all timesteps. However, generative trajectories typically govern coarse layout in early steps and fine texture in late steps; adaptive step sizing based on flow divergence could yield further improvements.
  • Sensitivity to Initial Reference Direction: The reference vector is initialized using the first sampled perturbation at step \(0\). If an extreme noise vector is drawn initially, subsequent candidate filtering could experience a minor systemic bias. Initializing via multi-sample aggregation could enhance robustness in edge cases.
  • vs FlowEdit: FlowEdit pioneered inversion-free flow editing via instantaneous velocity differences, but its per-step noise resampling causes severe trajectory jitter and higher computational costs; V-HOLD adds candidate selection and directional holding, improving preservation while cutting compute by up to 74%.
  • vs FlowAlign: FlowAlign attempts to regularize trajectories by introducing additional optimization terms during sampling, which substantially inflates inference time (8.40s on SD3); V-HOLD regularizes dynamics passively by holding valid vectors, achieving much lower instability (0.090 vs 0.546) and much faster speeds (2.18s).
  • vs RFEdit / FireFlow / RF-Inversion: Inversion-based rectified flow methods necessitate bidirectional ODE integration and feature caching, which constrain semantic editing flexibility and slow down inference; V-HOLD operates completely inversion-free with minimal overhead.

Rating

  • Novelty: ⭐⭐⭐⭐ [Offers a fresh dynamical perspective on editing degradation, supported by thorough trajectory stability formulations]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Spans multiple backbones, benchmarks, trajectory metric analyses, computational complexity breakdowns, and dual MLLM evaluations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear presentation with rigorous mathematical formulations and compelling empirical analyses]
  • Value: ⭐⭐⭐⭐⭐ [Provides an elegant, zero-training, acceleration-capable solution ready for immediate adoption in production diffusion and flow editing pipelines]