OSVE: One Step Video Editing with One Step Diffusion Models¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/KU-VGI/OSVE
Area: Video Generation
Keywords: Video Editing, One-Step Diffusion Models, Inversion Encoder, Cross-Frame Attention, Temporal Consistency
TL;DR¶
The first One-to-One framework adapting one-step Text-to-Image diffusion models for text-guided video editing via a single-pass structure-aware inversion encoder and unified latent frame concatenation, achieving visual quality on par with multi-step baselines while running 155โ171ร faster at over 15.6 FPS.
Background & Motivation¶
Text-guided video editing with diffusion models holds vast potential across media production and interactive creation, yet mainstream methods are severely bottlenecked by extreme computational costs. Existing solutions predominantly follow a "multi-step inversion and multi-step denoising" paradigm (such as DDIM inversion paired with cross-frame attention control). Processing every video frame across tens or hundreds of sequential sampling steps causes latency to scale multiplicatively with both step count and frame count. For example, RAVEโreported as one of the fastest T2I-based editing baselinesโstill requires approximately 24 hours to edit a 5-minute, 30 FPS video on a high-end RTX 6000 Ada GPU, rendering real-time or streaming video editing practically impossible.
Directly replacing multi-step backbones with distilled one-step Text-to-Image models (such as DMD2) theoretically compresses the generative forward pass, but exposes three fundamental hurdles. First, conventional multi-step DDIM inversion is fundamentally mismatched with the one-step generative manifold, resulting in severe information loss, blurred textures, and collapsed dynamics. Second, existing structure-preserving controls (like Prompt-to-Prompt or ControlNet) fail under single-step constraints due to an "over-steering" phenomenon: in multi-step generation, control signals are gradually absorbed and corrected across many steps, whereas concentrating the entire control force into a single pass deprives the model of any corrective feedback, causing external guidance to overwhelm the output. Third, one-step image models possess no temporal awareness, and frame-wise editing inevitably triggers rampant temporal flickering and inconsistency.
To resolve these tensions, this paper introduces a fundamental paradigm shift: under the one-step regime, structural preservation cannot be enforced via mid-generation intervention, but must instead be pre-encoded into the initial noise latent prior to generation. Core idea: establish a "One-to-One" (one-step inversion followed by one-step generation) video editing framework, using prompt perturbation to construct generator-aligned paired data for training a lightweight inversion encoder with a Structure-Aware Editing (SAE) loss, while spatially concatenating frame latents into a unified tensor to naturally trigger global cross-frame self-attention in a single forward pass.
Method¶
Overall Architecture¶
The OSVE framework consists of two core phases: single-step structure-aware latent inversion and Unified-Frame Editing (UFE) generation. For an input sequence of \(K\) frames, each frame is mapped to latent space via a pre-trained VAE encoder and then processed independently through a learnable inversion encoder to predict its terminal noise latent in a single forward pass. Next, all frame latents are concatenated along the spatial width dimension into a unified wide latent map, which is fed into the frozen one-step diffusion generator. The native spatial self-attention layers inherently perform cross-frame correspondence matching across all frames simultaneously. For extended video sequences, an anchor frame selected via DINO feature medoid is combined with a sliding window to maintain both local continuity and long-term global stability.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Video Frames<br/>V0 = [z0_0, ..., z0_K-1]"] --> B["Structure-Aware One-Step Inversion<br/>Forward encoder predicts noise latent zT"]
B --> C["Unified-Frame Latent Concatenation<br/>Concatenate along width to form Z_UFE"]
C --> D["Anchor-Guided Sliding Window<br/>DINO medoid anchor + window slicing"]
D --> E["One-Step Diffusion & Attention Sharing<br/>One-step generator Gฮธ aligns features"]
E --> F["Output Edited Video<br/>Discard anchor & stitch central frames"]
Key Designs¶
1. Structure-Aware One-Step Inversion: Pre-encoding geometric layout into initial noise
To circumvent the inversion-generation mismatch and eliminate the "over-steering" artifacts caused by mid-generation attention manipulation, the framework trains a dedicated inversion encoder \(I_\phi\). Initialized from the pre-trained generator's U-Net to inherit strong latent space priors, the encoder directly predicts the initial noise latent \(\hat{z}_T = I_\phi(z_0, c)\) from the source frame VAE latent \(z_0\) and source text embedding \(c\). To train this encoder without requiring external editors (which produce out-of-distribution targets that disrupt training gradients), the authors design a "prompt perturbation" data synthesis scheme. By sampling prompts from LAION and JourneyDB, adding Gaussian noise to their text embeddings \(c_{\text{edit}} = c + \mathcal{N}(0, \lambda_{\text{noise}}^2 I)\), and feeding both prompts with shared initial noise into the frozen one-step generator \(G_\theta\), they collect structurally aligned paired images \((z_0, z_0^{\text{edit}})\) strictly bounded within the generator's reachable manifold.
The encoder is optimized via a composite objective combining reconstruction error and a novel Structure-Aware Editing (SAE) loss:
with \(\lambda_{\text{sae}} = 1.0\). The reconstruction term enforces faithful recovery of the original frame, while the SAE term forces the decoded output of the same inverted latent under the perturbed prompt \(c_{\text{edit}}\) to match the reference target \(z_0^{\text{edit}}\). Because the reference target shares layout by construction, this objective trains the inverted latent to inherently retain structural geometry under semantic variation, requiring zero external guidance at test time.
2. Unified-Frame Latent Concatenation: Reusing spatial self-attention for cross-frame alignment
To achieve temporal consistency without adding temporal convolution or temporal attention modules, OSVE introduces Unified-Frame Editing (UFE). Taking advantage of the fact that diffusion U-Nets naturally support flexible spatial dimensions, the inverted noise latents of all \(K\) frames \(\hat{z}_{T, k} \in \mathbb{R}^{C \times H \times W}\) are concatenated along the horizontal width dimension:
When \(Z_T^{\text{UFE}}\) is passed through the frozen one-step generator \(G_\theta\), its native self-attention mechanism operates globally across all spatial patches from all frames. Visual tokens corresponding to identical semantic parts (such as an animal's nose or eyes) directly query and attend to corresponding features across neighboring and distant frames. This input-level unification enables seamless appearance and motion alignment in a single forward pass without modifying any network parameters.
3. Anchor-Guided Sliding Window: Harmonizing local smoothness with long-range anchoring
Because the memory and computational complexity of global attention scale quadratically with the total concatenated width, directly processing very long videos is infeasible. OSVE resolves this with an overlapping sliding window guided by an anchor frame. The video's global anchor frame \(\hat{z}_T^A\) is identified by calculating the medoid in the DINO feature space, capturing the most representative visual context of the sequence. The video is then processed in overlapping windows of size \(w=7\) with stride \(s=5\). For every window, the anchor latent is prepended to the window's latents: \([ \hat{z}_T^A, \hat{z}_{T, t}, \dots, \hat{z}_{T, t+w-1} ]\). After the single-step generation pass, the anchor slice is discarded and only the central \(s\) frames are retained and stitched into the final output. The sliding window guarantees smooth local transitions between adjacent frames, while the persistent anchor frame anchors color palette, illumination, and subject identity across the entire duration, completely suppressing long-term temporal drift.
Key Experimental Results¶
Main Results¶
The method was evaluated on 60 real-world videos (51 short videos of 20 frames and 9 long videos of 90 frames) with five distinct text-editing prompts per video (encompassing object modification, deletion, and global stylization) using the comprehensive VBench benchmark. Metrics include temporal consistency (Subject Consistency SC, Background Consistency BC, Temporal Flickering TF, Motion Smoothness MS), single-frame perceptual quality (Aesthetic Quality AQ, Imaging Quality IQ), the composite Balanced Quality Score (BQS), and inference speed (FPS).
| Framework | Method | SC โ | BC โ | TF โ | MS โ | AQ โ | IQ โ | BQS โ | FPS โ |
|---|---|---|---|---|---|---|---|---|---|
| Multi-to-Multi (SD1.5) | FLATTEN | 0.965 | 0.970 | 0.964 | 0.972 | 0.625 | 0.639 | 0.611 | 0.072 |
| Multi-to-Multi (SD1.5) | TokenFlow | 0.983 | 0.976 | 0.985 | 0.991 | 0.668 | 0.680 | 0.663 | 0.075 |
| Multi-to-Multi (SD1.5) | FRESCO | 0.978 | 0.974 | 0.973 | 0.991 | 0.649 | 0.729 | 0.674 | 0.078 |
| Multi-to-Multi (SD1.5) | RAVE | 0.982 | 0.976 | 0.975 | 0.986 | 0.637 | 0.695 | 0.653 | 0.091 |
| Multi-to-Multi (SD1.5) | COVE | 0.983 | 0.976 | 0.984 | 0.989 | 0.645 | 0.655 | 0.639 | 0.061 |
| Multi-to-One (DMD2) | Prompt-to-Prompt | 0.915 | 0.945 | 0.965 | 0.978 | 0.587 | 0.566 | 0.548 | 0.898 |
| Multi-to-One (DMD2) | ControlNet (Depth) | 0.968 | 0.961 | 0.970 | 0.984 | 0.658 | 0.673 | 0.646 | 0.578 |
| Multi-to-One (DMD2) | Plug-and-Play | 0.953 | 0.971 | 0.995 | 0.995 | 0.486 | 0.220 | 0.345 | 0.398 |
| One-to-One (DMD2) | OSVE (Ours) | 0.983 | 0.977 | 0.978 | 0.991 | 0.678 | 0.703 | 0.678 | 15.625 |
On long videos (90 frames), OSVE maintains an inference speed of 15.793 FPS and attains a top BQS of 0.677, outperforming both Multi-to-Multi baselines (FRESCO 0.655, RAVE 0.659) and Multi-to-One adaptations, while operating approximately 155ร faster than RAVE and 188ร faster than TokenFlow.
Ablation Study¶
Validation of the inversion encoder training loss on the PIE-Bench benchmark (measuring structural distance and text alignment):
| Config | \(\mathcal{L}_{\text{mse}}\) | \(\mathcal{L}_{\text{sae}}\) | Struct. Dist. โ | CLIP Whole โ | CLIP Edit โ | Note |
|---|---|---|---|---|---|---|
| Reconstruction only | โ | โ | 0.087 | 21.797 | 19.884 | Inadequate structural preservation |
| SAE loss only | โ | โ | 0.074 | 22.349 | 19.863 | Structure partially retained but detail compromised |
| Full model | โ | โ | 0.064 | 22.329 | 20.416 | Optimal synergy between structure and editability |
Ablation of UFE components on 90-frame videos:
| Sliding Window | Anchor Frame | SC โ | BC โ | TF โ | MS โ | BQS โ |
|---|---|---|---|---|---|---|
| โ | โ | 0.931 | 0.948 | 0.972 | 0.980 | 0.671 |
| โ | โ | 0.954 | 0.963 | 0.977 | 0.988 | 0.676 |
| โ | โ | 0.943 | 0.950 | 0.976 | 0.985 | 0.669 |
| โ | โ | 0.958 | 0.965 | 0.978 | 0.989 | 0.677 |
Key Findings¶
- Synergistic loss formulation: Training solely with \(\mathcal{L}_{\text{mse}}\) encourages pixel-level copying, leaving the latent vulnerable to structural collapse under edited text; training with \(\mathcal{L}_{\text{sae}}\) alone lacks pixel-faithful reconstruction constraints. Combining both drives structural distance down to 0.064 while reaching the highest edit-region CLIP score of 20.416.
- Complementarity of window and anchor: The sliding window primarily elevates short-range temporal smoothness (MS increases from 0.980 to 0.988), while the global DINO anchor strongly anchors subject and background consistency against long-term drift (SC improves from 0.931 to 0.958).
- Computational breakthrough: By converting both inversion and generation into single-pass forward computations, OSVE reaches over 15.6 FPS, bridging the practical gap from minutes/hours per video to true sub-second execution.
Highlights & Insights¶
- Pre-encoding structural control bypasses over-steering: In single-step diffusion, mid-generation attention steering lacks iterative feedback and induces destructive artifacts; OSVE proves that structural control must be baked into the initial noise latent prior to denoising.
- Generator-aligned prompt perturbation prevents gradient conflict: Generating structurally similar pairs using perturbed prompt embeddings on the frozen generator itself ensures all supervisory signals reside strictly within the generator's native data manifold.
- Input-level width concatenation activates emergent temporal attention: Splicing frame latents horizontally allows 2D spatial self-attention to serve as cross-frame correspondence matching without adding a single temporal parameter.
Limitations & Future Work¶
- Constrained by base image generator capacity: OSVE is built atop the SD1.5-distilled DMD2 one-step generator, inheriting its native resolution and complex compositional understanding limits.
- Sensitivity of single anchor to dramatic viewpoint or scene cuts: The global anchor relies on a single DINO medoid frame; for videos containing abrupt camera cuts or extreme perspective changes, a single anchor cannot represent multiple distinct visual regimes (requiring multi-anchor or adaptive keyframe extensions).
- Scalability to native one-step Text-to-Video models: As one-step distilled T2V models mature and overcome current blurriness issues, the foundational concepts of pre-encoded structure inversion and unified latent processing can be extended to native 3D spatio-temporal backbones.
Related Work & Insights¶
- vs Multi-step video editors (TokenFlow / RAVE / FRESCO): Standard methods require 50โ100 iterative steps and complex optical flow or token-injection pipelines, running at sluggish speeds (0.06โ0.09 FPS); OSVE introduces the first purely single-step One-to-One pipeline, achieving matched or superior quality at two orders of magnitude higher speed (15.6+ FPS).
- vs Single-step image editing adaptations (Prompt-to-Prompt / ControlNet in single-step): Naive ports suffer from severe over-steering and distortion; OSVE shows that pre-generation structure encoding via SAE loss is vastly superior to runtime attention intervention.
- vs Conventional diffusion inversion (DDIM / Null-text Inversion): Iterative inversion trajectories collapse on one-step models; OSVE's parameterized encoder performs direct one-step noise estimation, setting a benchmark for extreme low-latency diffusion inversion.
Rating¶
- Novelty: โญโญโญโญโญ [Pioneers the adaptation of one-step diffusion to video editing via pre-encoded structure inversion and input concatenation]
- Experimental Thoroughness: โญโญโญโญโญ [Comprehensive benchmarking on 60 videos across short and long regimes, full component ablations, and a 26-person user study]
- Writing Quality: โญโญโญโญโญ [Exceptionally clear problem formulation pinpointing the root causes of single-step over-steering and trajectory mismatch]
- Value: โญโญโญโญโญ [Propels diffusion video editing into the 15+ FPS interactive realm, demonstrating immediate practical utility]