Skip to content

RhymeFlow: Training Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling

Conference: ECCV 2026
Paper: ECCV Official
Area: Video Generation
Keywords: Video Diffusion, Training-Free Acceleration, Asynchronous Denoising Scheduling, Latent Trajectory Projection, DiT

TL;DR

RhymeFlow breaks the conventional synchronous denoising paradigm of video diffusion models where every frame undergoes identical dense computation, introducing content-aware sequential keyframe selection, progressive asynchronous scheduling, and lightweight latent trajectory projection to achieve 1.53× to 1.78× training-free speedups with negligible perceptual quality degradation.

Background & Motivation

Diffusion Transformers (DiTs) such as Wan 2.1, CogVideoX, and HunyuanVideo have established new state-of-the-art benchmarks in photorealistic text-to-video generation. However, their real-world deployment is severely hampered by inference latency. To model spatial appearance alongside complex inter-frame motions, DiT architectures rely on full spatiotemporal 3D attention mechanisms, which suffer from quadratic computational complexity with respect to the total number of spatial pixels and temporal frames. Combined with dozens of sequential denoising steps, generating a short high-definition video requires prohibitive GPU memory bandwidth and computation. Existing training-free acceleration works focus predominantly on intra-step computational savings, such as attention sparsification based on heavy hitters or temporal patterns (e.g., SVG, SAP) and feature reuse across adjacent timesteps via KV-cache or layer skipping (e.g., FasterCache, DeepCache).

Nonetheless, all existing approaches strictly abide by the classical diffusion constraint: across every single timestep of the reverse sampling process, every frame in the target video sequence must undergo a full, dense, and synchronous network forward pass. This uniform computational allocation is fundamentally redundant for natural video data. Due to physical motion continuity and visual scene redundancy, adjacent frames exhibit strong spatiotemporal coherence. Once a sparse set of pivotal keyframes capturing critical semantic transitions or rapid structural motions is anchored with step-by-step full computation, the intermediate latent states of neighboring non-keyframes evolve along remarkably predictable, smooth trajectories.

Subjecting every frame to identical step-by-step dense computation is therefore computationally inefficient, opening a promising yet untapped orthogonal dimension of frame-specific heterogeneous scheduling. The core angle of attack is to decouple the denoising trajectories of individual frames according to their semantic importance, prioritizing keyframes for structural fidelity while allowing non-keyframes to skip steps. Core idea: introduce RhymeFlow, a training-free framework that identifies pivotal keyframes to undergo step-by-step dense denoising, assigns non-keyframes a progressive step-skipping schedule, and analytically reconstructs missing temporal context via latent trajectory projection, achieving compounded acceleration while maintaining global spatiotemporal coherence.

Method

Overall Architecture

The RhymeFlow pipeline operates across three primary phases: an initial synchronous warm-up, sequential content-aware keyframe identification, and alternating progressive asynchronous denoising flow scheduling. During the initial \(T_w\) steps of the reverse diffusion trajectory, all latent frames are updated synchronously with full 3D attention to establish foundational composition and motion priors. Next, using single-step clean latent estimates, semantic transition points are selected as keyframes. In the subsequent asynchronous stage, keyframes advance step-by-step to preserve fine details, while non-keyframes skip computation along multi-step strides. At intermediate timesteps where non-keyframes are skipped, their latent representations are analytically reconstructed via linear projection to provide full temporal context for keyframe attention. Finally, at periodic rhythmic points, all frames re-synchronize through a full 3D attention pass to bound error accumulation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Gaussian Noise<br/>N Latent Video Frames"] --> B["Initial Synchronous Warm-up<br/>First Tw steps full dense 3D denoising for all frames"]
    B --> C["Sequential Keyframe Selection<br/>Anchor semantic transitions via single-step clean latent estimates"]
    C --> D["Progressive Asynchronous Denoising Scheduling<br/>Step-by-step keyframes / Progressive multi-step skipping for non-keyframes"]
    D -->|Intermediate Asynchronous Steps| E["Latent Trajectory Linear Projection<br/>Analytic interpolation across endpoints for non-keyframe context"]
    E --> F["Layer-wise Rolling KV-Cache Management<br/>Store last two post-attention states, project transiently on-the-fly"]
    D -->|Periodic Rhythmic Synchronous Points| G["Global 3D Attention Re-synchronization<br/>Joint forward pass across all frames to calibrate accumulated errors"]
    F --> H["Converged Latent Representation<br/>3D-VAE decoding into high-fidelity video"]
    G --> H

Key Designs

1. Sequential Keyframe Selection: Content-Aware Semantic Shift Anchoring via Clean Latent Proxies

At early diffusion stages, noisy latents \(\mathbf{z}_t^{(i)}\) are heavily corrupted by isotropic Gaussian noise, obscuring genuine inter-frame semantic correlations. RhymeFlow overcomes this by applying a closed-form single-step denoising prediction at the end of the warm-up phase to derive clean latent proxies \(\hat{\mathbf{z}}_0^{(i)}\). Recognizing that the initial video frame serves as the foundational anchor for long-range composition and visual fidelity, the algorithm explicitly initializes the keyframe set with the first frame: \(\mathcal{K} = \{\hat{\mathbf{z}}_0^{(1)}\}\). Candidate frames are then evaluated sequentially along the temporal axis. For each candidate \(\hat{\mathbf{z}}_0^{(j)}\), its cosine similarity against the nearest preceding keyframe \(\hat{\mathbf{z}}_0^{(k)} \in \mathcal{K}\) is computed: $\(\psi_{\text{sim}}(\hat{\mathbf{z}}_0^{(j)}, \hat{\mathbf{z}}_0^{(k)}) = \frac{\hat{\mathbf{z}}_0^{(j)} \cdot \hat{\mathbf{z}}_0^{(k)}}{\|\hat{\mathbf{z}}_0^{(j)}\|_2 \|\hat{\mathbf{z}}_0^{(k)}\|_2}\)$ If the similarity drops below a predefined threshold \(\tau\), signaling a significant semantic or motion shift, the candidate is added to \(\mathcal{K}\) as a new anchor. This dynamic chronological selection ensures that the computational budget is allocated to genuine motion turning points rather than arbitrary uniform intervals.

2. Progressive Asynchronous Scheduling: Staged Step-Skipping with Periodic Rhythmic Synchronization

The reverse diffusion trajectory exhibits marked non-uniformity across timesteps. High-noise stages (large \(t\)) govern global structural composition and low-frequency shapes, where trajectories are sensitive and errors are irrecoverable. Conversely, low-noise stages (small \(t\)) refine high-frequency textures, where latent trajectories become smooth and highly predictable. A static skipping stride either disrupts early foundational structures or under-accelerates later stages. RhymeFlow resolves this via a piecewise progressive skipping schedule for non-keyframes: $\(n_{\text{skip}}(t) = \begin{cases} n_{\text{small}}, & \text{if } T_{\text{mid}} < t \le T - T_w \\ n_{\text{large}}, & \text{if } t \le T_{\text{mid}} \end{cases}\)$ where \(T_w\) is the warm-up duration, \(T_{\text{mid}}\) is the midpoint step (typically \(T/2\)), and \(n_{\text{small}} = 2, n_{\text{large}} = 3\). While non-keyframes leap across timesteps, keyframes advance strictly one step at a time (\(z_t^{(k)} \to z_{t-1}^{(k)}\)). To prevent long-range drift, timesteps where staggered strides coincide are designated as "Rhythmic Points." At these points, all frames execute a standard synchronous 3D attention pass, resetting the error accumulation of non-keyframes against the high-fidelity keyframe anchors.

3. Latent Trajectory Projection: Analytic Reconstruction of Transient Temporal Context

During intermediate asynchronous steps \(\tau \in [t - n_{\text{skip}} + 1, t - 1]\), keyframes must be updated using 3D spatiotemporal attention, which strictly requires contextual representations from all other frames at timestep \(\tau\). However, non-keyframes have been fast-forwarded to \(t - n_{\text{skip}}\), leaving intermediate latents physically uncomputed. Dropping non-keyframes degrades spatiotemporal continuity, whereas evaluating network forward passes defeats acceleration. Leveraging the near-linear nature of ODE trajectories in modern Rectified Flow models, RhymeFlow estimates the missing intermediate latents via closed-form linear interpolation between endpoints: $\(\hat{\mathbf{z}}_\tau^{(j)} = (1 - \alpha) \mathbf{z}_{t_{\text{start}}}^{(j)} + \alpha \mathbf{z}_{t_{\text{end}}}^{(j)}, \quad \alpha = \frac{t_{\text{start}} - \tau}{t_{\text{start}} - t_{\text{end}}}\)$ This closed-form calculation incurs virtually zero FLOPs. Transient Key and Value representations are computed on-the-fly from \(\hat{\mathbf{z}}_\tau^{(j)}\), enabling keyframes to attend to a seamless, temporally consistent sequence without executing full forward passes for non-keyframes.

4. Layer-wise Rolling KV-Cache Management: Post-Attention State Caching for Peak Memory Reduction

Naively caching full input Key-Value pairs across skipped steps would quickly exhaust GPU memory during long video generation. RhymeFlow implements a per-layer rolling cache that retains only the two most recent post-attention hidden states \(h_{t_1}^{(\ell)}\) and \(h_{t_2}^{(\ell)}\) (\(t_1 > t_2\)) for each non-keyframe. Post-attention representations already incorporate intra-frame spatial interactions and global context, yielding higher interpolation fidelity than raw inputs. When a keyframe requires context at an intermediate step \(\tau \in (t_1, t_2)\), the interpolated representation serves as a transient KV provider for the current attention operation and is immediately discarded. This design eliminates persistent intermediate storage, actually decreasing peak GPU VRAM from 44.3 GB (in the dense baseline) down to 42.6 GB.

A Worked Example

Consider a 50-step diffusion generation on 21 latent frames with parameters \(T_w=8, T_{\text{mid}}=25, n_{\text{small}}=2, n_{\text{large}}=3\). From step 50 down to 42, all 21 frames undergo dense synchronous denoising. At step 42, single-step clean estimates identify 4 pivotal keyframes (e.g., frames 1, 7, 14, 21). Entering the \(n_{\text{small}}=2\) stage, keyframes advance sequentially \(42 \to 41 \to 40\), while non-keyframes fast-forward directly from 42 to 40. At step 41, when keyframes run attention, missing non-keyframe states at \(t=41\) are linearly interpolated from their step 42 and step 40 states (\(\alpha=0.5\)). At step 40 (a rhythmic point), all 21 frames join in a full 3D attention update. Once the diffusion process crosses step 25, the skip stride widens to \(n_{\text{large}}=3\), further accelerating generation without structural loss.

Loss & Training

RhymeFlow is an entirely training-free and fine-tuning-free inference acceleration framework, requiring no auxiliary distillation losses or model parameter updates. All hyper-parameters (\(\tau, T_w, T_{\text{mid}}, n_{\text{skip}}\)) are architectural configurations. In orthogonal combination scenarios, RhymeFlow's frame-level mask \(\mathbf{M}_{\text{Rhyme}}\) integrates seamlessly with token-level intra-step sparse attention masks (e.g., \(\mathbf{M}_{\text{SAP}}\)) via Hadamard product: \(\mathbf{M}_{\text{combined}} = \mathbf{M}_{\text{Rhyme}} \odot \mathbf{M}_{\text{SAP}}\), unlocking compounded hierarchical acceleration.

Key Experimental Results

Main Results

Evaluation was conducted on Wan 2.1 (1.3B, 81 frames at 720p, 21 latent frames) and CogVideoX-v1.5 (81 frames at 720p, 11 latent frames) using a single NVIDIA A800 GPU:

Base Model Acceleration Method PSNR ↑ SSIM ↑ LPIPS ↓ Subject Consistency ↑ Imaging Quality ↑ Latency (s) ↓ Speedup ↑
Wan 2.1 Dense Baseline - - - 0.9102 0.6946 993.5 1.00×
Wan 2.1 SpargeAttn [45] 20.399 0.613 0.393 0.8632 0.7118 719.7 1.38×
Wan 2.1 SVG [36] 22.419 0.694 0.290 0.8758 0.6913 708.0 1.40×
Wan 2.1 SAP [40] 24.454 0.730 0.223 0.8789 0.6837 608.5 1.63×
Wan 2.1 Ours (RhymeFlow) 26.291 0.783 0.168 0.8831 0.6706 650.4 1.53×
Wan 2.1 Ours + SAP 24.586 0.737 0.221 0.8792 0.6806 596.8 1.66×
CogVideoX-v1.5 Dense Baseline - - - 0.987 0.624 625.0 1.00×
CogVideoX-v1.5 MInference [15] 22.490 0.743 0.264 0.874 0.589 422.3 1.48×
CogVideoX-v1.5 PAB [47] 23.230 0.782 0.145 0.978 0.573 443.3 1.41×
CogVideoX-v1.5 SVG [36] 24.130 0.811 0.171 0.982 0.597 385.8 1.62×
CogVideoX-v1.5 Ours (RhymeFlow) 26.890 0.852 0.142 0.986 0.623 351.1 1.78×
CogVideoX-v1.5 Ours + SAP 25.574 0.821 0.157 0.972 0.609 323.8 1.93×

On HunyuanVideo benchmarks, RhymeFlow also established superior performance: reaching 26.34 PSNR and 0.918 SSIM at 2.26× speedup (2939s vs. dense baseline), significantly outperforming caching methods like EasyCache (23.51 PSNR) and DiCache (23.54 PSNR). In combination with SAP, it achieved a 2.60× acceleration ratio (2555s).

Ablation Study

Extensive ablations on Wan 2.1 validated the critical necessity of each algorithmic component:

Experiment Group Setup / Variant PSNR ↑ SSIM ↑ LPIPS ↓ Peak Memory (GB) ↓ Latency (s) ↓ Speedup ↑
Core Architecture Full Model (Ours) 26.291 0.783 0.168 42.6 650.4 1.53×
Core Architecture w/o Progressive Scheduling 25.399 0.753 0.172 - 708.0 1.40×
Core Architecture w/o Trajectory Projection 20.630 0.525 0.383 - 622.2 1.60×
Keyframe Selection Random Selection 20.630 0.525 0.383 - 650.0 1.53×
Keyframe Selection First-Frames Only 19.220 0.515 0.402 - 651.0 1.53×
Keyframe Selection Uniform Spacing 24.293 0.643 0.183 - 649.0 1.53×
Keyframe Selection Content-Aware Sequential (Ours) 26.291 0.783 0.168 - 650.4 1.53×
KV-Cache Management Native Skipping (w/o Cache) 26.291 0.783 0.168 49.9 653.6 1.52×
KV-Cache Management Layer-wise Rolling Cache (Ours) 26.291 0.783 0.168 42.6 650.4 1.53×

Key Findings

  • Trajectory projection is essential for temporal stability: Removing latent trajectory projection (w/o Projection) causes PSNR to drop drastically from 26.291 to 20.630 and SSIM to collapse to 0.525. Latent error analysis reveals a 10× increase in keyframe error (\(0.0019 \to 0.0190\)), confirming that attention sequences truncated by step-skipping induce severe visual breakdown.
  • Content-aware keyframes outperform uniform allocation: At identical latency (~650s), selecting keyframes via clean latent proxies improves SSIM by 0.140 (\(0.643 \to 0.783\)) compared to uniform sampling, demonstrating the necessity of anchoring genuine semantic transition points.
  • Pareto trade-off of warm-up and keyframe budget: Fixing \(T_w=8\) while increasing keyframes \(M\) from 3 to 5 boosts PSNR from 23.742 to 27.707 but lowers speedup from 1.60× to 1.39×; \(T_w=8, M=4\) offers the optimal sweet spot (1.53× speedup with minimal degradation).
  • Human preference shows no statistical degradation: In a double-blind user study with 82 participants, RhymeFlow significantly outperformed SVG and SAP (\(p < 0.05\)). Compared against the dense baseline, differences were statistically insignificant across all metrics (all \(p > 0.10\)), with tie rates exceeding 53% to 58%.

Highlights & Insights

  • Orthogonal inter-step scheduling compounding with intra-step sparsity: Rather than squeezing efficiency inside individual 3D attention matrices, RhymeFlow controls which frames participate across timesteps. Because the acceleration dimensions are orthogonal, compounding RhymeFlow with SAP yields a near 2× speedup (1.93× on CogVideoX-v1.5).
  • Capitalizing on Rectified Flow ODE linearity for near-zero-cost projection: Observing that flow matching trajectories in modern DiTs are approximately linear, the method estimates missing latent states via scalar-weighted linear interpolation, delivering complete temporal context without neural network evaluations.
  • Post-attention cache bounding VRAM footprint: Caching post-attention representations and discarding intermediate projections immediately prevents memory bloat, reducing peak VRAM from 44.3 GB to 42.6 GB while preserving acceleration gains.

Limitations & Future Work

  • Author-admitted limitations: The skipping stride \(n_{\text{skip}}\) and rhythmic synchronization intervals are governed by piecewise constant schedules, which do not continuously adapt to varying scene dynamics or motion intensity.
  • Potential caveats: For videos containing sudden scene cuts or high-frequency occlusions, linear interpolation across disparate endpoints may introduce brief transient blur near cut boundaries.
  • Future directions: Designing a lightweight controller to predict local ODE trajectory curvature dynamically; and extending the asynchronous flow paradigm to autoregressive video generation architectures.
  • vs SVG [36] / SAP [40] (Intra-step Sparse Attention): SVG and SAP prune or permute tokens within a single attention step based on head patterns. RhymeFlow optimizes along the temporal diffusion axis by skipping entire forward passes for non-keyframes. Their combination forms a hierarchical pruning pipeline.
  • vs DeepCache [27] / FasterCache [25] / DiCache [4] (Feature Caching): Prior caching methods reuse deep features across timesteps, which often leads to motion lag in dynamic video scenes. RhymeFlow preserves step-by-step updates on keyframes and uses ODE-guided linear interpolation for non-keyframes rather than static feature copying, ensuring smooth dynamic transitions.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneers asynchronous frame-level scheduling in video diffusion, breaking the traditional uniform full-frame denoising constraint]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluated across Wan 2.1, CogVideoX, and HunyuanVideo, detailed ablations on all components, and an 82-participant double-blind study]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation grounded in ODE trajectory linearity, well-structured formulation and self-consistent narrative]
  • Value: ⭐⭐⭐⭐⭐ [Completely training-free, memory-friendly, highly compatible with orthogonal sparse attention baselines, strong engineering impact]