Skip to content

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models

Conference: NeurIPS2026
arXiv: 2605.19398
Area: Video Generation
Keywords: image-to-video, reference-frame dominance, attention modulation, motion control, training-free inference

TL;DR

DyMoS subtracts an adjustable bias from non-reference-query-to-reference-key self-attention logits in the text-conditioned branch during early image-to-video denoising, raising Wan 2.2's VBench Dynamic Degree from 51.7 to 64.8, although stronger dynamics do not automatically imply more realistic motion or higher reference fidelity.

Background & Motivation

Image-to-video (I2V) models commonly add image conditioning to a text-to-video (T2V) backbone, using the input image as the generated video's first frame. The reference constrains subject appearance, background, and colors, but that constraint can become excessive: even when a prompt requests jumping or a moving vehicle, subsequent frames remain close to the first frame, producing an almost static sequence. Earlier methods attribute this behavior to conditional image leakage and inject noise into the image condition, change latent conditioning, or apply early-stage low-pass filtering as in ALG. These approaches may require training or trade appearance information for motion.

This paper asks whether the problem lies less in an overly detailed input image than in how the model repeatedly uses it. The authors construct 50 paired generations with CogVideoX-5B: generate a T2V video, feed its first frame to I2V, and retain the prompt, random seed, and sampling configuration. Aggregating token self-attention into frame-to-frame distributions shows that non-reference I2V frames allocate more attention to reference-frame keys, particularly during the initial 10% of denoising. This reference-frame dominance can over-propagate appearance through time and crowd out interactions among generated frames. However, the paired backbones still differ in their parameters, so this observation does not establish a unique explanation for every static-generation failure.

The authors also intervene in both directions: strengthening the pathway produces more static videos, whereas weakening it increases dynamics and moves frame-to-frame attention closer to paired T2V generation. Core Idea: retain the input image and model weights, but reduce generated frames' excessive reliance on the reference during the early sampling stage that establishes global motion, allowing existing video priors to contribute rather than first damaging the image condition and then recovering appearance.

Method

Overall Architecture

DyMoS (Dynamic Motion Slider) modifies attention during sampling; it is not an additional trained network. Inputs remain the reference image, text prompt, and random video latent. The reference is encoded normally, and each backbone retains its original image-conditioning pathway. Every sampling step still computes null-text and text-conditioned predictions, combines them using the original classifier-free guidance (CFG), and updates the video latent.

The added operation occurs inside the text-conditioned prediction: during early sampling, a bias is subtracted from the non-reference-video-query-to-reference-video-key block in all relevant self-attention layers before softmax. Reference-frame query rows, other video keys, and text tokens in MM-DiT are outside this target pathway. Once the early window ends, attention reverts to its original computation, and the existing VAE ultimately decodes the video.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Reference image, text<br/>current video latent"] --> B["Early-Window Gating"]
    B -->|Early: modify text-conditioned branch| C["Targeted Logit Bias Subtraction"]
    B -->|After window: original attention| D["Conditional-Branch CFG"]
    C --> D
    A -->|Original null-text prediction, still image-conditioned| D
    D --> E["Original sampler update<br/>VAE decoding after loop"]

Here, null text does not mean absence of image conditioning: both CFG branches continue to receive the same reference image. The diagram depicts inference data flow only, with no training supervision, additional motion teacher, or reward optimization stage.

Key Designs

1. Early-Window Gating: alter motion layout before restoring detail refinement

Early denoising establishes coarse motion and scene evolution, whereas later steps focus more on appearance and temporal detail. Suppressing reference information throughout sampling could remove not only static bias but also the subject's appearance anchor. DyMoS uses the sampling-step ratio \(\lambda\) to restrict modulation to an initial interval, then returns entirely to the original attention computation. The mechanism analysis aggregates the initial 10% of steps; this is not a universal intervention ratio for the main experiments, where Wan 2.2 uses \(\lambda=0.2\).

This gate concerns progress through sampling, not the first few output video frames. Query location along latent frames is a separate dimension controlled by the frame-wise schedule below. The appendix algorithm expresses the window using \(t=i/N\) and \(t<\lambda\), whereas the background describes flow-matching time along denoising from noise to a clean latent. These conventions should not be conflated. An implementation should gate by sampling progress and check the boundary rather than treat the algorithm's \(t\) directly as the backbone's native noise time.

2. Targeted Logit Bias Subtraction: reduce reference-reading probability without deleting reference information

Ordinary self-attention converts queryโ€“key similarities into logits and applies softmax to allocate reading weights. DyMoS does not modify the image, queries, keys, or values. Before softmax, it modifies only the matrix block through which generated video-frame queries read reference video-frame keys. With \(\mathcal{I}_{f_0}\) denoting reference-latent-frame token indices and \(f(i)\) the query token's latent-frame position, the central mechanism is the paper's Equation (6):

\[ \tilde{\mathcal{L}}[i,j]=\mathcal{L}[i,j]-\gamma\,\phi(f(i))\,\mathbf{1}[j\in\mathcal{I}_{f_0}]\,\mathbf{1}[i\notin\mathcal{I}_{f_0}]. \]

Here, \(\gamma\) is the motion slider and \(\phi\) is the frame-position schedule. \(\gamma>0\) lowers the targeted logits, \(\gamma=0\) recovers the baseline, and \(\gamma<0\) strengthens reference-frame dominance. The two indicator functions are essential: subtracting a constant from the entire attention row would cancel under softmax. Lowering only the reference-key group instead raises other keys' relative shares. The targeted keys' unnormalized weights are multiplied by \(e^{-\gamma\phi(f(i))}\), while relative weights within that group remain unchanged; this is a mathematical interpretation of the equation, not an additional author-proposed formula.

The operation neither zeros reference attention nor changes reference-frame query rows, so reference information remains accessible to generated frames. All backbones use the uniform schedule \(\phi=1\) in the main experiments. The appendix also tests linear and logarithmic schedules, \(f(i)/(F-1)\) and \(\log(1+f(i))/\log F\), applying weaker modulation near the reference frame and stronger modulation later. Uniform modulation performs better, so increasing modulation with frame distance should not be presented as the default method.

3. Conditional-Branch CFG: change only the text-conditioned prediction while retaining the image anchor

Changing both the null-text and text-conditioned branches would also change the guidance direction. The authors instead use modulated logits only in the text-conditioned branch, allowing prompt-described motion to influence generation more strongly. The null-text branch retains its original attention and still receives the original reference image. Equation (7) combines the predictions as follows:

\[ \tilde{\bm v}_\theta=\bm v_\theta(\bm z_t,t,\varnothing,\bm z_{\mathrm{ref}})+\omega\left[\bm v_\theta^{\tilde{\mathcal L}}(\bm z_t,t,\bm c,\bm z_{\mathrm{ref}})-\bm v_\theta(\bm z_t,t,\varnothing,\bm z_{\mathrm{ref}})\right]. \]

The parameter \(\omega\) remains the backbone's original CFG scale, not a new motion parameter. DyMoS preserves the original sampling loop and number of forward passes, fusing score modification into the attention kernel with PyTorch FlexAttention. Thus, training-free does not mean zero cost or no inference implementation changes, and the slider cannot necessarily be inserted into a closed model API.

A Worked Example

Consider Wan 2.2-14B generating a video whose prompt requests a moving vehicle, with 40 sampling steps, \(\gamma=0.6\), \(\lambda=0.2\), and \(\phi=1\). The reference image and random latent enter the existing pipeline. During roughly the initial fifth of sampling, every generated-frame query's logit for a reference-frame key is reduced by 0.6 in the text-conditioned branch. Other keys receive no such subtraction, and neither does the null-text branch.

The method specifies neither vehicle speed nor an explicit trajectory. It instead reduces repeated copying of the initial vehicle position, allowing the prompt and existing video prior to produce motion. After the early window, both branches use original attention and retain the original refinement process. Multiplying 40 by 0.2 gives a nominal eight-step window, but the appendix starts at \(i=1\) and uses a strict inequality; its pseudocode therefore does not guarantee exactly eight modified steps. The actual boundary requires implementation verification.

Loss & Training

DyMoS introduces no new loss, training data, or weight updates. The flow-matching and CFG background describes existing backbones, not modules retrained by this work. Although a single \(\gamma\) acts as the user-facing motion slider, reproducing experiments also requires backbone-specific \(\lambda\), sampler, CFG scale, and resolution.

Backbone Video size (framesร—widthร—height) / FPS Steps / CFG Sampler \(\gamma\) \(\lambda\)
Wan 2.2-14B 81ร—832ร—480 / 16 40 / 3.5 UniPC 0.6 0.20
Wan 2.1-14B 81ร—832ร—480 / 16 40 / 5.0 UniPC 1.0 0.25
HunyuanVideo-1.5 81ร—832ร—480 / 16 50 / 6.0 FlowMatch-Euler 1.0 0.20
CogVideoX-5B 49ร—720ร—480 / 8 50 / 6.0 DDIM 0.8 0.24

These configurations come from appendix Table 3; every backbone uses \(\phi=1\). Cross-backbone effectiveness therefore supports transferability of the intervention pathway, not untuned transfer of one hyperparameter configuration.

Key Experimental Results

Main Results

The main evaluation uses 246 VBench-I2V imageโ€“caption pairs, excluding background-quality and camera-motion-instruction subsets. Additional evaluations use 200 PVD video-first-frameโ€“caption pairs and 500 VidProM prompts paired with Flux.1-dev-generated reference images. Within each backbone, paired random seeds and GPU type are fixed when comparing the original baseline, official ALG, and DyMoS.

VBench Dynamic Degree uses RAFT optical flow and the highest 5% of flows to assess whether videos contain sufficiently large motion. VideoScore Dynamic Degree is predicted by a video evaluator trained on human feedback. Both assess dynamics rather than fully guaranteeing physical realism. ViCLIP measures videoโ€“text semantic similarity, and VisionReward is a learned proxy for human preference.

The following table preserves the values in the paper's Table 1. DD denotes Dynamic Degree and VQ denotes the authors' Video Quality. Columns have different scales and should not be compared numerically with one another.

Backbone Method DD (VBench) DD (VideoScore) VQ ViCLIP VisionReward
Wan 2.2-14B Baseline 51.7 2.83 58.0 0.2614 0.143
Wan 2.2-14B ALG 55.3 2.83 57.7 0.2616 0.140
Wan 2.2-14B DyMoS 64.8 2.84 57.7 0.2621 0.143
Wan 2.1-14B Baseline 48.7 3.21 57.7 0.2601 0.132
Wan 2.1-14B ALG 52.5 3.18 57.2 0.2615 0.127
Wan 2.1-14B DyMoS 54.5 3.25 57.4 0.2618 0.126
HunyuanVideo-1.5 Baseline 48.0 2.87 58.2 0.2624 0.140
HunyuanVideo-1.5 ALG 46.3 2.87 58.2 0.2629 0.141
HunyuanVideo-1.5 DyMoS 55.3 2.88 58.1 0.2653 0.141
CogVideoX-5B Baseline 30.9 2.90 57.8 0.2612 0.140
CogVideoX-5B ALG 52.0 2.98 57.1 0.2647 0.127
CogVideoX-5B DyMoS 52.9 2.92 57.1 0.2652 0.131

Wan 2.2 gains 13.1 percentage points in VBench DD, approximately a 25.3% relative increase; its VQ changes from 58.0 to 57.7 rather than remaining completely unchanged. On CogVideoX, DyMoS exceeds ALG in VBench DD, but its VideoScore DD of 2.92 is below ALG's 2.98. Wan 2.1's VisionReward of 0.126 is below both the baseline's 0.132 and ALG's 0.127, so the prose summary should not be read as evidence that every preference metric is stable or improved.

The authors define their quality aggregate as:

\[ \mathrm{VideoQuality}=0.1\,\mathrm{TF}+0.25\,\mathrm{MS}+0.1\,\mathrm{AQ}+0.25\,\mathrm{IQ}. \]

TF, MS, AQ, and IQ denote Temporal Flickering, Motion Smoothness, Aesthetic Quality, and Imaging Quality. The published weights sum to 0.70 and exclude dynamics, subject consistency, and background consistency; they are not renormalized or corrected here. Consequently, this quality score does not directly establish unchanged reference-subject or background fidelity.

Ablation Study

Appendix Table 6 fixes the main-experiment \(\gamma\) and \(\lambda\) on Wan 2.2 and varies only the frame-wise schedule.

Frame-wise schedule DD (VBench) DD (VideoScore) VQ ViCLIP VisionReward
Uniform 64.8 2.84 57.7 0.2621 0.143
Linear 56.1 2.82 57.9 0.2598 0.143
Log 60.2 2.80 57.7 0.2597 0.141

Uniform modulation exceeds the linear schedule by 8.7 percentage points in VBench DD and the logarithmic schedule by 4.6 points. The linear schedule has slightly higher VQ but weaker dynamics, showing that increasingly suppressing reference attention with frame distance is not automatically preferable.

The main text also tests \(\gamma\in\{0.4,0.6,0.8,1.0\}\) and reports an 11.7% relative DD gain from 0.4 to 0.6. Absolute graph values are absent from the cached text and are not reconstructed. Increasing \(\lambda\) from 0.24 to 0.30 instead reduces DD, so a longer early window is not necessarily better.

Key Findings

  • The mechanism distance \(D(\gamma)\) is defined as the mean Jensenโ€“Shannon divergence between I2V and T2V frame-to-frame attention row distributions over non-reference query frames. It reaches its minimum around \(\gamma=0.6\) in the paired analysis, alongside T2V-like dynamics. This supports attention-structure similarity, not a universally quality-optimal setting.
  • Wan 2.2's VBench DD increases from 79.0 to 85.0 on PVD and from 54.2 to 67.0 on VidProM. However, PVD ViCLIP changes from 0.2147 to 0.2143, so not every metric improves monotonically across datasets either.
  • The MTurk study uses 25 VBench-I2V imageโ€“prompt pairs and collects 30 responses per question, comparing motion, reference fidelity, text alignment, and overall preference. The paper reports that DyMoS is selected most often on all four criteria. Exact percentages from Figure 5(c) are absent from the cache, so no numerical win rates are supplied.
  • Single-A100-80GB timings show DyMoS overheads of 1.2% for Wan 2.2, 1.9% for Wan 2.1, 3.5% for HunyuanVideo-1.5, and 4.3% for CogVideoX, with total runtimes of 841, 749, 503, and 242 seconds respectively. The 1.2% figure should not be generalized to all backbones.

Highlights & Insights

  • The intervention changes repeated access to image information rather than the image itself. Releasing motion while preserving the condition suggests that excessive conditioning may arise from internal routing, not simply input-signal strength.
  • Positive and negative biases serve both mechanism testing and user control: stronger dominance yields more static output, while weaker dominance yields more dynamic output. This bidirectional evidence strengthens the pathwayโ€“motion association beyond a one-sided improvement demonstration, but does not remove confounders such as backbone parameter differences.
  • The early window and conditional-branch choice keep the modification localized. The existing sampler, null-text image condition, and later refinement remain active, allowing motion tendencies to change without retraining.

Limitations & Future Work

  • The authors acknowledge that excessively large \(\gamma\) harms appearance consistency and temporal stability. Larger optical flow or dynamics scores are not sufficient evidence of more realistic motion, and the slider should not be interpreted as a reliable speed controller.
  • Each backbone is tuned separately, with fixed parameters rather than prompt-, subject-, or motion-dependent adaptation. Attention-share-triggered control is a possible research direction, not an implemented module in this paper.
  • Reference fidelity is supported mainly by qualitative examples and a small user study, while aggregate VQ explicitly excludes subject/background consistency. Future evaluation should add those dimensions, physical plausibility, failure rates, and uncertainty across random seeds.
  • The source contains inconsistent wording: Appendix C.1 says to โ€œadd a bias \(\gamma\),โ€ whereas Equation (6) subtracts a positive bias under the main-text sign convention. This note follows the equation and retains the discrepancy. The appendix algorithm's time convention and strict window boundary also require implementation verification.
  • vs ALG: ALG uses early low-pass guidance on high-frequency image content and adds a forward pass during its active interval. DyMoS modifies internal attention logits, retains the input image, and adds no model forward passes. Its dynamics advantage depends on the metric: CogVideoX VideoScore DD does not exceed ALG.
  • vs conditional image leakage mitigation: Zhao et al. perturb image conditioning with time-dependent noise, while FlashI2V redesigns condition injection using Fourier-guided latent shifting. This paper instead changes routing in an already trained model, making it relevant when retaining the original weights is important.
  • Transferable insight: Other strongly reference-conditioned generation tasks can first inspect which queries excessively read reference keys and test targeted intervention instead of immediately weakening the entire condition. Target-token identification, early-window timing, and conditional-branch semantics must be revalidated rather than copied directly from I2V configurations.

Rating

  • Novelty: 4/5 โ€” Reference-frame-dominance analysis motivates a localized attention intervention with a simple but specific mechanism.
  • Experimental Thoroughness: 4/5 โ€” Four backbones, three datasets, schedule ablations, and a user study provide broad coverage, but realism and fidelity quantification remain limited.
  • Writing Quality: 3/5 โ€” The main narrative is clear, but metric summaries, appendix bias signs, and sampling-time conventions need sharper distinctions.
  • Value: 4/5 โ€” No weight updates and modest overhead make it useful for motion adjustment on open backbones, subject to per-backbone tuning and degradation checks.