MVI2V: Human Centric Image to Video Generation with Multiview Consistent Appearance¶
Conference: ECCV 2026
Paper: ECCV Official Page
PDF: ECCV Paper PDF
Area: Video Generation
Keywords: Multiview Image-to-Video, Appearance Consistency, Human Video Generation, Multi-stream Diffusion Architecture, E-commerce Applications
TL;DR¶
MVI2V addresses the severe clothing and texture hallucination of unseen angles (e.g., garment back) in human-centric image-to-video generation by introducing an independent lightweight LoRA reference stream, cross-stream bidirectional self-attention, an inpainting training subtask, and an angular-filtered data curation pipeline, achieving high multiview appearance fidelity with only ~2.1% extra parameters.
Background & Motivation¶
Human-centric image-to-video (I2V) synthesis has recently made significant strides across entertainment, animation, and personalized content creation. However, in high-stakes commercial scenarios such as e-commerce garment showcase and virtual apparel presentation, existing models confront a fundamental bottleneck. Conventional I2V diffusion models condition strictly on a single initial frame. When the subject executes dynamic body reorientations, full turns, or complex camera sweeps, regions unobserved in the first frame (such as garment back cuts, rear prints, or side contours) lack explicit visual grounding. Consequently, models fall back on statistical hallucinations learned from general pretraining, yielding back views and textures that sharply diverge from actual clothing items.
Directly extending single-image paradigms to multiview conditions is non-trivial. Standard human animation frameworks focus on pose-driven single-character animation, whereas video virtual try-on operates by transferring flat garment layouts onto existing video sequences; neither maintains end-to-end temporal coherence while binding multi-angle physical appearance. Furthermore, naively concatenating multiview reference tokens with video tokens creates severe modality conflict: clean reference latents act as deterministic trajectory targets, whereas noisy video tokens represent intermediate, noise-corrupted flow states along the diffusion path. In addition, the strong visual prior of the initial frame easily traps the model in a shortcut: generating plausible video dynamics while ignoring new reference views altogether.
To resolve these tensions, this paper proposes an architectural and training framework that upgrades pretrained single-stream or dual-stream DiT backbones without retraining from scratch. Core idea: equip the model with an independent, structurally identical yet decoupled reference forward stream interacting via cross-stream bidirectional self-attention, complemented by a random inpainting training subtask and a large-angle rotation curation pipeline to suppress initial-frame shortcuts and enforce multiview appearance consistency at minimal parameter cost.
Method¶
Overall Architecture¶
MVI2V is built upon the Flow Matching framework and aims to synthesize temporally coherent human videos conditioned on text prompts, an initial frame, and multiple reference images of the person or garment. Treating clean reference images as an independent visual modality, MVI2V augments the pretrained video diffusion backbone (e.g., single-stream Wan2.1 or dual-stream MMDiT structures) with an identical dedicated reference stream. Visual features interact across modalities bidirectionally solely during the self-attention stage, preventing representation collapse caused by direct sequence concatenation. The training process integrates a specialized data pipeline filtering large angular rotations with an inpainting subtask that randomly masks the first frame's subject to force reference utilization.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
DataPipeline["Angular & Garment Curation Pipeline<br/>Single-person filter + HMR pose clustering + garment segmentation"] --> Input["Input Preparation<br/>First frame c_img + text prompt + multiview reference c_ref"]
Input --> InpaintingSubtask["Random Inpainting Training Subtask<br/>Randomly mask first-frame human box to eliminate shortcuts"]
LoRAAdaptation["Lightweight LoRA Adaptation<br/>Add low-rank branch with only 2.1% extra parameters"] -.->|Parameter-efficient scaling| CrossStreamAttention["Cross-Stream Bidirectional Self-Attention<br/>Bidirectional global interaction between video and reference tokens"]
InpaintingSubtask --> CrossStreamAttention
CrossStreamAttention --> Output["Video Latent Decoding<br/>Output human video with multiview appearance consistency"]
Key Designs¶
1. Angular & Garment Curation Pipeline: Mining complementary multiview human and clothing priors High-quality videos featuring significant body reorientations occupy the long tail in general video corpora. To prevent the model from degenerating into single-view generation due to trivial view angles, the authors constructed an automated multi-stage curation pipeline. First, a human detector filters out videos containing zero, multiple, or heavily truncated subjects, retaining only centered single-person sequences. Next, Human Mesh Recovery (HMR) estimates per-frame body orientation angles, discarding clips whose angular coverage falls below a strict threshold. From each qualifying video, the pipeline extracts canonical frontal and back frames with maximum visual coverage, followed by K-Means clustering over remaining rotation angles to select three centroid frames. Finally, Gemini and Sa2va segment the upper and lower garments from frontal and back views, yielding an augmented set of nine complementary reference images (2 canonical views + 3 clustered views + 4 garment parts) per video.
2. Random Inpainting Training Subtask: Breaking the local optima caused by strong first-frame priors When finetuning pretrained I2V models, the initial frame provides an overwhelming identity and appearance prior. In early diffusion steps, the network can readily synthesize temporally plausible frames solely from the first image, creating an optimization shortcut that entirely bypasses the reference token stream. To break this dependency, MVI2V introduces an auxiliary inpainting objective during training: using the human bounding box, the subject's entire rectangular region in the first frame is randomly masked out. Without first-frame appearance clues, the model is compelled to retrieve clothing colors, structural cuts, and texture patterns strictly from the external multiview reference stream. This masking mechanism is exclusively active during training and completely removed at inference, dramatically boosting reference cross-attention weights at zero runtime cost.
3. Lightweight LoRA Adaptation: Parameter-efficient expansion of the dedicated reference stream Processing clean reference tokens through an independent forward stream requires dedicated parameter capacity, but duplicating the entire video backbone doubles GPU memory consumption and training overhead. MVI2V resolves this via Low-Rank Adaptation (LoRA), augmenting original projection matrices with paired low-rank factorizations: $\(W'_S = W_S + \alpha A_S B_S, \quad W'_C = W_C + \alpha A_C B_C, \quad W'_F = W_F + \alpha A_F B_F\)$ where \(W_S, W_C, W_F\) represent pretrained weights for self-attention, cross-attention, and FFN layers, \(A\) and \(B\) are down- and up-projection matrices, and \(\alpha\) denotes the scaling constant. The reference stream directly inherits pretrained visual representations while updating only low-rank deltas. On the 14B Wan2.1-I2V backbone, the LoRA layers add merely 0.3B parameters (~2.1% overhead), preserving the generative prior while ensuring rapid convergence.
4. Cross-Stream Bidirectional Self-Attention: Preventing feature interference across distinct modalities Concatenating reference tokens directly into the primary sequence (Token-Concat) causes severe modality entanglement: reference images are pristine latent features, whereas video tokens represent noise-perturbed states along the flow trajectory. Merging them prematurely degrades anatomical structure and creates motion artifacts. MVI2V encodes reference images into token sequence \(y\) and keeps cross-attention and FFN layers strictly segregated within each stream. The two modalities interact bidirectionally only within the self-attention module of each DiT block: $\([x_1, y_1] = \mathbf{SelfAttention}(x \oplus y; M_S, M'_S, W_S, W'_S)\)$ This design enables video tokens to globally query fine-grained texture details across all reference perspectives, while shielding clean reference tokens from temporal noise corruption.
Loss & Training¶
The framework is optimized under the Flow Matching objective. The video tensor is mapped to latent state \(x_1\) via a VAE encoder, and linearly combined with standard Gaussian noise \(x_0 \sim \mathcal{N}(0, I)\) at timestep \(t \in [0, 1]\) to obtain \(x_t = t x_1 + (1 - t) x_0\). The model predicts velocity field \(u\) against target \(v_t = x_1 - x_0\), minimizing the mean squared error conditioned on text \(c_{\text{txt}}\), first frame \(c_{\text{img}}\), and reference views \(c_{\text{ref}}\): $\(\mathcal{L} = \mathbb{E}_{x_0, x_1, c_{\text{txt}}, c_{\text{img}}, c_{\text{ref}}, t} \left\| u(x_t, c_{\text{txt}}, c_{\text{img}}, c_{\text{ref}}, t; \theta) - v_t \right\|^2\)$ During training, the base video stream weights are largely preserved, and only the reference LoRA layers and interaction projection matrices are actively optimized.
Key Experimental Results¶
Main Results¶
The method was evaluated on a curated benchmark of 300 e-commerce test cases covering multi-angle person and independent garment references. Metrics assess garment consistency (GPT-4 VQAScore-guided grading and DINO feature cosine similarity) and overall video quality (Subject Consistency, Background Consistency, and Motion Smoothness from VBench).
| Model | GPTperson | GPTgarment | DINOperson | DINOgarment | Subject Consistency | Background Consistency | Motion Smoothness |
|---|---|---|---|---|---|---|---|
| Wan2.1 (Baseline) | 0.752 | 0.409 | 0.667 | 0.582 | 0.916 | 0.928 | 0.986 |
| Token-Concat | 0.797 | 0.512 | 0.715 | 0.620 | 0.902 | 0.922 | 0.982 |
| Concat-ID | 0.841 | 0.542 | 0.719 | 0.613 | 0.918 | 0.936 | 0.988 |
| MVI2V-Wan2.1 (Ours) | 0.862 | 0.585 | 0.728 | 0.656 | 0.927 | 0.938 | 0.988 |
| In-House (Dual-stream Baseline) | 0.692 | 0.416 | 0.700 | 0.601 | 0.926 | 0.934 | 0.993 |
| MVI2V-In-House (Ours) | 0.868 | 0.607 | 0.776 | 0.633 | 0.930 | 0.934 | 0.993 |
Ablation Study¶
The incremental impact of each proposed component built on the Wan2.1 backbone is presented below:
| Config | GPTperson | GPTgarment | DINOperson | DINOgarment | Subject Consistency | Background Consistency | Motion Smoothness | Note |
|---|---|---|---|---|---|---|---|---|
| Wan2.1 Baseline | 0.752 | 0.409 | 0.667 | 0.582 | 0.916 | 0.928 | 0.986 | No multiview reference ability |
| + MVI2V Architecture (Random Selection) | 0.784 | 0.470 | 0.744 | 0.631 | 0.905 | 0.923 | 0.982 | Multi-stream added, but data lacks orientation diversity |
| + Reference Selection (Angular Clustering) | 0.853 | 0.562 | 0.743 | 0.638 | 0.916 | 0.932 | 0.986 | Complementary views yield substantial gains |
| + Inpainting Subtask (Full Model) | 0.862 | 0.585 | 0.728 | 0.656 | 0.927 | 0.938 | 0.988 | Masked first frame suppresses shortcut learning |
Furthermore, evaluation on the standard VBench I2V benchmark confirms that foundational single-image animation capabilities are preserved: MVI2V-Wan2.1 slightly improves Subject Consistency from 0.938 to 0.953 and Motion Smoothness from 0.980 to 0.987 over the original Wan2.1 baseline.
Key Findings¶
- Data orientation filtering drives the largest fidelity jump: Adding the curated reference selection strategy boosted GPTgarment by +0.092 (0.470 to 0.562). Redundant reference images resembling the initial frame contribute little, whereas orthogonal and rear angles provide crucial discriminative signals.
- Inpainting subtask cures texture abandonment: Without the masking objective, models frequently blur or miss complex back patterns during 180-degree turns. Randomly obscuring the first frame forces the network to route information through the LoRA reference stream, pushing GPTgarment to a peak of 0.585.
- Viewpoint saturation curve: Consistency gains are steepest when moving from 1 to 3 reference views. In practical deployments, providing just two complementary perspectives (front and back) covers the vast majority of human self-occlusion scenarios.
Highlights & Insights¶
- Modality-isolated multi-stream architecture: Segregating clean static references and noisy dynamic video latents into dedicated streams while restricting their cross-talk to self-attention effectively avoids anatomical distortion and preserves native motion priors.
- Inpainting as an anti-shortcut catalyst: Introducing random subject masking during training elegantly breaks first-frame over-reliance, providing an effective blueprint for multi-condition generative systems.
- Minimal parameter footprint: With an addition of only 2.1% trainable parameters via LoRA, the framework adapts flexibly to both single-stream (Wan2.1) and hybrid dual-stream (In-House) backbones with high engineering feasibility.
Limitations & Future Work¶
- Author-admitted limitations: Under extreme, rapid body deformations or non-rigid dynamic accessories (e.g., dangling ribbons, sheer fabrics, dynamic unbuttoning), subtle cross-frame texture drift can still occur.
- Pipeline dependencies: The curation pipeline relies on 2D segmentation models (Gemini and Sa2va) and HMR orientation estimation; segmentation artifacts in preprocessing could propagate into generation.
- Future directions: Integrating explicit 3D geometry anchors (such as SMPL-X or 3D Gaussian Splatting guidance) into the cross-attention mechanism could further anchor spatial texture correspondence under severe non-rigid motions.
Related Work & Insights¶
- vs Wan2.1 / CogVideoX (Standard I2V): Traditional I2V models infer unobserved viewpoints purely through unguided statistical hallucination; MVI2V introduces an explicit reference stream to ground unseen textures in physical ground truth.
- vs Animate Anyone / MimicMotion (Character Animation): Character animation frameworks rely on single frontal images with skeletal poses and lack multiview consistency mechanisms; MVI2V targets multiview identity and apparel fidelity across dynamic views.
- vs CatV2TON / MagicTryOn (Video Virtual Try-On): Video try-on models focus on clothing warping and garment transfer onto an existing model video; MVI2V synthesizes entirely new dynamic motions conditioned on multi-angle references and text prompts.
Rating¶
- Novelty: โญโญโญโญโ [Pioneers a systematic multi-stream architecture and inpainting objective tailored for multiview-consistent video generation, addressing a critical commercial pain point]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive comparisons against concatenation baselines and proprietary backbones across custom 300-case e-commerce benchmarks and standard VBench]
- Writing Quality: โญโญโญโญโญ [Clear structural organization, thorough motivation and methodology breakdown, self-consistent data, and detailed ablation analyses]
- Value: โญโญโญโญโญ [Highly relevant to digital human and e-commerce virtual showcase; lightweight design facilitates direct production integration]