MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Video Generation
Keywords: Multi-View Video Generation / Autoregressive Diffusion / Spatio-Temporal Self-Forcing / 4D Geometric Prior / Explicit 3D Reconstruction
TL;DR¶
Addressing geometric inconsistency and compounding exposure bias in long-horizon multi-view video synthesis of dynamic scenes, MV-Forcing unifies temporal and view-wise autoregression in a single few-step diffusion model, leveraging a recurrent 4D reconstruction model as an accumulated geometric bridge alongside spatio-temporal self-forcing distillation to eliminate multi-axis error drift.
Background & Motivation¶
Simultaneously generating long, temporally coherent, and 3D geometrically consistent video across multiple camera viewpoints is a foundational capability for immersive virtual reality, cinematic video production, and dynamic embodied simulation. However, jointly modeling high-fidelity video distributions across unbounded temporal durations while enforcing strict physical geometry across arbitrary camera trajectories presents a formidable technical challenge.
Recent advances in video diffusion have evolved along two orthogonal paradigms: on one hand, autoregressive video diffusion models have pushed temporal duration to minute-long horizons by conditioning on causal temporal context; on the other hand, multi-view models (such as SynCamMaster and CVD) achieve cross-camera synchronization for dynamic scenes using bidirectional attention across the entire time-view grid. Yet, these paradigms cannot be straightforwardly unified. Existing multi-view approaches rely heavily on dense all-to-all bidirectional attention across both temporal frames and viewpoints; this scales quadratically with respect to frame and view counts, completely preventing streaming inference and confining generation to short, fixed-length windows (e.g., 81 frames, 2 views). Conversely, directly stringing temporal autoregressive models into sequential view generation induces severe exposure bias; without an explicit geometric anchor, sequential view generation quickly drifts along the camera chain, causing catastrophic geometric tearing and structural degradation.
Confronting this core tension between the non-scalability of dense spatio-temporal attention and the severe geometric drift of ungrounded autoregression, this paper introduces a new formulation: decoupling time and view autoregression while adopting a feed-forward dynamic 3D reconstruction model as an online geometric bridge. The core idea is to compose temporal and view-sequential autoregression within a single causal few-step diffusion model, using an autoregressive 4D reconstruction model to accumulate persistent scene geometry and render target-view priors, optimized via spatio-temporal self-forcing distillation to close both temporal and view-wise exposure bias gaps for unbounded, geometrically consistent multi-view video synthesis.
Method¶
Overall Architecture¶
The core architecture of MV-Forcing integrates a causal few-step student generator, a recurrent 4D geometric bridge, and spatio-temporal self-forcing distillation. Given a text prompt \(P\) and a sequence of \(N\) camera trajectories \(\text{cam}_0, \dots, \text{cam}_{N-1}\), the system generates video sequentially across viewpoints and temporal blocks. When generating view \(k\), previously generated views are decoded to pixel space via a 3D VAE decoder and fed into the recurrent dynamic 3D reconstruction model CUT3R, which incrementally updates a persistent latent state \(\mathcal{S}\). Querying this state with the target camera trajectory renders a geometric prior and a per-pixel confidence map, which are projected via a zero-initialized 3D convolution layer and added as a residual to the noisy latents of view \(k\). The causal student employs causal temporal attention to maintain temporal continuity and multi-view synchronization (MVS) cross-view attention to attend to the clean preceding view. The entire model is supervised via Distribution Matching Distillation (DMD) with dual-axis self-forcing rollouts to ensure robustness to self-generated imperfections.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Text Prompt + Multi-View Camera Sequences"] --> B["View-Sequential Autoregression & Joint Denoising<br/>Clean preceding view condition + noisy current view input"]
B --> C["4D-Grounded Recurrent Geometric Prior Bridge<br/>CUT3R online state accumulation + raymap query rendering"]
C --> D["Spatio-Temporal Self-Forcing Distillation<br/>Dual-axis unrolled generation + bidirectional teacher DMD scoring"]
D --> E["Dual-Axis Free Traversal Streaming Inference<br/>Temporal KV cache progression + view-wise chained scaling"]
Key Designs¶
1. View-Sequential Autoregression & Joint Denoising: Bypassing Attention Bottlenecks Standard multi-view models must hold all views in memory simultaneously for bidirectional interaction, resulting in quadratic computation overheads. MV-Forcing proposes view-sequential autoregression, generating views \(k = 0, 1, \dots, N-1\) one by one. To synthesize view \(k\), the student initializes its latent from pure Gaussian noise \(z_k^{\tau_Q}\) and denoises it over a \(Q\)-step schedule, while the preceding view \(z_{k-1}\) is kept fully clean as conditioning. Cross-view synchronization is handled by lightweight MVS cross-view attention layers where spatial tokens in the noisy view attend to corresponding tokens in the clean preceding view, alongside blockwise causal temporal masking. To enable the same architecture to generate the initial view from text alone without cross-view conditioning, the training pipeline incorporates joint view denoising: with probability \(p\), both view slots are initialized from pure noise, training the network for independent text-to-video generation; with probability \(1-p\), one view slot is conditioned on a clean preceding view, training the view-sequential conditioning pathway.
2. 4D-Grounded Recurrent Geometric Prior Bridge: Anchoring Multi-View Consistency Relying solely on implicit cross-view attention over a sequential chain causes minor pixel errors to compound, leading to catastrophic 3D structural collapse. MV-Forcing introduces the feed-forward dynamic 3D reconstruction network CUT3R as an explicit geometric bridge. CUT3R maintains a persistent latent state \(\mathcal{S}\) that encodes the 4D geometry of the scene. Each time a video chunk is decoded into pixel space \(V_{k-1} = \mathcal{D}(z_{k-1})\), the state is updated recurrently: $$ \mathcal{S}' = \text{CUT3R_update}(\mathcal{S}, V_{k-1}) $$ When synthesizing view \(k\), rather than rendering explicit point clouds, a pixel-wise raymap \(\mathcal{R}_k\) is constructed from view \(k\)'s camera parameters to query state \(\mathcal{S}\), decoding a rendered RGB image \(\hat{V}_k\) alongside a per-pixel confidence map \(\hat{C}_k\). The rendered frames are encoded into VAE latents and concatenated channel-wise with the confidence map to form a conditioning tensor \(g_k \in \mathbb{R}^{T \times D' \times H \times W}\), which is added to the patch tokens via a zero-initialized 3D convolution layer: $$ \tilde{x}_k = x_k + \text{Conv3d}(g_k) $$ Zero-initialization prevents the geometric prior from disrupting initial diffusion features, allowing the model to smoothly learn reliance on geometry during distillation. Critically, because state accumulation spans all prior views and temporal blocks, synthesizing view \(k+1\) benefits from the global 4D geometry accumulated so far, effectively eliminating drift.
3. Spatio-Temporal Self-Forcing Distillation: Eliminating Dual-Axis Exposure Bias While view-sequential autoregression solves memory scaling, at test time the model conditions on its own imperfect historical predictions rather than clean ground truth. This exposure bias rapidly causes generation quality to degrade. MV-Forcing generalizes Self-Forcing into both temporal and view dimensions. During distillation training, autoregressive rollouts are executed along both axes: starting from the ground-truth first view \(z_0^{gt}\), the causal student sequentially unrolls generation to produce views \(\hat{z}_1\) and \(\hat{z}_2\) using its few-step schedule. A frozen bidirectional SynCamMaster teacher acts as the data distribution score function \(s_{\text{data}}\), scoring consecutive generated pairs \((\hat{z}_{k-1}, \hat{z}_k)\) through the Distribution Matching Distillation (DMD) loss: $$ \nabla_\phi \mathcal{L}{\text{DMD}} \approx -\mathbb{E}\tau \left[ \left( s_{\text{data}}(\hat{z}\tau, \tau) - s}}(\hat{z\tau, \tau) \right) \frac{\partial \hat{z}\tau}{\partial \phi} \right] $$ This objective forces the student to denoise and correct its own imperfect generation distribution, aligning training with inference. Furthermore, to bridge domain shifts for real-world scenarios where no native multi-view text-to-video teacher exists, MV-Forcing adopts ReCamMaster (a video-to-video re-rendering model) as the real-world score teacher, finetuning the student over limited iterations to achieve robust zero-shot open-domain generalization.
Loss & Training¶
The student generator is trained through a multi-stage curriculum: 1. ODE Trajectory Pre-alignment: The causal student is first trained on a small set of teacher ODE solution pairs to warm up the few-step generation trajectory; 2. Spatio-Temporal DMD Distillation: The base temporal layers (initialized from pretrained Self-Forcing) remain frozen, while MVS layers and the Conv3d projection are trained on the synthetic SynCamVideo dataset under spatio-temporal self-forcing unrolling; 3. Real-World Domain Adaptation: The student is finetuned for \(N_{ft}\) iterations on the Mixkit real-world video subset from Open-Sora using ReCamMaster as the score teacher (\(p=0\) since reference conditioning is available), while the initial view is generated via standard single-view Self-Forcing, delivering an end-to-end real-world multi-view pipeline.
Key Experimental Results¶
Main Results¶
The framework is rigorously evaluated across two benchmarks: a short sequence setting (2 views, 81 frames) compared against the bidirectional SynCamMaster teacher, and a long sequence setting (3 views, 162 frames) compared against composite baselines (SF+ReCamMaster and SF+ReCamMaster+SF) across both real-world (Real) and synthetic (Synth.) distributions.
1. Short Sequence Evaluation (2 views, 81 frames)
| Method | FID ↓ | FVD ↓ | CLIP-T ↑ | CLIP-F ↑ | RotErr ↓ | TransErr ↓ | Mat. Pix.(K) ↑ | FVD-V ↓ | CLIP-V ↑ |
|---|---|---|---|---|---|---|---|---|---|
| SynCamMaster (Teacher) | 166.57 | 1451.44 | 30.02 | 99.32 | 3.83 | 8.83 | 236.72 | 1697.37 | 90.13 |
| MV-Forcing (Ours) | 167.90 | 1468.84 | 29.67 | 99.21 | 3.64 | 8.26 | 251.13 | 1691.05 | 91.81 |
In the short sequence setting, MV-Forcing requires only a few sampling steps yet outperforms the bidirectional teacher across all camera accuracy metrics (RotErr drops to 3.64, TransErr drops to 8.26) and view synchronization metrics (matching pixels increase from 236.72K to 251.13K, CLIP-V improves to 91.81), demonstrating that explicit 4D geometric grounding provides stronger geometric alignment than implicit full self-attention.
2. Long Sequence Evaluation (3 views, 162 frames)
| Setting | Method | FID ↓ | FVD ↓ | CLIP-T ↑ | CLIP-F ↑ | RotErr ↓ | TransErr ↓ | Mat. Pix.(K) ↑ | FVD-V ↓ | CLIP-V ↑ |
|---|---|---|---|---|---|---|---|---|---|---|
| Real | SF+ReCamMaster | 157.57 | 1397.42 | 32.27 | 98.22 | 4.74 | 10.12 | 146.81 | 1759.82 | 87.26 |
| Real | SF+ReCamMaster+SF | 156.78 | 1363.23 | 32.48 | 98.67 | 4.89 | 10.61 | 127.48 | 1771.34 | 86.61 |
| Real | MV-Forcing (Ours) | 153.27 | 1309.68 | 32.74 | 99.03 | 3.88 | 8.78 | 239.37 | 1554.23 | 90.83 |
| Synth. | SF(ft)+ReCamMaster | 192.34 | 1576.83 | 29.07 | 98.03 | 4.32 | 10.27 | 143.77 | 1956.38 | 87.19 |
| Synth. | SF(ft)+ReCamMaster+SF | 191.26 | 1592.78 | 29.26 | 98.52 | 4.67 | 10.73 | 121.29 | 1987.43 | 86.27 |
| Synth. | MV-Forcing (Ours) | 186.73 | 1560.54 | 29.54 | 99.17 | 3.72 | 8.41 | 243.63 | 1729.35 | 89.88 |
Ablation Study¶
Ablations on synthetic data at 3 views and 162 frames examine key components: removing view-sequential self-forcing unrolling (w/o View Unrolling), removing the CUT3R prior entirely (w/o CUT3R), removing state accumulation across views (w/o Accumulation), and replacing learned raymap querying with classical z-buffer splatting (Manual Rendering).
| Config | FID ↓ | FVD ↓ | CLIP-T ↑ | CLIP-F ↑ | RotErr ↓ | TransErr ↓ | Mat. Pix.(K) ↑ | FVD-V ↓ | CLIP-V ↑ | Note |
|---|---|---|---|---|---|---|---|---|---|---|
| Full Model (Ours) | 186.73 | 1560.54 | 29.54 | 99.17 | 3.72 | 8.41 | 243.63 | 1729.35 | 89.88 | full model |
| w/o View Unrolling | 212.89 | 1791.40 | 27.74 | 97.66 | 4.72 | 10.45 | 159.73 | 2318.94 | 86.41 | Severe exposure bias causes catastrophic multi-view collapse |
| w/o CUT3R | 199.10 | 1613.21 | 28.87 | 98.21 | 4.56 | 9.83 | 173.75 | 2109.82 | 87.12 | Removing explicit 3D prior sharply degrades cross-view geometry |
| w/o Accumulation | 193.66 | 1576.84 | 29.26 | 98.78 | 3.97 | 9.04 | 229.63 | 1984.35 | 88.27 | Conditioning only on single preceding view loses global context |
| Manual Rendering | 191.14 | 1571.44 | 29.11 | 99.10 | 3.93 | 8.93 | 232.71 | 1963.42 | 88.36 | Forward splatting suffers from disocclusion holes and ghosting |
Key Findings¶
- View-sequential unrolling is critical for long view chains: Disabling view unrolling causes FID to surge from 186.73 to 212.89 and FVD-V to degrade from 1729.35 to 2318.94, confirming that training-inference distribution drift is the most catastrophic failure mode in sequential multi-view synthesis.
- Explicit 4D geometric grounding beats implicit attention: Omitting CUT3R drops matching pixels from 243.63K to 173.75K, verifying that implicit cross-view attention cannot guarantee strict 3D physical constraints over extended sequences.
- Robust scaling along view and time axes:
- View scaling: Increasing views from 2 to 5 at 81 frames exhibits almost negligible degradation (matching pixels remain steady from 251.13K to 250.95K, CLIP-V stays at 91.81 to 91.73), validating that persistent state accumulation prevents drift across progressive viewpoints.
- Temporal scaling: Extending sequence length from 81 frames to 648 frames (\(8\times\) horizon) at 2 views keeps cross-view geometric metrics virtually flat (RotErr stays at 3.64-3.67, TransErr at 8.26-8.29), with only minor drops in frame consistency CLIP-F (99.21 to 96.23) due to standard temporal autoregressive error accumulation.
Highlights & Insights¶
- Persistent dynamic 3D reconstruction as an online generative bridge: Rather than treating 3D reconstruction as an offline post-processing step, MV-Forcing embeds the feed-forward recurrent reconstruction model CUT3R directly into the diffusion loop, utilizing raymap rendering to decouple camera trajectory representation from rendering artifacts.
- Extending Self-Forcing into spatio-temporal dimensionality: The paper highlights that view-sequential autoregression suffers from an identical exposure bias problem as temporal autoregression; unrolling generation along both dimensions during distillation effectively resolves multi-axis drift.
- Flexible dual-axis traversal during inference: Because both temporal and spatial transitions rely solely on causal KV caches and accumulated geometric states, inference can freely advance along either the temporal axis or the view axis, offering unparalleled deployment flexibility.
Limitations & Future Work¶
- Pairwise supervision limits from the teacher model: Because the underlying SynCamMaster teacher model only denoises two views jointly, the DMD distillation loss supervises pairwise consistency rather than multi-view cycle closure, leaving potential closure discrepancies across full \(360^\circ\) orbits.
- Artifacts under extreme dynamic motions: CUT3R may produce uncertain or blurred confidence maps when encountering fast non-rigid deformations, mirror reflections, or large disocclusions, occasionally causing subtle texture flickering during diffusion refinement.
- Future directions: Incorporating multi-view cycle distillation across triplets or closed-loop camera trajectories, and integrating next-generation dynamic representations such as streaming 4D Gaussian Splatting into the persistent memory state.
Related Work & Insights¶
- vs SynCamMaster: SynCamMaster employs dense bidirectional spatio-temporal attention, which is effective for short clips but prohibits streaming inference and explodes quadratically; MV-Forcing distills it into a causal few-step student with an explicit 4D prior, outperforming the teacher in camera accuracy while scaling to 648+ frames and 5+ views.
- vs Self-Forcing: Self-Forcing mitigates exposure bias exclusively in single-view temporal autoregression; MV-Forcing is the first to extend the self-forcing principle to the spatio-temporal domain, conquering spatial drift across views.
- vs ReCamMaster: ReCamMaster operates as a video-to-video re-rendering system without text-to-video capability, and independent per-chunk re-rendering leads to temporal discontinuities; MV-Forcing utilizes ReCamMaster strictly as a real-world score teacher, enabling native text-to-long-multi-view generation.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneering integration of spatio-temporal autoregression with an online recurrent 4D reconstruction prior for multi-view video diffusion]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Comprehensive evaluations across synthetic and real datasets, scaling analyses up to 5 views and 648 frames, and detailed ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Exceptionally clear problem framing, mathematically rigorous formulation, and coherent narrative structure]
- Value: ⭐⭐⭐⭐⭐ [Establishes a highly practical and scalable foundation for 3D-consistent long video synthesis in VR, gaming, and robotic simulation]