Novel View Synthesis as Video Completion¶
Conference: ECCV 2026
Paper: ECCV Official
Project: https://frame-crafter.github.io
Area: 3D Vision
Keywords: Novel View Synthesis / Video Diffusion Models / Permutation Invariance / Plücker Ray / LoRA Fine-Tuning
TL;DR¶
FrameCrafter casts sparse novel view synthesis as low-frame-rate video completion, converting pretrained video diffusion models into strictly permutation-invariant multi-view generators via per-view independent VAE encoding, query-centered Plücker ray conditioning, and temporal RoPE removal—outperforming image diffusion baselines trained on massive multi-view datasets using only 1K scenes for fine-tuning.
Background & Motivation¶
Sparse novel view synthesis (NVS) requires synthesizing photorealistic imagery from a novel target camera pose given only a small, sparse collection of input views (e.g., \(K \approx 5\)) with calibrated camera parameters. Classical feed-forward regression methods based on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) rely heavily on dense multi-view geometric overlap; when given sparse observations or wide baselines, they frequently suffer from severe blurring, floaters, and incomplete geometry. Recently, generative approaches such as SEVA and EscherNet leverage pretrained 2D image diffusion priors to hallucinate unobserved regions. However, because models trained purely on single 2D images fundamentally lack multi-view spatial awareness, they require immense pose-annotated multi-view datasets (such as hundreds of thousands of synthetic 3D assets and 80K real-world scenes) to learn cross-view spatial consistency from scratch.
This paradigm overlooks a far more natural and scalable source of 3D priors: web-scale video data. Natural video sequences continuously capture camera transitions and smooth viewpoint shifts over time; as a result, large-scale video diffusion models (such as Wan2.1 and CogVideoX) implicitly develop rich 3D structural awareness and spatio-temporal consistency during pretraining. Nevertheless, directly repurposing video diffusion models for sparse NVS faces an architectural roadblock: modern video architectures are fundamentally engineered around temporal causality. Their 3D VAEs apply causal temporal striding (typically 4:1 compression) and feature caching that enforce strict frame ordering, while the Diffusion Transformer (DiT) uses 3D Rotary Positional Embeddings (RoPE) that index frames chronologically. In contrast, sparse NVS inputs form an unordered set with large baselines and no intrinsic temporal order; running them through a standard causal video pipeline severely degrades spatial fidelity through temporal blending and causes predictions to vary wildly with input permutations.
To reconcile video priors with unordered multi-view synthesis, this paper avoids training dedicated multi-view architectures from scratch. Instead, it introduces minimal architectural modifications that allow a pretrained video model to forget temporal sequence ordering. Core idea: formulate sparse novel view synthesis as low-frame-rate video completion, using per-view independent VAE encoding to eliminate causal compression, query-centered Plücker ray conditioning via pixel unshuffle for lossless geometric injection, and temporal RoPE removal to build a strictly permutation-invariant multi-view generator.
Method¶
Overall Architecture¶
The FrameCrafter pipeline comprises a frozen video VAE encoder-decoder and a video Diffusion Transformer (DiT) adapted with Low-Rank Adaptation (LoRA). Given \(K\) sparse, unordered input images \(\{ \mathbf{I}_k \}_{k=1}^K\) with known poses \(\{ \boldsymbol{\pi}_k \}_{k=1}^K\) and a target query pose \(\boldsymbol{\pi}_{\text{tgt}}\), the model first encodes each view independently into the latent space while mapping camera parameters to normalized Plücker ray maps. Context latents with binary masks and ray maps are concatenated along channels and fed into the DiT, where self-attention aggregates cross-view geometry to predict the velocity field under flow matching. Finally, the denoised target latent is decoded back to image space using the independent single-frame VAE decoder.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
IN["Sparse Unordered Input Images + Target Camera Pose"] --> D1["Per-View Independent VAE Encoding<br/>Single-frame decoupled forward pass, blocking causal compression and feature caching"]
D1 --> D2["Query-Centered Plücker Ray Conditioning<br/>Target view as world origin, lossless channel injection via Pixel Unshuffle"]
D2 --> D3["Zero-Temporal Positional Encoding & Target-Only Supervision<br/>Zero out temporal RoPE to remove order bias, supervised solely on target flow matching"]
D3 --> OUT["Predicted Novel View Latent & Independent Decoding<br/>Output high-fidelity, geometrically consistent target view image"]
Key Designs¶
1. Per-View Independent VAE Encoding: Eliminating Causal Compression and View Ordering Bias Standard 3D causal VAEs in modern video models (e.g., Wan2.1) encode the first frame independently while compressing every subsequent 4 frames via temporal strided 3D convolutions with feature caching. Passing sparse, wide-baseline input views through this pipeline blends distinct viewpoints into shared latent tokens, destroying fine-grained high-frequency spatial detail and introducing severe left-to-right temporal dependency. FrameCrafter completely decouples view encoding by processing each input view \(\mathbf{I}_k\) independently as a single-frame "video": $\(\mathbf{z}_k = \mathcal{E}(\mathbf{I}_k) \in \mathbb{R}^{d_z \times 1 \times h \times w}, \quad k = 1, \dots, K\)$ The query view is represented by encoding a zero placeholder \(\mathbf{z}_{\text{tgt}} = \mathcal{E}(\mathbf{0})\), and all latents are concatenated along the temporal dimension into \(\mathbf{z}_{\text{all}} = [\mathbf{z}_1; \dots; \mathbf{z}_K; \mathbf{z}_{\text{tgt}}] \in \mathbb{R}^{d_z \times (K+1) \times h \times w}\). At test time, the predicted target latent is decoded independently: \(\hat{\mathbf{I}}_{\text{tgt}} = \mathcal{D}(\mathbf{z}_{\text{tgt}})\). This architectural change ensures that latent representations are mathematically invariant to input permutation and maintain full spatial resolution without temporal compression artifacts.
2. Query-Centered Plücker Ray Conditioning: Lossless Geometry Injection and Coordinate Invariance Prior NVS methods typically align world coordinates to the first input camera, which inherently breaks permutation invariance whenever the input order is reshuffled. FrameCrafter instead anchors the world coordinate frame to the target query camera, transforming all input poses into this query-centric coordinate system and normalizing the mean camera distance to unit length. For each pixel \((u, v)\), back-projected normalized ray directions \(\hat{\mathbf{d}}\) and camera centers \(\mathbf{o}\) form a 6D Plücker coordinate \(\mathbf{p}(u, v) = [\hat{\mathbf{d}}; \mathbf{o} \times \hat{\mathbf{d}}]\). Rather than using bilinear interpolation that blurs sharp ray directions, the model applies Pixel Unshuffle with factor \(f_s\), folding spatial ray coordinates into \(6 f_s^2\) channels (e.g., \(6 \times 8^2 = 384\) channels) at latent resolution \(h \times w\). This preserves exact ray geometry and provides rich geometric alignment when concatenated with noisy latents \(\mathbf{z}_t\), context visual features, and binary target masks.
3. Zero-Temporal Positional Encoding & Target-Only Supervision: Enforcing Pure Geometric Reasoning Video DiT architectures standardly apply 3D Rotary Positional Embeddings (RoPE) along the time axis, which injects strong sequential order bias. To enforce strict permutation invariance across the \((K+1)\) views, FrameCrafter zeros out the temporal component of RoPE, treating all view tokens as an unordered set in self-attention. This forces the transformer to infer cross-view correspondence strictly from Plücker camera rays rather than temporal indices. Furthermore, during training, Flow Matching supervision is applied exclusively to the target view: $\(\mathcal{L} = \mathbb{E}_{t, \epsilon, \mathbf{z}_0} \left\| \epsilon_\theta(\mathbf{z}_t, t, \mathbf{c}) - (\epsilon - \mathbf{z}_0) \right\|_{2, \text{tgt}}^2\)$ where \(\mathbf{c}\) denotes all conditioning inputs. Supervising context frames dilutes the gradient signal by biasing the network toward trivial reconstruction of visible frames; target-only supervision concentrates the entire optimization capacity on synthesizing occluded and unobserved geometry.
Loss & Training¶
FrameCrafter adopts Wan2.1-I2V-14B as its primary backbone. LoRA modules with rank \(r=32\) are inserted into the attention and feed-forward layers of the DiT, and the patch embedding convolution is re-initialized and trained with full gradients to support the expanded input channels, while all other pretrained weights are frozen (training only \(\sim 1\%\) of total parameters). Training follows a two-stage resolution curriculum: initial warmup at \(192 \times 336\) followed by fine-tuning at \(480 \times 832\). The model is trained on only 1,000 scenes from DL3DV-10K using a probabilistic frame sampling scheme (80% uniform random sampling across the full clip and 20% local window sampling to balance wide-baseline extrapolation and close-view interpolation).
Key Experimental Results¶
Main Results¶
Under the standard 6-view sparse NVS setting, FrameCrafter is evaluated against feed-forward regression models (LVSM, E-RayZer) and state-of-the-art generative diffusion baselines (EscherNet, Aether, SEVA) across DL3DV-Benchmark and Mip-NeRF 360. All predictions are evaluated at a unified \(480 \times 480\) resolution.
| Dataset | Method | Backbone | #Scenes | PSNR ↑ | SSIM ↑ | LPIPS ↓ | DreamSim ↓ |
|---|---|---|---|---|---|---|---|
| DL3DV | LVSM | – | – | 17.09 | 0.478 | 0.333 | 0.204 |
| E-RayZer | – | – | 16.85 | 0.442 | 0.455 | 0.254 | |
| EscherNet | SD1.5 | 10K | 12.07 | 0.251 | 0.484 | 0.227 | |
| Aether | CogVideoX-5B | – | 12.66 | 0.258 | 0.469 | 0.140 | |
| SEVA | SD2.1 | 80K | 16.15 | 0.470 | 0.253 | 0.088 | |
| Ours | CogVideoX-5B | 1K | 13.65 | 0.272 | 0.389 | 0.113 | |
| Ours | Wan2.1-14B | 1K | 17.18 | 0.445 | 0.223 | 0.066 | |
| Mip360 | LVSM | – | – | 15.25 | 0.317 | 0.609 | 0.577 |
| E-RayZer | – | – | 16.56 | 0.343 | 0.621 | 0.340 | |
| EscherNet | SD1.5 | 10K | 11.14 | 0.126 | 0.540 | 0.315 | |
| Aether | CogVideoX-5B | – | 12.60 | 0.220 | 0.651 | 0.334 | |
| SEVA | SD2.1 | 80K | 14.59 | 0.294 | 0.372 | 0.137 | |
| Ours | CogVideoX-5B | 1K | 13.09 | 0.209 | 0.540 | 0.204 | |
| Ours | Wan2.1-14B | 1K | 15.64 | 0.279 | 0.365 | 0.111 |
Ablation Study¶
Ablation experiments conducted on DL3DV-Benchmark at \(192 \times 336\) resolution highlight the individual impact of each design component:
| Config | PSNR ↑ | SSIM ↑ | LPIPS ↓ | DreamSim ↓ | Note |
|---|---|---|---|---|---|
| Full model | 15.69 | 0.326 | 0.246 | 0.114 | Best overall perceptual quality and consistency |
| ✗ Per-View Encoding | 11.50 | 0.147 | 0.676 | 0.656 | Joint causal VAE collapses multi-view details (+0.43 LPIPS) |
| ✗ Pixel Unshuffle | 13.53 | 0.212 | 0.359 | 0.129 | Standard spatial downsampling degrades ray precision |
| ✗ Target-Only Supervision | 15.08 | 0.287 | 0.270 | 0.116 | Full-sequence loss dilutes query prediction focus |
| w/ PRoPE | 14.91 | 0.289 | 0.315 | 0.174 | Relative frustum attention requires heavier full retraining |
| w/ Original RoPE | 15.30 | 0.300 | 0.265 | 0.122 | Temporal order bias harms generalization |
Furthermore, assessing permutation invariance across 10 random orderings of input views demonstrates: - First-view coordinate normalization: PSNR \(15.67 \pm 0.4730\), LPIPS \(0.262 \pm 0.0203\) (high variance due to arbitrary origin choice); - With original temporal RoPE: PSNR \(15.61 \pm 0.3106\), LPIPS \(0.254 \pm 0.0212\); - Ours (Query-centered + Zero Temporal RoPE): PSNR \(16.21 \pm 0.0081\), LPIPS \(0.224 \pm 0.0003\) (near-zero standard deviation, demonstrating true mathematical invariance).
Key Findings¶
- Per-view independent encoding is non-negotiable: Replacing independent encoding with standard causal 3D VAE joint encoding causes DreamSim error to surge from 0.114 to 0.656 and drops PSNR by over 4 dB, proving that causal temporal compression destroys sparse-view geometric reasoning.
- Extreme data efficiency over image diffusion: In data scaling analyses, FrameCrafter trained on merely 20 scenes (PSNR \(\approx 14.2\) dB) already surpasses EscherNet trained on 10,000 scenes (12.07 dB), indicating an efficiency advantage of over \(500\times\) owing to rich video foundation priors.
- Permutation invariance acts as implicit combinatorial data augmentation: On temporally ordered evaluation sequences, the permutation-invariant model (PSNR 15.69 / LPIPS 0.246) clearly outperforms the model explicitly trained with ordered inputs (PSNR 15.04 / LPIPS 0.276). Treating inputs as an unordered set naturally exposes the model to \(K!\) equivalent input permutations during training.
Highlights & Insights¶
- Repurposing video foundation priors for 3D vision: Rather than spending vast resources annotating 3D pose data for image diffusion, this work demonstrates that web-scale video models already possess cross-view consistency, needing only lightweight adaptation.
- Clean and targeted surgery: Without heavy 3D geometric machinery (such as explicit point cloud rendering or epipolar cost volumes), FrameCrafter achieves permutation invariance via three clean steps: single-frame VAE decoupling, zeroing temporal RoPE, and query-centered Plücker ray maps with pixel unshuffle.
- Direct backbone scalability: Moving from CogVideoX-5B to Wan2.1-14B yields an immediate LPIPS improvement from 0.389 to 0.223 on DL3DV, confirming that the framework effortlessly inherits performance gains as video foundation models scale up.
Limitations & Future Work¶
- High inference latency of 14B DiT: Utilizing large video diffusion backbones incurs noticeable GPU memory and compute overhead, hindering real-time interactive rendering.
- Hallucinations under extreme sparsity or deep occlusions: When inputs collapse to 1–2 views with substantial occlusion, the diffusion model may generate plausible yet hallucinated textures that deviate from the ground truth scene geometry.
- Future directions: Integrating FrameCrafter with feed-forward 3D Gaussian Splatting by using synthesized multi-view latents to initialize and regularize explicit 3D Gaussians for real-time downstream rendering.
Related Work & Insights¶
- vs SEVA: SEVA adapts SD2.1 using 80K scene-level and 340K object-level samples; FrameCrafter adapts a video diffusion backbone on only 1K scenes and surpasses SEVA in perceptual fidelity (LPIPS 0.223 vs. 0.253), validating video pretraining as a superior 3D prior.
- vs Aether: Aether generates continuous trajectory video frames under temporal order; FrameCrafter completely decouples time to perform permutation-invariant completion over sparse, discrete views, avoiding causal video compression artifacts.
- vs EscherNet: EscherNet relies on first-view relative camera conditioning in image diffusion; FrameCrafter uses query-centered coordinates and lossless Plücker ray unshuffle to prevent coordinate drift.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Elegant reframing of sparse NVS as video completion while removing causal video inductive biases.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive cross-dataset benchmarks, rigorous permutation variance stress testing, and data scaling analyses.
- Writing Quality: ⭐⭐⭐⭐⭐ Lucid motivation, coherent architectural narrative, and well-supported empirical claims.
- Value: ⭐⭐⭐⭐⭐ Unlocks a highly scalable, data-efficient blueprint for leveraging video foundation models in 3D vision.