Skip to content

GeoV2V: Geometry-Grounded Video Diffusion Model for Driving Scene Generation

Conference: ECCV 2026
Paper: ECCV Official
Project Page: https://xiyuche.github.io/GeoV2V/
Area: Video Generation
Keywords: novel view synthesis, video diffusion model, autonomous driving, geometric priors, lane-shift extrapolation

TL;DR

Built upon the Wan 2.1 backbone, GeoV2V presents a geometry-grounded video-to-video diffusion framework and a synchronized multi-lane synthetic dataset Para4D, achieving high-fidelity, geometrically consistent, and dynamic lighting-preserving driving video synthesis under large lateral lane shifts without requiring 3D bounding box annotations.

Background & Motivation

Synthesizing driving videos from novel viewpoints along unseen trajectories is a foundational capability for autonomous driving simulation and closed-loop world modeling. In contrast to conventional viewpoint interpolation, simulating lane shifts or lateral maneuvers requires extrapolating several meters off the recorded path while maintaining the underlying 3D spatial layout and the temporal dynamics of surrounding traffic. Existing approaches generally fall into two paradigms: reconstruction-based methods and reconstruct-then-restore generative pipelines. Classic reconstruction methods such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) rely heavily on densely overlapping views; while they achieve high pixel accuracy on interpolation benchmarks, their visual fidelity degrades sharply under lateral extrapolation due to missing geometry and severe perspective occlusions. Conversely, reconstruct-then-restore methods (e.g., FreeVS and StreetCrafter) condition single-frame generative models on projected colored LiDAR point clouds; although individual frames look plausible, they struggle with long-range video-level consistency, resulting in structural jitter and dynamic hallucinations. Moreover, real-world driving environments feature intricate non-rigid lighting dynamics and optical media effectsโ€”such as turn-signal blinking, wet-ground specular reflections, and lens water dropletsโ€”which remain fundamentally difficult to model through explicit geometric primitives alone.

A primary bottleneck preventing models from mastering cross-lane transformations is the complete absence of synchronized multi-trajectory training supervision in real-world driving datasets. Popular benchmarks like Waymo and nuScenes record sequences along a single vehicle path, offering no simultaneous observations from parallel adjacent lanes within the same dynamic scene. Consequently, existing models are forced to rely on unaligned surrounding cameras or adjacent frames, introducing parallax errors and temporal mismatches that destabilize geometric learning.

This paper tackles the challenge by unifying high-quality multi-lane supervision with large video diffusion transformers. The authors introduce Para4D, a CARLA-based synthetic dataset capturing synchronized multi-lane driving streams with exact extrinsics and lateral shifts up to 4 meters. On top of this, they build GeoV2V using the 1.3B Wan 2.1 DiT architecture. Core idea: decouple single-frame dense depth reprojections and sparse LiDAR point clouds into hierarchical geometric priors injected across distinct DiT depths, while conditioning on full source video latents via RoPE frame concatenation to jointly preserve geometric structure and implicit complex lighting dynamics without 3D annotations.

Method

Overall Architecture

Given a source-view driving video \(V_{src} = \{I_t\}_{t=1}^T\) and a target camera extrinsic trajectory, GeoV2V directly synthesizes the novel-view video \(V_{tgt}\) observed from the alternative path. The pipeline comprises two cooperative components: a multi-modal 3D geometric prior preprocessing branch and a geometry-guided video-to-video diffusion transformer backbone. In the geometric branch, per-frame global 3D representations are constructed by fusing depth-completed reprojected point clouds with sparse LiDAR points, followed by a depth-aware morphological hole-filling operation on target-view projections. In the diffusion backbone, Wan 2.1 DiT takes full source-video latents concatenated along the frame dimension to seamlessly transfer long-range visual appearance and dynamics via native Rotary Position Embeddings (RoPE), while two lightweight ControlNets hierarchically inject dense depth and sparse LiDAR cues into early and late layers, respectively.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Source Video & Target Trajectory"] --> B["Multi-modal 3D Priors & Morphological Infilling<br/>Dense Depth Reprojection + Sparse LiDAR Separation"]
    B --> C["Spatio-Temporal Full-Video Conditioning<br/>RoPE Frame Concatenation + Camera Pose Embeddings"]
    C --> D["Hierarchical Decoupled Geometry Injection<br/>Dense Depth in First 15 Layers + Sparse LiDAR in Last 15"]
    D --> E["Parameter-Efficient Tuning & Two-Stage Training<br/>Frozen DiT FFN Layers + Independent LiDAR Branch Convergence"]
    E --> F["Output: Consistent Lane-Shifted Video"]

Key Designs

1. Multi-modal 3D Priors & Morphological Infilling: Balancing dense global coverage with fine structural precision

To equip the diffusion model with structural awareness without impairing its capacity to model non-rigid dynamics, GeoV2V constructs 3D representations strictly at the single-frame level, avoiding explicit inter-frame motion constraints and eliminating the need for manual 3D bounding boxes. Because LiDAR point clouds are inherently sparse, the pipeline estimates dense depth maps via Video Depth Anything and applies depth completion to form dense reprojected pseudo-images, keeping them separate from sparse LiDAR projections. However, projecting single-view data into laterally shifted views produces grid-like disocclusion holes and disparity cracks. Naive nearest-neighbor interpolation causes foreground colors to bleed into the distant background. To resolve this, the authors introduce a depth-constrained morphological filling strategy:

Given the binary projection validity mask \(M\), candidate void regions are identified via morphological closing \(\Omega = ((M \oplus \mathcal{S}) \ominus \mathcal{S}) \setminus M\). For each hole pixel \(u \in \Omega\), its Euclidean nearest valid pixel is determined by \(u_n = \arg\min_{v \in M} \|u - v\|_2\). The depth-consistent neighborhood is then defined as \(\mathcal{N}_d(u) = \{ v \in M \mid \|u - v\|_2 < r, |D(v) - D(u_n)| < \tau \}\), and the infilled color is obtained via spatial Gaussian distance weighting:

\[ \hat{I}(u) = \frac{\sum_{v \in \mathcal{N}_d(u)} w(u, v) I(v)}{\sum_{v \in \mathcal{N}_d(u)} w(u, v)}, \quad w(u, v) = \exp\left(-\frac{\|u - v\|_2^2}{\sigma^2}\right) \]

This formulation confines color propagation to surfaces with depth differences strictly within threshold \(\tau\), suppressing foreground boundary bleeding and providing clean, reliable structural priors.

2. Spatio-Temporal Full-Video Conditioning: Long-range dynamics transfer without explicit attention overhead

In novel-view driving simulation, the source-view video contains all essential temporal transitions, fine surface textures, and transient illumination, but lacks strict pixel-level correspondence with the shifted camera view. Adopting heavy cross-attention mechanisms across 3D video sequences would introduce prohibitive memory and compute overhead. Adopting the conditioning philosophy of ReCamMaster, GeoV2V injects source conditioning across temporal and spatial dimensions. Temporally, it leverages the Rotary Position Embedding (RoPE) built into Wan 2.1 by concatenating source latents and target noisy latents along the frame dimension. This design ensures that time-synchronized source-target frame pairs maintain an invariant relative distance in sequence space, allowing the standard self-attention blocks to retrieve appearance and dynamic motion across viewpoints without bespoke cross-attention layers. Spatially, per-frame camera extrinsics are encoded into vector embeddings and added directly to frame latents, granting the model explicit perspective awareness.

3. Hierarchical Decoupled Geometry Injection: Isolating coarse and fine guidance to prevent signal competition

Camera pose embeddings provide view direction but cannot prevent geometry collapsing during drastic lateral shifts. Thus, GeoV2V introduces both dense depth-completed images and sparse LiDAR projections as explicit guidance. Rather than naively fusing both representations into a single feature stream, the model adopts a hierarchical injection schedule: within the 30-layer DiT backbone, the first 15 layers receive dense depth priors, while the subsequent 15 layers receive sparse LiDAR priors. Guided by findings that early DiT layers govern global scene layout while later layers refine detailed textures, introducing dense depth first establishes scene layout and coarse camera perspective; following up with sparse LiDAR featuresโ€”encoded through a specialized ConvNeXt branchโ€”supervises sharp physical boundaries. Both control paths are parameterized as lightweight ControlNets (61M parameters each, negligible compared to the 1.3B backbone), preserving generative richness while strictly enforcing 3D geometric fidelity.

4. Parameter-Efficient Tuning & Two-Stage Training: Mitigating sensor-specific overfitting for robust transfer

To protect the foundational spatiotemporal priors acquired during large-scale video pretraining, the feed-forward network (FFN) layers across all DiT blocks remain completely frozen during downstream driving adaptation. Only the self-attention blocks and camera embedding modules are updated. Furthermore, because physical LiDAR scan patterns and beam geometries vary substantially across datasets (e.g., distinct beam counts and reflection profiles between Waymo and nuScenes), jointly training the entire network risks overfitting to a single sensor pattern. GeoV2V resolves this through a two-stage training scheme: the first stage optimizes the video backbone and dense depth control branch until convergence, establishing robust general perspective transformation capabilities. The second stage freezes the primary network and trains only the ConvNeXt encoder and sparse LiDAR ControlNet. This modular training enables the model to transfer zero-shot from CARLA synthetic data to real-world datasets with different sensor configurations.

Key Experimental Results

Main Results

Evaluation was conducted on Waymo, nuScenes, and the synthetic Para4D benchmark. For real-world datasets lacking multi-lane ground-truth videos, lateral lane-shift generation (at 1m, 2m, and 4m lateral offsets) was assessed using FID for frame quality and FVD for temporal video consistency. For Para4D, synchronized parallel ground-truth streams enabled full pixel-level evaluation across PSNR, SSIM, LPIPS, and FID.

Dataset Shift Offset Metric Ours Prev. SOTA Prev. SOTA Method Gain
Waymo (Real) 1m FID (โ†“) 17.53 18.69 StreetCrafter -1.16
Waymo (Real) 1m FVD (โ†“) 159.00 212.56 StreetGS -53.56
Waymo (Real) 2m FVD (โ†“) 232.06 272.85 StreetCrafter -40.79
Waymo (Real) 4m FVD (โ†“) 328.69 369.64 StreetCrafter -40.95
Para4D (Synthetic) 1m PSNR (โ†‘) 25.18 24.34 StreetGS +0.84 dB
Para4D (Synthetic) 1m LPIPS (โ†“) 0.040 0.172 StreetGS -0.132
Para4D (Synthetic) 2m PSNR (โ†‘) 23.88 23.00 StreetGS +0.88 dB
Para4D (Synthetic) 2m LPIPS (โ†“) 0.051 0.181 StreetGS -0.130
Para4D (Synthetic) 4m PSNR (โ†‘) 21.31 19.85 StreetGS +1.46 dB
Para4D (Synthetic) 4m LPIPS (โ†“) 0.080 0.291 StreetGS -0.211

Ablation Study

Ablations on core architectural choices evaluated on the Para4D benchmark under a 2m lateral shift:

Config PSNR (dB) SSIM LPIPS Note
Full priors (full model) 23.88 0.697 0.051 Complete pipeline with dense depth, sparse LiDAR, video conditioning & Para4D
w/o video (omit source video) 23.28 0.695 0.052 Small metric change, but qualitative vehicle structures distort from depth noise
w/o geometry (dense) 17.18 0.504 0.124 Removing dense depth drops PSNR by 6.70 dB; model fails to shift viewpoints
w/o geometry (sparse) 22.86 0.668 0.058 Removing LiDAR degrades fine geometric boundaries (-1.02 dB PSNR)
w/o geometry (both) 16.45 0.490 0.128 Complete geometric collapse without explicit spatial priors
w/o Para4D (trained on misaligned views) 20.92 0.593 0.102 Weakly aligned side cameras cause severe parallax mismatch (-2.96 dB PSNR)

Key Findings

  • Dense geometric guidance is vital for lateral extrapolation: Removing the dense depth prior (w/o geometry (dense)) triggers a catastrophic collapse in PSNR (from 23.88 to 17.18 dB) and more than doubles LPIPS. Qualitative samples reveal that without dense depth, camera pose embeddings alone fail to elicit correct perspective parallax shifts.
  • Synchronized multi-lane supervision provides crucial inductive bias: Replacing Para4D parallel supervision with unaligned surround-view camera configurations (w/o Para4D) leads to an immediate 2.96 dB drop in PSNR and degrades LPIPS from 0.051 to 0.102, demonstrating that accurate learning of lateral camera translation requires time-synchronized parallel ground truths.
  • End-to-end video diffusion naturally retains transient dynamics: Qualitative evaluations (Fig. 1, Fig. 2, and Fig. 7 in the paper) confirm that while explicit 3DGS and per-frame restoration methods erase dynamic turn signals, brake lights, and rain streaks on windshields, GeoV2V seamlessly preserves these temporal optical variations without explicit physical rendering terms.

Highlights & Insights

  • Hierarchical DiT prior separation: Leveraging the structural nature of DiT layersโ€”where coarse view structure crystallizes in early blocks while high-frequency textures settle in late blocksโ€”the decoupled 15/15 layer injection avoids conflicting multimodal supervision.
  • Zero-overhead cross-view condition retrieval: Concatenating source and target latents along the frame dimension under native RoPE positions allows standard attention blocks to handle cross-trajectory synthesis without introducing heavy cross-attention modules.
  • Effective synthetic-to-real zero-shot transfer: Even when trained exclusively on the synthetic Para4D dataset, GeoV2V achieves superior FVD and realistic dynamic preservation on real-world Waymo sequences, outperforming baselines directly fitted on real driving data.

Limitations & Future Work

  • Hallucinations under extreme lateral offsets: When lateral shifts exceed 4 meters, regions hidden entirely from the source camera must be hallucinated by the video diffusion prior, which occasionally introduces plausible but non-factual background objects.
  • Absence of metrics sensitive to dynamic optical transients: Standard frame metrics (PSNR, SSIM) are largely insensitive to transient lighting events like turn signals or windshield droplets. The authors emphasize the critical need to design quantitative benchmarks targeting transient optical realism in closed-loop simulation.
  • vs Explicit Reconstruction (StreetGS / EmerNeRF / OmniRe): Traditional methods mandate labor-intensive 3D bounding boxes to segregate static and dynamic scene primitives and fail under lateral extrapolation; GeoV2V operates without bounding boxes and models dynamic vehicles end-to-end within the video diffusion prior.
  • vs Reconstruct-then-Restore (FreeVS / StreetCrafter): Prior restoration approaches denoise individual frames conditioned on projected colored LiDAR, yielding high temporal flicker; GeoV2V preserves video continuity through full-video latent conditioning and hierarchical injection, slashing 4m lane-shift FVD on Waymo from 445.36 to 328.69.

Rating

  • Novelty: โญโญโญโญโญ [Presents hierarchical decoupled prior injection, full-video RoPE conditioning, and the first synchronized multi-lane 4D driving dataset]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive evaluations across Waymo, nuScenes, and Para4D covering both interpolation and 1m/2m/4m lane shifts with dynamic lighting ablations]
  • Writing Quality: โญโญโญโญโญ [Clear structural organization with rigorous mathematical formulation of depth-aware morphological infilling]
  • Value: โญโญโญโญโญ [Offers a robust, annotation-free, scalable generative foundation for autonomous driving simulation and world models]