GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/SenseTime-FVG/OpenDWM
Area: Video Generation
Keywords: Autonomous Driving Video Generation, Flow Matching, Geometry-Aligned Prior, Spatially-Adaptive Noise Injection, Few-Step Sampling
TL;DR¶
Addressing the high latency and spatiotemporal inconsistency of diffusion and flow matching models initialized from standard Gaussian noise, GeoFlow leverages multi-view metric geometry and spatially-adaptive noise injection to construct a geometry-aligned prior as the source distribution, straightening and shortening the transport path to surpass 40-step baseline quality in only 8 sampling steps.
Background & Motivation¶
The scalable development and closed-loop evaluation of autonomous driving systems heavily rely on massive, diverse, and photorealistic driving video data covering rare corner cases and safety-critical scenarios. Generative world models based on diffusion probabilistic models and flow matching (such as MagicDrive and DriveDreamer) have emerged as powerful engines for multi-view controllable simulation. However, these models inherently depend on iteratively solving ordinary differential equations (ODEs) or stochastic differential equations (SDEs), requiring dozens to hundreds of denoising steps during inference. This introduces steep computational latency and cost, severely hindering real-time interaction and data generation throughput.
A fundamental driver of this inefficiency lies in the ubiquitous assumption that the source distribution is standard Gaussian white noise. In autonomous driving contexts, video sequences exhibit strong physical determinism and spatiotemporal coherence: static backgrounds (roads, buildings, vegetation) dominate the field of view and evolve continuously under camera ego-motion, while dynamic traffic participants move along constrained trajectories. Forcing the generative model to regenerate deterministic scene structures already present in historical frames from pure noise is computationally redundant and, in few-step regimes, frequently leads to texture flickering, geometric drift, and visual collapse.
Rather than synthesizing future scenes from pure randomness, an intuitive direction is to initialize generation from a coarse geometric projection of past context. However, directly warping history frames via camera poses produces depth inaccuracies, occlusions, and misplaced dynamic objects, which lead to severe artifact overfitting if used as a rigid deterministic source. Core idea: construct a Geometry-Aligned Prior (GAP) distribution by combining multi-view metric depth warping with a spatially-adaptive continuous noise injection mask, preserving reliable static geometry while injecting stochastic fallback into occluded, dynamic, and uncertain regions to fundamentally straighten and shorten the optimal transport trajectory.
Method¶
Overall Architecture¶
GeoFlow operates within a conditional flow matching framework. Conditioned on a past reference frame \(I_{ref}\) and future control conditions \(C\) (including camera parameters \(C_{cam}\), dynamic 3D bounding boxes \(C_{obj}\), and road map \(C_{map}\)), the goal is to generate future multi-view video chunks \(V\). Instead of initializing from pure noise \(x_0 \sim \mathcal{N}(0, I)\), GeoFlow extracts geometry in the VAE latent space, warps the features to future target coordinates, and injects localized noise through a three-factor uncertainty mask to form the geometry-aligned prior \(\hat{x}_0\). The generative model then resolves the residual velocity field with only a handful of numerical ODE integration steps toward clean video latents \(x_1\).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Reference Frame & Control Conditions<br/>I_ref, C_cam, C_obj, C_map"] --> B["Latent Geometry Extraction & Feature Splatting<br/>Zref unprojected to 3D point cloud & rendered to Zwarp"]
B --> C["Spatially-Adaptive Noise Injection<br/>Blend occlusion/dynamic/uncertainty masks into GAP x0_hat"]
C --> D["Short-Horizon Autoregressive Flow Matching<br/>Few-step ODE integration to future video chunk"]
Key Designs¶
1. Latent Geometry Extraction and Feature Splatting: Building a Future-Aligned Coarse Canvas
To eliminate redundant regeneration of existing static backgrounds without incurring heavy pixel-space rendering costs, GeoFlow performs 3D geometric projection directly within the VAE latent space. A pretrained metric depth estimator predicts the metric depth map \(D_{ref}\) and uncertainty map \(M_{unc}\) from \(I_{ref}\), while the VAE encoder produces latent features \(Z_{ref} = \mathcal{E}(I_{ref})\). Using camera intrinsic parameters, \(Z_{ref}\) is unprojected into a 3D feature point cloud \(P_{ref}\). Given the relative pose transformation \(T_{rel} \in \mathrm{SE}(3)\) from vehicle planning or control inputs, points are transformed into future target coordinates: $\(P_{target} = T_{rel} \cdot P_{ref}\)$ When projecting 3D feature points back onto the 2D latent plane, multiple points often map onto the same pixel location. GeoFlow employs a Z-buffer feature splatting mechanism: for each target pixel coordinate \(u=(u,v)\), the warped feature \(Z_{warp}(u)\) is assigned the feature vector of the candidate point closest to the camera in camera depth, establishing a structurally coherent canvas for the future viewpoint.
2. Spatially-Adaptive Noise Injection: Mitigating Geometric Artifacts While Preserving Stochastic Diversity
Using deterministic warped features \(Z_{warp}\) directly as source distribution \(\hat{x}_0\) causes two major issues: flow matching models require stochasticity to explore the target data manifold and prevent inference error accumulation, and imperfect depth estimation produces distorted stretched textures and out-of-view voids. Uniform Gaussian noise injection fails because it indiscriminately corrupts well-aligned static regions. GeoFlow introduces spatially-adaptive noise injection governed by a continuous pixel-wise reliability mask \(M\), computed as the element-wise maximum across three geometric uncertainty sources: $\(M = \max(M_{occ}, M_{dyn}, M_{unc})\)$ Here, \(M_{occ}\) captures occluded and disoccluded out-of-view regions identified by the Z-buffer; \(M_{dyn}\) is derived by projecting 3D bounding boxes of dynamic agents to mask out regions where static warping improperly misplaces moving objects; and \(M_{unc}\) represents the normalized depth uncertainty map covering ambiguous regions such as sky or reflective surfaces. The Geometry-Aligned Prior source distribution is formulated via spatially modulated interpolation: $\(\hat{x}_0 = (1 - M) \odot Z_{warp} + M \odot \epsilon, \quad \epsilon \sim \mathcal{N}(0, I)\)$ Reliable static regions strictly retain the geometric prior, while invalid and dynamic areas smoothly fall back to Gaussian noise for generative inpainting.
3. Short-Horizon Autoregressive Adaptation and Flow Matching Optimization: Low-Cost Residual Refinement
To ensure valid visual overlap across frames during re-projection, GeoFlow adopts a short-horizon autoregressive chunking strategy with window size \(L=6\) (1 reference frame conditioning 5 generated frames). Training simplifies from generating from scratch to learning a residual refinement field. The intermediate states along the optimal transport path are sampled via linear interpolation \(x_t = (1-t)\hat{x}_0 + t x_1\), and network \(F_\theta\) is trained to predict the vector field matching analytical velocity \(v_t = x_1 - \hat{x}_0\): $\(\mathcal{L} = \mathbb{E}_{t, x_1, \hat{x}_0, C} \left\| F_\theta(x_t, t, C) - v_t \right\|^2\)$ Because baseline models already possess strong denoising and image restoration capabilities from pretraining, GeoFlow adapts efficiently with fewer than 10,000 fine-tuning iterations (costing under 30 H100 GPU hours), delivering rapid plug-and-play acceleration without architectural modifications.
Key Experimental Results¶
Main Results¶
Evaluated on the NuScenes validation set across 150 scenes generating six-view 16-frame driving videos (\(256 \times 448\) resolution), benchmarked against state-of-the-art models and the OpenDWM baseline:
| Method | Steps | FID โ | FVD โ |
|---|---|---|---|
| MagicDrive-V2 | 30 | 20.9 | 94.8 |
| DreamForge | 30 | 14.6 | 103.6 |
| Drive-WM | 50 | 15.2 | 122.7 |
| DriveDreamer-2 | - | 11.2 | 55.7 |
| UniMLVG (3 ref frames) | 50 | 5.8 | 36.1 |
| OpenDWM Baseline | 40 | 6.8 | 38.8 |
| GeoFlow (Ours) | 15 | 6.8 | 32.5 |
| GeoFlow (Ours) | 8 | 8.3 | 38.6 |
Across varied sampling steps, GeoFlow demonstrates consistent performance advantages over the OpenDWM baseline:
| Sampling Steps | OpenDWM FID โ | OpenDWM FVD โ | GeoFlow FID โ | GeoFlow FVD โ |
|---|---|---|---|---|
| 5-step | 20.4 | 123.5 | 11.7 | 49.2 |
| 8-step | 14.7 | 77.4 | 8.3 | 38.6 |
| 10-step | 12.6 | 66.1 | 7.5 | 35.0 |
| 15-step | 10.1 | 52.8 | 6.8 | 32.5 |
| 20-step | 8.7 | 45.1 | 6.8 | 32.6 |
| 40-step | 6.8 | 38.8 | 6.9 | 34.0 |
Across diverse base architectures, GeoFlow applied to OpenDWM-tvae, OpenDWM-vae, and UniMLVG yields relative FVD reductions of 62.5%, 59.6%, and 38.6% under 5-step generation, respectively.
Ablation Study¶
Ablation of each component in the Noise Injection Mask \(M\) under a fixed 10-step inference budget:
| Config | \(M_{occ}\) | \(M_{unc}\) | \(M_{dyn}\) | FID โ | FVD โ | Note |
|---|---|---|---|---|---|---|
| Naรฏve Warping | - | - | - | 8.2 | 43.6 | Overfits rendering artifacts and texture stretching |
| Occlusion Mask | โ | - | - | 7.9 | 40.0 | Inpaints out-of-view regions cleanly |
| Occlusion + Uncertainty | โ | โ | - | 7.6 | 40.2 | Refines ambiguous depth regions like sky |
| Full Model | โ | โ | โ | 7.5 | 35.0 | Eliminates dynamic ghosting, best temporal coherence |
Key Findings¶
- Straighter transport path yielding \(5\times\) step acceleration: GeoFlow achieves an FVD of 38.6 at only 8 sampling steps, outperforming the baseline at 40 steps (FVD 38.8), and reaches its peak FVD of 32.5 at 15 steps.
- Fast adaptation with negligible training overhead: Casting video synthesis as residual refinement allows GeoFlow to saturate within 10,000 iterations (under 30 H100 GPU hours), avoiding prohibitive distillation pipelines.
- Minimal geometric latency overhead: Profiling on an NVIDIA L20 GPU indicates that geometric reconstruction (0.43s, 2.84%) and feature rendering (0.49s, 3.24%) account for only ~6% of single-chunk latency, whereas ODE integration accounts for 91.65% (13.83s), resulting in a net end-to-end speedup of \(4.2\times\).
Highlights & Insights¶
- Orthogonal paradigm shift in generative acceleration: While existing acceleration predominantly focuses on solver redesign or model distillation, GeoFlow shifts the source distribution itself, directly shortening trajectory length while remaining compatible with existing solvers.
- Parameter-free adaptive error handling: Unifying occlusion geometry, 3D dynamic bounding boxes, and metric depth uncertainty into an unparameterized continuous mask prevents artifact propagation without adding trainable parameters.
- Zero-shot depth estimator interchangeability: Demonstrates strong robustness when swapping the depth estimator at test time to newer architectures (e.g., MapAnything-v1.1 or DepthAnything-3) without retraining.
Limitations & Future Work¶
- Reliance on camera overlap: The short-horizon chunking window (\(L=6\)) limits single-pass synthesis under sharp turns or extreme vehicle velocity where field-of-view overlap drops substantially.
- Absence of object-level motion priors: Dynamic object regions are reset to pure Gaussian noise rather than warped along estimated 3D object velocity vectors, which leaves dynamic agent motion solely to generative priors.
Related Work & Insights¶
- vs MagicDrive / DriveDreamer: Prior driving world models inject 3D geometry via attention mechanisms while initializing from pure Gaussian noise; GeoFlow shows that incorporating physical geometry directly into the source distribution provides dramatic efficiency gains.
- vs Video Bi-flow: Video Bi-flow pioneered non-Gaussian source initialization using noisy past frames, but lacks 3D camera geometry and multi-view warping; GeoFlow handles rigorous 3D spatial transformations and multi-source uncertainty explicitly.
Rating¶
- Novelty: โญโญโญโญโ (Redefines source distribution via geometry-aligned priors, offering an elegant acceleration path)
- Experimental Thoroughness: โญโญโญโญโญ (Exhaustive comparisons across sampling steps, base models, latency profiles, and mask ablations)
- Writing Quality: โญโญโญโญโญ (Clear logical narrative, comprehensive experimental support, and rigorous formulation)
- Value: โญโญโญโญโญ (Highly practical and plug-and-play for real-time autonomous driving simulation and data generation)