Skip to content

DreamWorld: Geometry-Grounded Video Diffusion for 3D-Consistent World Modeling

Conference: ECCV 2026
Paper: ECCV 2026
Project: https://yanghb22-fdu.github.io/DreamWorld
Area: Video Generation
Keywords: World Modeling, Video Diffusion Model, 3D Geometry, Geometry Diffusion, Decoupled Generation

TL;DR

DreamWorld presents a geometry-grounded decoupled world modeling framework that distills rich spatial priors from 3D foundation models into a geometry video diffusion model to predict complete 3D structural representations, which subsequently guide an appearance video diffusion model to synthesize photorealistic, cross-view consistent 3D world exploration videos.

Background & Motivation

Camera-controlled video diffusion models (VDMs) have rapidly emerged as a promising paradigm for constructing interactive 3D world models, allowing users to navigate and explore dynamic scenes along customized camera trajectories across applications in embodied AI, virtual reality, and cinematic production. Nonetheless, most existing camera-guided video models rely heavily on implicit spatiotemporal latent representations or numeric extrinsic parameter embeddings without explicit 3D geometric grounding. Such geometry-agnostic modeling frequently leads to severe structural collapse, distorted object shapes, and prominent cross-view spatial inconsistencies, especially under large-angle camera rotations and substantial viewpoint translations.

Introducing explicit geometric conditions into video generation poses an intrinsic tension. While point cloud based back-projection warping provides accurate camera motion cues, raw warped renderings are inevitably plagued by substantial disocclusion holes, missing peripheral regions, and projection distortion artifacts. Directly feeding these degraded warped frames into a video diffusion model forces the network to simultaneously tackle two entangled objectives: inferring extensive missing geometric structures while hallucinating fine-grained textures and correcting projection noise. This task entanglement induces severe optimization ambiguity, often resulting in over-smoothed visual details or erratic geometric drift. Furthermore, while emerging 3D foundation models (e.g., Depth Anything 3 and VGGT) encapsulate robust multi-scale 3D structural and spatial layout priors across diverse viewpoints, they operate on geometric feature manifolds that are structurally and semantically disconnected from the compressed latent space of video Variational Autoencoders (VAEs), making direct integration challenging.

The key insight of this paper is to rethink the role of 3D geometry in video diffusion models: instead of treating geometry as an implicit byproduct of appearance generation, geometry is elevated to an explicit intermediate "structural pivot" through a decoupled "geometry-then-appearance" generative paradigm. Core idea: distill high-level structural priors from a pretrained 3D foundation model into a geometry video diffusion model to first predict complete and multi-view consistent 3D latent representations from partial warped observations, and then condition an appearance video diffusion model on this geometric scaffold to synthesize photorealistic novel views.

Method

Overall Architecture

DreamWorld establishes a two-stage decoupled generation pipeline. In the first stage, given an initial observation image \(I_0\) and a user-defined camera trajectory \(\pi\), point cloud forward warping generates an incomplete view sequence from which a 3D foundation model extracts multi-scale geometry features. A geometry video diffusion model is trained via feature-level flow matching to complete these sparse features into full, view-consistent 3D structural representations. In the second stage, an appearance video diffusion model takes the completed 3D structural features as physical scaffolding via channel concatenation with noisy VAE latents, synthesizing high-fidelity, geometrically consistent RGB frames guided by text prompts and the initial reference image.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Initial Observation I0<br/>& Camera Trajectory π"] --> B["Point Cloud Warping & Extraction<br/>Partial Views + 3D Foundation Model"]
    B --> C["Geometry Video Diffusion Model<br/>Feature-level Flow Matching Completion"]
    C --> D["Geometry Projection Module M<br/>Multi-scale Alignment & KL Regularization"]
    D --> E["Appearance Video Diffusion Model<br/>Channel Concatenation of Geometry Scaffold"]
    E --> F["High-Fidelity 3D-Consistent Video"]

Key Designs

1. Geometry Video Diffusion Model: Completing Structural Pivots on Latent Manifolds To circumvent the visual artifacts, severe disocclusions, and projection holes caused by direct pixel-level point cloud warping, DreamWorld conducts geometric completion purely in a compact structural latent space. The initial reference \(I_0\) is unprojected into a local point cloud and reprojected along the camera trajectory \(\pi\) into partial warped frames \(x_{warp}\). A frozen 3D foundation model (e.g., DA3) extracts multi-scale geometric features across network depths, which are mapped to incomplete geometry conditions \(z^{3d}_{warp}\). The geometry video diffusion model \(F_{geo}\), parameterized as a Diffusion Transformer (DiT), operates under the rectified flow formulation using the ground-truth video's 3D features \(z^{3d}_0\) as the regression target: $$ \mathcal{L}{geo} = \mathbb{E}}^{3d0, \epsilon, t} \left| v}(\mathbf{z}^{3dt, t, \mathbf{z}^{3d}_0) \right|_2^2 $$ By isolating structural completion from high-frequency RGB texture synthesis, the geometry diffusion branch focuses purely on learning spatial layouts and cross-view physical coherence, generating dense, artifact-free 3D scaffolds.}, I_0, \mathcal{T}) - (\epsilon - \mathbf{z}^{3d

2. Geometry Projection and Regularization Module: Bridging 3D Representations and VAE Space Multi-scale features from 3D foundation models possess high dimensionality and spatial characteristics that do not natively match video VAE latent distributions. DreamWorld designs a lightweight MLP-based projection module \(M(\cdot)\) that maps shallow surface normal details and deep semantic layout embeddings into compact diffusion-friendly tokens \(z^{3d}_0\). To align feature distributions and promote smooth diffusion interpolation, a Gaussian Kullback-Leibler (KL) divergence penalty is enforced during optimization: $$ \mathcal{L}{\text{KL}} = D)\right) $$ This projection module serves as a calibrated translation bridge, standardizing high-dimensional spatial priors into a compact conditioning manifold for generative networks.}}\left(\mathcal{N}(\mu(z^{3d}), \Sigma(z^{3d})) \parallel \mathcal{N}(0, \mathbf{I

3. Appearance Video Diffusion Model: High-Fidelity Rendering Grounded on Structural Anchors Once provided with predicted complete geometry features \(\hat{z}^{3d}_{tgt}\), the appearance video diffusion model \(F_{app}\) is liberated from the burden of inferring occluded depths or spatial arrangements. The projected geometry scaffold is concatenated directly along the channel dimension with noisy VAE latents \(z_t\), while the initial observation \(I_0\) and text prompt \(\mathcal{T}\) guide cross-attention layers. The appearance model optimizes a rectified flow matching objective regularized by the KL loss: $$ \mathcal{L}{app} = \mathbb{E}0, \epsilon, t} \left| v}(\mathbf{zt, t, \mathbf{z}^{3d}_0, I_0, \mathcal{T}) - (\epsilon - \mathbf{z}_0) \right|_2^2 + \lambda \mathcal{L} $$ To eliminate train-test distribution shift—where ground-truth 3D features are available during training but generated features are fed at inference—controlled Gaussian perturbations are injected into the geometry features during stage-one optimization, significantly improving the appearance network's robustness against minor geometric inaccuracies.}

Loss & Training

DreamWorld follows a progressive two-stage optimization strategy based on the Wan 2.1 video diffusion foundation. In the first stage, the 3D foundation model remains frozen while the geometry projection module \(M\) and the appearance video diffusion model \(F_{app}\) are jointly trained with small feature perturbations to establish the visual mapping. In the second stage, \(M\) is frozen and the geometry video diffusion model \(F_{geo}\) is trained to master feature-level distillation and structural completion. Both stages utilize the AdamW optimizer with a constant learning rate of \(5 \times 10^{-5}\), a global batch size of 32, and 15,000 optimization iterations, trained on curated 81-frame video clips at a spatial resolution of \(832 \times 480\).

Key Experimental Results

Main Results

On zero-shot novel view synthesis benchmarks RealEstate10K and Tanks-and-Temples across Easy and Hard splits, DreamWorld is compared against implicit camera parameter injection methods (MotionCtrl, CameraCtrl) and explicit warped view conditioning models (ViewCrafter, Gen3C), along with evaluations on the WorldScore benchmark.

Table 1: Quantitative comparison of single-image novel view synthesis on RealEstate10K and Tanks-and-Temples.

Dataset Difficulty Method PSNR ↑ SSIM ↑ LPIPS ↓ FID ↓ Rotation Error Rdist ↓ Translation Error Tdist ↓
RealEstate10K Easy Set MotionCtrl 14.12 0.538 0.413 88.13 4.372 8.778
RealEstate10K Easy Set CameraCtrl 19.09 0.706 0.325 42.18 1.254 2.796
RealEstate10K Easy Set ViewCrafter 19.21 0.719 0.315 39.75 0.713 2.086
RealEstate10K Easy Set Gen3C 20.73 0.726 0.279 33.97 0.556 1.582
RealEstate10K Easy Set DreamWorld (Ours) 23.15 0.781 0.249 25.18 0.301 1.112
RealEstate10K Hard Set MotionCtrl 13.05 0.495 0.453 97.66 7.116 9.484
RealEstate10K Hard Set CameraCtrl 17.15 0.554 0.395 75.76 1.875 3.104
RealEstate10K Hard Set ViewCrafter 17.76 0.677 0.376 52.86 1.161 2.682
RealEstate10K Hard Set Gen3C 18.04 0.681 0.334 46.67 0.838 2.347
RealEstate10K Hard Set DreamWorld (Ours) 20.94 0.739 0.267 31.16 0.521 1.318
Tanks-and-Temples Easy Set MotionCtrl 14.88 0.554 0.497 80.73 7.447 8.603
Tanks-and-Temples Easy Set CameraCtrl 15.79 0.587 0.396 71.31 1.656 2.913
Tanks-and-Temples Easy Set ViewCrafter 18.11 0.671 0.327 45.67 0.966 2.175
Tanks-and-Temples Easy Set Gen3C 18.95 0.683 0.308 41.95 0.884 1.677
Tanks-and-Temples Easy Set DreamWorld (Ours) 21.66 0.721 0.256 29.85 0.377 1.214
Tanks-and-Temples Hard Set MotionCtrl 13.45 0.524 0.514 89.55 8.731 9.402
Tanks-and-Temples Hard Set CameraCtrl 14.01 0.532 0.493 82.01 1.975 3.124
Tanks-and-Temples Hard Set ViewCrafter 17.41 0.645 0.353 57.76 1.437 2.891
Tanks-and-Temples Hard Set Gen3C 17.95 0.659 0.331 48.42 1.266 2.510
Tanks-and-Temples Hard Set DreamWorld (Ours) 19.59 0.704 0.283 36.95 0.584 1.337

Table 2: Quantitative comparison on the WorldScore benchmark.

Method 3D Consist. ↑ Photo Consist. ↑ Style Consist. ↑ Cam. Ctrl. ↑ Obj. Ctrl. ↑ Cont. Align. ↑ Subj. Qual. ↑ Avg. ↑
WonderWorld 82.85 67.86 55.79 92.32 47.63 79.09 69.03 70.65
AETHER 79.84 58.68 72.09 57.44 52.26 28.06 41.11 55.64
Uni3C 78.59 85.48 88.32 62.94 45.83 47.40 57.00 66.51
Voyager 56.00 80.68 72.89 45.92 57.69 48.36 44.74 58.04
FantasyWorld 83.31 86.11 94.22 57.05 34.46 38.45 57.40 64.43
DreamWorld (Ours) 84.96 92.70 92.98 81.97 63.11 50.37 59.17 75.04

Ablation Study

Ablations on RealEstate10K Easy Set examine condition space choices and the impact of the decoupled training strategy.

Table 3: Ablation study on core design choices in DreamWorld (RealEstate10K Easy Set).

Config Condition Space (VAE / 3D) Geometry & Appearance (Joint / Decoupled) PSNR ↑ SSIM ↑ LPIPS ↓ Note
Baseline VAE Latent Geometry-agnostic / Direct Conditioning 19.47 0.721 0.287 Warped views encoded into VAE latents directly
w/o Geometry Diffusion 3D Geometry Raw Incomplete 3D Feature Conditioning 20.05 0.724 0.283 Bypasses diffusion completion; uses raw sparse features
w/o Geo-App Decoupled 3D Geometry Joint End-to-End Multitask Optimization 21.87 0.752 0.256 Single model jointly predicting 3D and RGB
DreamWorld (Full Model) 3D Geometry Decoupled: Geometry-then-Appearance 23.15 0.781 0.249 Complete geometry diffusion + appearance synthesis

Key Findings

  • Decoupled generation avoids optimization interference: Compared to joint multitask training (21.87 PSNR), the decoupled two-stage architecture delivers a +1.28 dB gain in PSNR, confirming that segregating structural reasoning from texture synthesis substantially eases network convergence.
  • Latent geometry diffusion is essential: Omitting the geometry diffusion module causes PSNR to drop sharply from 23.15 to 20.05 (-3.10 dB), indicating that explicitly completing disoccluded 3D features is critical for eliminating synthesis ambiguity.
  • Robustness under extreme camera transformations: On the Hard Set, DreamWorld outperforms the previous SOTA Gen3C by +2.90 dB PSNR on RealEstate10K and +1.64 dB on Tanks-and-Temples, while cutting rotation and translation errors by more than 30%, showing strong resilience across complex camera trajectories.

Highlights & Insights

  • Repurposing 3D Foundation Models as Generative Anchors: Rather than using 3D models solely for downstream geometry recovery, this work distills their rich structural priors into generative diffusion, instilling strict spatial awareness into video synthesis models.
  • Decoupling Geometry from Appearance: Eliminates the traditional burden of having a single diffusion model handle geometric reasoning and appearance rendering concurrently, resolving task entanglement.
  • Feature Perturbation for Inference Drift: Injecting subtle noise into geometry representations during stage-one appearance training prevents error accumulation from imperfect geometry completion during inference.

Limitations & Future Work

  • Static Scene Assumption: DreamWorld assumes rigid point cloud projection and static 3D environments; dynamic foreground objects and non-rigid physical deformations lead to projection ghosting and structural smearing.
  • Computational Overhead: The two-stage pipeline chains two sequential DiT diffusion models, doubling the denoising sampling steps and limiting real-time interactive simulation.
  • Future Directions: Exploring 4D spatiotemporal geometry feature representations and diffusion distillation techniques to enable real-time interactive world modeling.
  • vs CameraCtrl / MotionCtrl: Implicit parameter-conditioned models lack explicit 3D grounding, resulting in severe warping and spatial drift under large view changes; DreamWorld anchors video synthesis on explicit 3D foundation model features.
  • vs ViewCrafter / Gen3C: Warping-based methods feed artifact-prone partial images directly into video diffusion, amplifying blurriness and distortion; DreamWorld conducts warping in latent geometry feature space and fills missing structures before rendering.
  • vs FantasyWorld: FantasyWorld models video and 3D objectives jointly within a unified network, creating task conflict; DreamWorld decouples geometry completion and appearance synthesis into two specialized stages, achieving superior consistency and photorealism.

Rating

  • Novelty: ⭐⭐⭐⭐☆ An elegant formulation distilling 3D foundation models into a geometry diffusion branch as intermediate structural pivots for world modeling.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluations across indoor and outdoor NVS benchmarks and the rigorous WorldScore suite with detailed ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear motivation, logical narrative, clean mathematical formulation, and consistent diagrams.
  • Value: ⭐⭐⭐⭐⭐ Establishes a solid benchmark for physical consistency in generative world models and embodied AI simulators.