Skip to content

LivingWorld: Interactive 4D World Generation with Environmental Dynamics

Conference: ECCV 2026
Paper: ECCV Official
Code: Project Page & Code
Area: 3D Vision
Keywords: Interactive 4D World Generation, Environmental Dynamics, Global Motion Fields, 3D Scene Flow Alignment, 4D Gaussian Splatting

TL;DR

To tackle the static nature of existing interactive 3D world outpainting, LivingWorld models continuous environmental dynamics via geometry-aware Kabsch alignment and a compact hash-based 3D motion field, achieving globally coherent 4D world generation with bidirectional motion propagation in just 12 seconds per step.

Background & Motivation

Recent breakthroughs in 3D Gaussian Splatting (3DGS) and generative diffusion models have enabled interactive large-scale 3D environment creation from single images or text prompts (e.g., WonderWorld, Text2Room, LucidDreamer). In these interactive systems, users can steer virtual cameras and trigger outpainting to expand explorable environments. However, environments synthesized by these approaches remain fundamentally static: they concentrate on extending geometric surfaces and appearance textures, leaving natural phenomena such as flowing rivers, ocean waves, drifting clouds, smoke, and flames completely frozen. In the physical world, environmental dynamics are not merely isolated moving objects; rather, they are continuous physical evolutions intrinsically coupled with the scene's large-scale geometry. Omitting these dynamics prevents environments from faithfully reflecting real-world behavior, which critically hampers applications in embodied intelligence, robotic perception, and interactive simulation.

Incorporating temporally coherent environmental dynamics into interactive world expansion introduces severe technical tensions between geometric consistency and real-time interactive feedback. Existing controllable video generation models (e.g., Veo 3.1, CogVideoX, Tora) can synthesize visually compelling dynamic videos, but they lack an explicit, persistent 3D world representation. Consequently, when the scene expands or is revisited under wide viewpoint trajectories, they inevitably suffer from multi-view geometric distortions and temporal drift. Conversely, dynamic 4D scene generation methods (e.g., 4DGS-Cinemagraphy, PerpetualWonder) rely heavily on iterative video-driven appearance reconstruction or physics engine simulators. These optimization pipelines incur massive computational overhead (often thousands of seconds per step), eliminating any possibility of interactive visual feedback. Furthermore, Lagrangian formulations that tie deformation parameters directly to individual Gaussian primitives face explosive optimization costs as the scene scales up, and prolonged integration causes particle drift that tears dynamic regions into unfillable holes.

The angle of attack in this work is to bypass iterative video-driven distillation and directly construct decoupled continuous Eulerian motion fields over explicit 3D geometry. Core idea: decouple environmental dynamics into a continuous global 3D motion field, resolve cross-view scene flow ambiguities via closed-form geometry-aware Kabsch alignment, and achieve low-latency interactive 4D world generation with seamless bidirectional motion propagation.

Method

Overall Architecture

Starting from a single input image, LivingWorld operates via a progressive two-stage interactive workflow: it first builds an initial 4D scene and then iteratively expands the 4D world guided by user-specified camera trajectories and outpainting text prompts. At each expansion step, user clicks or prompts produce a sparse dynamic region mask, from which a 2D Eulerian flow predictor estimates instantaneous velocities that are lifted into 3D using predicted metric depth. Next, a geometry-aware alignment module aligns the newly lifted sparse 3D scene flow with accumulated historical flows through rigid rotation and uniform scale calibration. A continuous global 3D motion field is then fitted using multi-resolution hash encoding and a lightweight MLP. Finally, during rendering, Gaussians are advected along bidirectional trajectories and blended via a time-dependent opacity scheduler, producing seamless, hole-free 4D dynamic sequences.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Single Input Image / Outpainted View"] --> B["2D Eulerian Motion Estimation<br/>Click mask + instantaneous flow prediction"]
    B --> C["Geometry-Aware Alignment Module<br/>3D scene flow lifting + closed-form Kabsch alignment"]
    C --> D["Hash-based Global Motion Field Learning<br/>Multi-resolution hash table + MLP continuous field regression"]
    D --> E["Bidirectional Motion Propagation & Rendering<br/>Forward/backward Euler advection + temporal opacity blending"]
    E --> F["Interactive 4D Dynamic Scene Output"]

Key Designs

1. Geometry-aware alignment module: resolving directional and scale ambiguities across views When independent 2D Eulerian flow estimators predict flow fields across different outpainted views and unproject them into 3D space using monocular depth, severe directional discrepancies and magnitude inconsistencies emerge in overlapping regions due to depth scale drift and perspective deformation. Prior approaches (such as 3D-MOM) optimize cross-view consistency through random initialization and multi-view 2D flow reprojection, which is prohibitively slow (hundreds of seconds) and prone to divergence in weakly overlapping regions. LivingWorld establishes spatial correspondences \(\mathcal{M}\) in 3D overlapping space, formulating the alignment between the newly lifted scene-flow samples \(\mathbf{S}_i\) and previously accumulated samples \(\mathbf{S}_{\text{prev}}\) directly as an optimization of optimal rotation \(\mathbf{R} \in \mathrm{SO}(3)\) and uniform scale \(s\):

\[\arg\min_{\mathbf{R}, s} \sum_{k \in \mathcal{M}} \left\| \mathbf{S}_{\text{prev}}^{(k)} - s \mathbf{R} \mathbf{S}_i^{(k)} \right\|_2^2\]

This problem is solved in closed form via the Kabsch algorithm using singular value decomposition (SVD) for rotation, alongside a 1D least-squares solution for scale \(s\). Because the Kabsch algorithm has linear computational complexity with respect to the number of point correspondences (\(\mathcal{O}(N)\)), coarse alignment executes within milliseconds. A lightweight gradient refinement step then cleans up residual discrepancies, delivering geometrically consistent 3D scene flow supervision within just 3 seconds.

2. Compact hash-based continuous motion field: primitive-decoupled Eulerian dynamics Conventional dynamic Gaussian Splatting methods predominantly adopt Lagrangian formulations, dedicating trajectory embeddings or deformation MLPs to individual Gaussians. This tightly couples optimization cost with the primitive count, making scaling across expansive environments computationally intractable. LivingWorld adopts a 3D Eulerian formulation, learning a continuous velocity field \(F_\theta: \mathbb{R}^3 \to \mathbb{R}^3\) independent of specific Gaussian primitives. To overcome the lack of regular grid structures in unstructured 3D space and the sparsity of flow supervision, the motion field is parameterized using multi-resolution spatial hash encoding combined with a lightweight MLP. For any coordinate \(\mathbf{x} \in \mathbb{R}^3\), spatial hashing maps it into hash tables of size \(T = 2^{19}\) using large prime numbers and XOR operations:

\[h_\ell(\mathbf{x}) = \left( \bigoplus_{j=1}^3 x_j^\ell \cdot \pi_j \right) \bmod T\]

Interpolated features across all resolution levels are concatenated and passed through a shallow MLP to regress the instantaneous 3D velocity \(F_\theta(\mathbf{x})\). This design allows any Gaussian to instantly query its velocity based on spatial coordinates, eliminating per-primitive backpropagation and naturally providing spatial smoothness and dynamic extrapolation to unseen neighboring regions.

3. Bidirectional motion propagation and opacity scheduling: preserving Gaussian density and coverage In the absence of multi-view video re-optimization, advecting Gaussians forward along an Eulerian velocity field inevitably causes primitives to drift downstream, leaving severe density holes and tearing artifacts at flow sources. LivingWorld introduces a symmetric bidirectional Euler integration scheme. For each Gaussian primitive \(g\) with center \(\mathbf{p}_g\) at time \(t-1\) and per-axis step size vector \(\boldsymbol{\psi}\), forward and backward trajectories are computed via:

\[\mathbf{p}_g^f(t) = \mathbf{p}_g^f(t-1) + \boldsymbol{\psi} \odot F_\theta\big(\mathbf{p}_g^f(t-1)\big)\]
\[\mathbf{p}_g^b(t) = \mathbf{p}_g^b(t-1) - \boldsymbol{\psi} \odot F_\theta\big(\mathbf{p}_g^b(t-1)\big)\]

During rendering, forward and time-reversed backward trajectories are merged with static primitives. To eliminate abrupt boundary transitions and form a closed-loop animation cycle, a bidirectional opacity scheduler modulates primitive opacities using linear temporal blending \(w(t) = t/T\):

\[\tilde{\alpha}_g^f(t) = \big(1 - w(t)\big) \alpha_g^f, \quad \tilde{\alpha}_g^b(t) = w(t) \alpha_g^b\]

As time progresses, the opacity contribution smoothly transitions from the forward to the backward trajectory, seamlessly looping back to the initial scene configuration within \(T\) steps and preserving complete Gaussian spatial coverage without density holes.

Loss & Training

The continuous global motion field is trained purely from the aligned sparse 3D scene flow samples \(\{(\mathbf{x}_i, \mathbf{s}_i)\}_{i=1}^N\) using a direct \(L_1\) regression loss:

\[\mathcal{L}_{\text{motion}} = \sum_{i=1}^N \|F_\theta(\mathbf{x}_i) - \mathbf{s}_i\|_1\]

No expensive photometric rendering losses are needed. On a single NVIDIA RTX 5090 GPU, geometric scene outpainting takes ~9 seconds, while 2D flow estimation, closed-form Kabsch alignment, and hash motion field training collectively take only 3 seconds, capping total expansion latency at approximately 12 seconds per step.

Key Experimental Results

Main Results

The evaluation benchmark consists of 60 natural scenes collected from Pexels and Unsplash across four environmental dynamics categories: clouds, water, smoke/fog, and fire (15 scenes each). Evaluations compare leading controllable video generation models and explicit 4D scene generation frameworks across VBench metrics (Imaging Quality, Aesthetic Quality, Motion Smoothness, Temporal Flickering), GPT-5.5-based physical plausibility (PhysReal), and per-step runtime.

Category Method VBench Imaging (↑) VBench Aesthetic (↑) VBench Motion (↑) VBench Flicker (↑) PhysReal (↑) Runtime (s) (↓)
Video Gen. Veo 3.1 0.694 0.625 0.992 0.979 0.622 140
Video Gen. CogVideoX 0.677 0.611 0.991 0.983 0.575 1510
Video Gen. Tora 0.649 0.609 0.992 0.976 0.571 550
4D Scene 4DGS-Cinemagraphy 0.637 0.604 0.996 0.988 0.605 1980
4D Scene PerpetualWonder 0.553 0.553 0.979 0.972 0.554 3580
4D Scene LivingWorld (Ours) 0.673 0.639 0.995 0.989 0.655 12

In the 2AFC human preference study, LivingWorld achieves substantial win rates against all baselines: 72%(±8.7) on Motion and 75%(±8.4) on Flicker versus Veo 3.1; over 68%–75% across all metrics versus 4DGS-Cinemagraphy; and 87%(±6.6) on Aesthetic and 92%(±5.4) on Flicker versus PerpetualWonder.

Ablation Study

Ablation experiments evaluate the contribution of the geometry-aware alignment module compared to naive scene flow accumulation and optimization-based 3D-MOM alignment. Metrics include local consistency (Motion Correlation Accuracy MCA, Flow Magnitude Variance FMV) and global consistency over distant points (Cosine similarity, Magnitude Ratio).

Alignment Configuration Runtime (s) (↓) Local MCA (↑) Local FMV (↓) Global Cosine (↑) Global Mag. Ratio (↑) Note
WonderWorld + Naive Scene Flow 9.3 0.0550 1.91 0.54 0.72 Unaligned flows cause explosive magnitude drift
WonderWorld + 3D-MOM Alignment 600.0 0.0597 1.66 0.66 0.76 Reprojection optimization is slow and distorts new areas
LivingWorld Alignment (Ours) 12.1 0.0742 0.29 0.91 0.84 Closed-form Kabsch + light refinement; fast and coherent

Ablations on motion field architecture confirm: 1. Without hash motion field (advecting Gaussians with static unprojected flow vectors): Primitives accumulate identical velocities, inflating displacement and producing severe visual tearing. 2. Without bidirectional propagation (forward-only advection): Primitives permanently drift downstream, leaving catastrophic density voids and unrendered black holes at flow origins.

Key Findings

  • The closed-form Kabsch alignment reduces alignment runtime from 600s (3D-MOM) to milliseconds, while elevating global flow cosine similarity from 0.66 to 0.91 and slashing flow magnitude variance (FMV) by 82.5% (from 1.66 down to 0.29).
  • LivingWorld attains the highest PhysReal score (0.655), surpassing top-tier generative video models like Veo 3.1 (0.622) and physics-guided PerpetualWonder (0.554), demonstrating that 3D Eulerian representations capture fluid-like environmental dynamics with superior physical fidelity.

Highlights & Insights

  • Formulating 3D cross-view flow alignment as a closed-form Kabsch problem: By recognizing that independent monocular flow errors primarily manifest as global spatial rotations and scale shifts, SVD-based rigid alignment replaces black-box backpropagation, resolving the core bottleneck of dynamic 4D scene expansion.
  • Symmetric bidirectional advection for fluid density preservation: Instead of incorporating particle emitter simulators to counteract Eulerian advection voids, the symmetric forward-backward loop with opacity decay achieves self-closing cyclical motion over static Gaussian geometry.
  • Native composability with localized object-centric motions: The global environmental field coexists seamlessly with localized rigid perturbations (e.g., swaying branches or waving flags), providing a dynamic substrate for interactive world creation.

Limitations & Future Work

  • Constrained support for complex non-rigid articulated motions: The framework is primarily tailored for fluid-like environmental phenomena and currently does not handle multi-joint articulated human bodies or complex rigid-body collisions.
  • Dependence on 2D flow and monocular depth quality: Failure cases in the 2D flow estimator (e.g., specular water surfaces or low-texture regions) or scale distortions in monocular depth estimation can propagate errors into Kabsch alignment.
  • Future directions: Integrating lightweight physical incompressibility and momentum conservation priors into the motion field loss, and coupling LivingWorld with embodied AI simulator platforms.
  • vs 3D-MOM / 4DGS-Cinemagraphy: 4DGS-Cinemagraphy optimizes dynamic Gaussians via per-frame multi-view reprojection taking thousands of seconds, and lacks global consistency in newly revealed areas. LivingWorld achieves 100× speedups and seamless expansion via closed-form 3D Kabsch alignment.
  • vs PerpetualWonder / WonderPlay: Physics-simulation hybrids require heavy MPM engines and diffusion video distillation, restricting scalability in open environments. LivingWorld extracts continuous 3D velocity fields directly from visual prompts with instant responsiveness.
  • vs Veo / CogVideoX / Tora: Pure video generative models lack persistent 3D geometry and multi-view consistency across expanded exploration trajectories, whereas LivingWorld provides an explicit, explorable 4D world.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Pioneers continuous global Eulerian motion fields and closed-form Kabsch alignment for interactive 4D scene expansion.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 60 scenes, VBench, GPT-5.5 PhysReal, human studies, and flow consistency ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear exposition, elegant mathematical formulation, and tight coupling between narrative and empirical validation.
  • Value: ⭐⭐⭐⭐⭐ Provides a fast, robust 4D environmental dynamics paradigm for embodied simulation, gaming, and virtual world creation.