Skip to content

UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation

Conference: ECCV 2026
Paper: ECCV Official Page
Area: Video Understanding
Keywords: Event-RGB fusion, future 4D scene generation, diffusion models, perceptual dynamics space, joint optical flow and depth

TL;DR

Addressing the lack of motion priors and susceptibility to motion blur in single-frame driving extrapolation, UniDynamics introduces a diffusion-based framework that leverages a single event-RGB pair to simultaneously generate future RGB video, 3D metric depth, and optical flow fields via event latent enhancement and a decoupled perceptual dynamics space.

Background & Motivation

In autonomous driving and robotics navigation, inferring the spatiotemporal evolution of surrounding environments from current observations is vital for downstream trajectory planning, obstacle avoidance, and risk-aware decision-making. Conventional generative driving world models heavily rely on continuous multi-frame histories paired with explicit control signals such as trajectories, velocities, or steering prompts to enforce physical controllability. However, in open-road deployment, consistent long-horizon video streams and structured kinematics are frequently unavailable. In practical single-frame prediction regimes, observing instantaneous velocity and motion vectors from a single static RGB frame is fundamentally ill-posed; moreover, high-speed camera motion introduces severe motion blur that degrades spatial textures, causing baseline video prediction models to suffer from catastrophic structure collapse and drift. Furthermore, previous predictive pipelines decouple static 3D geometry from temporal motion fields, lacking explicit dynamic physics to constrain spatiotemporal consistency across long extrapolation horizons.

Event cameras, with their microsecond-level temporal resolution and high dynamic range, continuously capture brightness changes triggered by relative motion with minimal latency. They preserve sharp moving edges and directional motion vectors even under extreme illumination conditions and high-speed motion blur where conventional RGB frames severely degrade. Consequently, event streams serve as an ideal, blur-resilient motion prior to complement single-frame RGB observations. In parallel, anticipating future environments requires harmonizing multi-view representations: realistic appearance rendering must stay consistent with underlying 3D metric depth and inter-frame optical flow fields.

The angle of attack in this work is to exploit asynchronous event streams as an alternative dynamic prior for single-image future prediction, while jointly modeling future appearance, 3D geometry, and motion fields within a unified latent domain. Core idea: encode asynchronous event streams into diffusion-injectable soft momentum priors via an Event Latent Enhancement (ELE) module, and embed a Perceptual Dynamics Space (PDS) within the multi-scale U-Net to decouple and bidirectionally interact geometric and motion latents, feeding them back into diffusion decoders via zero-convolutions for unified future 4D dynamic scene generation.

Method

Overall Architecture

UniDynamics builds upon the latent diffusion architecture of Stable Video Diffusion (SVD). Given a single RGB frame at timestamp \(T\) and the asynchronous event stream over the temporal window \(T-1 \to T\), the objective is to jointly forecast the future \(N\)-frame sequence of appearance RGB images, metric depth maps, and forward optical flow fields. To eliminate the computational overhead of training separate multi-modal autoencoders and ensure seamless latent alignment, the framework utilizes the latent space of a pretrained image VAE encoder as a shared representation domain. Diffusion denoising operates directly on image latents, while depth and flow latents are initialized from image latents and progressively updated and decoded through dedicated U-Net dynamics modules.

The pipeline comprises two foundational innovations: First, the event stream is converted into temporal voxel grids and processed by the Event Latent Enhancement (ELE) module with bidirectional cross-attention against RGB latents, producing an event latent that is injected into the diffusion input as a soft pseudo-historical state prior. Second, a Perceptual Dynamics Space (PDS) is embedded across the multi-scale stages of the U-Net backbone to decouple features into spatial depth and spatiotemporal flow subspaces; after gated bidirectional interaction, the resulting depth and flow representations are fed back into the spatial and temporal transformer layers of the U-Net decoder via zero-initialized convolutions, guiding the final prediction of future 4D scenes.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    InRGB["Single-Frame RGB Observation<br/>Frame at timestamp T"]
    InEvt["Asynchronous Event Stream<br/>Events over window T-1 to T"]
    InRGB --> ELE["Event Latent Enhancement (ELE)<br/>Bidirectional Cross-Attention & Fusion"]
    InEvt --> ELE
    ELE --> DiffIn["Diffusion Sequence Construction<br/>Soft Momentum Prior Injection"]
    InRGB --> DiffIn
    DiffIn --> UNet["SVD Spatiotemporal Diffusion U-Net<br/>Multi-scale Denoising Backbone"]
    UNet <--> PDS["Perceptual Dynamics Space (PDS)<br/>Decoupled Adapters & Gated Interaction"]
    PDS --> Feedbk["Dynamic Feedback Injection<br/>Depth to Spatial / Flow to Temporal"]
    Feedbk --> UNet
    UNet --> Out4D["Unified Shared VAE Decoding<br/>Future 4D Predictions (RGB + Depth + Flow)"]

Key Designs

1. Event Latent Enhancement: extracting diffusion-injectable soft momentum priors from asynchronous spikes

Raw event camera data consists of sparse, discrete, and asynchronous spike streams, which cannot be directly consumed by continuous latent diffusion networks. UniDynamics accumulates events within a \(100\,\text{ms}\) window and applies linear interpolation to discretize them into a voxel grid with \(B=15\) temporal bins, followed by a 2D residual encoder to produce event features. Because event signals degrade in stationary or textureless regions, the ELE module deploys bidirectional cross-attention (BiCA) layers, alternating event features and projected RGB latents as queries, keys, and values to achieve cross-modal semantic alignment. A subsequent convolutional fusion block and appearance-guided cross-attention step infuse RGB structural textures into the event representation, producing the enhanced latent \(Z(E)\). Critically, directly concatenating event latents along the temporal dimension as a rigid condition across all future frames causes severe transient bias, structural artifacts, and long-horizon drift. UniDynamics addresses this with a soft conditioning scheme: \(Z(E)\) is treated as the latent of a pseudo-historical state at \(T-1\), corrupted with noise jointly with future target frames during training and iteratively denoised. This formulation allows the U-Net to internalize event-driven initial momentum without corrupting pretrained temporal attention priors.

2. Decoupled Adapters with Gated Bi-interaction: disentangling geometry and motion in Perceptual Dynamics Space

Scene depth encapsulates static 3D geometric boundaries and surface occlusions, whereas optical flow delineates dynamic pixel correspondences and continuous temporal displacements. Forcing both modalities to share an identical feature space triggers gradient tug-of-war during multi-task learning, leading to structural artifacts where noisy flow vectors distort crisp object boundaries. PDS embeds dedicated adapters across U-Net stages (\(1\), \(1/2\), and \(1/4\) resolution scales), deploying 2D convolutions in the depth adapter to model spatial geometry \(F_d\) and 3D convolutions in the flow adapter to capture spatiotemporal trajectories \(F_f\). Recognizing that real-world motion boundaries are physically constrained by depth discontinuities, PDS incorporates a gated bidirectional interaction mechanism: concatenated features \(F_{fu}\) pass through dual convolutional branches with Sigmoid activations to produce normalized spatial gating maps, which modulate cross-modal residual updates:

\[\bar{Z}(D) = \text{Sigmoid}(\text{Conv}_f(F_{fu})) \odot F_f + F_d, \quad \bar{Z}(F) = \text{Sigmoid}(\text{Conv}_d(F_{fu})) \odot F_d + F_f\]

This formulation enables the flow branch to sharpen motion boundaries using geometric silhouettes, while the depth branch leverages instantaneous motion cues to segment moving foreground objects, achieving adaptive physical coordination without parameter entanglement.

3. Physics-guided Dynamic Feedback Injection: progressive multi-modal regularization via zero-convolutions

To ensure that the geometric and kinematic representations learned in PDS actively regularize visual synthesis, UniDynamics establishes a feedback loop routing PDS latents into the U-Net decoder. Depth latents, which contain structural boundaries and surface orientations, are injected directly into the spatial transformer layers of the U-Net. In contrast, optical flow latents, which govern frame-to-frame displacement and occlusion dynamics, are routed into the temporal transformer layers. To prevent unrefined geometry and motion latents from destroying the generative priors of pretrained SVD weights during early training steps, all injection channels utilize zero-initialized convolutions (ZeroConv):

\[Y_d = \text{ZeroConv}(X_d), \quad Y_f = \text{ZeroConv}(X_f)\]

At initialization, the outputs of the zero convolutions are strictly zero, preserving baseline diffusion capabilities. Throughout training, the network gradually strengthens the injection weights, providing a progressive transition from weak to strong physical constraints that force RGB appearance to remain spatiotemporally coherent and physically plausible.

Loss & Training

The entire pipeline is trained end-to-end using a multi-task composite loss spanning appearance generation, depth estimation, optical flow regression, and event latent grounding. For RGB, a latent denoising loss \(L(I)\) enforces reconstruction accuracy, dynamic enhancement, and structural fidelity. For depth, a latent MSE loss \(L(D)\) is combined with a scale- and shift-invariant loss \(L_{\text{SSI}}\) in pixel space. For optical flow, a latent MSE loss \(L(F)\) is combined with a pixel-level \(L_1\) endpoint loss. Furthermore, an MSE alignment loss \(L(E)\) between \(Z(E)\) and the actual image latent at timestamp \(T-1\) stabilizes ELE training. The overall objective function is formulated as:

\[L = L(I) + L(D) + L(F) + L(E) + \lambda_1 L_{\text{SSI}} + \lambda_2 L_1\]

Hyper-parameters are set to \(\lambda_1 = 2.0\) and \(\lambda_2 = 1.0\). The model is trained on 2 \(\times\) A100 GPUs for 16k steps using AdamW with an initial learning rate of \(5 \times 10^{-5}\) and batch size 1. A \(15\%\) random condition dropout is applied to enable classifier-free guidance, complemented by exponential moving average (EMA) weight smoothing and DeepSpeed ZeRO-2 optimization.

Key Experimental Results

Main Results

Quantitative evaluations are conducted on the synthetic autonomous driving benchmark E-VKitti2, evaluating the prediction of future \(N=9\) frames across video generation fidelity (FID, FVD), metric depth accuracy (AbsRel, \(\delta_n\) threshold accuracy, RMSE), and optical flow accuracy (EPE, AE, nPE outlier percentages). As detailed in the table below, UniDynamics outperforms task-specific baselines across all three modalities simultaneously.

Category Model FID ↓ FVD ↓ Depth AbsRel ↓ Depth \(\delta_1\) ↑ Flow EPE ↓ Flow AE ↓ Flow 1PE ↓
Generation Only SimVP 137.8 1301.7 Unsupported Unsupported Unsupported Unsupported Unsupported
Generation Only SVD 114.5 1244.4 Unsupported Unsupported Unsupported Unsupported Unsupported
Generation Only Vista 101.8 1364.6 Unsupported Unsupported Unsupported Unsupported Unsupported
Perception Only FLODCAST Unsupported Unsupported 0.57 21.4% 11.7 55.7 76.4%
Perception Only SVD-Depth Unsupported Unsupported 0.39 52.8% Unsupported Unsupported Unsupported
Perception Only SVD-Flow Unsupported Unsupported Unsupported Unsupported 8.5 40.3 88.8%
Joint Gen. & Single Perc. UniFuture-Depth 44.0 584.1 0.38 56.8% Unsupported Unsupported Unsupported
Joint Gen. & Single Perc. UniFuture-Flow 46.9 615.3 Unsupported Unsupported 5.9 33.8 78.8%
Unified 4D Gen. & Perc. UniDynamics (w/o events) 46.8 577.1 0.36 56.0% 5.3 34.1 79.2%
Unified 4D Gen. & Perc. UniDynamics (Full Model) 39.1 223.6 0.27 67.1% 3.0 21.0 58.1%

Ablation Study

Extensive ablation experiments on E-VKitti2 evaluate the individual contributions of event inputs, conditioning mechanisms, modality branches, and the PDS adapter architecture.

Configuration Variant FID ↓ FVD ↓ Depth AbsRel ↓ Depth \(\delta_1\) ↑ Flow EPE ↓ Flow AE ↓ Note
UniDynamics (w/o events) 46.8 577.1 0.36 56.0% 5.32 34.1 Loss of initial momentum leads to massive temporal degradation (FVD +353.5)
UniDynamics (hard in) 43.3 417.6 0.33 60.2% 4.63 28.2 Rigid channel concatenation induces conditioning bias and drift
UniDynamics-Depth (w/o flow) 42.0 264.2 0.27 68.6% Unsupported Unsupported Absence of explicit motion field degrades video FVD to 264.2
UniDynamics-Flow (w/o depth) 44.1 268.3 Unsupported Unsupported 3.64 24.8 Lack of geometric constraints elevates flow EPE to 3.64
UniDynamics (w/o adapters) 50.8 384.9 0.30 62.6% 4.60 30.3 Direct joint 3D convolutions trigger severe gradient interference
UniDynamics (Full Model) 39.1 223.6 0.27 67.1% 3.00 21.0 Optimal spatiotemporal and physical consistency across all metrics

Robustness to Motion Blur and Real-World Zero-Shot Transfer

  1. Motion Blur Robustness: Under simulated high-speed motion blur on E-VKitti2, single-frame RGB baselines suffer dramatic breakdowns (UniFuture-Depth scores FID 56.1, FVD 668.7, \(\delta_1\) 57.9%). In stark contrast, UniDynamics leverages blur-invariant event streams to achieve superior performance (FID 53.6, FVD 272.8, \(\delta_1\) 68.1%, and flow EPE 2.86 vs UniFuture-Flow's 5.67), establishing the event camera as an indispensable perceptual modality in dynamic settings.
  2. Zero-shot Cross-Dataset Generalization: When evaluated zero-shot on real-world datasets without fine-tuning, UniDynamics delivers strong temporal coherence on DSEC (\(FVD = 224.0\) vs UniFuture-Depth's 285.7) and halves optical flow error (\(EPE = 6.8\) vs 12.4). On the challenging outdoor benchmark M3ED, it achieves FID 68.5 and FVD 508.3 (outperforming UniFuture baselines of 88.6 and 705.4), proving that unified latent space modeling yields outstanding sim-to-real transferability.

Key Findings

  • Event streams unlock single-frame temporal extrapolation: Removing event inputs results in a catastrophic increase in FVD from 223.6 to 577.1. Without historical video context, the model cannot infer scene velocity vectors, making event streams the essential physical carrier for initial momentum.
  • Mutual synergy between geometry and motion: Omitting the depth branch increases optical flow EPE from 3.00 to 3.64, demonstrating that accurate motion estimation requires stable static 3D geometric boundaries as spatial anchors.
  • Decoupled interaction mitigates gradient conflict: Replacing task-specific adapters with naive joint 3D convolutions causes substantial degradation in both generation and perception metrics, confirming that decoupled feature modeling with gated interaction is necessary to prevent multi-task gradient tug-of-war.

Highlights & Insights

  • Treating event latents as soft pseudo-historical frames: Formulating event latents as historical frame representations at \(T-1\) within the joint diffusion process injects crucial momentum priors while avoiding the severe structural distortions caused by hard conditioning.
  • Perceptual Dynamics Space (PDS) for decoupled multi-task dynamics: Separating 2D spatial depth and 3D spatiotemporal flow while dynamically exchanging boundary features through gated cross-talk establishes an elegant bridge between diffusion priors and physical scene dynamics.
  • Shared latent representation across heterogeneous modalities: Deriving depth, optical flow, and visual appearance from a single weight-shared VAE latent space completely eliminates redundant encoders, minimizes GPU memory usage, and enforces intrinsic multi-modal alignment.

Limitations & Future Work

  • Sensitivity to raw event noise: In high-vibration scenarios or adverse weather (such as heavy rain or fog), real event sensors exhibit high-frequency background noise and hot pixels; explicit noise-filtering mechanisms for event streams remain to be integrated.
  • Implicit 4D representations requiring post-processing: The current framework outputs 2D depth and flow maps to infer 3D point clouds and scene flow; directly generating native 4D Gaussian Splatting (4D-GS) or continuous neural radiance fields remains an open challenge.
  • Future directions: Integrating predicted 4D dynamic fields directly into end-to-end driving planners to establish closed-loop predictive control and trajectory optimization.
  • vs UniFuture: While UniFuture pioneered single-frame driving forecasting, it strictly requires explicit velocity and trajectory controls and completely fails under high-speed motion blur; UniDynamics operates autonomously without control priors by fusing passive event streams.
  • vs Vista / SVD World Models: Conventional video world models synthesize only 2D RGB frames without explicit geometric or kinematic representations; UniDynamics introduces a unified 4D paradigm combining depth geometry and optical flow dynamics.
  • vs Event-RGB Perception Methods (e.g., DCEIFlow, SRFNet): Existing fusion approaches are limited to current-frame perception; UniDynamics elevates event-RGB fusion into generative future forecasting.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering framework for unified future 4D dynamic scene generation from a single event-RGB pair; soft momentum injection and PDS are conceptually innovative.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarks across synthetic and three real-world datasets (VKitti2, DSEC, MVSEC, M3ED) with rigorous motion-blur and ablation evaluations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear theoretical framing, rigorous mathematical formulations, and seamless alignment between architectural diagrams and methodological explanations.
  • Value: ⭐⭐⭐⭐⭐ Highly impactful for extreme-condition autonomous driving world models, motion deblurring, and embodied robotics scene forecasting.