DIVER: Disentangling Camera–Object and Active–Passive Motion for Video Generation¶
Conference: ECCV 2026
Paper: ECCV Official
Project: https://research.nvidia.com/labs/sil/projects/moright
Area: Video Generation
Keywords: video generation, motion disentanglement, camera control, causal reasoning, flow matching
TL;DR¶
A unified flow-matching video generation framework that decouples camera from object motion via a canonical static-view dual-stream architecture and models action-consequence causality through active-passive motion dropout, enabling arbitrary viewpoint exploration and bidirectional causal interaction reasoning.
Background & Motivation¶
Controllable video generation serves as a cornerstone for interactive visual reasoning, physical world simulators, and embodied AI. Users naturally seek to direct scene dynamics through intuitive inputs—such as sparse strokes or trajectories on an initial frame—while rendering the synthesized interaction under freely specified camera viewpoints. However, prevailing trajectory-conditioned video generators directly condition diffusion models on 2D pixel-space displacements. Under dynamic viewpoints, camera movements alter all pixel coordinates simultaneously, severely entangling camera motion with intrinsic object dynamics and forcing the generator to resolve ambiguous spatial assignments without explicit geometric guidance.
Beyond geometric entanglement, existing models treat user-specified trajectories merely as kinematic displacements, overlooking physical action-consequence dependencies (motion causality). In real-world interactions, active manipulation inevitably triggers cascading passive reactions—pushing a teapot causes it to slide, tilt, and pour water. Demanding dense trajectories across all interacting objects is impractical, while relying on external physics simulators or two-stage LLM planners introduces severe error accumulation, unnatural cross-modal translations, and rigid domain constraints.
This work builds upon a key physical insight: object motion is unambiguous and self-contained when expressed within a canonical static camera frame, and interactive physical causality can emerge directly from video data under asymmetric supervision. Core idea: construct a shared-weight dual-stream generation pipeline that anchors object trajectories in a canonical static view and transfers dynamics to arbitrary target camera poses via temporal cross-view self-attention, while leveraging active-passive motion decomposition and asymmetric dropout to learn forward consequence prediction and inverse action reasoning.
Method¶
Overall Architecture¶
The framework takes a single reference frame \(I\), a collection of canonical-view object trajectories \(\mathcal{T} = \{\tau_i\}_{i=1}^T\), and a target camera trajectory \(\{C_i\}_{i=1}^T\) to generate a dynamic video \(x^{\text{tar}}\) adhering to both controls. The architecture comprises a canonical stream and a target stream operating in parallel within a DiT backbone. The canonical stream conditions on an identity camera pose and reprojected trajectories, acting as an unperturbed dynamic anchor. The target stream conditions on warped camera pose features while keeping trajectory inputs empty. Both streams share DiT weights and exchange geometric and dynamic representations at each transformer block through temporal cross-view self-attention.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Single Reference Image + Target Camera Poses + Canonical Trajectories"] --> B["Condition Encoding & Feature Injection<br/>Gen3C-style camera warping + trajectory temporal embeddings"]
B --> C["Dual-Stream Generation Architecture<br/>Shared-weight canonical static stream & target dynamic stream"]
C --> D["Temporal Cross-View Self-Attention<br/>Progressive motion transfer from canonical anchor to target view"]
D --> E["Active–Passive Causal Reasoning<br/>Asymmetric motion dropout inducing forward/inverse interaction dynamics"]
E --> F["Output: Physically Coherent Multi-View Interaction Video"]
Key Designs¶
1. Canonical Dual-Stream View Disentanglement: Eliminating Camera-Motion Geometric Entanglement
To prevent viewpoint changes from corrupting user-drawn trajectories, the framework establishes a canonical coordinate system on the first frame. Using estimated per-frame depth maps \(D_i\) and camera poses \(C_i\) from foundation perception models (e.g., ViPE), any 2D trajectory point is unprojected into 3D world space and reprojected onto the initial reference frame: $\(\tau^{\text{can}}_i = \pi \bigl(K,\,C_0\,C_i^{-1}\,\pi^{-1}(K,\,\tau_i,\,D_i)\bigr)\)$ In the DiT backbone, the canonical stream receives conditions \(\mathbf{c}^{\text{can}} = \{I,\, C_1,\, \tau^{\text{can}}\}\), whereas the target stream receives \(\mathbf{c}^{\text{tar}} = \{I,\, \{C_i\},\, \emptyset\}\). Because the canonical camera is fixed to identity, this stream solely models object deformability and trajectory dynamics. During test-time sampling, the canonical stream serves as a virtual anchor; only the target latent representation is reconstructed by the VAE decoder \(\hat{x} = \mathcal{D}(\hat{z}^{\text{tar}}_0)\), yielding orthogonal control over viewpoints and object movements.
2. Condition Injection and Cross-View Self-Attention: Seamless Latent Dynamics Transfer
Camera and trajectory signals are translated into dense feature grids matching the latent video resolution. Camera poses warp the first frame using estimated depth and are encoded into \(z^{\text{cam}} \in \mathbb{R}^{\hat{T} \times \hat{H} \times \hat{W} \times d}\) via the VAE encoder. Trajectories are encoded into a continuous temporal-correspondence map \(e^{\text{trk}} \in \mathbb{R}^{\hat{T} \times \hat{H} \times \hat{W} \times d}\). Inside every transformer block, projected conditions are injected additively into the latent tokens: $\(\mathbf{f}^i \leftarrow \mathbf{f}^i + W_{\text{cam}} \mathbf{z}^{i,\text{cam}} + W_{\text{trk}} \mathbf{e}^{i,\text{trk}}, \quad i \in \{\text{can},\, \text{tar}\}\)$ Tokens from both streams are concatenated along the temporal axis and processed by joint self-attention: $\([\mathbf{f}^{\text{can}};\; \mathbf{f}^{\text{tar}}] := \text{SelfAttn}\bigl([\mathbf{f}^{\text{can}};\; \mathbf{f}^{\text{tar}}]\bigr)\)$ Target-view tokens directly attend to motion-conditioned canonical tokens, allowing the network to progressively propagate the canonical physical motion into the target perspective across layers.
3. Active–Passive Motion Decomposition & Dropout: Inducing Action–Consequence Causality
To enable genuine physical reaction rather than passive trajectory tracking, foreground motion is segmented into active motion \(\tau^{\text{act}}\) (the initiating action, e.g., a hand pushing) and passive motion \(\tau^{\text{pas}}\) (the resulting reaction, e.g., an object tumbling). During training, asymmetric motion dropout masks out one category with probability \(p\): $\(\tilde{\tau}_i := \begin{cases} \tau_i^{\text{act}}, & \xi < p \\ \tau_i^{\text{pas}}, & \text{otherwise} \end{cases}\)$ The model is supervised against the full unmasked video containing complete scene dynamics. This forces the model to synthesize the missing cause or effect, natively unlocking two complementary capabilities at test time: forward reasoning (predicting environmental consequences from specified active actions) and inverse reasoning (inferring plausible driving manipulations that yield an intended passive outcome).
Loss & Training¶
The framework is optimized under the Flow Matching objective. Given latent video \(z_0\) and Gaussian noise \(\epsilon \sim \mathcal{N}(0, I)\), the linear interpolation path \(z_t = (1-t)z_0 + t\epsilon\) is regressed by velocity model \(\mathcal{G}_\theta\): $\(\mathcal{L}_{\text{FM}} = \mathbb{E}_{t, z_0, \epsilon} \bigl[ \|\mathcal{G}_\theta(\mathbf{z}_t, t, \mathbf{c}) - (\epsilon - \mathbf{z}_0)\|^2 \bigr]\)$ Training uses a three-tier hybrid supervision strategy: 1. Paired Multi-View Synthetic Data: Curating static-camera segments from real videos and generating corresponding dynamic camera versions via a video-to-video diffusion model; 2. Single-View Real-World Hybrid Training: For static-camera real videos, duplicating the stream to train condition transfer; for dynamic real videos, computing loss exclusively on the target stream; 3. Multi-Granularity & Occlusion Dropout: Trajectories are randomly pooled over spatial patches to simulate coarse user strokes, and segments are randomly masked to ensure robustness against occlusions and tracking noise.
Key Experimental Results¶
Main Results¶
Evaluation is conducted on DynPose-100K (in-the-wild dynamic camera sequences) and Cooking (complex hand-object physical manipulation). Baselines include unguided Wan2.1, camera-controlled Gen3C, and trajectory-guided models MP, ATI, and WanMove. Notably, tracking baselines require full future-frame pixel trajectories across both foreground and background (Full info), whereas this method requires only first-frame reprojected trajectories and camera poses.
| Dataset | Method | Full info | PSNR ↑ | SSIM ↑ | Rot (°) ↓ | Trans ↓ | EPE ↓ |
|---|---|---|---|---|---|---|---|
| DynPose-100K | Wan2.1 | × | 11.23 | 0.435 | - | - | - |
| Gen3C* | × | 12.45 | 0.507 | 5.46 | 4.09 | - | |
| MP* | ✓ | 11.72 | 0.455 | 6.76 | 6.04 | 7.56 | |
| ATI | ✓ | 13.18 | 0.493 | 5.62 | 6.54 | 8.43 | |
| WanMove | ✓ | 13.91 | 0.521 | 4.12 | 3.56 | 8.05 | |
| Ours | × | 12.30 | 0.457 | 4.55 | 4.61 | 7.64 | |
| Cooking | Wan2.1 | × | 14.23 | 0.527 | - | - | - |
| Gen3C* | × | 15.37 | 0.613 | 1.97 | 10.03 | - | |
| MP* | ✓ | 15.68 | 0.564 | 2.50 | 12.24 | 4.25 | |
| ATI | ✓ | 15.93 | 0.582 | 4.25 | 16.94 | 5.87 | |
| WanMove | ✓ | 16.42 | 0.589 | 2.93 | 13.27 | 5.47 | |
| Ours | × | 16.44 | 0.594 | 2.16 | 10.11 | 4.27 |
*Note: * indicates authors' reimplementation; Full info specifies whether the model receives ground-truth future pixel tracks.
On physical commonsense (WISA) and interactive manipulation (Cooking) benchmarks, baselines receive full text prompts describing both actions and outcomes, whereas this method receives only the active action description:
| Dataset | Method | Full Prompt | FID ↓ | FVD ↓ | Physical Commonsense (PC) ↑ | Semantic Adherence (SA) ↑ |
|---|---|---|---|---|---|---|
| WISA | MP* | ✓ | 57.29 | 975.94 | 0.75 | 0.82 |
| ATI | ✓ | 69.80 | 990.82 | 0.75 | 0.83 | |
| WanMove | ✓ | 61.34 | 1088.23 | 0.73 | 0.83 | |
| Ours | × | 52.95 | 876.03 | 0.76 | 0.82 | |
| Cooking | MP* | ✓ | 43.49 | 759.53 | 0.87 | 0.89 |
| ATI | ✓ | 55.80 | 881.94 | 0.85 | 0.90 | |
| WanMove | ✓ | 53.51 | 882.90 | 0.84 | 0.87 | |
| Ours | × | 39.94 | 730.46 | 0.88 | 0.89 |
Ablation Study¶
Ablations on the Cooking benchmark examine architectural designs, causal training, data mixtures, and motion input granularities:
| Configuration | FID ↓ | FVD ↓ | PSNR ↑ | SSIM ↑ | Rot ↓ | Trans ↓ | EPE ↓ | PC ↑ | SA ↑ |
|---|---|---|---|---|---|---|---|---|---|
| Cascaded (2-stage) | 41.74 | 728.80 | 15.98 | 0.569 | 2.69 | 11.50 | 5.05 | 0.87 | 0.89 |
| w/o Fixed-View Branch | 51.17 | 997.83 | 14.15 | 0.515 | 3.36 | 14.57 | 14.30 | 0.87 | 0.89 |
| w/o Motion Reasoning | 44.04 | 784.19 | 15.55 | 0.562 | 2.88 | 12.49 | 5.05 | 0.87 | 0.88 |
| w/o Mixed Training | 41.94 | 808.96 | 16.29 | 0.583 | 2.22 | 12.80 | 4.09 | 0.87 | 0.89 |
| Ours (Coarse Input) | 39.83 | 725.88 | 16.45 | 0.594 | 2.21 | 10.98 | 4.37 | 0.88 | 0.88 |
| Ours (Passive Input) | 44.20 | 838.67 | 15.99 | 0.588 | 2.21 | 11.04 | 7.27 | 0.87 | 0.88 |
| Ours (Active Input - Full) | 39.94 | 730.46 | 16.44 | 0.594 | 2.16 | 10.11 | 4.27 | 0.88 | 0.89 |
Key Findings¶
- Crucial Role of Canonical Anchoring: Removing the canonical fixed-view stream (w/o fixed view) degrades EPE from 4.27 to 14.30 and severely impairs camera pose tracking, validating that canonical space is indispensable for resolving geometric ambiguity.
- Joint Optimization vs. Cascaded Stacking: The cascaded baseline accumulates rendering errors between stages, yielding higher control errors and worse perceptual fidelity compared to joint dual-stream denoising.
- Emergent Causal Dynamics: Active-passive decomposition and asymmetric dropout enhance physical commonsense (PC) and distribution realism (FVD/FID), demonstrating that the model learns physical interaction rules rather than simplistic texture dragging.
Highlights & Insights¶
- Geometric Elegance: By reprojecting 2D tracks onto the canonical first-frame plane and coordinating via cross-view attention, the system sidesteps cumbersome 3D neural fields and achieves lightweight, precise viewpoint-motion decoupling.
- Bidirectional Causality: First work to introduce bidirectional action-consequence reasoning in video diffusion, empowering the model to deduce downstream physical chain reactions or reconstruct plausible driving forces.
- Robust Granularity Flexibility: Smoothly transitions between fine-grained dense point tracks and coarse user-drawn sketches, offering practical usability for interactive UI design and human-in-the-loop animation.
Limitations & Future Work¶
- Degradation Under Severe Viewpoint Changes: Rapid camera trajectories and extensive occlusions can cause reprojected tracks to land out of bounds, introducing geometric blur and peripheral artifacts.
- Long-Horizon Object Disappearance: During intricate contact dynamics over extended horizons, unconstrained deformation can occasionally trigger physical hallucinations or disappearing objects.
- Sampling Latency: Denoising dual concatenated latent streams increases compute, requiring roughly 15 minutes per video on an A100 GPU and motivating future research into flow distillation.
Related Work & Insights¶
- vs Motion-Conditioned Baselines (MP, ATI, WanMove): Existing trajectory models require per-pixel tracks across all frames and entangle camera transformations; this method operates purely on first-frame canonical projections with decoupled camera poses.
- vs Physics-Grounded Simulators (PhysGen, WonderPlay): Simulation pipelines depend on rigid physics engines and explicit force parameters; this work learns unconstrained, multi-material physical causality directly from video data.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Decouples camera and object dynamics using canonical dual streams and pioneer bidirectional active-passive causality.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Validated across DynPose-100K, WISA, and Cooking, supported by extensive ablations and human evaluations.
- Writing Quality: ⭐⭐⭐⭐⭐ Problem motivation is exceptionally clear, technical explanations are mathematically rigorous and well-structured.
- Value: ⭐⭐⭐⭐⭐ Establishes a foundational paradigm for interactive visual simulation, world modeling, and embodied AI.