Skip to content

LooseControlVideo: Directorial Video Control using Spatial Blocking

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Video Generation
Keywords: Video Generation & Editing, Directorial Control, Spatial Blocking Proxy, Oriented 3D Bounding Boxes, DNOCS Representation

TL;DR

LooseControlVideo adopts the cinematic concept of spatial blocking by employing coarse, oriented 3D bounding boxes as sparse proxies rendered into depth-modulated normalized object coordinate space (DNOCS), allowing a frozen video DiT backbone to decouple directorial trajectory choreography from fine-grained morphological deformations and achieve high-precision 3D trajectory control and localized video editing.

Background & Motivation

Modern video diffusion models such as Wan 2.2, Sora, and Veo have demonstrated remarkable photorealism and physical understanding. However, orchestrating complex narratives involving intricate spatio-temporal synchronization and multi-object interactions remains severely constrained under current control interfaces. Existing control modalities force an impractical trade-off: natural language prompts are inherently too ambiguous to dictate exact physical trajectories in space and time, whereas dense per-frame structural guidance—such as dense depth maps or detailed skeletal motion sequences—is prohibitively labor-intensive for users to author from scratch for dynamic scenes involving deformable entities (e.g., an eagle stooping to catch a dodging rabbit).

The fundamental difficulty stems from the fact that traditional dense structural signals conflate two distinct control axes that ought to be decoupled: the high-level choreography (camera trajectory, spatial layout, macro-motion, and object collision paths) versus the low-level execution (fine-grained articulation, muscular deformation, and secondary dynamics). Forcing human creators to author frame-accurate dense depth maps not only erects an insurmountable authoring barrier but also underutilizes the rich physical animation and deformation priors already mastered by large generative video backbones.

Drawing inspiration from the "blocking" phase in professional cinematography, where directors orchestrate scene layouts using coarse spatial proxies to define movement and pacing, the core idea of this work is to explicitly decouple macro-level choreography from micro-level geometric deformation by using sparse, oriented 3D bounding boxes as intuitive spatial blocking proxies, rendering them via virtual camera projection into a hybrid DNOCS (Depth-modulated Normalized Object Coordinate Space) representation that jointly encodes 6-DOF orientation and global depth ordering, thereby empowering a frozen video DiT to autonomously infer realistic deformations, mutual occlusions, and physical interactions.

Method

Overall Architecture

LooseControlVideo (LCV) is structured around a three-stage pipeline: lightweight 3D proxy authoring, image-space geometric projection with hybrid DNOCS shading, and conditional residual injection into a video DiT backbone. The system accepts as input a text prompt \(p\), a time-varying sequence of oriented 3D bounding boxes \(b = \{b_t\}_{t=1}^T\) (each characterized by center, scale, and 3D rotation), and a virtual camera trajectory \(c = \{c_t\}_{t=1}^T\). In video editing scenarios, a reference source video \(v\) can also be supplied.

Rather than passing raw 3D coordinates into the latent space via MLP tokenization—which forces 2D diffusion models to mentally reconstruct complex perspective projections and typically leads to training failure—LCV renders the 3D proxies onto a 2D control video \(v_{\text{ctrl}}\) using ray casting. This control video explicitly exposes perspective geometry, surface orientation, and depth ordering. Keeping the Wan 2.2 DiT backbone completely frozen, LCV introduces lightweight WAN-VACE condition modules tuned with LoRA, injecting spatio-temporal control features into the backbone to synthesize photorealistic videos adhering strictly to the directorial blocking.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["User Inputs<br/>Text Prompt + 3D Oriented Box Trajectory + Camera Path"] --> B["DNOCS Virtual Rendering<br/>Ray Casting + Orientation Hue / Inverse Depth Brightness"]
    B --> C["Multi-Modal Control Video Assembly<br/>Pure Generation / Segment Holding / Removal / Spatial Mix"]
    C --> D["Frozen DiT Backbone & VACE Conditioning<br/>Wan 2.2 Frozen + LoRA Fine-Tuned VACE Adapter"]
    D --> E["Generated Video<br/>Spatio-Temporal Alignment + Occlusion + Realistic Deformation"]

Key Designs

1. DNOCS Virtual Rendering: Decoupling 3D Projection and Jointly Encoding Orientation with Global Depth

A naive strategy of encoding 3D bounding box coordinates (center \(o_t\), scale \(s_t\), rotation \(R_t\)) via MLPs and appending them as tokens into DiT attention layers fails in practice because 2D video models lack native 3D geometry awareness and struggle to infer projective depth and occlusions. LCV circumvents this by pre-rendering boxes into image space. Given camera intrinsics \(K\), extrinsics \((R_{cw}, t_{cw})\), and a box with center \(C_w\), rotation \(R_{wb}\), and half-extents \(h\), rays are cast per pixel and transformed into the box-local coordinate frame. A standard slab intersection test yields the hit mask \(m(u,v)\), local intersection point \(p_b(u,v)\), and camera-space depth \(z(u,v)\).

To encode 6-DOF pose and distance within a single 3-channel frame, DNOCS maps the normalized local surface vector \(n(u,v) = p_b(u,v) / \|p_b(u,v)\|\) to a spherical color wheel with constant luminance (\(L=0.55, a=0.35\)), providing orientation hue \(\mathrm{rgb}_{\text{orient}}(u,v)\). Simultaneously, a normalized inverse depth signal \(d(u,v) \in [0, 1]\) is computed from the 2nd and 98th percentile depths \((z_{\min}, z_{\max})\) of the active boxes: $\(d(u,v) = 1 - \mathrm{clip}\left(\frac{z(u,v) - z_{\min}}{z_{\max} - z_{\min}}, 0, 1\right)\)$ Depth is converted into multiplicative brightness with an exponential decay and floor: $\(b(d) = \beta_{\min} + (1 - \beta_{\min})\exp\left(-k(1 - d)\right)\)$ with default parameters \(\beta_{\min}=0.08\) and \(k=2.0\). The final pixel representation is \(\mathrm{rgb}_{\text{DNOCS}}(u,v) = \mathrm{rgb}_{\text{orient}}(u,v) \odot b(d(u,v))\). Closer objects appear vivid and bright while distant ones fade, and physical overlaps naturally resolve depth ordering at the pixel level, presenting the generative model with unambiguous topological depth cues.

2. Multi-Modal Control Video Assembly: Unifying Video Generation and Trajectory Retargeting Editing

To enable both pure text-to-video synthesis and localized video motion editing within a single architecture, LCV introduces an occlusion-aware compositing pipeline. For pure generation, \(v_{\text{ctrl}}\) consists solely of black-background DNOCS renderings. In video editing or motion retargeting scenarios, users seek to alter the trajectory of a specific subject while preserving global environment consistency.

By estimating monocular depth across the source video using Video Depth Anything (VDA), the system composites newly positioned 3D DNOCS proxies into the input video in a depth-consistent manner. Individual frames of \(v_{\text{ctrl}}\) can be authored as: (i) pure DNOCS frames; (ii) unedited source video segments for strict temporal anchoring; (iii) empty black frames for unrestricted generative inpainting; or (iv) spatially composited mixtures of source backgrounds and rendered DNOCS proxies. In addition, users can paint gray masks over regions where the original subject should be neutralized, enabling the model to cleanly erase original subjects and re-synthesize high-speed drifts or intricate path weaving along with secondary physics effects like tire skid marks and dust plumes.

3. Automated Data Pipeline and Dual-Regime Training: Mining 3D Supervision from In-The-Wild Video

Because manual 3D bounding box annotations for video datasets are virtually nonexistent at scale, LCV builds an automated curation pipeline over ~10,000 in-the-wild stock videos. For each sequence, GroundingDINO and SAM-2 extract instance tracking masks, followed by per-object point cloud reconstruction using monocular depth estimation from VDA. Oriented 3D bounding boxes are fitted per frame and refined via 3D Kalman filtering to enforce smooth, temporally consistent 6-DOF trajectories.

During training, to prevent the network from overfitting strictly to clean synthetic renders or becoming overly reliant on source video frames, control inputs are sampled randomly: 70% as pure DNOCS renders and 30% as mixed composited editing signals. Built upon the WAN-VACE framework, LCV freezes the base Wan 2.2 DiT model completely and only fine-tunes the VACE conditioning adapter modules using LoRA (rank 64) for 10,000 iterations on 4 H100 80GB GPUs, isolating the conditioning representation and proving that high directorial control can be acquired with modest compute.

A Worked Example

Consider orchestrating a dynamic scene of an eagle stooping to catch a fleeing rabbit: 1. Directorial Blocking: The creator places two primitive 3D rectangular boxes in a 3D editor. Box A (eagle) is keyed along a descending parabolic arc across 3 keyframes, pitched downward 45 degrees along its velocity vector; Box B (rabbit) is placed on a ground plane executing an evasive zig-zag. 2. DNOCS Projection: The ray caster projects both boxes onto the virtual camera plane. Box A is shaded with orientation hues that gradually brighten exponentially as it plunges closer to the camera. At the moment of interception, Box A's pixels directly occlude Box B's pixels on the projection canvas. 3. DiT Synthesis: Wan 2.2 receives the composite DNOCS stream through VACE adapters. Rather than rendering rigid cubes, the diffusion backbone populates Box A with an eagle spreading its wings and extending its talons to brake, Box B with a sprinting rabbit kicking up dirt, and automatically synthesizes cast shadows and motion blur that respect the geometric trajectory.

Key Experimental Results

Main Results

Evaluation is conducted across three diverse benchmarks with ground-truth 3D annotations: the large-scale autonomous driving dataset nuScenes (urban traffic with multi-agent interactions), and the articulated interaction benchmarks HO-3D (hand-object manipulation with rapid 3D rotations) and BEHAVE (full-body human interactions with large dynamic objects). Metrics include object mask Containment (Contain ↑), Trajectory Error (TrajErr ↓ in pixels), Global Overlap Winner (GOW ↑), Global Motion Field Agreement (GMFA ↓), Rigid Motion Consistency (RMC ↓), Occlusion Accuracy (OcclAcc ↑ with threshold NDR > 0.55), and VBench visual Quality.

The table below summarizes performance on the nuScenes validation split (Table 1) and HO-3D / BEHAVE interaction benchmarks (Table 2):

Dataset Method Input Modality Contain (%) ↑ GMFA ↓ RMC ↓ TrajErr (px) ↓ OcclAcc (%) ↑ Quality
nuScenes Control-free Baseline GT First + Last frames 10.22 0.828 0.863 90.12 41.45 76.45
nuScenes VACE 2D Flow 21.23 0.135 0.566 7.86 73.91 73.90
nuScenes VACE ft 2D Flow 22.45 0.093 0.528 6.78 79.32 75.50
nuScenes VACE ft 2D Bounding Boxes 96.33 0.232 0.735 16.66 42.45 66.34
nuScenes LCV (Ours) Rendered Oriented 3D Boxes (DNOCS) 87.93 0.066 0.318 5.79 92.69 74.45
HO-3D Control-free Baseline GT First + Last frames 46.8 0.440 0.362 38.5 53.4 76.4
HO-3D VACE ft 2D Flow 69.1 0.071 0.192 5.4 84.2 73.6
HO-3D VACE ft 2D Bounding Boxes 97.9 0.126 0.181 9.7 55.1 72.4
HO-3D LCV (Ours) Rendered Oriented 3D Boxes (DNOCS) 91.3 0.045 0.122 3.9 94.1 72.9
BEHAVE Control-free Baseline GT First + Last frames 42.3 0.611 0.490 54.8 48.6 76.2
BEHAVE VACE ft 2D Flow 63.8 0.098 0.318 7.6 78.5 75.8
BEHAVE VACE ft 2D Bounding Boxes 95.8 0.238 0.412 14.9 49.7 69.3
BEHAVE LCV (Ours) Rendered Oriented 3D Boxes (DNOCS) 88.6 0.062 0.207 5.8 90.2 75.0

When evaluated against the 3D-aware baseline 3DTrajMaster (restricted to scenes with \(\le 3\) entities on HO-3D and BEHAVE), LCV outperforms 3DTrajMaster by \(10\times\) to \(12\times\) in Trajectory Error (dropping from 47.2 to 3.9 on HO-3D, and from 52.6 to 5.8 on BEHAVE) and gains over 25 percentage points in Occlusion Accuracy (improving from 67.5% to 94.1% on HO-3D, and from 64.8% to 90.2% on BEHAVE).

Ablation Study

The ablation study on nuScenes isolates the representation choices and the depth modulation rate \(k\) within the DNOCS encoding:

Configuration / Variant Depth Modulation \(k\) TrajErr (px) ↓ RMC ↓ OcclAcc (%) ↑ Note & Analysis
Depth-only guidance - 7.42 0.471 81.34 Lacks orientation hue; struggles with axial rotations
Standard NOCS (no depth mod.) \(k = 0\) 6.18 0.402 73.56 Missing depth variation; occlusion accuracy drops by ~19.1 pp
DNOCS mild modulation \(k = 1\) 6.05 0.371 88.21 Moderate depth falloff restores relative depth ordering
DNOCS balanced (Ours) \(k = 2\) 5.79 0.318 92.69 Optimal trade-off between orientation hue and depth contrast
DNOCS heavy modulation \(k = 4\) 6.34 0.395 89.47 Overly rapid depth decay washes out orientation hues in mid/far range

Key Findings

  • Depth modulation is indispensable for occlusion reasoning: Omitting depth modulation (\(k=0\)) degrades OcclAcc from 92.69% down to 73.56% (a drop of 19.13 pp), proving that hue alone is insufficient for resolving overlapping multi-object interactions. However, over-modulating (\(k=4\)) diminishes visibility of orientation hues in mid-to-far ranges.
  • 2D boxes create an illusion of control: While 2D axis-aligned boxes achieve the highest 2D containment (96%+), they fail in 3D physical grounding (TrajErr 14-16 px, OcclAcc ~42-55%), causing severe drifting, unnatural deformation, and perspective violations during rotations.
  • High resilience to sparse keyframe authoring: Downsampling the 3D control boxes to every 8th frame (sparse keyframing) incurs less than a 2% increase in TrajErr on nuScenes, confirming that the video generative model smoothly interpolates 3D trajectory dynamics across unannotated intermediate frames.

Highlights & Insights

  • Solving 3D problems in the 2D image domain: Rather than forcing a 2D DiT to learn 3D graphics rendering through raw token projections, LCV leverages classic ray casting to project 3D proxies into an image-space DNOCS map, amortizing the 3D-to-2D rendering burden and preserving the pre-trained generative prior.
  • Cinematic blocking meets generative priors: By decoupling high-level directorial choreography from micro-level morphology, human creators are liberated from drawing impossible frame-by-frame deformations, while diffusion models are fully utilized to infer secondary physics and natural articulations.
  • Efficient plug-and-play conditioning: Using a frozen Wan 2.2 backbone with a rank-64 LoRA VACE adapter trained on 4 H100 GPUs for only 10,000 steps demonstrates an accessible and highly scalable path toward production-ready video directorial controls.

Limitations & Future Work

  • Absence of explicit identity-to-box binding: The current architecture does not explicitly anchor specific subject appearances to individual 3D boxes. In crowded scenes with visually similar entities crossing paths, residual features can occasionally bleed. Future work should integrate multi-view identity reference latents assigned to specific 3D anchors.
  • Manual keyframe timing dependency: Although local shape deformation is automated, pacing and timing of interactions remain governed by user-specified keyframes. Developing models that infer physical acceleration and timing from sparse path splines represents a promising future avenue.
  • vs ControlNet / ControlVideo: Dense edge/depth conditioners require tedious frame-accurate authoring and constrain natural object deformation; LCV uses loose 3D boxes to provide macro guidance while letting the generative model synthesize natural motion.
  • vs Boximator / GLIGEN / Ctrl-V: These systems rely on 2D screen-space boxes that collapse under 3D rotations and depth occlusions; LCV leverages 6-DOF oriented 3D boxes and DNOCS rendering to preserve true 3D spatial relationships.
  • vs 3DTrajMaster: 3DTrajMaster injects sparse coordinates into attention tokens and is hard-capped at 3 entities per clip; LCV renders 3D proxies into image space, scaling naturally to scenes with 10+ interacting agents like nuScenes traffic.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Adapting directorial spatial blocking via DNOCS-rendered 3D oriented bounding boxes introduces an elegant framework that effectively decouples trajectory planning from object deformation.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Evaluated across nuScenes, HO-3D, and BEHAVE with comprehensive structural metrics (Contain, GOW, GMFA, RMC) and a 2-AFC user preference study.
  • Writing Quality: ⭐⭐⭐⭐⭐ The tension between dense guidance and creative authoring is clearly articulated, with rigorous mathematical formulations and compelling qualitative demonstrations.
  • Value: ⭐⭐⭐⭐⭐ Offers a practical, accessible bridge between high-level human creative intent and foundation video generation models for cinematography and video post-production.