Skip to content

MotionSplicer: Part-Based Motion Editing for 4D Volumetric Videos

Conference: ECCV 2026
Paper: ECCV 2026
Project Page: ivl.cs.brown.edu/research/motionsplicer
Area: 3D Vision
Keywords: 4D Gaussian Splatting, Part-Based Motion Editing, Linear Blend Skinning, Temporal Difference Segmentation, Neural Feature Grid

TL;DR

Addressing the inability to interactively manipulate motions in 4D Gaussian volumetric videos with multiple objects and non-rigid parts, MotionSplicer introduces an annotation-free, template-free part-based editing framework that discovers 3D motion parts via temporal difference prompts and adaptive 3D unify-and-split, while learning spatiotemporally coherent Gaussian skinning weights through a time-invariant neural coordinate feature grid and temporal consistency refinement.

Background & Motivation

4D volumetric video representation powered by 3D Gaussian Splatting (3DGS) has emerged as a cornerstone for immersive real-world digitization, offering real-time rendering speeds and free-viewpoint navigation for applications in virtual reality, world model simulation, and interactive digital entertainment. However, conventional dynamic Gaussian radiance fields typically model scene dynamics as a single unconstrained deformation field, lacking explicit decomposition into semantically and kinematically distinct components. Consequently, users cannot interactively alter, isolate, retime, or recompose individual parts and objects once the 4D capture is completed.

Prior efforts in 3D and 4D dynamic editing predominantly focus on either isolated single objects or strictly rigid articulated objects governed by 1-DoF revolute or prismatic joints. They heavily rely on pre-existing structural priors such as parametric body templates (e.g., SMPL), CAD meshes, or URDF kinematic trees, or require labor-intensive manual keyframe annotations. While recent template-free methods have begun exploring unannotated dynamic scenes, they are mostly confined to tabletop setups with a single object and struggle with non-rigid boundaries and multi-object interactions. Furthermore, naively lifting 2D foundational segmentation models (e.g., SAM or SAM 2) across multi-view videos exposes a fundamental domain gap: these models are trained to detect 2D static visual semantics rather than 4D physical motion boundaries, inevitably resulting in severe over-segmentation of static textures, under-segmentation of interacting movable components, and multi-view flickering.

To resolve the core tension between static visual semantics and genuine physical motion boundaries without relying on manual supervision, this paper presents a motion-centric part discovery paradigm. Instead of parsing static frames, it extracts initial motion cues from reference temporal difference images, lifts them into 3D Gaussians, and adaptively refines them through a geometric 3D unification and splitting mechanism. To track non-rigid dynamics robustly without per-primitive noise, it anchors linear blend skinning (LBS) weights onto a time-invariant coordinate feature grid and regularizes continuous streaming optimization via historical frame re-evaluation. Core idea: automate 3D motion part discovery via temporal difference prompts coupled with adaptive 3D unify-and-split, and learn spatiotemporally coherent Gaussian skinning weights through a time-invariant coordinate feature grid with temporal consistency refinement, enabling unannotated, template-free part-level 4D motion editing across complex multi-object scenes.

Method

Overall Architecture

MotionSplicer takes synchronized multi-view RGB video streams as input and yields an editable 4D volumetric video representation parameterized by sparse 6-DoF motion handles and continuous skinning weights over 3D Gaussian primitives. The pipeline comprises three cooperative stages: first, at the canonical reference frame (\(t=0\)), a temporal difference image prompts a foundational segmentation model to identify dynamic regions across multi-view views, which are lifted to 3D and calibrated by an adaptive 3D Unification & Splitting (U&S) algorithm into clean motion handles; second, a continuous spatial skinning weight field is instantiated via a multi-resolution hash feature grid, mapping canonical 3D coordinates to normalized skinning weights to govern forward linear blend skinning (LBS); finally, a streaming spatio-temporal optimization updates frame-wise 6-DoF handle transformations alongside the shared weight grid, regularized by a temporal consistency refinement loss that randomly re-evaluates historical frames, enabling downstream creative motion manipulations including speed re-scaling, echoing, ghosting, freezing, and cross-scene compositing.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Synchronized Multi-View Dynamic Videos<br/>(t = 0 to T)"] --> B["Temporal Difference & 3D Unify-and-Split Part Discovery<br/>Difference Image Prompt + SAM 2 Propagation + Adaptive U&S"]
    B --> C["Time-Invariant Feature Grid Skinning Binding<br/>Instant-NGP Continuous Coordinates + Forward LBS Mapping"]
    C --> D["Streaming Temporal Consistency Optimization & Editing<br/>Multi-View Photometric Loss + Historical Frame Replay + Interactive Editing"]
    D --> E["Free-Viewpoint 4D Part-Based Motion Rendering<br/>(Speed Control / Echo / Ghosting / Cross-Scene Compositing)"]

Key Designs

1. Temporal Difference & 3D Unify-and-Split Part Discovery: bridging the 2D semantic to 4D motion domain gap

Directly applying 2D foundational segmentation models to static scene images produces excessive semantic clutter (e.g., clothing patterns, floor tiles, background furniture) that obscures physical kinematic boundaries. To circumvent this issue, MotionSplicer feeds reference temporal difference images \(I_d = \|I_{t=0} - I_{t=\bar{t}}\|\) (where \(\bar{t}\) is an anchor frame with pronounced motion) into SAM. The temporal difference suppresses static geometry and highlights moving regions, allowing SAM to generate clean dynamic masks that are sequentially tracked across spatially sorted multi-view cameras via SAM 2. All unmasked pixels are assigned to a unified background handle.

These 2D masks supervise an auxiliary logit vector \(m_i \in \mathbb{R}^{P'}\) attached to each 3D Gaussian primitive \(g_i\). Because lifted 2D predictions inevitably suffer from multi-view occlusions and inaccurate part boundaries, the method performs a 3D Unification and Splitting (U&S) procedure during the optimization of \(t=1\): over-segmented groups exhibiting minimal relative translation and rotation are merged, whereas under-segmented clusters accumulating large gradient magnitudes in the skinning weight grid are bisected along their principal geometric axis estimated via Principal Component Analysis (PCA). This process condenses initial candidate parts \(P'\) into \(P\) coherent physical handles, with canonical handle centers \(c_p \in \mathbb{R}^3\) placed at the geometric centroid of assigned primitives.

2. Time-Invariant Feature Grid Skinning Binding: enforcing smooth non-rigid deformation fields

Optimizing independent skinning weight vectors \(w_i\) for millions of discrete Gaussian primitives tends to fall into noisy local minima, producing spatial disconnections and flying floaters across time. To guarantee spatial regularity and smooth deformation across soft, non-rigid boundaries, MotionSplicer parameterizes the skinning field using an Instant-NGP multi-resolution hash feature grid paired with a compact 2-layer MLP. Taking the canonical mean coordinate \(\mu_i\) of each Gaussian primitive as input, the network outputs continuous, normalized skinning weights \(w_i(\mu_i) \in \mathbb{R}^P\) satisfying \(\sum_{p=1}^P w_{ip} = 1\).

At any timestep \(t > 0\), the \(P\) sparse handles undergo 6-DoF rigid transformations \(h_t^p = (R_t^p, T_t^p) \in SE(3)\) relative to the canonical frame. The position \(\mu_i^t\) and orientation quaternion \(q_i^t\) of each Gaussian primitive are transformed via Linear Blend Skinning (LBS): $$ \mu_i^t = \sum_{p=1}^P w_{ip} \left( R_t^p (\mu_i - c_p) + c_p + T_t^p \right), \quad q_i^t = \text{quat}\left( \sum_{p=1}^P w_{ip} R_t^p R(q_i) \right) $$ where \(R(\cdot)\) maps a quaternion to its rotation matrix. By anchoring discrete Gaussian transformations onto a continuous neural coordinate field, non-rigid tissues and flexible joints (such as a crawling turtle's shoulder or bending limbs) smoothly interpolate between adjacent handles, avoiding the artificial seam tears common in strict rigid-body articulation.

3. Streaming Temporal Consistency Optimization & Editing: preventing long-term drift and unlocking versatile 4D edits

The framework optimizes parameters sequentially from frame \(t = 1\) to \(T\). For each frame, rendered multi-view images \(I\) are evaluated against ground truth views \(\hat{I}_t^v\) using a composite photometric loss \(\mathcal{L} = \lambda \|I - \hat{I}\|_1 + (1 - \lambda)(1 - \text{SSIM}(I, \hat{I}))\) to back-propagate gradients into handle transformations \(H_t\) and the shared feature grid. However, sequential streaming optimization is susceptible to temporal drift when encountering disocclusion or transient lighting changes. To enforce long-term fidelity, the model incorporates a Temporal Consistency Refinement loss: at frame \(t\), it randomly samples an earlier optimized frame \(t' < t\), freezes its previously converged handle transformation \(H_{t'}\), re-renders the multi-view output at \(t'\), and updates exclusively the time-invariant skinning feature grid. This replay mechanism forces the neural grid to retain valid skinning representations across diverse historical deformation configurations.

At test time, because the 4D volumetric scene is decoupled into time-invariant Gaussian skinning assignments and sparse handle trajectories \(H_t\), users can perform granular motion editing with high responsiveness: altering part velocity (Speed), extending repetitive gestures (Loop), isolating specific dynamic agents while muting others (Ghosting/Isolation), creating delayed motion echoes (Echo), and transferring segmented Gaussians into novel scenes (Compositing), as well as applying novel 6-DoF spatial rotations and scaling to individual parts.

Key Experimental Results

Main Results

The method is comprehensively evaluated across synthetic isolated objects (Sketchfab), synthetic multi-object scenes (OmniGibson BEHAVIOR-1k), and real-world multi-view captures (Panoptic Studio and custom multi-camera rigs). Baselines include sparse-controlled dynamic splatting (SC-GS), articulated digital twin modeling (ArtiGS), hierarchical object rigging (RigGS), and monocular manipulation imitation (RSRD). Metrics evaluate novel-view reconstruction fidelity (PSNR, SSIM, LPIPS) and motion editing fidelity: temporal smoothness via T-LPIPS, perceptual realism via DINOv2-based Kernel Video Distance (KVD), optical flow consistency via Flow-Warp error, and user intent alignment rated by GPT-4o on a 1โ€“5 scale (GPT-EE).

Benchmark / Setting Method PSNR โ†‘ SSIM โ†‘ LPIPS โ†“ T-LPIPS โ†“ KVD โ†“ Flow-Warp โ†“ GPT-EE โ†‘
Synthetic Objects SC-GS 40.04 0.993 0.0166 0.0079 0.9493 0.9611 3.7
ArtiGS 27.07 0.896 0.0921 0.0052 0.8716 0.3217 3.9
RigGS 31.83 0.983 0.0473 0.0023 0.7444 2.8741 4.8
RSRD 27.26 0.949 0.0398 0.0161 0.7973 0.0000 4.5
MotionSplicer (Ours) 36.52 0.985 0.0311 0.0109 0.6111 0.6723 4.8
Synthetic Scenes SC-GS 32.30 0.959 0.1386 0.0033 1.0945 0.7440 1.0
ArtiGS 26.72 0.930 0.1646 0.0052 0.7836 0.2647 2.8
RigGS 29.90 0.971 0.2629 0.0142 0.9728 2.2541 4.5
RSRD 32.25 0.948 0.0891 0.0167 0.9274 0.4389 3.0
MotionSplicer (Ours) 34.49 0.967 0.0428 0.0022 0.7007 0.4407 4.7
Real Scenes SC-GS 22.05 0.730 0.3990 0.0125 1.0039 29.1976 2.0
ArtiGS 16.36 0.679 0.5268 0.0494 1.0734 17.4413 1.7
RigGS 26.40 0.853 0.3299 0.0810 0.9895 31.1659 1.7
RSRD 19.21 0.677 0.3856 0.0436 1.0770 26.5958 1.8
MotionSplicer (Ours) 27.36 0.892 0.1683 0.0064 0.8086 31.5803 4.7

Ablation Study

A systematic ablation isolates the impact of each core component: relying solely on SAM 2 masks without motion guidance, removing temporal difference initialization (\(I_d\)), disabling the neural feature grid (optimizing per-primitive weights directly), and removing temporal consistency refinement (TC refine).

Config PSNR โ†‘ SSIM โ†‘ LPIPS โ†“ T-LPIPS โ†“ KVD โ†“ Flow-Warp โ†“ GPT-EE โ†‘ Note
SAM2 only 30.40 0.9265 0.1530 0.0088 0.6658 0.329 3.5 Lacks dynamic cues; static semantic boundaries cause severe under/over-segmentation
w/o \(I_d\) (no difference image) 32.81 0.9588 0.0681 0.0094 0.6726 0.329 4.8 Clutters initial state with static semantic parts, burdening convergence
w/o feature grid 34.00 0.9583 0.0681 0.0088 0.6462 0.325 4.5 Discrete weights lack spatial smoothness, creating noisy boundary tears
w/o TC refine 35.55 0.9681 0.0479 0.0092 0.6425 0.313 4.8 Suffers from long-term temporal drift across streaming frames
Full Model 36.54 0.9729 0.0345 0.0014 0.6693 0.303 4.8 Best overall synthesis; T-LPIPS drops by nearly an order of magnitude

Key Findings

  • Superiority in complex real scenes: As evaluations transition from isolated synthetic objects to complex real-world scenes, MotionSplicer's advantage widens dramatically. On real captures, it attains 27.36 dB PSNR versus SC-GS's 22.05 dB and RSRD's 19.21 dB, halving the perceptual error (LPIPS: 0.1683 vs. 0.3299โ€“0.5268), demonstrating that single-axis models and rigid tree skeletons collapse when confronted with multi-object interactions.
  • Critical role of temporal refinement: Integrating historical frame re-evaluation (TC refine) slashes T-LPIPS temporal jitter from 0.0092 to 0.0014 (an 84.8% reduction), effectively eliminating trailing floaters and flickering artifacts during long-term video streaming.
  • The zero-motion artifact in flow warping: On synthetic objects, RSRD achieves an apparently flawless 0.0000 Flow-Warp error because monocular optimization fails to register movement, leaving the object entirely stationary. Furthermore, occlusion masking in bidirectional optical flow filters out complex non-rigid motion regions. This underlines the necessity of holistic metrics like GPT-EE and perceptual KVD alongside user studies.

Highlights & Insights

  • Temporal difference as a lightweight dynamic filter: Computing \(\|I_{t=0} - I_{t=\bar{t}}\|\) to prompt SAM focuses segmentation exclusively on active kinematic regions, eliminating tens of thousands of static background components and avoiding expensive \(O(N \times P)\) parameter explosions.
  • Self-correcting 3D Unify-and-Split: Rather than blindly trusting 2D prompts, the framework evaluates geometric gradients and motion clustering in 3D during initial streaming, splitting merged kinematic clusters and unifying fragmented static patches in a closed-loop fashion.
  • Hybrid discrete-continuous parameterization: Coupling discrete explicit 3D Gaussians with a continuous neural hash grid achieves the ideal balance: maintaining real-time rendering speed while enforcing physically plausible smooth skinning transitions across soft non-rigid boundaries.

Limitations & Future Work

  • Reliance on synchronized multi-view capture: The framework presumes calibrated multi-camera rigs to resolve depth ambiguities, limiting immediate deployment to in-the-wild monocular or single mobile phone recordings.
  • Inadequate modeling of extreme topological dynamics: Linear blend skinning relies on localized rigid and elastic transformations; it struggles to represent extreme topological mutations such as fluid splashing, fabric tearing, smoke dispersion, or fracturing.
  • Per-scene optimization overhead: Optimizing feature grids and handle trajectories per scene incurs non-negligible training time on extended sequences, highlighting the need for generalizable, feed-forward 4D part manipulation foundation models.
  • vs SC-GS (Huang et al., CVPR 2024): SC-GS drives deformation through dense control points suitable for holistic continuous deformation, but lacks kinematic semantic separation; when applied to scenes with multiple independently rotating doors, parts become entangled in local minima.
  • vs ArticulatedGS (Guo et al., 2025): ArticulatedGS predicts a single joint rotation axis on static assets, completely failing in environments with multiple moving bodies or non-rigid organic motions.
  • vs RSRD (Kerr et al., 2024) / POD (Wu et al., 2025): RSRD and POD depend on monocular DINO clustering on tabletop objects, exhibiting high sensitivity to camera viewpoints and lacking streaming multi-view temporal regularization.

Rating

  • Novelty: โญโญโญโญโญ Seamlessly unifies temporal difference prompts, 3D unify-and-split, and neural skinning grids for unannotated 4D multi-object editing.
  • Experimental Thoroughness: โญโญโญโญโญ Rigorous evaluation across synthetic objects, multi-object simulation, and real 360 captures with deep metric analysis and user studies.
  • Writing Quality: โญโญโญโญโญ Well-structured, lucid exposition connecting theoretical design choices to empirical failure modes.
  • Value: โญโญโญโญโญ Sets a benchmark for interactive 4D content creation, XR video manipulation, and embodied robotics simulation.