MoBa-GS: Learning a Spatially-Varying Motion Basis over a Dynamic Canonical Space for 4D Reconstruction¶
Conference: ECCV 2026
Paper: CVF Open Access
Code: https://github.com/tgy1221/MoBa-GS
Area: 3D Vision
Keywords: 4D Reconstruction, 3D Gaussian Splatting, Motion Basis Factorization, Dynamic Canonical Space, Non-rigid Deformation
TL;DR¶
MoBa-GS resolves 4D dynamic reconstruction bottlenecks by introducing a structural inversion from global temporal trajectories to a spatially-varying motion basis, factorizing non-rigid deformations into a low-frequency dynamic canonical space and sparse linear combinations of local motion primitives, achieving state-of-the-art fidelity from random initialization without SfM point cloud priors alongside real-time rendering (>163 FPS) and 16x faster training via time-invariant caching.
Background & Motivation¶
Novel view synthesis and 4D spatiotemporal reconstruction of complex dynamic scenes represent a core frontier for immersive technologies, augmented reality, and embodied spatial intelligence. While 3D Gaussian Splatting (3DGS) has revolutionized static scene reconstruction with photorealistic rendering fidelity at interactive framerates, lifting this explicit primitive formulation into the four-dimensional continuum introduces fundamental optimization and representation dilemmas. Prevailing dynamic 3DGS frameworks predominantly follow one of two paradigms: the first relies on monolithic deformation multilayer perceptrons (MLPs) that map spatiotemporal coordinates directly to per-Gaussian spatial displacements, rotations, and scaling factors. This formulation inherently entangles underlying 3D geometry with dynamic motion, creating a notorious "moving target" optimization conflict where canonical shape and deformation fields destructively interfere, resulting in severe geometric overfitting. The second paradigm attempts to regularize the motion field via factorization using globally shared time-basis trajectories; however, this strategy implicitly assumes the scene consists of rigid or semi-rigid components riding on a finite set of global temporal paths. When encountering heterogeneous non-rigid dynamics—such as wrinkling cloth, flowing fluids, or independent local deformations—global temporal bases suffer a catastrophic explosion in basis rank and fail to span the divergent kinematic space.
Beyond these structural limitations, existing dynamic reconstruction approaches suffer from an operational vulnerability: an extreme dependency on pre-computed Structure-from-Motion (SfM) sparse point clouds (e.g., COLMAP) for geometric initialization. In real-world dynamic scenes, rapid motion and continuous non-rigid deformation systematically cause feature correspondence matching to fail across consecutive frames, leaving moving subjects severely under-reconstructed or plagued with outliers in the SfM point cloud. Lacking calibrated geometric point priors, existing methods frequently suffer topological collapse or degenerate into floating artifacts. Furthermore, evaluating dense monolithic deformation networks on every Gaussian at runtime incurs heavy computational overhead, preventing sustained high-framerate rendering on resource-constrained platforms.
To overcome these intertwined bottlenecks, this paper builds upon a critical physical insight: while high-frequency non-rigid motion appears overwhelmingly complex across the entire scene, it is strictly low-dimensional within any localized spatial neighborhood, where kinematic capabilities are inherently governed by local spatial coordinates and material properties. Core idea: introduce a structural inversion from global time-bases to a Spatially-Varying Motion Basis, where a lightweight network first establishes a low-frequency dynamic canonical space to absorb coarse global drift, while static spatial coordinates predict a localized physical motion dictionary linearly combined with dynamic temporal weights to construct an implicit neural scaffold that dispenses with SfM point cloud priors.
Method¶
Overall Architecture¶
MoBa-GS implements a hierarchical, spatially factorized 4D reconstruction architecture that decouples coarse motion from fine-grained non-rigid kinematics. Given an input sequence of monocular or multi-view RGB video frames alongside time step \(t\), the system outputs time-deformed 3D Gaussians that render photorealistic novel views. The pipeline coordinates three primary stages: first, the Base Motion Network functions as a temporal low-pass filter, mapping low-frequency temporal embeddings to coarse base displacements that transform static canonical positions into an adaptive dynamic canonical space, neutralizing camera ego-motion and scene-wide rigid translation; second, the Basis Predictor takes solely the static canonical coordinates to generate a localized, time-invariant motion basis dictionary, which is cached in GPU memory once geometry stabilizes; third, the Weight Predictor takes the dynamic canonical coordinates and high-frequency temporal embeddings to predict sparse dynamic blending weights that linearly scale the local basis vectors to synthesize high-frequency residual displacements. Throughout optimization, the framework integrates motion-guided densification and positional annealing to concentrate geometric capacity on dynamic regions and eliminate late-stage geometric overfitting.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Canonical Gaussians p_canon & Timestamp t"] --> B["Dynamic Canonical Space Construction<br/>Low-frequency time embedding predicts coarse base motion"]
B --> C["Spatially-Varying Motion Basis Decoupling<br/>Static positions predict local bases & dynamic states predict weights"]
C --> D["Time-Invariant Basis Caching & Residual Synthesis<br/>Cached basis projection matrix & sparse linear combination"]
D --> E["Spatiotemporal Synergistic Optimization<br/>Motion-guided densification & decoupled alternating annealing"]
E --> F["Output: High-fidelity deformed Gaussians & real-time rendered views"]
Key Designs¶
1. Dynamic Canonical Space Construction: Decoupling Low-Frequency Global Drift to Stabilize the Reference Frame
To resolve the optimization instability caused by the "fixed origin" assumption without polluting fine geometric gradients, the framework first constructs a coarse dynamic canonical reference frame. Directly forcing a high-frequency residual network to absorb global rigid translation along with intricate non-rigid deformations causes destructive interference during gradient backpropagation. The Base Motion Network \(D_{\text{base}}\) is explicitly designed as a temporal low-pass filter, implemented as a lightweight 4-layer MLP (hidden dimension 128) conditioned on a low-frequency temporal positional embedding with maximum frequency \(L=4\). The network predicts a coarse base displacement \(\Delta \mathbf{p}_{\text{base}}\), transforming the static canonical position \(\mathbf{p}_{\text{canon}}\) into a dynamic canonical space:
This dynamic reference frame digests shared low-frequency trajectories (such as global translation and residual camera ego-motion), insulating the subsequent high-frequency residual modules from low-frequency drift and constraining residual learning strictly to localized non-linear variations.
2. Spatially-Varying Motion Basis Decoupling: Formulating Localized Low-Rank Physical Dictionaries
To overcome the inability of global temporal trajectories to model heterogeneous non-rigid dynamics and to prevent unconstrained deformation regression, the paper introduces a Spatially-Varying Motion Basis. The high-frequency residual deformation is factorized into a time-invariant spatial basis field and a time-varying coefficient activation field. The Basis Predictor \(B\) (a 2-layer MLP with hidden dimension 256) takes solely the positional encoding of the static canonical coordinate \(\mathbf{p}_{\text{canon}}\) as input, producing \(K\) local 3D basis vectors that form a local projection matrix:
Simultaneously, a higher-capacity Weight Predictor \(W\) (an 8-layer MLP trunk with hidden dimension 256) conditions on the dynamic canonical coordinate \(\mathbf{p}'(t)\) and high-frequency temporal embedding \(\gamma_{L=10}(t)\) to predict a sparse set of \(K\) scalar blending weights \(\mathbf{w}(\mathbf{p}', t) = [w_1, \dots, w_K]^\top \in \mathbb{R}^K\). The final residual 3D displacement \(\Delta \mathbf{p}_{\text{res}}\) is synthesized via linear projection:
Auxiliary prediction heads on the Weight Predictor concurrently output residual transformations for quaternion rotation (\(\Delta \mathbf{q}\)) and scaling (\(\Delta \mathbf{s}\)). By restricting spatiotemporal trajectories to a sparse combination of learned spatial bases, the optimization is inherently bounded within a locally low-dimensional manifold of natural kinematics. This architecture functions as an Implicit Neural Scaffold: rigid regions naturally learn isotropic, dispersed bases for general transformation, while deforming regions establish directionally aligned motion primitives. Crucially, this intrinsic manifold prior eliminates the need for handcrafted smoothness heuristics, allowing the model to recover high-fidelity geometry and motion entirely from random point cloud initialization without SfM priors.
3. Time-Invariant Basis Caching & Residual Synthesis: Eliminating Runtime MLP Overhead
In dynamic 3DGS synthesis, evaluating deep MLPs for hundreds of thousands of Gaussians at every frame creates severe latency bottlenecks. Because the Basis Predictor \(B\) is conditioned strictly on static canonical coordinates \(\mathbf{p}_{\text{canon}}\), the predicted local basis matrix \(\mathbf{V}(\mathbf{p}_{\text{canon}})\) is completely time-invariant. Once scene geometry stabilizes during late-stage training, all basis tensors across all Gaussians are pre-computed in a single forward pass and cached directly in GPU memory. During test-time rendering, the execution of the Basis Predictor MLP is completely bypassed; runtime deformation reduces to lightweight weight inference followed by a vectorized matrix-vector dot product, cutting inference time by an order of magnitude and achieving over 163 FPS.
4. Spatiotemporal Synergistic Optimization: Motion-Guided Densification & Positional Annealing
To suppress mutual interference between evolving geometry and non-rigid deformation networks and prevent capacity waste on high-frequency static textures, MoBa-GS introduces a two-pronged training curriculum: - Motion-Guided Densification: Standard 3DGS densification relies exclusively on 2D view-space projection gradients, blindly over-densifying static textures with high visual contrast. MoBa-GS formulates densification as spatiotemporal importance sampling. A Gaussian \(G_i\) is eligible for cloning or splitting if and only if both its view-space gradient exceeds \(\tau_{\text{grad}}\) and its predicted temporal motion magnitude exceeds \(\tau_{\text{motion}}\):
This dynamic feedback loop suppresses densification across static backgrounds, concentrating Gaussian geometric budget onto complex deforming elements. - Decoupled Alternating Optimization & Positional Annealing: To break the unstable feedback loop where canonical geometry and deformation networks serve as moving targets for one another, training alternates across phases: Phase 1 freezes Gaussian geometry and trains only deformation networks for \(N_{\text{def}}=2\) steps; Phase 2 freezes deformation networks and optimizes canonical parameters for \(N_{\text{geo}}=1\) step. Once training reaches iteration \(i_{\text{anneal}}\), Positional Annealing gradually freezes the learning rate \(\eta_{\mathbf{p}}(i)\) of canonical coordinates, compelling the motion networks to absorb all residual spatiotemporal variations and preventing late-stage geometric overfitting on training views.
Loss & Training¶
The framework is optimized end-to-end using self-supervised photometric consistency against multi-view or monocular RGB supervision, requiring no external depth or optical flow priors. The total reconstruction objective combines an \(L_1\) color loss with the D-SSIM structural dissimilarity loss:
All models are trained on a single NVIDIA A100 GPU with default basis rank \(K=8\). In addition, a dynamic acceleration penalty is applied to regularize Gaussian trajectories, imposing rigid preservation on static regions while decaying exponentially for rapidly moving Gaussians to safeguard non-rigid kinematic fidelity.
Key Experimental Results¶
Main Results¶
MoBa-GS was evaluated extensively on the real-world NeRF-DS benchmark across seven challenging non-rigid scenes against current dynamic 3DGS baselines, as reported in Table 1. MoBa-GS achieves the highest mean PSNR (24.14 dB) and SSIM (0.857) across all scenes, outperforming Deformable-GS, SC-GS, and 4DGS by a decisive margin. While D-MiSo achieves lower LPIPS scores, recent findings demonstrate that VGG-based perceptual metrics unduly reward motion-blurred, over-smoothed predictions; qualitative analysis confirms that MoBa-GS reconstructs crisp boundaries and fine structures without blur artifacts.
Table 1: Quantitative comparison on NeRF-DS dataset (7 real-world dynamic scenes)
| Method | As (PSNR/SSIM) | Basin (PSNR/SSIM) | Bell (PSNR/SSIM) | Cup (PSNR/SSIM) | Plate (PSNR/SSIM) | Press (PSNR/SSIM) | Sieve (PSNR/SSIM) | Mean PSNR↑ | Mean SSIM↑ | Mean LPIPS↓ |
|---|---|---|---|---|---|---|---|---|---|---|
| 3D-GS [18] | 22.69 / 0.802 | 18.42 / 0.717 | 21.01 / 0.789 | 21.71 / 0.830 | 16.14 / 0.697 | 22.89 / 0.816 | 23.16 / 0.820 | 20.29 | 0.782 | 0.292 |
| Deformable-GS [50] | 26.22 / 0.881 | 19.65 / 0.791 | 25.08 / 0.842 | 24.66 / 0.887 | 20.29 / 0.808 | 25.35 / 0.858 | 25.08 / 0.869 | 23.76 | 0.848 | 0.180 |
| SC-GS [16] | 24.47 / 0.847 | 19.30 / 0.806 | 21.98 / 0.801 | 20.08 / 0.740 | 19.92 / 0.814 | 24.89 / 0.870 | 25.12 / 0.896 | 22.25 | 0.824 | 0.203 |
| D-MiSo [36] | 26.06 / 0.881 | 19.79 / 0.784 | 24.75 / 0.851 | 24.36 / 0.885 | 20.50 / 0.805 | 25.71 / 0.864 | 26.12 / 0.885 | 23.90 | 0.851 | 0.151 |
| 4DGS [43] | 25.68 / 0.885 | 19.26 / 0.817 | 23.22 / 0.853 | 24.04 / 0.893 | 18.55 / 0.760 | 24.13 / 0.822 | 22.88 / 0.830 | 22.54 | 0.837 | 0.212 |
| Motion-GS [56] | 25.97 / 0.868 | 19.80 / 0.779 | 23.99 / 0.809 | 24.36 / 0.860 | 20.41 / 0.793 | 25.86 / 0.861 | 25.62 / 0.847 | 23.71 | 0.831 | 0.240 |
| MoBa-GS (Ours) | 26.48 / 0.885 | 19.80 / 0.798 | 25.23 / 0.847 | 24.79 / 0.894 | 20.86 / 0.823 | 25.96 / 0.877 | 25.86 / 0.878 | 24.14 | 0.857 | 0.197 |
On the HyperNeRF benchmark exhibiting aggressive camera motion and extreme topological deformation, MoBa-GS consistently establishes SOTA metrics, demonstrating pronounced gains on intricate deformation cases such as 'Banana' and 'Broom' (see Table 2).
Table 2: Quantitative comparison on HyperNeRF benchmark
| Method | 3D Printer (PSNR/SSIM) | Chicken (PSNR/SSIM) | Broom (PSNR/SSIM) | Banana (PSNR/SSIM) | Mean PSNR↑ | Mean SSIM↑ |
|---|---|---|---|---|---|---|
| SC-GS [16] | 19.33 / 0.60 | 22.46 / 0.63 | 20.26 / 0.49 | 21.63 / 0.80 | 20.92 | 0.63 |
| D-MiSo [36] | 20.13 / 0.66 | 23.66 / 0.70 | 20.86 / 0.31 | 25.22 / 0.80 | 22.47 | 0.62 |
| Deformable-GS [50] | 20.18 / 0.64 | 22.68 / 0.60 | 20.54 / 0.34 | 24.85 / 0.78 | 22.06 | 0.59 |
| MotionGS [56] | 20.47 / 0.61 | 22.23 / 0.61 | 19.64 / 0.25 | 21.51 / 0.51 | 20.96 | 0.50 |
| MoBa-GS (Ours) | 20.51 / 0.67 | 23.77 / 0.70 | 21.03 / 0.32 | 25.62 / 0.80 | 22.73 | 0.63 |
Ablation Study¶
Systematic component ablations and basis rank sensitivity evaluations on the complete NeRF-DS dataset are summarized in Table 3.
Table 3: Ablation studies and basis count sensitivity analysis on NeRF-DS
| Model Configuration | PSNR↑ | SSIM↑ | LPIPS↓ | Empirical Finding / Degradation Note |
|---|---|---|---|---|
| MoBa-GS (Full model, \(K=8\)) | 24.14 | 0.857 | 0.197 | Full framework achieves highest reconstruction fidelity |
| w/o Positional Annealing | 23.87 | 0.850 | 0.210 | Persistent canonical learning overfits training views (-0.27 dB PSNR) |
| w/o Motion Densification | 23.83 | 0.849 | 0.207 | Time-agnostic densification starves dynamic regions (-0.31 dB PSNR) |
| w/o Motion Basis | 23.48 | 0.841 | 0.221 | Replaced with standard Base+Residual MLP; massive drop of -0.66 dB PSNR |
| Basis Count: \(K = 2\) | 23.54 | 0.840 | 0.222 | Low rank underfits high-frequency non-rigid residual dynamics (-0.60 dB PSNR) |
| Basis Count: \(K = 4\) | 23.62 | 0.847 | 0.216 | Sub-optimal kinematic capacity (-0.52 dB PSNR) |
| Basis Count: \(K = 16\) | 23.54 | 0.845 | 0.216 | Excess capacity overfits training views (-0.60 dB PSNR) |
| Basis Count: \(K = 32\) | 23.44 | 0.841 | 0.218 | Further capacity bloat confirms symmetric U-shaped capacity curve |
Table 4: Computational training efficiency and inference footprint (NeRF-DS on NVIDIA A100)
| Scene | D-MiSo [36] Training Time | MoBa-GS Training Time | MoBa-GS Rendering Speed (FPS)↑ | #Gaussians (K)↓ | Model Size (MB)↓ |
|---|---|---|---|---|---|
| As | 2h 47m | 13m | 181 | 28 | 8.73 |
| Basin | 2h 38m | 15m | 146 | 41 | 11.96 |
| Bell | 3h 14m | 18m | 103 | 61 | 16.63 |
| Cup | 4h 00m | 11m | 150 | 37 | 10.92 |
| Plate | 2h 47m | 9m | 149 | 35 | 10.42 |
| Press | 1h 54m | 12m | 174 | 31 | 9.56 |
| Sieve | 3h 52m | 15m | 237 | 37 | 10.88 |
| Mean | ~3h 27m | ~13m (>16x speedup) | 163 FPS | 39K | 11.3 MB |
Key Findings¶
- Spatially-varying basis factorization is the core performance driver: replacing the local basis with a monolithic Base+Residual MLP results in a severe 0.66 dB drop in PSNR, proving that unconstrained MLPs fail to regularize non-rigid 4D deformation fields without explicit geometric priors.
- Basis rank \(K=8\) forms a symmetric U-shaped capacity optimum: smaller basis ranks (\(K=2,4\)) underfit complex kinematics, while higher ranks (\(K \ge 16\)) induce view overfitting, validating \(K=8\) as the optimal structural regularizer.
- Remarkable training and inference efficiency: facilitated by time-invariant basis caching, MoBa-GS reduces convergence time from over 3.4 hours (D-MiSo) to just 13 minutes (a 16x speedup) while achieving a real-time rendering speed of 163 FPS with an average of only 39K Gaussians and 11.3 MB total storage.
Highlights & Insights¶
- Structural Inversion of Motion Factorization: Reversing the established convention from global temporal trajectories to a spatially localized basis dictionary grounds motion modeling directly into spatial coordinates, enabling independent kinematic expression across diverse materials.
- Implicit Neural Scaffold Eliminating SfM Priors: Formulating deformation as a sparse projection onto local spatial bases enforces an intrinsic manifold constraint that guides high-fidelity geometric recovery entirely from random initialization, removing the long-standing dependency on SfM feature matching.
- Synergistic Spatiotemporal Annealing & Caching: Coupling time-invariant basis caching with motion-guided densification and positional annealing provides a robust blueprint for preventing geometric drift and optimizing dynamic explicit radiance representations.
Limitations & Future Work¶
- Extreme topological tearing: The continuous Lagrangian deformation formulation assumes homeomorphic temporal continuity, which may struggle with fluid splashes or abrupt topological fragmentation.
- Disentanglement of complex illumination: Time-varying specularities and cast shadows are currently absorbed by color spherical harmonics and motion weights, which can occasionally induce minor geometric compensation.
- Future directions: Integrating feed-forward foundation model priors (such as DUSt3R or VGGT) with the spatially-varying motion basis represents a promising pathway toward zero-shot generalizable 4D dynamic reconstruction.
Related Work & Insights¶
- vs DynMF [19]: DynMF employs globally shared temporal trajectory bases for all points, risking rank explosion on heterogeneous non-rigid scenes; MoBa-GS learns localized spatial motion bases tailored to local coordinate geometry.
- vs Deformable-GS [50] / SC-GS [16]: These methods rely on unconstrained monolithic MLPs that suffer from severe geometric overfitting and require SfM point clouds; MoBa-GS provides intrinsic manifold scaffolding, enables random initialization, and renders at >163 FPS via caching.
- vs D-MiSo [36]: D-MiSo requires >3.4 hours of training and yields over-smoothed boundaries; MoBa-GS trains in under 18 minutes (16x faster) and preserves sharp high-frequency geometry with higher PSNR and SSIM.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Elegant structural inversion paradigm resolving the global time-basis rank explosion and eliminating SfM initialization dependency.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across NeRF-DS and HyperNeRF, detailed ablation on basis rank \(K\), training duration, and storage footprint.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulation, clear physical intuition, and compelling visualization of learned motion primitives.
- Value: ⭐⭐⭐⭐⭐ Fast convergence (<18 min), compact storage (~11 MB), and high framerate (>163 FPS) establish an outstanding practical standard for 4D dynamic reconstruction.