Minute4D: Training High-Fidelity 4D Gaussian Splatting in One Minute¶
Conference: ECCV 2026
Paper: ECCV 2026
Area: 3D Vision
Keywords: 4D Gaussian Splatting, Dynamic View Synthesis, Monocular Dynamic Reconstruction, Density Control, 3D Point Tracking
TL;DR¶
To tackle camera-object motion ambiguity, view coverage scarcity, and prolonged optimization in casual monocular dynamic 3D reconstruction, Minute4D combines foundation-model-driven segmentation-tracking geometric completion with multi-view photometric loss-guided 4D Gaussian density control, compressing 4D Gaussian Splatting optimization time to ~40 seconds (a 5x+ speedup) while establishing state-of-the-art rendering fidelity and memory efficiency.
Background & Motivation¶
Reconstructing 4D dynamic scenes and synthesizing high-fidelity novel views from uncalibrated, casually captured monocular videos is an essential enabling technology for robotics, autonomous driving, virtual reality, and augmented reality. However, monocular dynamic reconstruction is inherently ill-posed: the coupling between camera trajectories and non-rigid object motions precludes traditional Structure-from-Motion (SfM) pipelines from reliably recovering camera poses. Furthermore, severe self-occlusions of dynamic subjects lead to pronounced geometric voids in per-frame back-projections, while existing implicit Neural Radiance Field (NeRF) variants and deformation-field-based 4D Gaussian pipelines require hours or even days to optimize. Even recent rapid systems like Instant4D still require multiple minutes and suffer from single-frame overfitting.
The core tension behind these limitations lies in the fact that geometric point clouds back-projected purely from per-frame depth maps lack cross-temporal correspondences. As a consequence, 4D Gaussian primitives fit independently to visible regions within individual frames at the expense of global consistency, producing severe geometric holes and tearing artifacts under novel viewpoints. Simultaneously, standard Gaussian densification and pruning heuristics depend strictly on single-view screen-space positional gradients, spawning massive numbers of redundant Gaussians that trigger GPU memory explosion and sluggish optimization convergence.
To overcome both geometric incompleteness and optimization inefficiency, this paper introduces a unified paradigm for ultra-fast monocular 4D Gaussian reconstruction. Core idea: integrate SAM2 semantic tracking with TAPIP3D persistent 3D point tracking to complete occluded dynamic structures across frames, coupled with a loss-guided density control strategy that uses multi-view photometric errors and temporal marginal weights to prune Gaussian redundancy, achieving sub-minute high-fidelity 4D Gaussian Splatting convergence.
Method¶
Overall Architecture¶
Minute4D takes an unconstrained monocular casual video sequence \(I = \{I_i\}_{i=1}^N\) as input, with the objective of building an explicit 4D Gaussian scene representation capable of real-time photorealistic novel view synthesis at arbitrary timestamps \(t^*\) and viewpoints \(P^* \in SE(3)\). The overall pipeline comprises three serial stages: visual SLAM geometric initialization, the Segmentation-Tracking Enhancement module, and Loss-Guided Density Control during Gaussian optimization.
Initially, the framework employs MegaSaM, a deep visual SLAM system, to jointly estimate camera poses, intrinsics, temporally consistent depth maps, and motion probability maps. Unprojecting per-frame depth maps followed by voxel grid hash downsampling yields an initial point cloud. Next, the Segmentation-Tracking Enhancement module refines dynamic masks via SAM2 and tracks dynamic points across the entire video using TAPIP3D, infilling occluded regions from cross-frame observations to establish a geometrically complete enhanced point cloud. Finally, during Gaussian optimization, the framework periodically samples multi-view ground-truth and rendered frames, constructs normalized photometric error maps, and weights Gaussian densification and pruning using temporal Gaussian marginal distributions, thereby eliminating redundancy and achieving rapid convergence under minimal GPU memory.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Monocular Dynamic Video Input<br/>Uncalibrated/Casual Sequential Frames"] --> B["MegaSaM Geometric Initialization<br/>Jointly Estimates Poses, Depths & Motion Maps"]
B --> C["Segmentation-Tracking Enhancement<br/>SAM2 Mask Refinement + TAPIP3D Trajectory Completion"]
C --> D["4D Gaussian Scene Construction<br/>Initializes 4D Covariances & 4D Spherindrical Harmonics"]
D --> E["Loss-Guided Density Control<br/>Multi-View Error Maps + Temporal Gaussian Weighting"]
E --> F["High-Fidelity 4D Scene Representation<br/>~40s Ultra-Fast Training / >1200 FPS Real-Time Rendering"]
Key Designs¶
1. Segmentation-Tracking Enhancement: Cross-frame correspondence and occluded dynamic geometry completion Dynamic foreground objects in monocular casual videos experience frequent self-occlusions and view-dependent omissions. Directly initializing Gaussians from per-frame depth maps without correspondence leads to catastrophic frame-wise overfitting, where primitives cannot form cohesive 3D structures. To establish clean dynamic separation and complete geometry, this design first thresholds the motion probability maps from MegaSaM (\(\tau_{motion}=0.7\)) to extract coarse dynamic masks. The foreground pixels of the first frame initialize SAM2 prompts for instance tracking across all frames. To capture new dynamic instances emerging in later frames, coarse masks are uniformly sampled every \(B\) frames, the currently tracked mask is subtracted, and the residual dynamic regions are segmented by SAM2 to generate refined dynamic masks \(M^r\).
Following dynamic-static disentanglement, the framework employs the pre-trained 3D point tracking model TAPIP3D (\(f_\phi\)) to track 3D dynamic trajectories over time. To preserve efficiency without dense query overhead, dynamic points are uniformly subsampled and queried in batches to predict full temporal trajectories \(\{y_i\}_{i=1}^N\): $$ {y_i}{i=1}^N = f\phi(X, y_{i^}), \quad i^ \in {1, \ldots, N} $$ Merging these predicted trajectories with the initial point cloud \(X\) yields an enhanced point cloud \(X' = \bigcup_y \{y_i\}_{i=1}^N \cup X\). By backfilling surface details visible in alternate frames into previously occluded views, the initial Gaussians obtain precise and structurally complete geometric priors, preventing tearing and hole artifacts during novel view synthesis.
2. Loss-Guided Density Control: Multi-view error aggregation and temporally weighted adaptive population control Standard 3D and 4D Gaussian Splatting relies on single-frame screen-space positional gradients to trigger densification, while pruning only targets primitives with near-zero opacity or excessive spatial scales. This gradient-based heuristic easily generates excessive redundant Gaussians, causing memory bloat and severely decelerating training. Minute4D replaces heuristic gradients with a multi-view photometric mean absolute error (MAE) formulation. During optimization, \(K=10\) frames are sampled to compute pixel-wise photometric error maps \(e_i^{u,v} = \mathbb{E}_c | \hat{I}_i^{u,v} - \bar{I}_i^{u,v} |\). These error maps are min-max normalized and filtered by a threshold \(\tau_{error}=0.1\) to suppress low-error background noise, producing refined error maps \(E_i^{refined} = E_i \odot \mathbb{I}(\text{Normalize}(E_i) > \tau_{error})\).
Each 4D Gaussian \(G_j\) with center \(\mu_j\) is projected onto the sampled frames, yielding 2D footprint footprints \(\Omega_i^j\). Because the marginal distribution of a 4D Gaussian along the temporal axis follows a 1D normal distribution \(p^j(i) = \frac{1}{\sqrt{2\pi}\sigma_t^j} \exp\left(-\frac{(i-\mu_t^j)^2}{2(\sigma_t^j)^2}\right)\), its visual influence varies across timestamps. The densification score \(s_{den}^j\) weights multi-view errors by temporal proximity to the temporal mean: $$ s_{den}^j = \frac{1}{K} \sum_{i=1}^K \sum_{q \in \Omega_i^j} p^j(i) \cdot E_i^{refined}(q) $$ Gaussians exceeding \(\tau_{den}=0.001\) are prioritized for densification. Conversely, for pruning, the overall photometric loss \(L^i = (1-\lambda) L_1^i + \lambda L_{SSIM}^i\) is incorporated as a frame degradation weighting factor to formulate the pruning score \(s_{pru}^j\): $$ s_{pru}^j = \text{Normalize}\left( \sum_{i=1}^K \left( \sum_{q \in \Omega_i^j} p^j(i) \cdot E_i^{refined}(q) \right) \cdot L^i \right) $$ When \(s_{pru}^j > \tau_{pru}=0.009\), the Gaussian is identified as contributing persistently to degraded multi-view photometric reconstruction and is pruned. This adaptive control concentrates primitives strictly on necessary geometric boundaries, delivering a 13x training speedup and reducing peak GPU memory by nearly 9x.
Loss & Training¶
During the Gaussian optimization phase, supervision is provided exclusively by the input monocular video frames. The optimization objective employs a composite photometric loss combining \(L_1\) color error and structural similarity (\(L_{SSIM}\)): $$ L_{total} = (1 - \lambda) L_1(\hat{I}i, \bar{I}_i) + \lambda L_i) $$ where }(\hat{I}_i, \bar{I\(\lambda\) is set to 0.2. The optimization schedule is compact: 5,000 iterations for the Dycheck iPhone and DAVIS datasets, and 1,500 iterations for the NVIDIA Dynamic dataset. All experiments run on a single NVIDIA A6000 GPU, achieving complete convergence in approximately 20 to 40 seconds without requiring expensive offline COLMAP bundle adjustment.
Key Experimental Results¶
Main Results¶
Quantitative evaluations are conducted on the handheld Dycheck iPhone dataset (featuring complex parallax and non-rigid motions) and the multi-camera NVIDIA Dynamic dataset. On the Dycheck iPhone benchmark under uncalibrated (pose-free) settings, Minute4D outperforms previous state-of-the-art Instant4D(Full) by +1.25 dB in average PSNR (25.77 dB vs. 24.52 dB). Training runtime is slashed from 0.12 hours (~7.2 minutes) to 0.008 hours (~28.8 seconds), while peak GPU memory drops from 8.0 GB to 1.0 GB.
On the NVIDIA Dynamic benchmark, Minute4D attains the highest rendering quality among all pose-free visual SLAM-based methods (PSNR 24.73 dB), while finishing training in merely 0.006 hours (~21.6 seconds) and rendering at an ultra-fast 1218 FPS.
| Method | Calibration | Dycheck PSNR (dB) โ | NVIDIA PSNR (dB) โ | Runtime (h) โ | Peak Mem (GB) โ | Rendering FPS โ |
|---|---|---|---|---|---|---|
| D-NeRF | Ground Truth | 21.50 | - | > 24 | 12 | < 1 |
| RoDynRF | Ground Truth | 16.80 | 25.89 (COLMAP) | 22 | 15 | 0.42 |
| 4D-GS | Ground Truth / COLMAP | 21.64 | 21.45 | 1.2 | 21 | 43 |
| Deform3D | Ground Truth | 22.63 | - | - | - | - |
| RoDynRF (w/o pose) | Uncalibrated (Joint Opt) | 14.90 | - | 22 | 15 | - |
| RoDyGS | Visual Geometry (MASt3R) | 17.37 | - | 1.0 | - | - |
| 4DGS* | Visual SLAM | - | 18.34 | 0.16 | - | 98 |
| InstantSplat* | Visual SLAM | - | 22.56 | 0.15 | - | 117 |
| Instant4D (Lite) | Visual SLAM (MegaSaM) | 23.02 | - | 0.03 | 1.1 | - |
| Instant4D (Full) | Visual SLAM (MegaSaM) | 24.52 | 23.99 | 0.12 / 0.02 | 8.0 | 822 |
| Minute4D (Ours) | Visual SLAM (MegaSaM) | 25.77 | 24.73 | 0.008 / 0.006 | 1.0 | 1218 |
Ablation Study¶
Ablation experiments on the Dycheck iPhone benchmark validate the distinct contributions of the proposed modules. Removing the Segmentation-Tracking Enhancement module (w/o Seg-Track) drops PSNR from 25.77 dB to 24.60 dB (-1.17 dB) and SSIM from 0.717 to 0.628, confirming that dynamic 3D point trajectory completion is vital for restoring occluded surfaces. Disabling Loss-Guided Density Control (w/o Loss-Guided) causes runtime to surge from 29.22 seconds to 397.73 seconds (>13x increase) and peak memory to jump from 1088 MB to 9406 MB (>8.6x increase), validating its primary role in suppressing redundant primitives.
| Config | PSNR (dB) โ | SSIM โ | Runtime (sec) โ | Peak Mem (MB) โ | Note |
|---|---|---|---|---|---|
| Full Settings | 25.77 | 0.717 | 29.22 | 1088 | Full model: optimal fidelity, 29s runtime, low memory |
| w/o Seg-Track | 24.60 | 0.628 | 31.78 | 1064 | Severe drop in quality (-1.17 dB); occlusions suffer holes and artifacts |
| w/o Loss-Guided | 25.74 | 0.694 | 397.73 | 9406 | Runtime surges by 13.6x, memory surges by 8.6x due to uncontrolled redundancy |
Key Findings¶
- Complementary System Design: Geometric enhancement dictates reconstruction fidelity and spatial completeness (elevating the performance ceiling), whereas loss-guided density control eliminates redundant primitives and enforces rapid convergence (slashing computational cost).
- Infilling Dynamic Occlusions: Qualitative comparisons across rotating windmills, walking camels, and human acrobatics show that while baseline methods produce torn or blurry surfaces due to missing single-frame observations, Minute4D faithfully recovers sharp geometric silhouettes.
- Extreme Inference Throughput: By suppressing redundant primitives, the optimized 4D Gaussian representation achieves 1218 FPS rendering speeds on consumer GPUs, unlocking low-latency interactive streaming for dynamic digital twins.
Highlights & Insights¶
- Decoupling Heavy Foundation Priors from Differentiable Rasterization: Rather than optimizing heavy neural deformation networks during training, Minute4D restricts foundation models (SAM2, TAPIP3D) to an offline preprocessing stage that constructs a robust initial scaffolding, preserving the raw rasterization speed of explicit Gaussians.
- Spatio-Temporal Error Projection via Marginal Probability Weighting: Incorporating 4D Gaussian temporal variance and mean into multi-view error back-projection gives densification and pruning physical meaning, offering a clean blueprint for general 4D dynamic primitive management.
Limitations & Future Work¶
- Dependency on Upstream SLAM Robustness: Reconstruction fidelity is bounded by MegaSaM's pose and depth accuracy. Extreme dynamic occlusions, severe motion blur, or textureless scenes that degrade visual SLAM tracking will inevitably impact downstream geometry and novel view rendering.
- Scalability to Long Sequences and Topological Mutations: While validated on casual clips with hundreds of frames, scaling to thousands of frames with complex non-rigid topological changes will require hierarchical temporal partitioning and dynamic sub-map blending.
Related Work & Insights¶
- vs Instant4D: Instant4D also leverages MegaSaM for fast initialization but relies solely on unprojected per-frame depth maps without cross-frame dynamic point tracking, leading to single-frame overfitting. Minute4D introduces SAM2+TAPIP3D trajectory completion and temporal-loss-guided density control, achieving a 5x+ training speedup and +1.25 dB higher PSNR.
- vs 4D-GS (Wu et al.) / Deformable 3DGS: Conventional deformation-based 4DGS approaches optimize continuous deformation fields (via MLPs or HexPlanes), requiring 1-2 hours of training alongside offline COLMAP camera calibration. Minute4D models dynamic scenes via explicit 4D Gaussians and multi-view photometric pruning, achieving convergence in tens of seconds without requiring calibrated poses.
Rating¶
- Novelty: โญโญโญโญโ Elegantly integrates foundation segmentation with 3D point tracking for occluded dynamic completion, combined with temporally weighted 4D error-driven density control.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive quantitative and qualitative evaluations across Dycheck iPhone, NVIDIA Dynamic, and in-the-wild DAVIS datasets.
- Writing Quality: โญโญโญโญโญ Clear exposition, self-consistent math notations, and informative qualitative visualizations.
- Value: โญโญโญโญโญ Compresses uncalibrated monocular dynamic 4D scene reconstruction to ~40 seconds while achieving real-time rendering at >1200 FPS, offering high practical utility for robotic vision and embodied spatial computing.