Tempo-SAM3D: Monocular Video to 4D via Temporal Memory-Guided Generation¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: Pending release
Area: 3D Vision
Keywords: monocular video to 4D generation, temporal memory guidance, layout-aware generation, spatial grounding, training-free
TL;DR¶
Tempo-SAM3D introduces the first training-free framework that lifts layout-aware single-image 3D generation to temporally coherent 4D generation from monocular video by synergizing dual-path temporal memory caches, solver-level velocity correction, and global HexPlane temporal smoothing.
Background & Motivation¶
Recent breakthroughs in feed-forward 3D generation and flow matching have enabled the synthesis of high-fidelity textured 3D assets from a single image in seconds. A particularly significant milestone is the advent of layout-aware 3D generation models such as SAM3D, which not only reconstruct the geometry and appearance of isolated objects but also recover their spatial arrangement within the physical scene—namely scale, rotation, and translation. This scene-level spatial understanding provides a crucial bridge between isolated asset creation and real-world dynamic scenarios such as robotics, autonomous driving, and augmented reality, where incoming visual observations naturally arrive as sequential video streams.
However, directly applying single-image generation models independently to each video frame yields severe temporal inconsistencies. Because each frame is processed in total isolation, the network is forced to hallucinate unobserved or occluded regions anew at every time step. This produces erratic structural discontinuities in the 3D latent space, glaring texture flickering across frames, and jittering spatial pose trajectories. Prior video-to-4D approaches largely sidestep this by adopting a canonical-plus-deformation paradigm, where a shared canonical 3D model is reconstructed and animated via a learned continuous deformation field. While naturally smooth, this paradigm inherently assumes fixed topology, struggles with large-scale non-rigid shape variations, and completely discards scene-level spatial grounding by operating in an isolated, normalized coordinate space.
Bridging the gap between per-frame generation quality and temporal coherence without forfeiting layout awareness requires addressing consistency across multiple complementary levels rather than relying on a single ad-hoc constraint. The core tension lies in injecting cross-frame continuity while preserving the powerful pre-trained spatial generative prior, all without expensive model retraining. Core idea: introduce a multi-level temporal consistency framework that injects dual-path memory caches into cross- and self-attention to propagate visual and structural memory, enforces trajectory smoothness and observation alignment via gradient-guided velocity correction at the ODE solver level, and regularizes residual high-frequency jitter through global HexPlane spatiotemporal decomposition.
Method¶
Overall Architecture¶
Tempo-SAM3D builds upon the layout-aware 3D generation architecture of SAM3D (grounded in TRELLIS). Given a monocular video sequence, estimated dynamic point maps (e.g., from MoGe or DUSt3R), and object segmentation masks, the system employs a two-stage flow matching diffusion transformer (DiT). Stage 1 generates a coarse sparse voxel structure along with the object's 3D layout pose (scale, rotation, translation) within the scene. Stage 2 refines this coarse voxel structure into high-resolution structured latents encoding sharp geometry and rich texture conditioned on image patch features.
During ODE sampling at step \(s\), the model calculates a lookahead clean state prediction \(\hat{x}_0 = x_s - s \cdot v_s\). Tempo-SAM3D establishes temporal coherence across three distinct, complementary tiers: internally at the attention level via Dual-Path Temporal Attention KV caching; at the solver level via gradient-guided velocity correction enforcing trajectory smoothness and geometric/rendering alignment; and post-generation via low-rank HexPlane decomposition and joint pose optimization.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
Input["Monocular Input Sequence<br/>RGB Video Frames + Point Maps + Segmentation Masks"] --> Feat["Two-Stage Flow Matching Backbone<br/>Stage 1 Sparse Voxel Structure / Stage 2 Structured Latents"]
Feat --> DPA["Dual-Path Temporal Attention<br/>Cross-frame KV caches: Visual Memory + Structural Memory"]
DPA --> GGC["Gradient-Guided Velocity Correction<br/>ODE Sampling: Latent Trajectory Smoothing + Multimodal Alignment"]
GGC --> Dec["Per-Frame Preliminary 4D Generation<br/>Decoded Coarse Geometry Voxels and Fine Textures"]
Dec --> HPS["HexPlane Temporal Smoothing & Global Pose Refinement<br/>Low-rank Spatiotemporal Factorization + Trajectory Optimization"]
HPS --> Out["Spatially-Grounded 4D Content<br/>Temporally coherent, geometrically detailed, and accurately placed"]
Key Designs¶
1. Dual-Path Temporal Attention: cross-frame propagation of visual and structural memory When generating video frames independently, object rotation causes previously observed regions to become occluded, forcing the model to hallucinate appearance from scratch and leading to texture drift and flickering. Meanwhile, unobserved back-facing geometry lacks cross-frame anchoring, resulting in severe structural jitter. To solve this, Dual-Path Temporal Attention maintains separate, importance-driven key-value (KV) caches in both cross-attention and self-attention modules:
In the cross-attention pathway, a KV pool \(\mathcal{P}_{\text{cross}}\) of bounded capacity \(N_{\text{cross}}\) stores 2D image patch features from historical frames. During cross-attention, cached historical patches are concatenated with the current frame's patches along the sequence dimension: $\(K = [K_{\text{img}}; K_{\text{pool}}], \quad V = [V_{\text{img}}; V_{\text{pool}}]\)$ This enables 3D query tokens to retrieve visual appearance that was clearly visible in earlier frames but is currently occluded. To prevent obsolete features from clogging the cache, patch eviction combines complementarity (dissimilarity to current patches) with usage (the actual attention weight received from 3D tokens during generation), while admission prioritizes current patches that maximize visual diversity relative to the pool.
Simultaneously, the cross-attention distribution provides an implicit confidence metric via normalized entropy: $\(H(l) = -\frac{1}{\log P}\sum_{p} a_l^{(p)} \log a_l^{(p)}\)$ where \(a_l^{(p)}\) denotes the cross-attention weight from 3D token \(l\) to patch \(p\). Directly visible tokens attend to specific local patches with sharp, low-entropy distributions, whereas occluded tokens diffuse attention across the entire image with high entropy. The observation confidence is defined as \(C(l) = 1 - H(l)\).
In the self-attention pathway, a structural KV pool \(\mathcal{P}_{\text{self}}\) of capacity \(N_{\text{self}}\) retains 3D latent tokens from historical frames. Cached structural tokens are concatenated with current tokens, allowing current queries to attend to historical 3D representations. Each candidate token is scored under a unified metric combining confidence \(C(l)\), popularity \(\text{Pop}(l)\) (total self-attention received from other tokens, highlighting geometric edges and articulation hubs), and exponential temporal decay: $\(S^{\text{self}}(l) = C(l) \cdot \text{Pop}(l) \cdot \exp(-\lambda_{\text{decay}}\cdot \Delta t)\)$ Retaining high-scoring structural anchor tokens provides a stable geometric prior that effectively suppresses arbitrary hallucinations in unobserved regions.
2. Gradient-Guided Velocity Correction: solver-level trajectory smoothing and multimodal observation alignment While attention-level memory softly biases generation toward temporal coherence, it offers only implicit coupling and cannot enforce rigid physical adherence to instantaneous observations. Tempo-SAM3D introduces test-time gradient guidance into the flow matching ODE solver. At each sampling step \(s\), the lookahead estimate \(\hat{x}_0 = x_s - s \cdot v_s\) is evaluated, and the velocity vector is steered along the gradient of guidance objectives: $\(v_{\text{guided}} = v_s - \alpha_s \cdot \nabla_{\hat{x}_0}\mathcal{L}_{\text{guide}}\)$ The total guidance loss combines three complementary terms: \(\mathcal{L}_{\text{guide}} = \lambda_{\text{traj}}\mathcal{L}_{\text{traj}} + \lambda_{\text{geo}}\mathcal{L}_{\text{geo}} + \lambda_{\text{render}}\mathcal{L}_{\text{render}}\).
First, Temporal Guidance \(\mathcal{L}_{\text{traj}} = \|\hat{x}_0^{(t)} - \hat{x}_0^{(t-1)}\|^2\) operates directly in latent space across all ODE steps, penalizing abrupt frame-to-frame trajectory deviations and acting as a continuous stabilizing anchor.
Second, Alignment Guidance anchors generation to physical observations at stage-specific intervals. In Stage 1, the lookahead state is decoded into coarse voxel geometry \(\tilde{\mathcal{V}}\) and spatial pose \(T_t = (s_t, R_t, t_t)\), and steered toward the reconstructed point map \(\mathcal{P}_t\) using a one-directional Chamfer distance: $\(\mathcal{L}_{\text{geo}} = \frac{1}{|\mathcal{P}_t|}\sum_{p \in \mathcal{P}_t} \min_{q \in T_t(\tilde{\mathcal{V}})} \|p - q\|^2\)$ A strictly one-directional distance from the point map to the generated structure is required because single-view point maps cover only the front-facing visible surface; a bidirectional distance would mistakenly penalize valid unobserved back-facing geometry. In Stage 2, the lookahead representation is rendered via a differentiable renderer, and perceptual loss (LPIPS) is computed against both the current input frame and the previous frame's rendering, ensuring sub-pixel color fidelity and perceptual continuity.
3. HexPlane Temporal Smoothing & Global Pose Refinement: low-rank spatiotemporal regularization and trajectory filtering Because online memory propagation and ODE guidance operate strictly causally, minor high-frequency temporal noise and slight frame-to-frame pose drift can persist across extended video sequences. A non-causal global post-processing stage is applied once all frame representations have been generated.
For latent feature regularization, the 4D sequence of per-frame latent feature volumes is factorized into six compact 2D feature planes via HexPlane decomposition: three spatial planes (\(XY, XZ, YZ\)) modeling static 3D structure and three spatiotemporal planes (\(XT, YT, ZT\)) capturing temporal dynamics. Because stochastic per-frame fluctuations cannot project coherently onto low-rank spatiotemporal planes, reconstructing features through this representation acts as an implicit spatiotemporal low-pass filter, suppressing high-frequency texture and structural jitter while preserving dynamic motion.
For spatial positioning, the sequence of estimated poses \(\{T_t\}_{t=1}^T\) is jointly optimized across the entire video. The objective balances instantaneous one-directional Chamfer alignment against the observed point maps with temporal smoothness penalties on consecutive scale, rotation, and translation changes. This global optimization resolves causal accumulated drift, resulting in physically stable, jitter-free 4D trajectories grounded in the scene frame.
Loss & Training¶
Tempo-SAM3D is an entirely training-free framework that leaves all pre-trained base model weights frozen. Experiments are conducted on a single NVIDIA A100 (80GB) GPU. Sampling schedules follow SAM3D defaults: 50 ODE steps for Stage 1 (sparse structure generation) and 25 ODE steps for Stage 2 (structured latent generation).
Key hyper-parameters include: - Self-attention KV cache capacity \(N_{\text{self}}\): 8,000 for Stage 1, 15,000 for Stage 2; - Cross-attention KV cache capacity \(N_{\text{cross}} = 1,024\); - Temporal decay rate in importance scoring \(\lambda_{\text{decay}} = 0.1\); - Guidance weights: \(\lambda_{\text{traj}} = 0.3\) with a linear decay schedule; \(\lambda_{\text{geo}} = 0.05\) (active during the final 50% of Stage 1 steps); \(\lambda_{\text{render}} = 0.1\) (active during the final 50% of Stage 2 steps); - HexPlane temporal resolution is set to 48 with a feature plane dimension of 32.
Key Experimental Results¶
Main Results¶
Quantitative evaluations strictly adhere to the V2M4 benchmark protocol, covering 40 animation sequences partitioned into Simple (Consistent4D, 20 sequences with subtle movements) and Complex (Mixamo and Sketchfab, 20 sequences with large-scale articulated motions). Metrics include CLIP visual similarity, LPIPS, DreamSim perceptual distance, and Fréchet Video Distance (FVD).
| Dataset Split | Method | CLIP ↑ | LPIPS ↓ | FVD ↓ | DreamSim ↓ |
|---|---|---|---|---|---|
| Simple (Consistent4D) | TRELLIS | 0.8905 | 0.1597 | 1342.66 | 0.1282 |
| DreamMesh4D | 0.8692 | 0.1019 | 914.28 | 0.0937 | |
| V2M4 | 0.9259 | 0.1017 | 825.59 | 0.0688 | |
| Tempo-SAM3D (Ours) | 0.9312 | 0.0992 | 841.27 | 0.0702 | |
| Complex (Mixamo / Sketchfab) | TRELLIS | 0.8887 | 0.1265 | 1216.19 | 0.1492 |
| DreamMesh4D | 0.8256 | 0.0804 | 1079.02 | 0.1850 | |
| V2M4 | 0.9008 | 0.0747 | 666.04 | 0.1220 | |
| Tempo-SAM3D (Ours) | 0.9085 | 0.0728 | 642.38 | 0.1165 |
To evaluate spatially-grounded generation in natural settings, the authors introduce the Tempo-SAM3D dataset (20 video sequences with complex dynamic backgrounds and reconstructed 3D/4D spatial reference data). Beyond rendering quality (LPIPS, PSNR, FVD), 3D spatial grounding accuracy is rigorously benchmarked via 3D bounding box IoU, Depth RMSE (D-RMSE), and pose accuracy within a 5 cm error margin (ACC5cm).
| Method | LPIPS ↓ | PSNR ↑ | FVD ↓ | IoU (%) ↑ | D-RMSE ↓ | ACC5cm (%) ↑ |
|---|---|---|---|---|---|---|
| TRELLIS | 0.165 | 17.82 | 1456.28 | — | — | — |
| V2M4 | 0.138 | 18.76 | 781.56 | — | — | — |
| SAM3D (Per-frame baseline) | 0.128 | 19.45 | 1218.63 | 73.5 | 0.028 | 90.3 |
| Tempo-SAM3D (Ours) | 0.119 | 20.12 | 732.45 | 83.8 | 0.018 | 95.8 |
Ablation Study¶
A progressive ablation study on the Tempo-SAM3D dataset dissects the exact contribution of each proposed module toward temporal consistency, visual fidelity, and spatial grounding accuracy.
| Config | FVD ↓ | LPIPS ↓ | PSNR ↑ | IoU (%) ↑ | D-RMSE ↓ | ACC5cm (%) ↑ | Note |
|---|---|---|---|---|---|---|---|
| (a) Baseline (SAM3D) | 1218.63 | 0.128 | 19.45 | 73.5 | 0.028 | 90.3 | Independent per-frame generation |
| (b) + Self-Attn Memory | 1076.52 | 0.131 | 19.32 | 73.1 | 0.029 | 89.8 | Structural memory reduces FVD (-142.1) |
| (c) + Temporal Guidance | 935.18 | 0.130 | 19.38 | 73.3 | 0.028 | 90.1 | Trajectory smoothness lowers FVD (-141.3) |
| (d) + Alignment Guidance | 871.25 | 0.121 | 20.18 | 82.6 | 0.019 | 95.2 | Spatial driver: +9.5% IoU, cuts D-RMSE |
| (e) + HexPlane (Full Model) | 732.45 | 0.119 | 20.12 | 83.8 | 0.018 | 95.8 | Suppresses high-frequency jitter; best overall |
Key Findings¶
- Superiority on large articulated motions: On complex sequences involving large non-rigid deformations (Mixamo and Sketchfab), canonical-deformation baselines (e.g., DreamMesh4D) often fail due to rigid topological assumptions. Tempo-SAM3D's per-frame generative paradigm anchored by temporal memory excels, achieving substantially lower FVD (642.38 vs 666.04) and sharper rendering fidelity.
- Alignment guidance is the core driver of spatial grounding: As revealed by the ablation study, adding attention memory and temporal trajectory guidance (stages b and c) primarily enhances temporal smoothness (FVD drops from 1218.63 to 935.18) but barely alters 3D IoU or depth error. The introduction of one-directional geometric alignment guidance (stage d) produces an immediate jump in 3D IoU (73.3% to 82.6%) and ACC5cm (90.1% to 95.2%), demonstrating that explicit geometric forces are indispensable for physical scene grounding.
- Qualitative necessity of cross-attention visual memory: Because appearance consistency in unobserved regions cannot be measured by input-view 2D metrics, a qualitative turning worker experiment demonstrates its impact: without cross-attention visual memory, front-facing jacket textures seen at \(T_0\) are completely forgotten once the worker turns away at \(T_1-T_3\), resulting in garbled hallucinations; with the memory pool, earlier front-view textures are accurately retrieved and rendered from novel views.
Highlights & Insights¶
- Full-stack, multi-level temporal synergy: Instead of undertaking costly 4D diffusion model pretraining, temporal consistency is cleanly decomposed into attention-level memory caches, solver-level ODE gradient guidance, and post-processing low-rank decomposition, delivering competitive 4D generation in a completely training-free manner.
- Principled one-directional geometric guidance: Recognizing that single-view point maps only cover the front-facing shell of an object, using a strictly one-directional Chamfer distance avoids penalizing plausible, unobserved back-facing geometry while firmly anchoring the visible surface to scene depth.
- Entropy- and popularity-driven memory eviction: Moving beyond naive FIFO caching, cross-attention entropy dynamically diagnoses observation confidence, while self-attention centrality identifies structural boundary hubs. This introspection-guided cache management provides a compelling blueprint for test-time scaling in multimodal generative models.
Limitations & Future Work¶
- Reliance on upstream monocular geometry and segmentation: Geometric alignment guidance depends on clean point maps and camera trajectories. If monocular depth or pointmap predictors (e.g., DUSt3R or MoGe) suffer severe scale drift or distortion under extreme dynamic occlusions, errors will propagate directly into the 4D generation process.
- Test-time gradient computational overhead: While avoiding offline training, running backpropagation through differentiable rendering and Chamfer distance during ODE sampling incurs non-trivial computation, currently precluding real-time interactive applications.
- Future directions: Integrating tightly coupled end-to-end dynamic geometry estimators (such as MonST3R or Flow3R) and reformulating HexPlane smoothing into a causal sliding-window low-rank filter could unlock real-time streaming 4D scene reconstruction from video.
Related Work & Insights¶
- vs V2M4: V2M4 relies on multi-stage mesh optimization pipelines that are slow and confined to normalized, isolated coordinates; Tempo-SAM3D operates via flow matching guidance, naturally producing spatially-grounded 4D assets positioned within physical scene coordinates with better temporal coherence on complex motions.
- vs DreamMesh4D: DreamMesh4D employs a hybrid Gaussian-mesh representation constrained by canonical deformation assumptions; Tempo-SAM3D generates assets per-frame with memory anchoring, naturally accommodating arbitrary topological changes and large deformations.
- vs SAM3D: SAM3D is restricted to static single-image inputs; Tempo-SAM3D successfully extends it across the temporal dimension, retaining its layout-awareness capabilities while eliminating frame-to-frame geometric jitter and appearance flickering.
Rating¶
- Novelty: ⭐⭐⭐⭐ [First training-free framework extending layout-aware single-image 3D diffusion to spatially-grounded 4D generation]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive quantitative benchmarks on V2M4 and the realistic Tempo-SAM3D dataset with detailed multi-level ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear structural organization and elegant mathematical formulation across attention, solver, and post-processing tiers]
- Value: ⭐⭐⭐⭐⭐ [Highly practical paradigm for dynamic scene understanding, embodied AI spatial grounding, and low-cost 4D asset creation]