Skip to content

MVFusion-GS: Motion-Variance Guided Temporal Attention for High-Quality Dynamic Gaussian Splatting

Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/toseeai-com/MVFusion-GS
Area: 3D Vision
Keywords: Dynamic Gaussian Splatting, Motion Variance, Temporal Attention, Dynamic-Static Decomposition, Distractor-free

TL;DR

To resolve background contamination and inaccurate foreground deformation caused by the lack of explicit motion awareness in dynamic 3D Gaussian Splatting, this paper introduces a plug-in refinement framework combining global motion-variance trajectory priors with local temporal cross-attention.

Background & Motivation

Extending 3D Gaussian Splatting (3DGS) to dynamic scenes via deformation fields has garnered intense attention, underpinning two representative tasks: high-fidelity dynamic scene reconstruction and distractor-free static scene reconstruction. Dynamic scene reconstruction strives to accurately map deformable foreground trajectories, whereas distractor-free reconstruction aims to eliminate transient occluders (such as walking pedestrians or passing vehicles) from casual video captures to restore a pristine static background. Decoupled dual-branch frameworks like DeGauss attempt to bridge these divergent goals by maintaining a static Gaussian set alongside a HexPlane-deformed dynamic Gaussian set, blending them through a learnable composition mask.

However, standard spatio-temporal deformation networks predict coordinate and attribute offsets purely based on local spatio-temporal coordinate grids, lacking explicit awareness of each Gaussian primitive's long-term motion magnitude or cross-frame temporal dependencies. When foreground objects undergo subtle, low-amplitude, or brief transient motions, the deformation network frequently underfits, failing to output sufficient offsets and mistaking dynamic components for stationary geometry. These pseudo-static foreground Gaussians remain trapped in the static branch, leaving severe ghosting artifacts in distractor-free backgrounds while causing blur and motion smearing in the reconstructed dynamic foreground.

This paper attacks the problem from a clear perspective: since local grid features cannot reliably differentiate subtle motion from static structure, deformation networks must be explicitly informed by long-term motion statistics and short-term temporal context. Core idea: extract a 13D global motion-variance trajectory signature via periodic offline sampling to guide coarse dynamic-static separation, combined with query-centered local temporal cross-attention to aggregate short-term motion context, refining deformation features in a plug-in manner.

Method

Overall Architecture

MVFusion-GS operates on top of a decoupled dynamic-static Gaussian Splatting framework, treating the baseline deformation network as a coarse predictor and refining its representations purely in feature space. The system integrates two complementary motion-aware branches: the Motion-Variance Guided Refinement (MVG) module periodically samples deformation trajectories across global timestamps to extract per-Gaussian displacement, rotation, and scale variances; the MotionFormer Temporal Attention (MFTA) module aggregates local motion context within a sliding temporal window via query-centered cross-attention. The resulting motion features are added directly to the baseline deformation feature before entering the unmodified deformation heads \(\Phi\), predicting refined offsets to update dynamic Gaussians before dual-branch composited rendering.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Input Dynamic Gaussians & Timestamp t"] --> BaseFeat["Baseline Deformation Network<br/>Extract Coarse Feature h_base"]
    BaseFeat --> MVG["Motion-Variance Guided Refinement (MVG)<br/>Global Trajectory Signature + Local Variance Dict"]
    BaseFeat --> MFTA["MotionFormer Temporal Attention (MFTA)<br/>Temporal Window Cross-Attention"]
    MVG --> FeatFusion["Feature-Space Residual Fusion<br/>h_base + h_MVG + h_MFTA"]
    MFTA --> FeatFusion
    FeatFusion --> HeadDecode["Unchanged Deformation Heads Φ<br/>Predict Refined Offset ΔG_d"]
    HeadDecode --> Render["Decoupled Rendering & Mask Composition<br/>Pristine Static Background & Sharp Dynamic Foreground"]

Key Designs

1. Motion-Variance Guided Refinement (MVG): Injecting explicit motion priors from global trajectory statistics

To address the underfitting of subtle motion in conventional spatio-temporal grids that falsely classifies moving Gaussians as static structures, MVG periodically samples timestamps during training without back-propagating gradients, evaluating the deformation network across all dynamic Gaussians to simulate full trajectories. It computes the mean and variance of spatial displacement, scalar rotation magnitude, and both isotropic and anisotropic log-scale deviations across time, constructing a compact 13D global trajectory signature along with a scalar motion intensity score:

\[s_i^{\text{intensity}} = 0.7\,\sigma_{\text{pos}}(i) + 0.2\,\sigma_{\text{scale}}(i) + 0.1\,\sigma_{\text{rot}}(i)\]

To complement this long-term signature with instantaneous variation cues, MVG also caches a sliding-window local displacement variance dictionary over sparse timestamps and interpolates it at arbitrary query times \(t\). Both descriptors are mapped by a lightweight MLP into a motion feature \(h_{\text{MVG}}\), providing unambiguous dynamic-static discriminative cues.

2. MotionFormer Temporal Attention (MFTA): Aggregating local temporal dependencies via cross-attention

While the global variance signature offers a robust summary of overall motion amplitude, it cannot capture frame-to-frame temporal coherence and instantaneous interactions. MFTA constructs a local temporal neighborhood around the current query timestamp \(t_i\). For the query frame and neighboring frames, composite temporal tokens are assembled from deformation hidden features, coarse deformation predictions, and instantaneous local variance values. Using the current frame's token as Query and neighboring tokens as Key and Value, query-centered cross-attention is performed:

\[\mathbf{a}_{i, t_i} = \mathrm{Attn}\left(\mathbf{z}_{i,t_i},\,\mathbf{z}_{i,t_{i-w}},\,\mathbf{z}_{i,t_{i+w}}\right)\]

This mechanism allows each Gaussian to dynamically attend to its short-term temporal context weighted by instantaneous motion energy, smoothing sudden trajectory jitters and resolving ambiguous motion directions into a temporal motion feature \(h_{\text{MFTA}}\).

3. Feature-Space Plug-in Residual Fusion: Non-invasive integration and progressive three-stage training

To maximize compatibility with existing dynamic Gaussian architectures (such as 4DGS or DeGauss), MVG and MFTA operate strictly within the latent feature space. The baseline deformation feature is fused additively with the motion-aware features and fed directly into the original deformation heads \(\Phi\), requiring no alterations to downstream rasterization or head architectures. A progressive three-stage training regime stabilizes optimization: Stage 1 optimizes base geometry and densification without deformation; Stage 2 activates the baseline deformation network to learn coarse motion fields and begins collecting trajectory statistics; Stage 3 activates zero-initialized MVG and MFTA branches for end-to-end refined optimization. At inference, cached trajectory descriptors are directly reused, and the temporal attention module can optionally be disabled for higher efficiency.

Loss & Training

The overall training objective combines photometric reconstruction losses with standard Gaussian geometric regularization:

\[\mathcal{L} = \lambda_{\text{rgb}}\mathcal{L}_{\text{rgb}} + \lambda_{\text{ssim}}\mathcal{L}_{\text{ssim}} + \lambda_{\text{reg}}\mathcal{L}_{\text{reg}}\]

where \(\mathcal{L}_{\text{rgb}}\) is the \(L_1\) photometric loss, \(\mathcal{L}_{\text{ssim}}\) enforces structural similarity, and \(\mathcal{L}_{\text{reg}}\) constrains Gaussian scale ratios and sizes. Optimization is carried out on a single NVIDIA RTX 3090 GPU using the Adam optimizer for 30k iterations. Trajectory statistics in MVG are updated every 2000 iterations over \(K=64\) sampled timestamps, and the temporal window size in MFTA is set to \(w=5\) frames.

Key Experimental Results

Main Results

The method was evaluated on both distractor-free static reconstruction benchmarks (NeRF On-the-go, evaluating novel views of clean static backgrounds) and full dynamic scene reconstruction benchmarks (Neu3D).

Table 1: Distractor-free static scene reconstruction on NeRF On-the-go

Method Mountain (PSNR / SSIM / LPIPS) Corner (PSNR / SSIM / LPIPS) Patio-High (PSNR / SSIM / LPIPS) Mean (PSNR↑ / SSIM↑ / LPIPS↓)
RobustNeRF 17.54 / 0.496 / 0.383 23.04 / 0.764 / 0.244 20.54 / 0.578 / 0.366 19.64 / 0.583 / 0.369
3DGS 19.40 / 0.638 / 0.213 20.90 / 0.713 / 0.241 17.29 / 0.604 / 0.363 19.30 / 0.668 / 0.253
WildGaussians 20.43 / 0.653 / 0.255 24.16 / 0.822 / 0.139 22.23 / 0.725 / 0.206 22.16 / 0.746 / 0.182
DeSplat 19.59 / 0.715 / 0.175 26.05 / 0.885 / 0.095 22.59 / 0.845 / 0.125 22.58 / 0.813 / 0.130
SpotlessSplats 21.64 / 0.725 / 0.195 25.77 / 0.877 / 0.117 22.98 / 0.808 / 0.155 23.42 / 0.813 / 0.145
DeGauss 22.31 / 0.746 / 0.163 25.94 / 0.869 / 0.078 23.35 / 0.799 / 0.124 23.91 / 0.819 / 0.113
MVFusion-GS (Ours) 22.44 / 0.755 / 0.127 25.58 / 0.868 / 0.061 23.20 / 0.805 / 0.097 23.94 / 0.826 / 0.085

Table 2: Dynamic scene reconstruction on Neu3D

Method Cut Beef (PSNR / SSIM / LPIPS) Cook Spinach (PSNR / SSIM / LPIPS) Flame Steak (PSNR / SSIM / LPIPS) Mean (PSNR↑ / SSIM↑ / LPIPS↓)
K-Planes 31.82 / 0.966 / 0.114 32.60 / 0.966 / 0.114 32.39 / 0.970 / 0.102 31.63 / 0.964 / 0.117
4DGS 32.66 / 0.946 / 0.053 32.46 / 0.949 / 0.052 32.75 / 0.954 / 0.040 31.12 / 0.937 / 0.058
MangoGS 33.28 / 0.949 / 0.043 32.83 / 0.947 / 0.042 34.11 / 0.958 / 0.035 31.89 / 0.940 / 0.052
DeGauss 32.56 / 0.957 / 0.042 32.61 / 0.950 / 0.041 32.75 / 0.955 / 0.034 31.52 / 0.942 / 0.047
MVFusion-GS (4DGS plug-in) 33.31 / 0.952 / 0.052 33.01 / 0.954 / 0.051 33.45 / 0.957 / 0.039 31.66 / 0.942 / 0.057
MVFusion-GS (Full Decoupled) 33.62 / 0.960 / 0.040 33.41 / 0.954 / 0.039 34.13 / 0.959 / 0.032 32.07 / 0.943 / 0.046

Ablation Study

Table 3: Ablation study of core components (Neu3D and NeRF On-the-go)

Config Neu3D PSNR↑ / SSIM↑ / LPIPS↓ NeRF On-the-go PSNR↑ / SSIM↑ / LPIPS↓ Note
w/o MVG & MFTA (Base deformation) 31.52 / 0.942 / 0.047 23.86 / 0.807 / 0.122 Baseline without any motion-aware refinement
w/o MVG (no motion statistics) 31.64 / 0.937 / 0.052 23.87 / 0.810 / 0.114 Relies only on MFTA temporal attention
w/ MVG (position variance only) 31.93 / 0.943 / 0.046 23.92 / 0.815 / 0.096 Uses 3D position variance instead of full 13D signature
w/o MFTA 31.78 / 0.942 / 0.047 23.89 / 0.820 / 0.089 Uses MVG statistics only; lacks local temporal aggregation
w/o Cross-attention (Self-attention) 32.01 / 0.943 / 0.046 23.93 / 0.822 / 0.090 Self-attention without query-centered aggregation
Full Model 32.07 / 0.943 / 0.046 23.94 / 0.826 / 0.085 Full combination of 13D MVG and MFTA cross-attention

Key Findings

  • Motion trajectory statistics (MVG) are critical for dynamic discrimination: introducing position variance alone boosts Neu3D PSNR by 0.41 dB over the baseline, and the full 13D signature brings further gains by incorporating rotation and scale dynamics vital for articulated structures.
  • Temporal cross-attention (MFTA) prevents abrupt deformation errors: removing MFTA incurs a 0.29 dB PSNR drop on Neu3D, causing fast-moving elements like flames and liquids to lose sharp boundary definitions.
  • In distractor-free reconstruction, perceptual quality (LPIPS) shows marked improvement (dropping from 0.113 to 0.085 on NeRF On-the-go), fully eradicating the ghostly pedestrian silhouettes and street smudges typical of DeGauss.

Highlights & Insights

  • Gradient-free offline trajectory sampling: Constructing the 13D statistics and sparse time dictionary via periodic forward evaluations avoids tracking memory-intensive computation graphs across hundreds of frames, providing global dynamic priors with negligible overhead.
  • Unified synergy between distractor removal and dynamic reconstruction: Proves that distractor elimination and dynamic reconstruction are two sides of the same coin; accurate motion awareness correctly re-attributes subtle movers to the dynamic branch, simultaneously purifying backgrounds and sharpening foregrounds.
  • Non-invasive feature-space plug-in interface: Easily integrates into both decoupled dual-branch frameworks (DeGauss) and single-branch architectures (4DGS), yielding steady performance gains across diverse baselines.

Limitations & Future Work

  • Reliance on initial deformation baseline: If the baseline deformation network severely underfits due to heavy occlusions or insufficient views, the sampled trajectory variance loses discriminative power, potentially leaving residual artifacts.
  • Ambiguity between creeping motion and lighting shifts: Extremely subtle motion (e.g. slight foliage rustle) or gradual illumination changes over time can sometimes be difficult to distinguish from genuine static geometry using displacement/scale/rotation metrics alone.
  • Future directions: Incorporating foundation model semantic/depth priors (such as DINO or monocular depth estimators) to guide explicit Gaussian reassignment under sparse or ill-posed observation conditions.
  • vs DeGauss: DeGauss proposed a decoupled static/dynamic dual-branch with mask blending, but its deformation network lacks explicit motion cues, leaving ghosting residuals in the static background; MVFusion-GS introduces global MVG statistics and local MFTA attention to cleanly re-attribute pseudo-static Gaussians.
  • vs 4DGS / SC-GS: 4DGS and SC-GS use HexPlanes or sparse control nodes to model continuous temporal deformations but lack explicit characterization of individual Gaussian motion histories; MVFusion-GS serves as a universal plug-in that boosts 4DGS performance (improving Neu3D mean PSNR by 0.54 dB).
  • vs SpotlessSplats / WildGaussians: These methods rely on 2D pre-trained feature clustering or pixel-space uncertainty masks; MVFusion-GS grounds motion modeling directly in 3D Gaussian trajectory variance, offering a lighter-weight, geometrically faithful solution.

Rating

  • Novelty: ⭐⭐⭐⭐☆ [Elegant synthesis of global motion trajectory variance statistics and temporal cross-attention to resolve dynamic-static decomposition ambiguity]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Rigorous evaluation across NeRF On-the-go, RobustNeRF, and Neu3D datasets covering static backgrounds, dynamic foregrounds, and ablations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, well-formulated technical mechanisms, and consistent experimental presentation]
  • Value: ⭐⭐⭐⭐☆ [A versatile plug-in module readily applicable to various dynamic 3DGS baselines for both academic exploration and practical pipelines]