Skip to content

StreetForward: Perceiving Dynamic Street with Feedforward Causal Dynamics

Conference: ECCV 2026
Paper: ECCV 2026
Area: Autonomous Driving
Keywords: Dynamic Street Reconstruction, Feedforward 3DGS, Causal Temporal Attention, Cross-Frame Rendering Consistency, Tracker-Free 4D Perception

TL;DR

StreetForward introduces a pose-free, tracker-free, and segmentation-free feedforward framework for 4D dynamic street reconstruction, leveraging causal masked attention to disentangle temporal directed motion alongside spatio-temporal rendering consistency to achieve high-fidelity spatio-temporal view synthesis and depth estimation.

Background & Motivation

Closed-loop simulation and verification in autonomous driving crucially depend on fast, highly accurate 4D digital reconstruction of dynamic real-world environments. Traditional dynamic neural rendering frameworks based on Neural Radiance Fields (NeRFs) or 3D Gaussian Splatting (3DGS)—such as StreetGS and OmniRe—rely almost exclusively on computationally expensive, per-scene offline optimization. This optimization paradigm becomes prohibitively slow when confronted with massive, diverse, and continuously updated real-world driving data streams, posing a severe bottleneck for scalable closed-loop testing.

To overcome the latency of per-scene optimization, recent efforts extend general 3D foundation models (e.g., DUSt3R, VGGT) to recover scene geometry and camera poses in a single feedforward pass. However, porting feedforward reconstruction to dynamic 4D urban scenes encounters two fundamental hurdles: first, geometry degradation under motion—the alternating attention (AA) backbone in models like VGGT was inherently designed for unordered multi-view static scenes, where global cross-frame attention treats tokens from all timestamps symmetrically, failing to resolve the directed temporal flow (source → target) and yielding severe geometric blur and drift on moving foregrounds; second, brittle tracking dependencies and supervisory bottlenecks—existing feedforward dynamic methods (such as DGGT) heavily depend on external 3D point trackers to supply candidate trajectories, which frequently collapse or drift into empty space under occlusion and long-term extrapolation, while direct 4D scene flow ground truth remains virtually unobtainable at scale.

This paper's angle of attack is to eliminate external trackers and segmentation priors entirely, reformulating the feedforward attention mechanism through an explicit causal temporal lens and exploiting the physical deformation consistency of 3DGS across adjacent frames. Core idea: equip visual geometry transformers with causal masked attention to explicitly model directional source-to-target temporal dynamics, and supervise dense pixel-wise velocity fields and 4D Gaussians in an entirely tracker-free, self-supervised manner via cross-frame rendering exclusion, forward-backward symmetry, and local rigidity regularization.

Method

Overall Architecture

StreetForward takes an unposed monocular vehicle video sequence as input and directly predicts per-frame camera intrinsics and extrinsics, dense depth maps, full-scene 3D Gaussian primitives, and dense forward/backward velocity fields in a single forward pass, supporting high-fidelity rendering at arbitrary novel camera poses and timestamps. The pipeline operates across four stages: feature extraction and geometric initialization, causal dynamics modeling, bidirectional velocity decoding with dynamic masking, and joint optimization under spatio-temporal consistency constraints.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Driving Video<br/>Unposed monocular multi-frame images"] --> B["Alternating Attention Geometry Backbone<br/>DINOv2 + interleaved frame/global AA"]
    B --> C["Base Geometry & Pose Decoding<br/>Estimate intrinsics/extrinsics/depth/3DGS"]
    B --> D["Causal Masked Temporal Attention<br/>Sinusoidal time embeddings + directional causal mask"]
    D --> E["Bidirectional Velocity & Mask Decoding<br/>Regress forward/backward velocity + dynamic score"]
    C & E --> F["Disentangled Cross-Frame Rendering & Consistency<br/>Global static fusion + cross-frame dynamic propagation"]
    F --> G["4D Dynamic Street Scene<br/>High-fidelity synthesis at novel poses & timestamps"]

Key Designs

1. Causal Masked Attention: Resolving Directional Motion Ambiguity The standard alternating attention layers in VGGT disperse attention symmetrically across all frames, unable to differentiate physical forward and backward causality. StreetForward resolves this by first concatenating a sinusoidal temporal embedding \(\tau_f \in \mathbb{R}^{d_t}\) to every image patch token in frame \(f\), forming augmented tokens \(\tilde{\mathbf{X}}\) that are flattened into \(\tilde{\mathbf{Z}} \in \mathbb{R}^{B \times (F \cdot P) \times D'}\). A structured binary causal attention mask \(\mathbf{M}\) is then enforced: $$ \mathbf{M}[b, h, i, j] = \begin{cases} 1, & \text{if } \mathrm{frame}(i) = f_{\mathrm{s}} \text{ and } \mathrm{frame}(j) = f_{\mathrm{t}} \ 0, & \text{otherwise} \end{cases} $$ where \(f_{\mathrm{t}} = f_{\mathrm{s}} + 1\) during forward motion modeling and \(f_{\mathrm{t}} = f_{\mathrm{s}} - 1\) during backward modeling. The mask enters the attention matrix via a logarithmic barrier: $$ \mathbf{A}^{(h)} = \mathrm{softmax}\left(\frac{\mathbf{Q}^{(h)}{\mathbf{K}^{(h)}}^\top}{\sqrt{d_h}} + \log \mathbf{M}\right) \mathbf{V}^{(h)} $$ This design forces cross-frame attention to aggregate visual tokens strictly along the temporal arrow, preserving static geometric representations while generating highly discriminative motion-aware latent features \(\mathbf{Y}\).

2. Disentangled Static-Dynamic Representation and Tracker-Free Cross-Frame Propagation Prior dynamic Gaussian methods frequently assign an ephemeral lifespan parameter to individual Gaussians, which commonly leads to flickering and sudden disappearance of static background structures during viewpoint interpolation. StreetForward completely discards the lifespan attribute, allowing static Gaussians to persist across the full duration of the sequence. A DPT-style motion decoder regresses bidirectional pixel-wise velocities \(\mathbf{v}_{f,u} \equiv [\mathbf{v}^+_{f,u}, \mathbf{v}^-_{f,u}] \in \mathbb{R}^6\) alongside a dynamic probability \(s_{f,u}\), which defines the static subset \(\mathcal{G}^{\text{static}}\) via thresholding \(\chi_{f,u} = \mathbb{I}[s_{f,u} \le \tau_{\text{dyn}}]\). Crucially, when rendering frame \(f\), the framework intentionally excludes the dynamic Gaussians parameterized at frame \(f\) itself, synthesizing the frame exclusively from the union of static Gaussians and dynamic Gaussians warped from neighboring frames: $$ \mathcal{G}f = \mathcal{G}^{\text{static}} \cup \bigcup $$ This proxy cross-frame rendering objective compels the model to explain the appearance of dynamic objects strictly through accurate motion displacement from adjacent frames, enabling fully self-supervised velocity learning without requiring dense optical flow, scene flow ground truth, or error-prone external trackers.} \mathcal{G}^{\text{dynamic}}_{f \leftarrow t

3. Spatio-Temporal Rigidity Regularization and Opacity Stabilization Unconstrained velocity estimation guided solely by RGB rendering losses is heavily under-constrained, easily resulting in disjoint structural floaters or degenerate opacity collapse. To prevent these failure modes, the framework incorporates three complementary regularizers: First, a piecewise local rigidity loss \(\mathcal{L}_{\text{rigid}} = \mathcal{L}_{\text{rigid-2D}} + \mathcal{L}_{\text{rigid-3D}}\) penalizes velocity discrepancies across local 2D pixel windows \(\mathcal{N}(u)\) and 3D spatial \(K\)-nearest neighborhoods \(\mathcal{N}_K(u)\), enforcing coherent rigid motion over vehicle bodies and eliminating floating artifacts. Second, a forward-backward symmetry loss: $$ \mathcal{L}_{\text{fb}} = \sum_u |\mathbf{v}^{f \to f+1}_u + \mathbf{v}^{f \to f-1}_u|_2^2 $$ imposes a constant-velocity prior over the brief duration \(\Delta t\), ensuring trajectory smoothness. Finally, an opacity stabilization regularizer \(\mathcal{L}_{\alpha} = \lambda_{\alpha} \sum_{f} \sum_{k \in \mathcal{V}_f} w_{f,k} \|1 - \alpha_k^f\|_2\) encourages visible splats to maintain opacities close to 1, preventing the optimization from settling into degenerate local minima where primitives fade to transparent to evade geometric alignment penalties, thereby directing gradient updates toward true 3D spatial refinement.

Loss & Training

The framework adopts a two-stage progressive optimization protocol: 1. Stage 1 (Single-Frame Geometry and Pose Warmup): The DINOv2 backbone is kept frozen while the final linear projection of the motion decoder is zero-initialized and disabled. Cross-frame dynamic aggregation is deactivated. The alternating-attention backbone, camera head, depth head, and Gaussian head are trained using RGB reconstruction loss \(\mathcal{L}_{\text{rgb}}\), opacity loss \(\mathcal{L}_{\alpha}\), and rasterized-predicted depth consistency loss \(\mathcal{L}_{\text{depth}}\) to establish high-fidelity static geometry. 2. Stage 2 (Joint Dynamic Optimization): The causal masked attention layers and motion head are activated, enabling cross-frame dynamic Gaussian warping and random temporal downsampling (\(1\times\) to \(4\times\) stride). The total training objective becomes: $$ \mathcal{L} = \mathcal{L}{\text{rgb}} + \mathcal{L}} + \mathcal{L{\text{depth}} + \lambda}}\mathcal{L{\text{rigid}} + \lambda $$ This staged schedule stabilizes the overall convergence, preventing noisy early motion estimates from disrupting 3D geometric initialization.}}\mathcal{L}_{\text{fb}

Key Experimental Results

Main Results

StreetForward was thoroughly evaluated on the Waymo Open Dataset (202 test scenes under diverse weather and lighting conditions). With 20-frame input video clips where only frames 1, 5, 10, and 15 serve as context views, the model is evaluated on original-view temporal interpolation (Dynamic Only vs. Full Image) and sparse LiDAR depth estimation (RMSE). Zero-shot transfer capabilities were additionally benchmarked on the CARLA simulation dataset.

Dataset / Evaluation Setting Metric StreetForward (Ours) DGGT (Prev. SOTA) STORM Gain / Margin
Waymo Dynamic Only PSNR ↑ 24.30 20.99 22.10 +3.31 dB (significant leap over prior art)
Waymo Dynamic Only SSIM ↑ 0.827 0.821 0.624 +0.006 (outperforming DGGT)
Waymo Full Image PSNR ↑ 27.01 27.41 26.38 Highly competitive full-frame fidelity
Waymo Full Image SSIM ↑ 0.818 0.846 0.794 Competitive structural fidelity
Waymo Sparse Depth (Dynamic RMSE) RMSE ↓ 3.45 6.37 7.50 45.8% error reduction
Waymo Sparse Depth (Full Image) RMSE ↓ 3.14 4.08 5.48 23.0% error reduction
CARLA Zero-Shot (Dynamic PSNR) PSNR ↑ 22.37 21.32 12.80 +1.05 dB (superior domain generalization)
CARLA Zero-Shot (Dynamic SSIM) SSIM ↑ 0.741 0.691 0.313 +0.050
CARLA Zero-Shot (Full Image PSNR) PSNR ↑ 24.62 23.91 18.23 +0.71 dB
CARLA Zero-Shot (Full Depth RMSE) RMSE ↓ 6.01 6.56 - Lowest error among all baselines

Ablation Study

A systematic ablation on the Waymo Open Dataset verifies the contribution of causal modeling, motion regularizers, and opacity stabilization across dynamic-only and full-frame rendering:

Config Dynamic Only PSNR ↑ Dynamic Only SSIM ↑ Full Image PSNR ↑ Full Image SSIM ↑ Note
Full model 24.30 0.827 27.01 0.818 Complete architecture with all regularizers
w/o causal attention 22.36 0.714 24.07 0.781 -1.94 dB dynamic PSNR; loss of directional motion
fixed causal horizon (\(k=1\)) 22.96 0.722 24.66 0.760 Inflexible horizon degrades multi-stride robustness
w/o \(\mathcal{L}_{\text{rigid}}\) 22.64 0.743 26.21 0.786 Disconnected motion fields and spatial floaters
w/o \(\mathbf{v}^-\) (forward-only) 23.26 0.786 26.88 0.809 Incomplete temporal fusion in occluded regions
w/o \(\mathcal{L}_{\text{fb}}\) 23.03 0.721 26.43 0.794 Trajectory inconsistency under temporal synthesis
w/o \(\mathcal{L}_{\alpha}\) 21.02 0.659 25.85 0.683 Catastrophic drop (-3.28 dB) due to opacity collapse

Key Findings

  • Crucial Role of Causal Masked Attention: Removing the causal mask incurs a steep drop of 1.94 dB on dynamic regions, confirming that unconstrained global attention creates temporal feature entanglement and fails to reliably separate dynamic motion from static backgrounds.
  • Opacity Stabilization Prevents Optimization Collapse: Removing \(\mathcal{L}_{\alpha}\) precipitates a catastrophic performance collapse (-3.28 dB in PSNR and -0.168 in SSIM). In multi-view rendering optimization, primitives readily shrink their opacity toward zero to hide geometric misalignments; the opacity regularizer successfully eliminates this degenerate shortcut.
  • Self-Supervised Dynamics Outperform External Trackers: Unlike DGGT which inherits tracking failures (e.g., drifting vehicles, false flying artifacts, incorrect associations) from off-the-shelf 3D trackers, StreetForward reliably maps parked vehicles to near-zero dynamic confidence and accurately synthesizes complex traffic without manual motion masks.

Highlights & Insights

  • Lightweight Directional Conditioning on Geometric Transformers: Instead of discarding pretrained geometric representations or deploying cumbersome video diffusion models, adding just 4 layers of causal masked attention on top of alternating attention provides an elegant and parameter-efficient mechanism to encode directed temporal dynamics.
  • Self-Supervised "Hold-Out" Dynamic Warping: Excluding the current frame's own dynamic Gaussians during training forces the network to explain dynamic foregrounds entirely through temporal propagation from neighboring frames, solving the longstanding supervision dilemma of 3D dynamic scene flow.
  • Lifespan-Free Static Persistence: Eliminating artificial Gaussian lifespans guarantees unbroken spatio-temporal coherence for static infrastructure, preventing sudden holes or popping artifacts during lateral lane-shift extrapolation.

Limitations & Future Work

  • Author-Acknowledged Limitations: The framework assumes piecewise constant velocity over short time steps \(\Delta t\). While effective for typical road traffic, it can experience trajectory approximation error during abrupt deceleration, tight emergency maneuvers, or complex non-rigid pedestrian articulated motions.
  • Sequence Length Boundaries: Current benchmarks focus on 20-frame sequences; extending feedforward 4D reconstruction to kilometer-scale continuous streaming without cumulative drift remains an open engineering challenge.
  • Future Directions: Integrating explicit physical bounding box dynamics as structural priors into the local rigidity loss, or coupling feedforward initialization with lightweight online memory banks for continuous streaming reconstruction.
  • vs VGGT / Pi3: VGGT and Pi3 excel at static multi-view reconstruction and camera pose estimation, but suffer sharp geometric degradation in dynamic settings (dynamic depth RMSE reaching 6.36m and 7.93m). StreetForward builds on their alternating-attention foundation but introduces causal dynamics to compress dynamic RMSE down to 3.45m.
  • vs DGGT: DGGT provides feedforward 4D driving reconstruction but relies on an external 3D tracker (TAPIR-3D), which prone to trajectory extrapolation failure and vehicle hallucination. StreetForward is entirely tracker-free and segmentation-free, outperforming DGGT by +3.31 dB on dynamic regions.
  • vs STORM: STORM models spatio-temporal outdoor scenes with discrete feature volumes, yielding blurred dynamic boundaries. StreetForward employs explicit 3D Gaussian primitives and bidirectional causal dynamics to deliver significantly sharper geometric and appearance fidelity.

Rating

  • Novelty: ⭐⭐⭐⭐ [Introduces causal masked attention into geometric transformers with a self-supervised cross-frame proxy rendering strategy for tracker-free dynamic perception]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations on Waymo and zero-shot CARLA across rendering metrics, sparse LiDAR depth, and rigorous ablations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear mathematical formulations, well-motivated structural design, and transparent ablation analysis]
  • Value: ⭐⭐⭐⭐⭐ [Provides a practical, highly efficient, and tracker-free foundation for closed-loop autonomous driving simulation and 4D digital twin generation]