Skip to content

Geometry-Aware Spatio-Temporal Context Modeling for 4D Occupancy Forecasting

Conference: ECCV 2026
Paper: ECCV Paper
Code: https://github.com/chenst27/GAST
Area: Autonomous Driving
Keywords: 4D occupancy forecasting, spatio-temporal context modeling, progressive explicit-implicit generation, dual-path spatio-temporal modeling, world models

TL;DR

GAST establishes an end-to-end continuous BEV paradigm for 4D occupancy forecasting that replaces discrete autoregressive tokenization with pose-driven progressive explicit-implicit generation and dual-path spatio-temporal context modeling, outperforming prior SOTA on Occ3D-nuScenes by 7.67% mIoU with a 2.84x speedup.

Background & Motivation

3D semantic occupancy grids discretize driving scenes into dense volumetric voxels labeled with semantic attributes, offering a comprehensive and holistic representation that surmounts the limitations of 3D bounding boxes and sparse LiDAR point clouds in describing irregular obstacles and fine-grained geometry. Extending this representation into the temporal domain, 4D occupancy forecasting aims to predict the future spatio-temporal evolution of 3D scene geometry and semantics given historical observations and ego-motion. This capability serves as an indispensable pillar for modern autonomous driving pipelines, facilitating long-horizon motion planning, corner-case simulation, and closed-loop synthetic validation.

However, prevailing state-of-the-art methods typically adopt a two-stage autoregressive world model paradigm (e.g., OccWorld, RenderWorld, Occ-LLM, I2-World). These approaches discretize continuous 3D occupancy into codebook tokens via VQ-VAE and autoregressively predict future token sequences using generative Transformers before decoding back into voxels. This paradigm suffers from three critical bottlenecks: first, static scene elements (such as roads, sidewalks, and buildings) physically adhere to rigid coordinate transformations governed by ego-motion, yet autoregressive token generation lacks explicit geometric constraints, leading to severe geometric distortion and spatial drift over extended horizons; second, the causal, step-by-step nature of autoregressive generation impedes global spatio-temporal context aggregation across timestamps, compounding prediction errors across frames; third, the disjoint two-stage training scheme limits the forecasting model's representational upper bound to the lossy reconstruction fidelity of the discrete tokenizer.

To overcome these structural limitations, the entry point of this paper is to return to continuous BEV space for direct end-to-end representation modeling: grounding static structures through explicit pose-driven geometric transformation, guiding dynamic evolution via motion-conditioned implicit modulation, and enforcing long-range consistency through dual-path spatio-temporal context aggregation in a unified world coordinate system. Core idea: discard discrete token autoregression and introduce an end-to-end geometry-aware spatio-temporal context modeling framework (GAST) that constructs per-frame future representations via progressive explicit-implicit generation, enforces cross-frame consistency through dual-path spatial-temporal modeling, and enables joint optimization of historical reconstruction and future forecasting.

Method

Overall Architecture

GAST takes historical 3D occupancy grids and ego-pose sequences as input, utilizes a lightweight occupancy encoder to extract compact bird's-eye-view (BEV) representations and predict future ego-poses, generates initial per-frame future BEV features via a progressive explicit-implicit generation module, reinforces cross-frame spatio-temporal coherence through a dual-path context modeling module, and decodes the unified representations into historical reconstruction and future 4D semantic occupancy grids.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Historical Occupancy & Poses<br/>O_{T-N_h+1:T}, P_{T-N_h+1:T}"] --> B["BEV Encoding & Ego-Pose Prediction<br/>Extract B_t & forecast future poses \hat{P}_t"]
    B --> C["Progressive Explicit-Implicit Generation<br/>Rigid spatial warping + Motion modulation + Local refinement"]
    C --> D["Global Spatial Context Aggregation<br/>World alignment + Multi-scale pyramid fusion"]
    C --> E["Temporal Dynamics Extraction<br/>ConvGRU sequential recurrence"]
    D --> F["Dual-Path Feature Fusion & Self-Attention<br/>Spatial alignment & temporal smoothness fusion"]
    E --> F
    F --> G["Joint Reconstruction & Forecasting Decoder<br/>Unified spatio-temporal voxel decoding"]
    G --> H["Output 4D Occupancy Grids<br/>Past reconstruction & future forecasting"]

Key Designs

1. Progressive Explicit-Implicit Generation: Rigid Geometric Warping and Dynamic Feature Refinement Purely implicit or autoregressive forecasters routinely produce blurred or distorted static geometry because they fail to explicitly exploit vehicle ego-motion as a strong physical prior. To tackle this, the progressive explicit-implicit generation (PEIG) module operates through three cascaded stages: explicit geometric transformation (EGT), implicit feature modulation (IFM), and feature refinement (FR). In the explicit stage, the current high-confidence BEV feature \(B_T\) is rigidly transformed into future ego-centric coordinate frames using predicted future absolute poses \(\hat{P}_t = (\hat{t}_t, \hat{q}_t)\) relative to current pose \(P_T\) via differentiable grid sampling: $\(B_t^{geo} = \text{Warp}(B_T;\, P_T^{-1} \circ \hat{P}_t)\)$ This parameter-free projection provides a geometrically faithful static scene canvas. Next, to model non-rigid traffic dynamics and vehicle interactions, implicit feature modulation concatenates the relative future translation deltas \(\Delta\hat{t}_t \in \mathbb{R}^3\) and rotation deltas \(\Delta\hat{q}_t \in \mathbb{R}^4\) into motion vectors, passes them through an MLP, and produces channel-wise scaling parameters \(\gamma_t \in \mathbb{R}^c\) and offset parameters \(\beta_t \in \mathbb{R}^c\) to apply an affine transformation: $\(B_t^{mod} = \gamma_t \odot B_t^{geo} + \beta_t\)$ Finally, to populate unobserved occluded areas and refine spatial details, feature refinement sets modulated features \(B_t^{mod}\) as queries, maps future BEV grid centers \(C_t\) back into current coordinate space via relative transformation \(\mathcal{T}_{t \to T}\) to obtain 2D reference points, and executes deformable cross-attention against current observation features \(B_T\) as keys and values: $\(B_t^{ref} = \text{DeformAttn}\big(B_t^{mod},\, \mathcal{T}_{t \to T}(C_t),\, B_T\big)\)$ This progressive flow equips each future frame with both rigid structural preservation and plausible semantic adaptability.

2. Global Spatial Context Aggregation: Multi-Scale Pyramid Fusion in Unified World Coordinates As an autonomous vehicle maneuvers across time, future BEV features generated in localized, shifting coordinate systems become spatially disjoint, hindering convolutional networks from capturing broader scene-level layouts. Global spatial context aggregation resolves this by reprojecting all historical features \(\{B_t\}_{t=T-N_h+1}^T\) and refined future features \(\{B_t^{ref}\}_{t=T+1}^{T+N_f}\) into a unified, fixed global world coordinate system using their respective poses, producing an aligned feature sequence \(\{B_t^{align}\}\). To capture both macro-scale road geometry and micro-scale object shapes, the aligned features are projected onto multi-scale BEV grids at \(K=3\) hierarchical resolutions (\(50 \times 50\), \(100 \times 100\), and \(200 \times 200\)). After localized convolutional context extraction at each scale, features are fused top-down in a coarse-to-fine manner via upsampling and residual convolution: $\(\tilde{F}_k = \text{Conv}\big(\text{Upsample}(\tilde{F}_{k-1}) \oplus F_k^{global}\big),\quad k=2,\dots,K\)$ The resulting full-resolution global context map \(F^{global} := \tilde{F}_K\) is mapped back into future ego-centric frames via differentiable grid sampling: $\(B_t^{spatial} = \text{GridSample}\big(F^{global},\, \phi(\hat{P}_t)\big)\)$ thereby grounding localized frame features within a consistent global geometric coordinate frame.

3. Temporal Dynamics Extraction: ConvGRU-Driven Motion Smoothness and Dual-Path Integration While world-coordinate reprojection enforces spatial rigidity, it does not explicitly capture the velocity, acceleration, and continuous trajectory kinematics of dynamic traffic participants. The temporal dynamics extraction module addresses this by concatenating historical BEV features and refined future features chronologically into a sequence and processing them with a lightweight ConvGRU encoder: $\(h_t = \text{ConvGRU}(\mathcal{B}_t,\, h_{t-1}),\quad t = T-N_h+1, \dots, T+N_f\)$ where \(h_{T-N_h}=0\). For future time steps (\(t > T\)), historical temporal momentum is propagated forward in a strictly causal fashion, yielding motion-aware representations \(B_t^{temp}\). Finally, the spatial context features \(B_t^{spatial}\) and temporal dynamics features \(B_t^{temp}\) are fused at each future step through channel concatenation and convolution: $\(\hat{B}_t = \text{Conv}(B_t^{spatial} \oplus B_t^{temp})\)$ followed by deformable self-attention to enrich intra-frame semantic correlations, achieving a balanced synergy between spatial structural fidelity and temporal dynamic smoothness.

4. Joint Reconstruction and Forecasting Decoder: End-to-End Regularization of Spatio-Temporal Volumes Conventional forecasting models optimize exclusively on future frames, often resulting in latent feature representations that overfit future objectives and degrade fundamental perceptual semantics. In GAST, historical BEV features and enhanced future features are concatenated along the temporal dimension into a unified 4D spatio-temporal volume. This volume is upsampled to the original voxel resolution via bilinear interpolation and processed by 3D residual convolution blocks to decode dense 4D semantic occupancy \(\hat{O} \in \mathbb{R}^{(N_h+N_f) \times S \times H \times W \times D}\). Crucially, the historical segment (\(t \le T\)) is explicitly supervised as a reconstruction target while the future segment (\(t > T\)) is supervised as the forecasting target. Historical reconstruction acts as an inductive regularizer, preventing representation drift and ensuring seamless continuity between observed past context and unobserved future predictions.

Loss & Training

The framework is optimized end-to-end via a combined loss encompassing 4D occupancy forecasting and ego-motion regression: $\(\mathcal{L}_{total} = \mathcal{L}_{occ}^{past} + \mathcal{L}_{occ}^{future} + \lambda_{trans}\mathcal{L}_{trans} + \lambda_{rot}\mathcal{L}_{rot}\)$ Occupancy prediction losses for both past reconstruction and future forecasting combine weighted cross-entropy loss and Lovász-Softmax loss to handle extreme foreground-background voxel class imbalances, with hyper-parameters \(\lambda_{wce}=\lambda_{lov}=1.0\). Future relative pose regression employs an \(L_2\) translation loss (\(\lambda_{trans}=0.01\)) and a quaternion-based angular distance loss (\(\lambda_{rot}=1.0\)). The model is trained on 4 NVIDIA RTX 3090 GPUs using AdamW for 30 epochs, with an initial learning rate of 0.001 decayed via cosine annealing and batch size 8.

Key Experimental Results

Main Results

On the Occ3D-nuScenes validation set, GAST is comprehensively compared against leading autoregressive and diffusion-based methods across 1s, 2s, 3s, and average horizons under diverse sensor input and trajectory conditions.

Input Modality Ego Trajectory Method mIoU (1s) mIoU (2s) mIoU (3s) mIoU (Avg.) IoU (1s) IoU (2s) IoU (3s) IoU (Avg.)
Surround Camera Predicted OccWorld-D 11.55 8.10 6.22 8.62 18.90 16.26 14.43 16.53
Surround Camera Predicted Occ-LLM 11.28 10.21 9.13 10.21 27.11 24.07 20.19 23.79
Surround Camera Predicted DFIT-OccWorld 13.38 10.16 7.96 10.50 19.18 16.85 15.02 17.02
Surround Camera Predicted GAST-STC (Ours) 19.32 13.96 10.62 14.84 27.97 23.35 20.22 23.87
Surround Camera Ground Truth DOME-STC 17.79 14.23 11.58 14.53 26.39 23.20 20.42 23.33
Surround Camera Ground Truth I2-World-STC 21.67 18.78 16.47 18.97 30.55 28.76 26.99 28.77
Surround Camera Ground Truth GAST-STC (Ours) 22.90 19.93 17.17 20.16 31.62 29.92 28.00 29.87
3D Occupancy Predicted OccWorld 25.78 15.14 10.51 17.14 34.63 25.07 20.18 26.63
3D Occupancy Predicted OccLLaMA 25.05 19.49 15.26 19.93 34.56 28.53 24.41 29.17
3D Occupancy Predicted RenderWorld 28.69 18.89 14.83 20.80 37.74 28.41 24.08 30.08
3D Occupancy Predicted COME 30.57 19.91 13.38 21.29 36.96 28.26 21.86 29.03
3D Occupancy Predicted DFIT-OccWorld 31.68 21.29 15.18 22.71 40.28 31.24 25.29 32.27
3D Occupancy Predicted GAST (Ours) 38.60 24.79 18.54 27.38 44.69 34.21 28.85 35.90
3D Occupancy Ground Truth DOME 35.11 25.89 20.29 27.10 43.99 35.36 29.74 36.36
3D Occupancy Ground Truth UniScene 35.37 29.59 25.08 31.76 38.34 32.70 29.09 34.84
3D Occupancy Ground Truth COME 42.75 32.97 26.98 34.23 50.57 43.47 38.36 44.13
3D Occupancy Ground Truth I2-World 47.62 38.58 32.98 39.73 54.29 49.43 45.69 49.80
3D Occupancy Ground Truth GAST (Ours) 55.54 46.33 40.18 47.40 60.77 55.97 51.96 56.24

Ablation Study

A step-by-step ablation study on Occ3D-nuScenes validation set (under GT ego-trajectory) reveals the individual contribution of each architectural component.

No. EGT IFM FR GSCA TDE mIoU (%) IoU (%) Note
1 - - - - - 18.31 29.01 Baseline independent convolutional forecaster
2 ✓ - - - - 31.78 41.65 Pose-driven explicit rigid spatial warping (+13.47% mIoU)
3 ✓ ✓ - - - 32.55 42.63 Motion-aware channel affine modulation (+0.77% mIoU)
4 ✓ ✓ ✓ - - 40.47 49.15 Full PEIG module with deformable cross-attention (+7.92% mIoU)
5 ✓ ✓ ✓ ✓ - 46.08 55.50 Adding multi-scale world spatial context aggregation (+5.61% mIoU)
6 ✓ ✓ ✓ ✓ ✓ 47.40 56.24 Full GAST model with dual-path spatio-temporal modeling (+1.32% mIoU)

Efficiency and Long-Horizon Forecasting

Evaluation Dimension Method #Params #FLOPs Latency Avg. mIoU (%) Avg. IoU (%)
Efficiency (RTX 3090) DOME (Diffusion) 444.07 M 2928.66 G 1899.15 ms 27.10 36.36
Efficiency (RTX 3090) COME (Diffusion + ControlNet) 692.97 M 4461.46 G 4377.67 ms 34.23 44.13
Efficiency (RTX 3090) I2-World (Two-Stage Autoregressive) 22.67 M 494.58 G 227.09 ms 39.73 49.80
Efficiency (RTX 3090) GAST (Ours) 12.09 M 416.27 G 80.03 ms 47.40 56.24
8-second Long-term (Avg.) DOME - - - 15.83 25.21
8-second Long-term (Avg.) COME - - - 19.07 29.96
8-second Long-term (Avg.) GAST (Ours) - - - 24.12 39.06

Key Findings

  • Explicit geometric warping unlocks massive performance gains: Adding explicit geometric transformation (EGT) single-handedly elevates mIoU from 18.31% to 31.78% (+13.47%) and IoU from 29.01% to 41.65% (+12.64%). Because background voxels represent the majority of driving environments, anchoring static structures with exact ego-motion transformations is far more effective than forcing neural networks to synthesize static spatial translations from scratch.
  • World-coordinate alignment prevents long-range structural drift: Incorporating global spatial context aggregation (GSCA) yields an additional 5.61% mIoU improvement. Hierarchical multi-scale resolution ablation confirms that fusing coarse (\(50 \times 50\)), medium (\(100 \times 100\)), and fine (\(200 \times 200\)) grids outperforms any single-scale implementation, validating the synergy between macro topological layout and fine geometric boundaries.
  • Substantial efficiency advantage: By eschewing sequential autoregressive token generation, GAST requires only 12.09M parameters and 416.27G FLOPs. Its per-frame inference latency on an RTX 3090 is 80.03 ms—delivering a 2.84x speedup over I2-World and running more than 54x faster than diffusion models like COME.
  • Robustness in long-term forecasting: Over an extended 8-second horizon, GAST maintains 15.96% mIoU and 30.87% IoU at the final 8th second, outperforming COME's 8-second average by 5.05% mIoU and 9.10% IoU.
  • Exceptional zero-shot cross-dataset generalization: Evaluated directly on Occ3D-Waymo without fine-tuning, GAST attains 57.47% mIoU at 10Hz and 46.70% mIoU at 2Hz, beating I2-World by 13.74% and 10.32% respectively.

Highlights & Insights

  • Decoupled static physics and dynamic learning: Explicitly handling deterministic static geometry via rigid transformations while delegating non-rigid dynamic evolutions to implicit modulation and attention provides a clean, physically grounded inductive bias.
  • Breaking the autoregressive dogma in world models: Demonstrates that continuous spatial BEV representations can drastically surpass discrete VQ-VAE autoregressive pipelines in both geometric precision and inference throughput, avoiding the error accumulation inherent to generative Transformers.
  • Unified world coordinate aggregation: Reprojecting cross-temporal historical and predicted features into a shared global reference frame resolves localized coordinate misalignments, offering an impactful design pattern for multi-agent cooperative perception and long-term 3D mapping.

Limitations & Future Work

  • Dependency on ego-pose accuracy: Rigid spatial warping relies heavily on accurate future ego-pose estimates; large ego-trajectory prediction errors in extreme handling or slipping scenarios may cause geometric projection artifacts.
  • Absence of explicit traffic topology priors: While ConvGRU captures temporal momentum, incorporating vectorized HD-map road geometry and traffic signal states could further enhance multi-agent interaction and turning intent forecasting.
  • Future directions: Integrating lane topology priors into implicit modulation and exploring uncertainty estimation within the spatial warping module to handle erratic ego-motions gracefully.
  • vs OccWorld / RenderWorld / Occ-LLM: Early occupancy world models quantize 3D scenes into discrete tokens for autoregressive generation; GAST retains continuous BEV features, runs ~3x faster, and avoids tokenizer compression artifacts.
  • vs I2-World: I2-World decouples scenes into intra-frame static and inter-frame dynamic discrete tokens across two stages; GAST achieves structural decoupling directly via explicit pose warping and lightweight convolution/attention in an end-to-end framework.
  • vs DOME / COME: Diffusion-based 4D forecasting produces continuous representations but requires multi-step iterative denoising with multi-second latency; GAST generates accurate 4D occupancy in a single forward pass within 80ms.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Elegant departure from discrete token autoregression, harmonizing physical geometric priors with dual-path spatio-temporal modeling.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Exhaustive evaluation across multiple input modalities, 8-second horizons, zero-shot Waymo generalization, and comprehensive ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous methodology, crisp motivation, and clear mathematical formulation.
  • Value: ⭐⭐⭐⭐⭐ Provides a foundational, highly efficient blueprint for real-time 4D occupancy forecasting in autonomous driving.