SEM-ROVER: Semantic Voxel-Guided Diffusion for Large-Scale Driving Scene Generation¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Autonomous Driving
Keywords: 3D Diffusion, Semantic Voxel Grid, Surface Tokens, Large-Scale Scene Generation, Deferred Rendering
TL;DR¶
SEM-ROVER introduces \(\Sigma\)-Voxfield, a discrete surface representation with fixed-cardinality surface samples, and scales local 1D transformer diffusion to 100,000 \(m^2\) driving scenes via progressive outpainting and feed-forward deferred rendering within 8 GB VRAM.
Background & Motivation¶
Synthesizing large-scale 3D outdoor driving scenes is foundational for simulation, corner-case synthetic training data, and controllable digital twin editing. However, existing paradigms suffer from a severe tripartite tension among multiview geometric consistency, spatial scalability, and photorealistic free-viewpoint rendering. Mainstream 2D or video diffusion models (e.g., DreamDrive, MagicDrive) produce visually appealing camera streams, but their outputs remain strictly tied to predefined inference trajectories and lack a persistent, unified 3D geometry in world coordinates, leading to conspicuous drift and geometric collapse under off-trajectory view perturbations. Conversely, occupancy-based or 3D Gaussian Splatting (3DGS) / NeRF optimization frameworks (e.g., Urban Architect, LSD-3D) demand time-consuming per-scene reverse optimization, while dense 3D voxel grids trigger cubic memory and computation growth with increasing spatial extent, rendering large-scale high-resolution urban generation intractable.
The core tension stems from the fact that jointly representing continuous 3D geometry and dense photometric appearance across expansive urban environments consumes prohibitive computational resources. Consequently, direct 3D diffusion has historically been confined to object-centric or room-scale bounds, whereas distilling 2D generative priors into 3D representations fundamentally damages geometric consistency and inherits strong viewpoint biases from front-facing training sequences.
This paper tackles the challenge by moving away from dense 3D volumetric convolutions or unstructured point clouds toward a compact, structured discrete surface representation. Core idea: represent the scene via a discrete \(\Sigma\)-Voxfield grid where each occupied voxel holds a fixed number of colorized surface samples, jointly diffuse geometry and photometry over local token neighborhoods using a 1D Diffusion Transformer, expand to arbitrary large-scale scenes via progressive boundary-overlapping outpainting with bounded 8 GB VRAM, and decode into photorealistic multiview frames via 2D Gaussian surrogates and a deferred diffusion renderer without per-scene optimization.
Method¶
Overall Architecture¶
The SEM-ROVER framework comprises three core stages: First, conditioned on a coarse semantic voxel grid (\(0.6\text{ m}\) voxel size), the continuous textured scene geometry is discretized into \(\Sigma\)-Voxfield units holding fixed-cardinality surface samples with local 3D coordinates and RGB values. Second, a 1D Diffusion Transformer (DiT) operating on local voxel subsets (\(\mathcal{X}_\xi \subset \mathcal{G}\)) jointly denoises geometric and photometric tokens guided by semantic labels and 3D sinusoidal positional encodings, while an iterative spatial outpainting strategy based on the RePaint scheduler expands the scene across overlapping boundaries to arbitrarily large scales under a constant compute budget. Third, the generated \(\Sigma\)-Voxfield is converted into surface-tangent 2D Gaussians and differentiably rasterized into a primary 3D buffer image, which feeds into a deferred diffusion rendering engine conditioned on visibility masks to synthesize sky, distant backgrounds, and photorealistic fine details without per-scene optimization.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Coarse Semantic Voxel Grid<br/>0.6m resolution layout prior"] --> B["ฮฃ-Voxfield Discrete Surface Field Construction<br/>fixed-cardinality surface samples"]
B --> C["Semantic & 3D Position-Guided Local Token Diffusion<br/>1D DiT denoising geometry and color"]
C --> D["Overlapping Neighborhood Progressive Spatial Outpainting<br/>boundary-conditioned iterative expansion"]
D --> E["2D Gaussian Surrogate-Guided Deferred Diffusion Rendering<br/>differentiable splatting and neural decoding"]
E --> F["Photorealistic Multiview Consistent Driving Scenes"]
Key Designs¶
1. ฮฃ-Voxfield Discrete Surface Field Construction: Discretizing Continuous Geometry and Photometry into Fixed-Length Tokens Outdoor driving scenes feature sparse spatial occupancy but intricate surface topology; dense 3D grids lead to volumetric explosion, while raw unstructured point clouds struggle to align with structured spatial priors. SEM-ROVER introduces the \(\Sigma\)-Voxfield grid, where the scene is partitioned into voxels of size \(v_s\), and empty voxels are discarded. Within each occupied voxel \(v_\Sigma\), a fixed cardinality of \(n\) points is sampled on the underlying surface, recording their local 3D offsets \((x^i, y^i, z^i)\) relative to the voxel center alongside their associated RGB colors \((r^i, g^i, b^i)\). To eliminate permutation ambiguity and establish deterministic tokenization, the \(n\) samples are ordered strictly by increasing Euclidean distance to the voxel center and stacked channel-wise into a flat vector:
This formulation allows every occupied voxel to act directly as a discrete token for sequential transformer processing, tightly coupling local micro-geometry with photometric appearance into a compact, unified representation.
2. Semantic & 3D Position-Guided Local Token Diffusion: Jointly Denoising Geometry and Color within Bounded Neighborhoods Diffusing full-scale urban scenes in one shot is computationally prohibitive. SEM-ROVER restricts generative diffusion to local subsets \(\mathcal{X}_\xi\) containing at most \(N_\xi \in [50, 150]\) adjacent occupied voxels (spanning roughly \(4 \times 4\text{ m}^2\)). Each token's world coordinate center \(x_{v_\Sigma}\) is mapped through a sinusoidal 3D positional encoding and a learnable projection layer, then injected into the noisy token to preserve spatial structure. Class categorical semantic labels \(s_{v_\Sigma}\) (e.g., road, sidewalk, building, vegetation) serve as conditioning vectors via cross-attention. Furthermore, masked self-attention is applied across tokens, strictly limiting attention to neighboring voxels within a 3-meter physical radius. The 1D DiT predicts the clean surface sample vector \(\widehat{\psi}(\mathcal{X}_{\xi, 0})\) under an \(\ell_2\) denoising objective, leveraging semantic constraints to prevent geometric hallucinations.
3. Overlapping Neighborhood Progressive Spatial Outpainting: Expanding to Infinite Extents under Constant Memory Budget Because diffusion operates solely on local neighborhoods, direct synthesis cannot cover large city blocks in a single pass. To scale up efficiently, SEM-ROVER employs an iterative spatial outpainting strategy powered by the RePaint scheduler. When expanding a new region, the local set \(\mathcal{X}_\xi\) is partitioned into a known portion \(\mathcal{X}_\xi^{\text{known}}\) (already generated) and an unobserved target portion \(\mathcal{X}_\xi^{\text{target}}\). During reverse diffusion sampling, \(\mathcal{X}_\xi^{\text{known}}\) is clamped to its known values at each step while only \(\mathcal{X}_\xi^{\text{target}}\) is updated, enforcing smooth surface elevation, edge alignment, and textural consistency across overlapping boundaries. By sequentially sliding the outpainting window, generation time scales linearly with total area while peak GPU memory remains strictly locked at 8 GB, effortlessly generating scenes exceeding 100,000 \(m^2\).
4. 2D Gaussian Surrogate-Guided Deferred Diffusion Rendering: Bridging Discretization Gaps without Per-Scene Optimization Directly projecting discrete surface points produces severe sparsity artifacts and ray penetrations. SEM-ROVER constructs an analytical surface-aligned 2D Gaussian at each surface point: PCA on spatial neighbors estimates the local surface normal, which initializes an SO(3) rotation matrix aligning the Gaussian with the tangent plane, paired with a fixed splat radius \(r = 0.04\text{ m}\). Differentiable 2D Gaussian Splatting rasterization renders this intermediate buffer into a primary view \(I_\Sigma\). However, \(I_\Sigma\) naturally lacks distant backgrounds, sky, and high-frequency specular reflections. A modified Stable Diffusion 1.5 model acts as a deferred neural renderer, conditioned on \(I_\Sigma\) latents and a visibility mask indicating uncovered regions. For multi-frame temporal consistency, the architecture incorporates Autoregressive Stable Diffusion (ASD, feeding the previous frame as conditioning) or Video Stable Diffusion (VSD, generating 12 frames jointly), achieving photorealistic novel-view synthesis feed-forwardly without per-scene 3DGS optimization.
Loss & Training¶
- 3D Diffusion Training: Standard DDPM objective trained for 1,000 denoising steps, minimizing the \(\ell_2\) loss between predicted \(\widehat{\psi}(\mathcal{X}_{\xi, 0})\) and ground-truth \(\psi(\mathcal{X}_{\xi, 0})\). Semantic labels are dropped with 10% probability for classifier-free guidance (CFG scale = 4.0 at inference). Optimized with Adam at learning rate \(5 \times 10^{-4}\) on \(2 \times 24\text{ GB}\) GPUs for 4 days.
- Deferred Rendering Training: Fine-tuned on pairs \(\{I_\Sigma, \widehat{I}_s\}\) rendered from static backgrounds of reconstructed multi-view scenes. Minimizes latent noise prediction error \(\mathcal{L}(\phi) = \mathbb{E}_{t, \epsilon}[\|R_\phi(x_t, x_\Sigma) - \epsilon\|_2^2]\) using Adam at learning rate \(5 \times 10^{-5}\) on 1 GPU for 4 days.
Key Experimental Results¶
Main Results¶
Evaluated on Waymo Open Dataset (WOD) across seen trajectories (ground-truth poses) and novel views (perturbed poses), comparing image synthesis metrics (FID / KID) and minimum inference VRAM requirements against state-of-the-art baselines:
| Method | FID (Seen) โ | FID (Novel) โ | KID (Seen) โ | KID (Novel) โ | Min. VRAM | Feed-Forward Pipeline |
|---|---|---|---|---|---|---|
| InfiniCube | 84.14 | 99.13 | 0.03 | 0.06 | 75 GB | Yes |
| GEN3C | 113.27 | 117.63 | 0.08 | 0.09 | 43 GB | Yes |
| SEM-ROVER (Ours, ASD) | 81.98 | 89.20 | 0.05 | 0.06 | 8 GB | Yes |
Ablation Study¶
Ablations evaluate the 3D diffusion components in feature space, the impact of deferred rendering conditioning, and sensitivity to voxelization hyper-parameters.
Table 1: Feature-space ablation of 3D diffusion model (PointNet++ feature extractor with F3D and MMD metrics)
| Config | F3D โ | MMD โ | Note |
|---|---|---|---|
| Full model (Ours) | 3.523 | 0.091 | Incorporates semantic conditioning & distance-based sorting |
| w/o ordering | 3.585 | 0.093 | Disabling point sorting impairs deterministic token alignment |
| w/o sem. cond. | 3.927 | 0.105 | Absence of semantic prior causes severe topological drift |
Table 2: Conditioning signals for deferred rendering
| Conditioning Signal | FVD โ | Note |
|---|---|---|
| Semantic maps (UniScene-style) | 143.62 | Lacks explicit geometry and color, inducing severe temporal flicker |
| Projected LiDAR (FreeVS-style) | 97.51 | Provides sparse depth but lacks dense surface coverage and texture |
| ฮฃ-Voxfield projection (Ours) | 53.01 | Joint geometry and color priors dramatically stabilize temporal rendering |
Table 3: Effect of surface points \(n\) and voxel size \(v_s\) on generation trade-offs
| Parameter Config | Speed / Metric | F3D / PSNR | MMD / SSIM | Empirical Conclusion |
|---|---|---|---|---|
| Points \(n = 5\) | 1.95ร speed | F3D: 4.11 | MMD: 0.08 | Undersampled surfaces degrade geometric fidelity |
| Points \(n = 20\) (Ours) | 1.00ร speed | F3D: 3.52 | MMD: 0.09 | Optimal Pareto frontier between quality and throughput |
| Points \(n = 70\) | 0.51ร speed | F3D: 3.50 | MMD: 0.09 | Marginal quality gains at 2ร computational penalty |
| Voxel size \(v_s = 0.3\text{ m}\) | - | PSNR: 15.20 | SSIM: 0.60 | Over-constrains conditioning, hurting geometric diversity |
| Voxel size \(v_s = 0.6\text{ m}\) (Ours) | - | PSNR: 15.50 | SSIM: 0.58 | Best compromise between layout control and surface realism |
| Voxel size \(v_s = 1.2\text{ m}\) | - | PSNR: 14.83 | SSIM: 0.54 | Coarse resolution blurs road boundaries and curbs |
Key Findings¶
- Superior Novel-View Robustness: Under trajectory perturbations on novel views, InfiniCube's FID degrades by 14.99 points (84.14 to 99.13), whereas SEM-ROVER only degrades by 7.22 points (81.98 to 89.20). Generating persistent 3D geometry directly avoids the off-trajectory distortions prevalent in view-distilled video models.
- Drastic Reduction in Resource Footprint: By combining local token diffusion with iterative spatial outpainting, SEM-ROVER slashes minimum VRAM to 8 GB (compared to 75 GB for InfiniCube and 43 GB for GEN3C), enabling generation on commodity GPUs.
- Efficacy of \(\Sigma\)-Voxfield Guidance: In deferred rendering, conditioning on \(\Sigma\)-Voxfield projections slashes FVD by 63.1% compared to pure semantic masks (143.62 \(\to\) 53.01), confirming that surface-aligned geometry and coarse colors are vital for video stability.
Highlights & Insights¶
- Fixed-Cardinality Surface Tokenization: Structuring irregular 3D surfaces into ordered, fixed-length \(\Sigma\)-Voxfield vectors bypasses 3D volumetric cubic scaling while allowing seamless integration into standard 1D Diffusion Transformers.
- RePaint-Style Spatial Outpainting in 3D: Decouples physical scene dimensions from GPU memory allocation; the model runs in a bounded \(4 \times 4\text{ m}^2\) local window while synthesizing continuous multi-kilometer urban environments.
- Feed-Forward 2D Gaussian Deferred Rendering: Translating discrete voxels into tangent-plane 2D Gaussians produces continuous intermediate buffers that neural decoders can render instantly, eliminating per-scene optimization entirely.
Limitations & Future Work¶
- Lack of Fine-Grained Photometric Control: Generation is guided primarily by semantics and geometry; atmospheric weather, solar angles, and fine material reflectances cannot be explicitly specified via text prompts.
- Static Scene Assumption: The current framework models only rigid background scenes, omitting dynamic moving cars, pedestrians, and evolving traffic dynamics. Extending tokens to 4D dynamic Gaussians remains an open frontier.
- Dependency on Reconstruction Pipelines: Constructing the ground-truth training tokens requires offline multi-view 3DGS (OmniRe) and marching cubes mesh extraction.
Related Work & Insights¶
- vs InfiniCube: InfiniCube employs an HDMap \(\to\) fine voxel \(\to\) video diffusion \(\to\) feedforward 3DGS pipeline, which struggles on side cameras and requires 75 GB VRAM; SEM-ROVER diffuses persistent 3D tokens directly, excelling across surround camera rigs with only 8 GB VRAM.
- vs GEN3C: GEN3C relies on 2D video diffusion conditioned on point cloud caches and suffers high FID on novel views (117.63); SEM-ROVER enforces rigorous physical 3D consistency, achieving 89.20 FID on novel views.
- vs LSD-3D & Urban Architect: Prior geometry-grounded models rely on heavy SDS distillation or per-scene Gaussian optimization taking hours; SEM-ROVER is fully feed-forward, synthesizing an entire large-scale scene in ~20 minutes.
Rating¶
- Novelty: โญโญโญโญ [Clever combination of fixed-cardinality surface tokens, local DiT diffusion, and 3D outpainting for large-scale driving environments]
- Experimental Thoroughness: โญโญโญโญ [Extensive validation across Waymo and PandaSet with 3D feature metrics, visual quality, and comprehensive hyper-parameter ablations]
- Writing Quality: โญโญโญโญโญ [Clear structural narrative, well-formulated methodology, and transparent exposition of trade-offs]
- Value: โญโญโญโญโญ [Provides a practical, low-compute blueprint for generating boundless, multiview-consistent 3D driving environments for AD simulation]