Map2World: Segment Map Conditioned Text to 3D World Generation¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Area: 3D Vision
Keywords: 3D World Generation, Structured Latent, 3D Gaussian Splatting, MultiDiffusion, TRELLIS
TL;DR¶
To overcome data scarcity and narrow domain restrictions in 3D world synthesis, Map2World introduces an arbitrary-shaped segment map conditioned framework that scales pretrained TRELLIS asset priors into complete, layout-controllable, scale-consistent, and detailed large-scale 3DGS worlds using 3D MultiDiffusion, spectral initial noise optimization, and a latent detail enhancer.
Background & Motivation¶
In large-scale three-dimensional world generation, high-fidelity virtual environments are crucial across gaming, immersive VR/AR content creation, autonomous driving simulation, and embodied AI. While 3D asset generation has achieved remarkable breakthroughs for isolated objects, extending generative models to unbounded, world-scale environments remains fundamentally bottlenecked by the extreme scarcity of high-quality, diverse 3D scene datasets. Existing paradigms largely fall into two camps: view-centric methods lift generated 2D images or video sequences into 3D via monocular depth estimation, which inherits rich 2D web priors but inevitably suffers from multi-view geometric inconsistency, catastrophic drift, and unobserved region occlusions; conversely, world-centric approaches (such as BlockFusion, CityDreamer, and InfiniCube) train 3D representations directly on domain-specific datasets, leaving each model tightly confined to a narrow domain (e.g., driving corridors or synthetic procedural city blocks) without generalizability across open-world domains.
To bypass the data bottleneck, an intuitive approach is to reuse off-the-shelf general 3D asset generators pretrained on massive object repositories, such as TRELLIS operating on structured 3D latents (SLAT). However, pretrained asset models are typically constrained to a fixed cubic voxel grid (e.g., \(64^3\)). Straightforward tiling schemes like SynCity divide scenes into uniform grids, generate assets independently per tile, and merge them afterwards. This rigid formulation fails to support arbitrary non-grid polygonal semantic layouts, produces stark boundary discontinuities between tiles, cannot synthesize continuous multi-tile geological terrain or large buildings, and frequently incurs severe global scale incoherence.
The central tension lies in seamlessly scaling fixed-resolution object-level 3D generative priors to arbitrary-shaped semantic layouts across large-scale physical scenes without requiring massive paired 3D world datasets, while maintaining global physical scale consistency and sharp microscopic surface fidelity. The core idea is to extrude an arbitrary-shaped 2D semantic layout into a 3D guidance mask, fuse multi-region denoising trajectories in the TRELLIS structured latent space via time-annealed Gaussian MultiDiffusion, enforce global scene scale constraints through spectral-domain initial noise optimization, and autoregressively refine local latents using a lightweight condition-aware detail enhancer.
Method¶
Overall Architecture¶
Map2World receives an arbitrary-shaped 2D semantic segmentation map along with user-specified text prompts assigned to each segment (e.g., "forest with tall green trees", "city consists of skyscrapers and roads"), and outputs a globally consistent, high-fidelity 3D Gaussian Splatting (3DGS) environment ready for real-time camera roaming. The overall pipeline operates in two major stages: large-scale spatial expansion in the structured latent (SLAT) space via rectified flow MultiDiffusion, followed by hierarchical local latent super-resolution via a condition-guided detail enhancer and fine-tuned SLAT decoding.
Specifically, the 2D layout is uniformly extruded along the height axis to form 3D semantic masks. The continuous 3D space is covered by overlapping cubic windows, where velocity field predictions from structure flow Transformers (\(\mathcal{G}_S\)) and latent flow Transformers (\(\mathcal{G}_L\)) are fused via Gaussian-weighted MultiDiffusion. Prior to denoising, a 3D FFT spectral parameterization aligns the initial noise to the desired scene scale manifold. The resulting coarse global structured latent is subsequently processed by a detail enhancer, which autoregressively upsamples each large latent cube into 8 fine-grained sub-cubes conditioned on adjacent and global context. Finally, a fine-tuned SLAT decoder reconstructs clean, artifact-free 3DGS primitives.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: 2D Segment Map<br/>+ Per-Segment Text Prompts"] --> B["Spectral Initial Noise Optimization<br/>Aligns Global Scale Manifold"]
B --> C["3D Latent Space MultiDiffusion Fusion<br/>Overlapping Windows & Dynamic Gaussian Mask"]
C --> D["Hierarchical Conditional Detail Enhancer<br/>1-to-8 Sub-Cube Autoregressive Refinement"]
D --> E["Fine-Tuned SLAT Decoder<br/>Outputs High-Fidelity 3DGS World"]
Key Designs¶
1. 3D Latent Space MultiDiffusion Fusion: Seamless Spatial Expansion on Arbitrary Semantic Layouts Because the active voxel space in pretrained TRELLIS is limited to a \(64^3\) cube, it cannot directly synthesize expansive scenes comprising ground planes, foliage, and architectural complexes. Map2World extends the 2D MultiDiffusion framework into 3D rectified flow latent space. The 3D scene volume is covered by overlapping \(64^3\) cubic windows \(\{\Omega_j\}\) with half-stride overlap. For any spatial coordinate \(\mathbf{x}\), its local latent velocity is computed by aggregating predictions from all active covering windows \(\mathcal{A}(\mathbf{x}) = \{ j \mid \mathbf{x} \in \Omega_j \}\) normalized by a 3D Gaussian weighting kernel \(W(\cdot)\) centered at \(\mathbf{c}_j\). To enable fine-grained layout control over arbitrary shapes, the 2D segmentation map is extruded into \(K\) 3D binary masks \(M_k(\mathbf{x})\). During diffusion step \(t\), a dynamic Gaussian kernel \(G(\sigma_t)\) spatially smooths the region boundaries:
The standard deviation \(\sigma_t\) is dynamically annealed with diffusion step \(t\), starting large to allow gentle semantic blending in early generative phases and sharpening toward zero near \(t=0\). This prevents hard seam tearing and geometric fragmentation along multi-segment borders.
2. Spectral Initial Noise Optimization: Enforcing Scale-Aware Manifold Consistency Rectified flow generative models exhibit high sensitivity to initial noise configurations: different Gaussian noise instances lead to drastically fluctuating macro-structures and divergent physical scales. The authors discover that sparse structures of similar physical scales (e.g., individual houses versus multi-block avenues) naturally cluster on scale-dependent sub-manifolds characterized by consistent ground planes and negative space patterns. To steer synthesis toward a prescribed scale, Map2World adopts an initial noise optimization strategy. Rather than backpropagating through the entire denoising trajectory, a single-step linear approximation \(S(t) \approx S_T + (1 - \frac{t}{T})[\mathcal{G}_S(S_T) - S_T]_{\mathrm{sg}}\) provides efficient gradients against a scale constraint target \(y\) over common ground/empty regions masked by \(\mathcal{M}\):
Because applying aggressive spatial-domain updates to sparse structure latents destabilizes the optimization trajectory and causes divergence, the feature tensor is parameterized in the spectral domain via a 3D Fast Fourier Transform (3D FFT). Updating frequency coefficients provides a smooth optimization surface, enabling a large learning rate of 9.0 to achieve convergence within just 5 steps.
3. Hierarchical Conditional Detail Enhancer and SLAT Decoder Fine-Tuning: Overcoming Capacity Bottlenecks A single global structured latent cannot encode dense microscopic geometries across an entire world due to bounded channel capacity. Training an end-to-end 3D super-resolution model directly on raw geometry is hindered by a lack of paired scene datasets. Map2World designs a hierarchical detail enhancer in the latent space. Using 17,500 large cubes extracted from Objaverse scenes, each large cube \(C_\mathcal{O}\) is subdivided into 8 identical smaller sub-cubes \(C_j\) (\(j=0,\dots,7\)). The detail enhancer autoregressively predicts sub-cube latents \(\mathbf{s}^j\) conditioned on the truncated regional coarse latent \(s^{\mathcal{O}|j}\) and previously generated adjacent sub-cube latents \(s^{\text{Adj}(j)}\). These conditions are channel-concatenated with noisy latents and passed through a lightweight shared MLP (\(\mathcal{F}_\theta\)) before feeding into the frozen flow Transformer:
\(\mathcal{F}_\theta\) is initialized as an identity diagonal mapping so that training begins smoothly from the base model behavior, with only \(\mathcal{F}_\theta\) updated under flow matching loss. Furthermore, because the original TRELLIS 3DGS decoder was trained strictly on closed single objects, decoding cropped partial open scenes introduces boundary blur and floating artifacts. Fine-tuning the latent decoder (\(D_L\)) on small scene cubes restores crisp geometric surfaces and refined textures.
Key Experimental Results¶
Main Results¶
Map2World is benchmarked against four representative baselines: InfiniCube (driving scenes), CityDreamer (urban environments), GaussianCube (unconditional 3D generation), and SynCity (a concurrent TRELLIS-based grid tiling method). Evaluation encompasses a GPT 5.3-based World Quality metric (WQ = \(0.15S + 0.45W + 0.25C + 0.15R\) measuring Sharpness \(S\), World Completeness \(W\), Coherence \(C\), and Realism \(R\)), distribution distances (FID / KID) measured against 35 ground-truth rendered scene meshes, and a pairwise double-blind human preference study across 19 participants on 21 evaluation pairs (with Map2World serving as reference).
| Method | S (Sharpness) โ | W (Completeness) โ | C (Coherence) โ | R (Realism) โ | WQ โ | FID (Incep.v2) โ | Human Preference (Map2World Win Rate) โ |
|---|---|---|---|---|---|---|---|
| InfiniCube [ICCV 2025] | 6.5 | 3.8 | 3.2 | 4.6 | 4.05 | 268.0 | 98.25% |
| CityDreamer [CVPR 2024] | 7.8 | 8.5 | 6.2 | 5.9 | 7.36 | 272.8 | 91.23% |
| GaussianCube [NeurIPS 2024] | 6.8 | 4.5 | 5.0 | 5.1 | 5.08 | 169.1 | 98.25% |
| SynCity [arXiv 2025] | 8.2 | 6.8 | 7.6 | 7.3 | 7.25 | 143.4 | 89.47% |
| Map2World (Ours) | 8.0 | 7.8 | 7.9 | 7.6 | 7.76 | 192.4 | Reference |
Ablation Study¶
Ablation experiments evaluate the architecture of the detail enhancer, the impact of Classifier-Free Guidance (CFG), and the necessity of SLAT decoder (\(D_L\)) fine-tuning. Evaluations are conducted on rendered views of test-set meshes using PSNR, LPIPS, and multi-model FID metrics (Inception v3, DINOv2, CLIP).
| Config | Enhancer Architecture | Use CFG | Decoder Fine-tuning | PSNR โ | LPIPS โ | FID (Incep.v3) โ | FID (DINOv2) โ | FID (CLIP) โ | Note |
|---|---|---|---|---|---|---|---|---|---|
| (a) Full Model (Ours) | Channel Concat MLP | No | Yes | 22.53 | 0.2137 | 16.98 | 32.67 | 11.79 | Best perceptual metrics and boundary continuity |
| (b) IP-Adapter | Cross-Attention Adapter | No | Yes | 20.28 | 0.2499 | 29.62 | 80.81 | 19.85 | Severe boundary disconnections and structural blur |
| (c) With CFG | Channel Concat MLP | Yes | Yes | 21.95 | 0.2174 | 19.06 | 38.15 | 21.32 | Oversaturated colors and pronounced geometric distortion |
| (d) No Decoder Fine-tuning | Channel Concat MLP | No | No | 22.08 | 0.2165 | 17.89 | 32.94 | 13.17 | Visible loss of fine microscopic surface sharpness |
Key Findings¶
- World completeness and seamless transitions dominate perceptual quality: While SynCity achieves a slightly higher sharpness score (\(S=8.2\)) by directly copying TRELLIS outputs into isolated grid cells, its completeness score drops to 6.8 due to conspicuous tile boundary artifacts. Map2World achieves a higher scene coherence of 7.9 and an overall WQ of 7.76, winning 89.47% of human votes against SynCity.
- CFG degrades performance in strong-condition latent super-resolution: Standard diffusion pipelines rely heavily on CFG. In latent detail enhancement, however, the unconditional trajectory drifts drastically from the conditional geometry; magnifying this difference induces geometric collapse and severe color over-saturation. Disabling CFG is critical for stable reconstruction.
- Spectral parameterization eliminates optimization divergence: In scale-aware noise optimization, spatial gradient descent with learning rate 9.0 diverges due to gradient spikes, whereas learning rate 1.0 requires excessive steps. 3D FFT parameterization smooths the landscape and reaches IoU/Dice > 0.90 in just 5 iterations.
Highlights & Insights¶
- Lifting 2D layout diffusion to 3D continuous rectified flow: Translating MultiDiffusion from 2D pixel grids into a continuous 3D flow-matching latent space with time-annealed Gaussian kernels enables seamless blending across non-convex, arbitrary-shaped semantic layouts.
- Spectral initial noise steering bypasses full trajectory backprop: Recognizing that scene scale correlates with specific noise sub-manifolds allows gradient optimization via a one-step linear approximation, while 3D FFT spectral parameterization guarantees stable and fast convergence.
- Data-efficient hierarchical latent super-resolution: By leveraging latent crops from only 35 high-quality Objaverse scenes and tuning only a lightweight MLP adapter, the framework achieves an effective 8-fold spatial resolution boost without expensive 3D scene re-training.
Limitations & Future Work¶
- Vertical height uniformity: The 3D semantic masks are currently extruded uniformly along the vertical axis from 2D floorplans, which limits control over complex vertically tiered environments (e.g., multi-story structures, underpasses, or tiered bridges).
- Asset distribution dependence: The synthesis diversity is tied to the underlying TRELLIS asset distribution, which may struggle with rare architectural typologies or dynamic interactive entities.
- Autoregressive inference overhead: The sequential 8-sub-cube generation introduces non-negligible inference latency when scaled to massive multi-kilometer terrains. Future work may explore parallel non-autoregressive latent diffusion.
Related Work & Insights¶
- vs SynCity [arXiv 2025]: SynCity tiles scenes into a rigid square grid using TRELLIS, suffering from boundary seams and inability to handle free-form layout boundaries. Map2World uses continuous 3D MultiDiffusion with soft dynamic Gaussian blending, supporting arbitrary layouts with smooth transitions.
- vs CityDreamer [CVPR 2024] / InfiniCube [ICCV 2025]: Prior scene generators depend on narrow domain datasets (driving videos, synthetic procedural cities), preventing cross-domain generalization. Map2World inherits open-domain 3D object priors and generalizes across natural landscapes, urban centers, and mixed environments.
- vs 2D MultiDiffusion [ICML 2023]: While 2D MultiDiffusion stitches planar windows, Map2World pioneers its adaptation to 3D volumetric rectified flows while resolving 3D-specific challenges including scale drift and latent capacity bottlenecks.
Rating¶
- Novelty: โญโญโญโญโ (Combines 3D MultiDiffusion, spectral noise optimization, and latent detail enhancement to enable arbitrary layout-conditioned 3D world generation)
- Experimental Thoroughness: โญโญโญโญโ (Extensive evaluations across multi-dimensional GPTScore, distribution distances, detailed ablations, and double-blind user studies)
- Writing Quality: โญโญโญโญโญ (Clean narrative arc, rigorous formulation of flow matching MultiDiffusion, and comprehensive ablation analysis)
- Value: โญโญโญโญโ (Presents a viable, data-efficient blueprint for general-purpose 3D world synthesis without requiring vast scene-scale 3D captures)