Skip to content

RoomPlanner: Reachability-Aware View Sampling for Text-to-Room 3D Gaussian Splatting

Conference: ECCV 2026
Paper: ECCV 2026 Poster 5856
Code: https://kaitlina-s.github.io/RoomPlanner/
Area: 3D Vision
Keywords: 3D Gaussian Splatting, text-to-3D scene generation, layout planning, reachability-aware view sampling, interval timestep flow sampling

TL;DR

Addressing the mismatch between spatial layout feasibility and camera view selection in text-to-room generation, RoomPlanner introduces hierarchical LLM planning with dual collision-and-reachability constraints, seamlessly bridging traversable layout maps to ReachView camera sampling and Interval Timestep Flow Sampling (ITFS) to synthesize physically plausible, decoupled, and editable 3D Gaussian Splatting indoor scenes in under 30 minutes from short text.

Background & Motivation

Synthesizing complete 3D indoor scenes directly from natural language text is a vital capability for embodied AI, gaming, virtual reality, and digital production. However, moving from isolated single-asset 3D generation to multi-object indoor environments introduces formidable challenges in geometric coherence and physical plausibility. Existing text-to-3D scene generation frameworks generally split into two paradigms: visual-guided methods and rule-based methods. Visual-guided methods leverage multi-view or panoramic 2D diffusion models to reconstruct 3D neural fields from generated intermediate images. In doing so, they typically treat foreground furniture and background room architecture as an indivisible continuous representation, sacrificing object-level decoupling and frequently generating distorted geometries, floating artifacts, or hazy boundaries. Rule-based and early LLM layout methods incorporate symbolic physical constraints, yet they either demand laborious manual layout authoring and expert intervention or rely on retrieving static, pre-existing CAD assets, which severely curtails the diversity and novelty of generated scenes.

The deeper, unaddressed tension lies in the disconnect between spatial layout generation and downstream neural rendering optimization. High-quality differentiable rendering—such as 3D Gaussian Splatting (3DGS)—relies heavily on comprehensive camera viewpoints and reliable traversal trajectories. However, conventional layouts derived from open-ended text or heuristic rules frequently block camera motion across the room, preventing close-up inspection of key objects. This mismatch starves object-centric views, degrades the multi-view observation set, and precipitates sticky geometry, opacity artifacts, and blurred surfaces during score-distillation optimization.

This paper tackles this issue by treating renderability and physical navigability as first-class citizens within the layout optimization loop. The core idea is to couple collision-free layout reasoning with an agent-based A* reachability constraint, directly projecting the resulting traversable room map into a hybrid ReachView camera trajectory and a mode-aware Interval Timestep Flow Sampling (ITFS) distillation scheme, thereby achieving physically rational, object-decoupled, and highly editable 3D indoor scene synthesis.

Method

Overall Architecture

RoomPlanner operates across three sequential stages: Hierarchical LLM Agents Planning, Layout Arrangement Planning, and Differentiable Scene Optimization. Given only an ambiguous, concise prompt (e.g., "A bedroom"), the system outputs a complete, high-fidelity 3DGS room representation \(\text{Scene}^\star\) featuring decoupled objects and clean room geometry.

The data flow progresses as follows: First, high-level and low-level LLM agent planners unpack the prompt into detailed object inventories \((O_i, S_i)\) endowed with real-world scales, styles, and materials. Second, a lightweight text-to-3D generator initializes coarse 3D point cloud assets. In the layout stage, the pipeline resolves spatial overlaps via collision constraints and subsequently enforces traversability through virtual-agent A* path planning. Finally, the resulting 2D traversable map directly guides ReachView camera trajectory sampling, which combines with Interval Timestep Flow Sampling (ITFS) in a single-stage optimization pass to distill rectified flow diffusion priors into 3D Gaussians.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400, 'subGraphTitleMargin': {'top': 8, 'bottom': 16}}}}%%
flowchart TD
    A["Concise User Prompt<br/>Short Prompt"] --> B["Hierarchical LLM Planning & Scale Alignment<br/>Hierarchical LLM Planning"]
    B --> C["Single-Asset 3D Point Cloud Initialization<br/>Point Cloud Assets Init"]
    C --> D["Dual Layout Arrangement Planning<br/>Dual Layout Constraints"]
    D --> E["ReachView Camera Trajectory Sampling<br/>ReachView Camera Sampling"]
    E --> F["Interval Timestep Flow Sampling Optimization<br/>Interval Timestep Flow Sampling"]
    F --> G["Decoupled High-Fidelity 3DGS Scene<br/>Output 3DGS Scene"]

Key Designs

1. Hierarchical LLM Planning & Scale Alignment: Grounding abstract text into physically grounded specifications

Prompting an LLM in a single step to place full 3D coordinates often causes floating objects and severe scale inconsistencies. RoomPlanner decouples reasoning into a high-level semantic planner and a low-level grounding planner. The high-level agent infers the room type and populates the scene with contextually congruent object categories \((O_i, S_i)\). The low-level agent enriches these assets with fine-grained style, texture, and material descriptors. Next, an off-the-shelf text-to-3D generator (e.g., Shap-E) produces initial point clouds \(p_i = (x_i, y_i, z_i, c_i)\) with candidate centers \((x_i^c, y_i^c, z_i^c)\). Crucially, a physical agent verifies canonical real-world dimensions against physical common sense, dynamically rescaling bounding boxes when mismatches are detected to yield a physically grounded layout specification \(\varphi_{\text{layout}} = (\hat{O}_i, \hat{S}_i, p_i)\).

2. Dual Layout Arrangement Planning: Enforcing collision-free placement and agent-based A* reachability

Within an enclosed room bounded by four walls, candidate placements are explored via depth-first search guided by symbolic rules (e.g., placing beds against walls, pairing nightstands beside beds). RoomPlanner guarantees spatial validity through two sequential hard constraints. First, a collision constraint penalizes 3D bounding box overlaps between objects and against walls: $\(R_{\text{coll}} = - \sum_{i=1}^{N}\sum_{j \neq i}^{N} \text{IoU}_{3D}(b_i, b_j) - \sum_{i=1}^{N}\sum_{k=1}^{4} \text{IoU}_{3D}(b_i, b_k^{\text{wall}})\)$ A placement is accepted only when \(R_{\text{coll}} = 0\) and containment within room boundaries holds. Second, to overcome the critical flaw of dense, trapped layouts that block access, a reachability constraint is introduced. Representing a virtual user or camera agent by bounding box \(b_{\text{agent}}\), the 3D layout is rasterized onto a 2D occupancy grid with obstacle clearance dilation. From a fixed room anchor (such as a door), A* shortest path search is conducted toward each object's contact area. The reachability reward is: $\(R_{\text{reach}} = \sum_{i=1}^{N} \mathbf{1}\left[\pi_i \neq \emptyset\right]\)$ where \(\pi_i\) is the valid path to object \(i\). Unreachable assets undergo constrained local grid resampling; if no collision-free, traversable spot is found within the search budget, the asset is omitted. This ensures an entirely connected, navigable indoor space.

3. ReachView Camera Trajectory Sampling: Direct reuse of traversable maps for unoccluded multi-view supervision

Prior 3D scene generators typically sample camera poses via unconstrained random spheres or fixed paths, frequently leading to camera-wall collisions, severe occlusions, or lost targets. ReachView resolves this by directly reusing the 2D traversable grid established during layout reachability validation. Camera locations are restricted to safe, navigable cells, filtering out any pose that breaches room boundaries, collides with furniture, or loses line-of-sight to the designated object. ReachView interleaves two complementary camera modes: a scene-centric global walk, which follows gain-aware A trajectories across nearest-neighbor objects from the room entrance to maintain holistic room context, and an object-centric orbit*, which executes close-range upper-hemisphere sweeps around specific assets to provide dense geometric supervision. Alternating between these modes creates a balanced "zoom-in and zoom-out" observation sequence that prevents both over-generalized blurry views and isolated single-asset over-fitting.

4. Interval Timestep Flow Sampling: Mode-aware timestep scheduling for rectified flow distillation

Conventional Score Distillation Sampling (SDS) optimized with single diffusion timesteps suffers from severe over-saturation and geometric collapse. RoomPlanner builds on Conditional Flow Matching (CFM) with Stable Diffusion 3.5, establishing a Flow Distillation Sampling (FDS) gradient. To simultaneously capture coarse layout coherence and high-frequency surface details, Interval Timestep Flow Sampling (ITFS) aggregates supervision over multiple sampled timesteps \(\{t_i\}_{i=1}^m\): $\(\nabla_\theta \mathcal{L}_{\text{ITFS}}(\theta) = \mathbb{E}_{\{t_i\}, \boldsymbol{\epsilon}} \left[ \sum_{i=1}^{m} w(t_i) \left( v_\phi(x_{t_i}; y, t_i) - (\boldsymbol{\epsilon} - x_0) \right) \times \frac{\partial g(\theta, c)}{\partial \theta} \right]\)$ Crucially, ITFS couples with the ReachView camera mode: during zoom-in object-centric orbits, timesteps are sampled from smaller, low-noise intervals (\(t \in [200, 600]\)) to enforce crisp surface boundaries and micro-geometry; during global-walk views, timesteps transition to higher noise intervals (\(t \in [400, 800]\)) to harmonize cross-object semantic alignment and global lighting consistency. This coordinated distillation shrinks convergence time to 20–30 minutes on a single consumer GPU.

Loss & Training

The 3D scene is parameterized using 3D Gaussian Splatting (3DGS). Differentiable rasterization is guided by a frozen 2D diffusion prior using Stable Diffusion 3.5 Medium. Optimization executes on a single NVIDIA RTX 4090 GPU (24GB VRAM) for 1,500 iterations at \(512 \times 512\) rendering resolution. The joint coupling of ReachView camera trajectories and ITFS avoids multi-stage warm-up or retraining, converging stably in a single pass.

Key Experimental Results

Main Results

The framework is evaluated across five representative indoor room categories: living room, kitchen, dining room, bedroom, and workplace. The evaluation features a double-blind user study involving 30 evaluators (professional interior designers, embodied AI researchers, and graduate students) rating Rationality and Quality on a 1–5 scale, alongside automated 2D-view metrics: Aesthetic Score (AS), ImageReward (IR), and CLIP Score (CS%).

Method Rationality ↑ Quality ↑ Aesthetic Score (AS ↑) ImageReward (IR ↑) CLIP Score (CS % ↑) Runtime
Set-the-Scene 1.13 2.80 4.83 41.10 26.05 ~1.5 h
GALA3D 2.61 3.22 5.07 42.67 25.40 ~3.5 h
DreamScene 3.10 3.15 5.13 43.85 26.77 ~1.0 h
RoomPlanner (Ours) 4.10 3.86 5.31 48.83 28.44 ~20 min

In addition, layout physical validity is evaluated via average colliding object pairs (Col), out-of-bounds objects (OOB), and the percentage of reachable objects via A* from the anchor (Reach %):

Layout Method Collisions (Col ↓) Out-of-Bounds (OOB ↓) Reachability (Reach % ↑)
LayoutVLM 11.09 7.53 86.2%
HOLODECK 0.00 0.87 97.3%
RoomPlanner (Ours) 0.00 0.20 100.0%

Ablation Study

Ablation experiments isolate the contributions of ReachView camera trajectory sampling and Interval Timestep Flow Sampling (ITFS), highlighting their specific roles in mitigating color over-saturation and eliminating geometric instability:

Config Distillation Mechanism Camera Trajectory Visual & Geometric Assessment
w/o ITFS & w/o ReachView Single-timestep vanilla SDS Standard unconstrained sampling Severe over-saturation and over-exposure; discontinuous geometry and surface blurring
ITFS & w/o ReachView Multi-interval rectified flow (ITFS) Standard unconstrained sampling Color saturation normalized, but missing viewpoints cause unstable object geometry and fuzzy door/armchair contours
Full Model (ITFS & ReachView) Mode-aware multi-interval ITFS A* traversable ReachView hybrid Optimal: sharp geometric contours, coherent textures, and complete absence of sticky multi-object artifacts

Key Findings

  • Reachability constraints eliminate visual observation blind spots: Resolving collisions alone (\(\text{Col}=0\)) is insufficient for robust 3DGS optimization. Without navigable clearance between furniture items, the camera cannot inspect side and rear surfaces, leading to gradient blind spots that cause objects to melt into adjacent walls or neighboring furniture. Achieving \(100\%\) reachability completely resolves these sticky geometry artifacts.
  • Mode-aware ITFS outperforms standard single-step SDS: Standard SDS over-saturates color channels during continuous Gaussian parameter updates. Allocating early timesteps \([200, 600]\) to close-up orbits and later timesteps \([400, 800]\) to global walks provides crisp edges locally while securing semantic harmony globally, boosting ImageReward to 48.83.
  • Order-of-magnitude acceleration in rendering convergence: By substituting standard diffusion sampling with rectified flow matching and training with hybrid trajectories, RoomPlanner achieves full convergence in 1,500 iterations, reducing runtime from 1–3.5 hours down to ~20 minutes on a single consumer RTX 4090.

Highlights & Insights

  • Downstream reuse of layout reachability for camera planning: Repurposing a 2D occupancy map and A* path solver from robot navigation into 3DGS camera trajectory generation is an elegant insight that completely prevents camera-wall collisions and occlusion at negligible computational overhead.
  • Coordinated pairing of camera scale and diffusion noise intervals: Synchronizing object-centric close-up views with low noise intervals and scene-centric global views with high noise intervals resolves the long-standing dilemma between local detail fidelity and global semantic unity.
  • Uncompromised object-level decoupling and user editability: By initializing individual 3D assets before compositional rendering rather than baking the entire scene into an inseparable volume, RoomPlanner retains complete post-generation editability, enabling rotation, translation, asset insertion/deletion, and style re-texturing.

Limitations & Future Work

  • GPU VRAM limits on late-stage Gaussian densification: Constrained by a 24GB VRAM ceiling on a single RTX 4090, Gaussian densification and splitting must be capped in later optimization rounds, which can slightly curtail fine-grained surface specularities and secondary light transport in large, sprawling rooms.
  • Absence of articulated assets: Generated objects are currently treated as rigid 3D point cloud assets. Extending the framework to incorporate articulated constraints (such as opening cabinet doors or pulling drawers) remains an exciting future avenue for embodied simulation.
  • vs DreamScene (ECCV 2024): While DreamScene pioneered 3DGS text-to-scene synthesis, it relies on standard pattern sampling, requiring ~1 hour and frequently producing floating opacity artifacts and coarse surfaces. RoomPlanner leverages ReachView and ITFS to slash generation time to 20 minutes while significantly enhancing surface fidelity and layout rationality.
  • vs HOLODECK (CVPR 2024): HOLODECK is a strong rule- and LLM-driven room layout generator, but it operates primarily by querying and placing pre-existing 3D CAD assets from a fixed repository. RoomPlanner generates custom 3D point cloud assets from scratch, providing unrestricted asset novelty while matching or exceeding HOLODECK's physical validity.
  • vs Pano2Room (SIGGRAPH Asia 2024): Pano2Room reconstructs 3D rooms from single panoramic images, intertwining foreground objects with walls into an inseparable mesh/radiance field. RoomPlanner maintains explicit object-level modularity, enabling downstream asset manipulation and flexible text-guided re-styling.

Rating

  • Novelty: ⭐⭐⭐⭐ [Clever coupling of agent-based A* reachability with 3DGS camera sampling, alongside mode-aware ITFS flow distillation]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across subjective user studies, multi-modal automated metrics, layout collision statistics, and ablation studies]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Structured and transparent narrative with rigorous mathematical formulation and clean algorithmic presentation]
  • Value: ⭐⭐⭐⭐⭐ [Substantially lowers the compute and time barriers for high-fidelity indoor 3DGS synthesis, providing strong utility for embodied AI and 3D content creation]