PolyLayout: Multi-room Manhattan Layout Estimation¶
Conference: ECCV 2026
Paper: ECCV Official Page
Project: https://ghanning.github.io/PolyLayout
Area: 3D Vision
Keywords: Multi-view Room Layout Estimation, Manhattan World Assumption, Featuremetric Alignment, Multi-room Joint Optimization, Adaptive Topology Refinement
TL;DR¶
PolyLayout parameterizes multi-room indoor layouts as 3D Manhattan polygons with shared orientation and floor/ceiling heights, formulating a differentiable objective over DINOv2 feature, edge, and confidence maps solved via coarse-to-fine Levenberg-Marquardt optimization with dynamic wall splitting and simplification.
Background & Motivation¶
Reconstructing 3D indoor room layouts from multi-view perspective images is a fundamental problem in computer vision, underlying autonomous indoor robotics, spatial computing, augmented reality (AR), and holistic scene understanding. Conventional approaches predominantly rely on a single perspective view or a single \(360^\circ\) panorama. However, monocular layout recovery remains inherently ill-posed due to unknown absolute scale, extensive occlusions, and severe domain-specific perspective distortions. While recent multi-view learning systems like Plane-DUSt3R fine-tune foundation models to directly predict structural planes, they exhibit catastrophic cross-dataset degradation. On the other hand, optimization-based frameworks like PixCuboid leverage multi-view featuremetric alignment to achieve high precision, yet they enforce a rigid single-room cuboid assumption that fundamentally fails on real-world multi-wall layouts such as L-shaped rooms, alcoves, and interconnected suites.
A critical structural tension exists in prior methodologies: individual rooms within a building inherently share architectural alignment—such as common Manhattan coordinate axes, continuous floor planes, and uniform ceiling heights—yet virtually all existing pipelines process individual rooms in isolation. This uncoordinated prediction not only accumulates independent angular drifts across rooms, but also discards strong structural co-visibility constraints across shared building boundaries. Furthermore, relying on unconstrained 3D point clouds or raw meshes introduces extreme vulnerability to featureless white walls, sensor noise, and clutter occlusions.
PolyLayout resolves this dilemma by decoupling neural feature scoring from explicit model-based geometric projection, expanding the parameterization to general 3D Manhattan polygons, and explicitly enforcing shared architectural parameters across rooms. Core idea: formulate multi-room layout estimation as a joint optimization of 3D Manhattan polygons with shared global orientation and floor/ceiling heights, minimizing a hybrid featuremetric, edge, and vanishing point objective over pre-trained DINOv2 representations with adaptive wall split-and-merge topology updates.
Method¶
Overall Architecture¶
PolyLayout takes as input a set of \(n\) perspective images \(\{I_i\}_{i=1}^n\) captured across one or multiple rooms with known camera poses \((R_i, t_i)\), intrinsic matrices \(K_i\), and room association indices. The goal is to recover a closed 3D Manhattan polygon \(\mathcal{P}\) for each room.
The pipeline comprises three major stages: first, a neural network featuring a pre-trained DINOv2 ViT-S/14 backbone and two lightweight convolutional decoder heads extracts dense feature maps \(F \in \mathbb{R}^{W \times H \times D}\), edge maps \(E \in \mathbb{R}^{W \times H}\), and corresponding confidence maps \(C_F, C_E \in \mathbb{R}^{W \times H}\) across three pyramid scales (\(1/16, 1/4, 1/1\)). Second, an initial Manhattan polygon \(\mathcal{P}_{\text{init}}\) is generated from the concave hull of camera positions combined with vanishing point alignment. Third, an unrolled coarse-to-fine Levenberg-Marquardt (LM) optimizer iteratively updates the layout parameters. Between optimization steps, self-intersections are repaired, redundant walls are merged using an adapted Visvalingam simplification algorithm, and under-parameterized walls are adaptively split, yielding the refined polygon \(\mathcal{P}_{\text{opt}}\).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multi-view Image Input<br/>Known camera poses and intrinsics"] --> B["Camera Concave Hull & VP Initialization<br/>Construct initial Manhattan polygon Pinit"]
B --> C["DINOv2 Multi-scale Feature & Edge Extraction<br/>Dense features F, edges E, and confidences CF, CE"]
C --> D["Featuremetric & Structural Objective Optimization<br/>Jointly minimize Efeat + Eedge + EVP + Eper"]
D --> E["Dynamic Polygon Topology Refinement<br/>Visvalingam simplification, repair, and wall splitting"]
E --> F["Multi-room Parameter Sharing<br/>Constrain shared orientation R and heights d1, d2"]
F --> G["Optimized 3D Manhattan Polygon Layout Popt"]
Key Designs¶
1. Manhattan Polygon Parameterization & Multi-room Parameter Sharing: Overcoming Cuboid Rigidity without Spatial Entanglement
Conventional bounding-box representations parameterize rooms via 3D center, rotation, and extents \((R, t, s_x, s_y, s_z)\), which entangles individual wall positions with room translation, making cross-room parameter sharing mathematically intractable. PolyLayout decouples global frame orientation from local axis-aligned plane offsets by representing each room layout via a global rotation \(R \in \mathrm{SO}(3)\) and a plane offset vector \(d = [d_1, d_2, \dots, d_p]^\top \in \mathbb{R}^p\) (\(p \ge 6\) and even).
In the local Manhattan frame, the horizontal floor and ceiling planes are specified by \(z = d_1\) and \(z = d_2\). The vertical walls alternate strictly between \(x\) and \(y\) axes: \(x = d_3\), \(y = d_4\), \(x = d_5\), continuing up to \(y = d_p\), forming an orthogonal closed boundary. In multi-room environments, all rooms within the scene explicitly share the identical global rotation matrix \(R\) and vertical floor/ceiling bounds \((d_1, d_2)\). This joint formulation dramatically reduces the degrees of freedom in the optimization landscape, suppresses individual room orientation drift, and transfers strong elevation priors from well-observed rooms to partially observed peripheral spaces.
2. Hybrid Learned-Geometric Objective: Integrating Dense Feature Warping with Model-Based Constraints
Instead of regressing layout coordinates through an uninterpretable black box, the optimization objective combines learned neural scoring with closed-form geometric projection:
Each constituent term addresses a distinct geometric ambiguity: - Featuremetric cost \(E_{\text{feat}}\): computes dense photometric feature consistency across views by warping sampled surface points \(x_{ik}\) on \(\mathcal{P}\) from view \(i\) to view \(j\) via \(\mathcal{W}_{i \to j}\). Residuals are weighted by interpolated confidences \(w_{ijk} = C_{F,i}[x_{ik}] C_{F,j}[\mathcal{W}_{i \to j}(x_{ik}, \mathcal{P})]\) and robustified via Barron's adaptive loss function \(\rho\) against foreground clutter. Self-occlusions inherent to non-convex polygons are explicitly resolved during ray casting. - Edge alignment cost \(E_{\text{edge}}\): uniformly samples 3D points \(X_k\) along all 3D polygon wireframe boundaries, projects them onto image planes via \(\Pi_i(X_k)\), and minimizes distance to predicted high-resolution edge maps \(E_i\). - Vanishing point cost \(E_{\text{VP}}\): aligns detected 2D line segments with the three orthogonal vanishing points defined by the columns of \(R_i R^\top = [v_{i,1}, v_{i,2}, v_{i,3}]\). - Perimeter shrinkage cost \(E_{\text{per}}\): regularizes 2D footprint perimeter via:
This term acts as a soft contracting regularizer, preventing unobserved walls from diverging outwards into unseen space.
3. Camera Concave Hull Initialization & Adaptive Topology Refinement: Dynamic Non-Rigid Polygon Evolution
Directly optimizing a high-degree polygon from scratch causes severe local minima traps. PolyLayout computes an initial camera up-vector from the mean camera vertical axes and extracts an \(\alpha\)-shape concave hull from camera center coordinates in the local \(xy\)-plane. Expanding this hull by margin \(\delta\) and rasterizing it onto an orthogonal grid produces an initial Manhattan polygon \(\mathcal{P}_{\text{init}}\) that naturally encloses the exploration trajectory.
During the multi-scale LM optimization iterations, the polygon topology is dynamically adjusted: - Camera containment enforcement: if any camera center drifts outside the polygon boundary after an LM step, the nearest wall is outward-projected past the camera position. - Manhattan Visvalingam simplification: to prevent over-segmentation and noisy micro-facets, the Visvalingam-Whyatt polyline simplification is adapted to the Manhattan domain by removing pairs of adjacent \((x, y)\) walls whose spanning rectangular area falls below an importance threshold, conditioned on wall convergence (\(|\Delta d_i| < \epsilon\)). - Self-intersection removal & wall splitting: polygon loops and self-intersections are resolved via the Shapely planar engine. At the conclusion of coarse and medium optimization scales, walls extending beyond a maximum geometric length are bisected, dynamically providing additional degrees of freedom to capture alcoves and structural recesses.
Loss & Training¶
The network weights and per-parameter LM damping factors \(\lambda\) are trained end-to-end via unrolled optimization. Given ground truth 2D-3D point pairs on walls, floor, and ceiling \(\{(x_{ik}^{\text{GT}}, X_{ik}^{\text{GT}})\}\), the layout loss is defined as:
Gradients propagate through unrolled LM iterations into the DINOv2 backbone and convolutional heads. To prevent over-smoothing of fine-scale feature representations, supervision on higher-resolution scales is applied only when the coarser scale successfully converges below an error threshold. Edge prediction heads are additionally pre-trained via weighted MSE against rendered ground truth wireframe lines.
Key Experimental Results¶
Main Results¶
PolyLayout was benchmarked on Aria Synthetic Environments (ASE), a newly annotated ScanNet++ v2 real multi-room dataset, and 2D-3D-Semantics. Baselines include point cloud structured model SceneScript, zero-shot unposed foundation model Plane-DUSt3R, multi-view cuboid alignment method PixCuboid, and floor plan extractor RoomFormer.
| Dataset | Method | 3D IoU (%) ↑ | Chamfer (m) ↓ | Wall Recall (%) ↑ | Room Recall (%) ↑ | Depth RMSE (m) ↓ | Normal Recall 10° (%) ↑ | Inference Time (s) ↓ |
|---|---|---|---|---|---|---|---|---|
| ASE (Synthetic multi-room) | SceneScript (sparse point cloud) | N/A | 2.44 | 16.7 | 9.1 | 1.18 | 64.5 | 7.45 |
| ASE | Plane-DUSt3R | N/A | 23.16 | 0.1 | 0.0 | 1.90 | 19.7 | 105.60 |
| ASE | PixCuboid | 68.1 | 0.93 | 54.9 | 45.5 | 0.57 | 90.0 | 1.01 |
| ASE | RoomFormer | 13.8 | 2.37 | 5.0 | 0.5 | 1.75 | 67.8 | 0.06 |
| ASE | PolyLayout (Ours) | 94.3 | 0.12 | 89.0 | 79.0 | 0.09 | 98.2 | 5.48 |
| ScanNet++ (Real multi-room) | SceneScript (COLMAP point cloud) | N/A | 1.80 | 11.9 | 0.3 | 0.80 | 55.8 | 6.16 |
| ScanNet++ | Plane-DUSt3R | N/A | 16.82 | 0.2 | 0.0 | 1.06 | 18.2 | 57.72 |
| ScanNet++ | PixCuboid | 78.8 | 0.36 | 58.9 | 26.1 | 0.26 | 87.8 | 0.87 |
| ScanNet++ | RoomFormer | 10.3 | 1.61 | 3.4 | 0.0 | 0.98 | 48.6 | 0.11 |
| ScanNet++ | PolyLayout (Ours) | 87.4 | 0.20 | 69.3 | 36.3 | 0.16 | 91.6 | 3.54 |
| 2D-3D-S (Cuboid rooms) | PixCuboid | 89.0 | 0.18 | 93.8 | 85.0 | 0.10 | 96.1 | 0.42 |
| 2D-3D-S | PolyLayout (Ours) | 90.0 | 0.15 | 92.2 | 80.6 | 0.10 | 96.4 | 1.11 |
Note: SceneScript and Plane-DUSt3R output open planar surfaces, precluding valid volumetric 3D IoU calculation; their Chamfer distances are evaluated after excluding floor and ceiling ground truth planes.
Ablation Study¶
Comprehensive ablations on the ASE test benchmark dissect the architectural design space across parameter sharing, visual backbones, loss terms, initialization geometries, and topology transformations:
| Ablation Dimension | Configuration | 3D IoU (%) ↑ | Chamfer (m) ↓ | Wall Recall (%) ↑ | Room Recall (%) ↑ | Depth RMSE (m) ↓ | Normal Recall 10° (%) ↑ |
|---|---|---|---|---|---|---|---|
| Parameter Sharing | None (independent rooms) | 93.5 | 0.13 | 88.5 | 78.5 | 0.11 | 97.9 |
| Parameter Sharing | Shared Orientation \(R\) only | 93.8 | 0.13 | 88.9 | 79.6 | 0.10 | 98.0 |
| Parameter Sharing | Shared Orientation & Floor/Ceiling | 94.3 | 0.12 | 89.0 | 79.0 | 0.09 | 98.2 |
| Visual Encoder | ResNet-101 (\(E_{\text{feat}}\) only) | 24.6 | 3.55 | 8.5 | 0.5 | 1.70 | 28.5 |
| Visual Encoder | DINOv2 ViT-S/14 (\(E_{\text{feat}}\) only) | 45.1 | 2.05 | 24.1 | 7.3 | 1.05 | 58.1 |
| Visual Encoder | ResNet-101 (Full model) | 82.7 | 0.43 | 66.8 | 43.2 | 0.29 | 93.9 |
| Visual Encoder | DINOv2 ViT-S/14 (Full model) | 94.3 | 0.12 | 89.0 | 79.0 | 0.09 | 98.2 |
| Loss Formulation | \(E_{\text{feat}} + E_{\text{edge}} + E_{\text{VP}}\) only | 93.5 | 0.19 | 89.3 | 79.1 | 0.11 | 98.1 |
| Loss Formulation | Add Perimeter Cost \(E_{\text{per}}\) | 94.3 | 0.12 | 89.0 | 79.0 | 0.09 | 98.2 |
| Initialization Geometry | Bounding Cuboid | 90.5 | 0.22 | 83.4 | 73.7 | 0.16 | 96.9 |
| Initialization Geometry | Camera Circle | 93.2 | 0.15 | 87.3 | 77.7 | 0.11 | 97.8 |
| Initialization Geometry | Camera Concave Hull | 94.3 | 0.12 | 89.0 | 79.0 | 0.09 | 98.2 |
| Topology Updates | No splitting | 93.6 | 0.14 | 89.2 | 82.2 | 0.10 | 98.0 |
| Topology Updates | No simplification | 92.9 | 0.16 | 86.9 | 72.5 | 0.11 | 97.5 |
| Topology Updates | Simplify & Split (Full model) | 94.3 | 0.12 | 89.0 | 79.0 | 0.09 | 98.2 |
Key Findings¶
- Architectural priors boost multi-room convergence: Sharing the global orientation matrix \(R\) consistently improves accuracy across all spatial metrics. Enforcing common floor and ceiling heights across rooms further elevates 3D IoU to 94.3% and reduces Chamfer distance to 0.12 m, confirming that structural coupling regularizes partially observed rooms.
- Superiority of self-supervised ViT features: Despite comparable parameter counts (~28.5M for DINOv2 ViT-S/14 vs ~30M for ResNet-101), DINOv2 features provide smooth, well-conditioned basins of attraction on textureless indoor walls, boosting isolated featuremetric alignment IoU from 24.6% to 45.1% and full-model IoU from 82.7% to 94.3%.
- Decisive role of polygon simplification: Ablating Visvalingam simplification degrades 3D IoU to 92.9% and room recall to 72.5%, demonstrating that pruning redundant planar degrees of freedom is essential to avoid overfitting noisy local image gradients.
Highlights & Insights¶
- Decoupled neural evaluation and analytical projection: Retaining an explicit camera projection and parameter-based polygon representation while delegating visual matching to pre-trained transformers avoids black-box coordinate regression and achieves out-of-the-box generalization to novel cameras and scene domains.
- Manhattan-constrained Visvalingam reduction: Extending classical polyline generalisation to orthogonal geometries by pruning adjacent orthogonal wall pairs based on bounded rectangular area establishes a principled mechanism for non-continuous topological simplification.
- Implicit shrinkage prior for unobserved views: The perimeter regularization term \(E_{\text{per}}\) injects a subtle inward contracting gradient, mathematically resolving planar divergence in unobserved blind spots without requiring manual scene bounding boxes.
Limitations & Future Work¶
- Orthogonal Manhattan boundary constraints: The formulation is strictly tied to axis-aligned Manhattan structures; non-orthogonal walls, slanted attic ceilings, or circular bay windows cannot be modeled without violating coordinate assumptions.
- Dependency on known poses and room clustering: The pipeline assumes pre-computed camera poses and prior knowledge of which views belong to which rooms. While the method exhibits robust tolerance when paired with visual odometry models like \(\pi^3\) (89.4% IoU), automating view-to-room clustering end-to-end remains an open direction.
- Absence of shared inter-room wall constraints: Although orientation and vertical elevations are shared across the building, common partition walls dividing adjacent rooms are currently optimized as separate plane offsets rather than enforcing zero-thickness coincidence.
Related Work & Insights¶
- vs PixCuboid: While both adopt multi-view featuremetric optimization, PixCuboid is fundamentally restricted to single 6-sided cuboids. PolyLayout generalizes to arbitrary Manhattan polygons, introduces camera concave hull initialization, supports dynamic topology evolution, and achieves superior accuracy even on purely cuboid spaces (90.0% vs 89.0% IoU on 2D-3D-S).
- vs Plane-DUSt3R: Plane-DUSt3R fine-tunes DUSt3R for unposed multi-view plane estimation but suffers from severe out-of-distribution failure, collapsing to near-zero recall on new datasets. PolyLayout maintains robust physical projection geometry and guarantees closed, valid 3D room volumes.
- vs SceneScript / Point-based Fitting: SceneScript and point cloud baselines require heavy pre-processing to compute dense point clouds via MVS, which frequently fail on untextured walls. PolyLayout operates directly on multi-view 2D images, bypassing dense 3D reconstruction and demonstrating higher resilience to sparse inputs.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ (Formulates multi-room layout estimation as a joint Manhattan polygon optimization with parameter sharing and adaptive topology updates)
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ (Curates new multi-room benchmarks on ASE and real ScanNet++, featuring exhaustive baseline comparisons and ablations)
- Writing Quality: ⭐⭐⭐⭐⭐ (Clear mathematical formulation, clean geometric reasoning, and precise exposition)
- Value: ⭐⭐⭐⭐⭐ (Offers an elegant, robust, and deployable framework for indoor 3D reconstruction, robotics navigation, and AR layout mapping)