Skip to content

Pixel-wise Planarity for High-Precision Monocular Plane Segmentation

Conference: ECCV 2026
Paper: ECCV 2026
Area: Segmentation
Keywords: monocular plane segmentation, pixel-wise planarity, geometric consistency, parallel region growing, monocular geometric foundation model

TL;DR

This paper presents a high-precision monocular plane segmentation framework that freezes a pretrained monocular geometric backbone, attaches a lightweight pixel-wise planarity prediction head, and performs GPU-parallel non-iterative region growing alongside geometry-consistent ground truth generation, dramatically suppressing false planar detections under strict distance thresholds.

Background & Motivation

Planar structures are ubiquitous in human-engineered indoor and outdoor environments, constituting dominant structures such as walls, floors, ceilings, pavements, and architectural facades. Robust monocular plane segmentation provides a structured, compact geometric prior that directly empowers downstream spatial AI applications, including visual SLAM, plane-guided room layout estimation, neural radiance fields (NeRF), and 3D Gaussian Splatting. In these applications, even minute segmentation deviations or spurious plane detections can propagate into severe reconstruction drift and non-planar distortions across multi-view fusion pipelines.

Prior monocular plane reconstruction frameworks have primarily evolved along two paths: instance-detection architectures (such as PlaneNet, PlaneRCNN, and PlanarRecon) and transformer-based query models (such as PlaneTR, PlaneRecTR, and Zeroplane). While these paradigms directly predict plane parameters and instance masks, they frequently produce over-segmentation and false positive detections in smooth, non-planar regions. More critically, existing benchmarks rely on coarse geometric fitting to generate ground-truth annotations, which suffer from severe coplanarity violations and outlier noise under tight distance thresholds. Consequently, models trained on such noisy annotations inherently lack the fine-grained supervision required to differentiate true physical planes from locally smooth curved surfaces.

To resolve this bottleneck, the paper decouples geometric representation learning from discrete instance clustering, reformulating planarity as a dense, pixel-wise physical property. Core idea: attach a lightweight pixel-wise planarity prediction head to a frozen monocular geometric foundation model and perform single-pass GPU-parallel region growing under geometry-consistent ground truth supervision to achieve high-precision planar segmentation with minimal computational overhead.

Method

Overall Architecture

The proposed pipeline comprises four tightly coupled stages: first, generating high-precision ground truth annotations via strict 3D distance, angular, and semantic consistency constraints; second, extracting dense metric depth, surface normals, and 3D point maps using a frozen MoGeV2 geometric backbone; third, estimating dense pixel-wise planarity confidence maps with a lightweight normal-initialized head; and fourth, grouping plane instances through a GPU-parallel non-iterative region-growing operator, followed by closed-form SVD parameter estimation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Monocular RGB Input"] --> B["Geometry-Consistent Ground Truth Generation<br/>Strict distance, normal, and semantic filtering"]
    B -->|BCE Supervision| C["Lightweight Pixel-wise Planarity Head<br/>Normal-initialized dense planarity confidence"]
    A --> D["Frozen MoGeV2 Backbone<br/>Predicts metric depth, normals, and 3D points"]
    D --> C
    D --> E["Geometry-Driven Parallel Region Growing<br/>Joint confidence, normal smoothness, and relative depth"]
    C --> E
    E --> F["Closed-Form SVD Plane Fitting<br/>Instance point cloud SVD analytical solve"]
    F --> G["High-Precision 3D Plane Segments and Parameters"]

Key Designs

1. Geometry-Consistent Ground Truth Generation: Reconstructing High-Fidelity Supervisions via Strict 3D Tolerances Standard plane datasets, such as the ScanNet annotations generated via PlaneRCNN, exhibit severe geometric inconsistencies under tight thresholds due to sensor depth noise and coarse fitting heuristics. At a strict 0.1 cm distance threshold, their inlier ratio drops drastically, frequently labeling curved sofas, bedding, and non-planar furniture as planar surfaces. To eliminate noise at the supervision source, the authors build a progressive 3D region-growing annotation pipeline starting from high-precision 3D meshes or clean depth maps. Backprojecting pixels into 3D coordinates \(P = d \cdot K^{-1} [u, v, 1]^\top\), the pipeline enforces strict dual thresholds: point-to-plane Euclidean distance \(|n^\top P - c| < \tau\) and angular normal deviations. Furthermore, semantic segmentation constraints ensure that a plane segment never crosses semantic class boundaries and filters out non-planar objects. On ScanNet++, this reconstructed ground truth boosts the F1 score at 0.1 cm from 42.8% to 50.8% and achieves 81.3% at 1.0 cm, providing clean physical supervision for planarity learning.

2. Lightweight Pixel-wise Planarity Head: Normal-Prior Warm-Starting and Dense Confidence Estimation From differential geometry, a planar surface maintains a spatially zero normal gradient, whereas curved or cluttered geometries produce rapid normal variations. Consequently, intermediate features learned by surface normal estimators inherently encapsulate the spatial differential cues needed for planarity estimation. Rather than training a heavy network end-to-end, the framework keeps the pretrained MoGeV2 backbone and its original heads completely frozen, mounting only a lightweight head of 4.7M parameters onto the shared representations. This head adopts the identical architecture of the pretrained surface normal decoder, replacing only the final layer with a single output channel followed by a Sigmoid activation, and is warm-started using the pretrained normal decoder weights. The head is supervised using a binary cross-entropy loss against ground-truth planarity maps: $\(\mathcal{L}_{\text{planarity}} = -\frac{1}{N} \sum_{i=1}^N \left[ y_i \log \hat{y}_i + (1 - y_i) \log (1 - \hat{y}_i) \right]\)$ where \(y_i \in \{0, 1\}\) represents ground-truth planarity and \(\hat{y}_i\) denotes predicted confidence. This design adds only ~10% latency overhead over the base backbone and converges within 2 epochs (12 hours) on a single NVIDIA RTX 3090.

3. Geometry-Driven Parallel Region Growing: Non-Iterative Clustering via Normal Smoothness and Scale-Adaptive Depth Traditional plane clustering relies either on fixed-query transformer decoders, which struggle with scene complexity variations, or on sequential BFS/DFS traversals on CPU, which cannot exploit GPU parallelism. The authors design a single-pass GPU-parallel region-growing operator. Pixels with planarity confidence above threshold \(\tau_p = 0.3\) are marked as planar candidates. For each candidate pixel, its \(5 \times 5\) neighborhood (24 neighbors) is evaluated simultaneously across three joint geometric criteria: (i) the neighbor is also a planar candidate (\(\ge \tau_p\));
(ii) the neighbor resides in a smooth normal field, satisfying a Sobel normal-gradient magnitude \(\|\nabla n\| \le \sqrt{2 - 2 \cos \theta}\) with \(\theta = 5^\circ\);
(iii) the depth difference satisfies a scale-adaptive relative depth tolerance \(|d_c - d_n| < \tau_{\text{rel}} \cdot d_c\) with \(\tau_{\text{rel}} = 0.025\).
A pixel is established as an active connected node if at least \(T = 8\) neighbors satisfy all three conditions. A standard single-pass connected components algorithm on GPU then extracts discrete plane segments without sequential loops, dynamically adapting to arbitrary plane counts while retaining sharp structural boundaries.

4. Closed-Form SVD Plane Fitting: Analytical Parameter Recovery Without Instance Bounds Direct regression of plane parameters via neural network branches frequently produces misalignments between algebraic plane equations and predicted 2D masks. In contrast, this pipeline adopts an analytical formulation based on the dense 3D point map predicted by MoGeV2. For each extracted segment \(S = \{P_i\}_{i \in S}\), plane normal and offset parameters are solved via least-squares minimization: $\(\hat{\mathbf{n}}, \hat{d} = \arg\min_{\|\mathbf{n}\|=1, d} \sum_{i \in S} |\mathbf{n}^\top P_i - d|^2\)$ This optimization is resolved in closed form via Singular Value Decomposition (SVD) on the zero-centered point cloud matrix, where the left-singular vector corresponding to the smallest singular value provides the optimal normal \(\hat{\mathbf{n}}\) and the centroid projection yields the offset \(\hat{d}\). This analytical fit guarantees exact 2D-3D consistency and executes in microseconds without dedicated regression loss tuning.

Loss & Training

The geometric backbone and its depth, normal, and 3D point heads remain frozen throughout training. Only the 4.7M parameter planarity head is trained using AdamW with an initial learning rate of \(1 \times 10^{-4}\), weight decay of \(1 \times 10^{-5}\), and batch size of 4. Input images are resized to \(476 \times 644\). Training is conducted for 2 epochs on a combined dataset of ScanNet++ (real indoor), Hypersim (synthetic indoor), Virtual KITTI 2 (synthetic outdoor), and SYNTHIA (synthetic outdoor), taking approximately 12 hours on a single NVIDIA RTX 3090 GPU.

Key Experimental Results

Main Results

On indoor benchmarks ScanNet++ and Hypersim, the proposed framework is evaluated against leading monocular plane reconstruction baselines across strict 3D inlier distance thresholds (with gate ratio \(\rho=0.9\)) and standard 2D segmentation metrics (RI, VOI, SC):

Dataset Method RI (↑) VOI (↓) SC (↑) [email protected] [email protected] [email protected] [email protected] [email protected] [email protected] [email protected] [email protected] [email protected]
ScanNet++ PlaneRCNN 0.78 2.28 0.54 32.3 26.0 28.8 60.8 48.8 54.1 72.3 57.7 64.2
PlaneTR 0.77 2.04 0.51 13.3 12.8 13.0 30.1 28.7 29.4 38.8 37.1 37.9
PlanarRecon 0.77 2.25 0.51 18.9 16.4 17.6 39.2 34.2 36.5 49.6 43.2 46.2
PlaneRecTR 0.77 2.06 0.52 11.6 8.6 9.9 28.2 21.6 24.4 38.9 29.9 33.8
Sequential-RANSAC 0.76 1.96 0.56 31.0 20.8 24.9 60.8 40.4 48.6 72.5 47.8 57.6
Zeroplane 0.81 2.04 0.58 28.3 25.8 27.0 58.8 53.6 56.1 70.4 64.1 67.1
Ours 0.83 1.81 0.66 62.1 41.8 50.0 88.8 60.1 71.7 92.8 62.8 74.9
Hypersim PlaneRCNN 0.74 2.99 0.43 41.3 26.8 32.5 45.9 29.4 35.9 50.8 32.3 39.5
PlaneTR 0.72 3.11 0.37 10.4 9.6 10.0 11.5 10.5 11.0 12.9 11.7 12.3
PlanarRecon 0.76 3.38 0.36 17.1 13.5 15.1 19.4 15.2 17.0 21.4 16.7 18.8
PlaneRecTR 0.76 3.07 0.39 11.8 8.8 10.1 13.2 9.8 11.2 14.9 11.1 12.7
Sequential-RANSAC 0.73 2.63 0.47 43.2 25.6 32.1 48.9 28.6 36.1 53.6 31.3 39.5
Zeroplane 0.86 2.49 0.52 35.4 31.1 33.1 39.1 34.3 36.5 43.2 37.9 40.3
Ours 0.80 2.68 0.54 78.4 46.8 58.6 83.1 49.4 62.0 85.3 50.6 63.5

Ablation Study

Ablation of geometric cues within the region growing procedure on ScanNet++: surface normals (N), depth continuity (D), and predicted planarity (P):

Normals (N) Depth (D) Planarity (P) RI (↑) VOI (↓) SC (↑) [email protected] [email protected] [email protected] [email protected] [email protected] [email protected] [email protected] [email protected] [email protected]
- ✓ - 0.33 2.32 0.31 0.6 0.6 0.6 1.5 1.5 1.5 2.1 2.1 2.1
✓ - - 0.81 2.32 0.58 43.0 36.4 39.5 72.9 61.6 66.7 81.1 68.4 74.2
✓ ✓ - 0.81 2.30 0.59 45.5 38.5 41.7 76.0 64.2 69.6 84.3 71.1 77.1
✓ ✓ ✓ 0.83 1.81 0.66 62.1 41.8 50.0 88.8 60.1 71.7 92.8 62.8 74.9

Key Findings

  1. Explicit Planarity Acts as an Indispensable Geometric Gate: Normals and depth alone achieve only 45.5% precision at 0.1 cm. Integrating predicted planarity boosts precision to 62.1% (+16.6% absolute gain) and SC from 0.59 to 0.66, proving that planarity confidence is critical to prevent smooth non-planar surfaces from leaking into planar segments.
  2. Depth Discontinuities Fail Without Directional Guidance: Relying solely on depth yields near-zero 3D precision ([email protected] of 0.6%), confirming that depth difference identifies step edges but cannot distinguish planar from curved geometry.
  3. High Efficiency and Modularity: On 1,000 ScanNet++ images, the model operates at 13.3 FPS with 335.6M parameters (only 5.6M over the base backbone), outperforming Zeroplane with DUSt3R (3.4 FPS, 621.4M params) and DINOv2 (4.5 FPS, 374.0M params) by 3-4x in speed while remaining much faster than Sequential-RANSAC (1.4 FPS).

Highlights & Insights

  • Normal-Prior Warm-Starting: Initializing the planarity head from the pretrained normal head leverages existing spatial differential representations, requiring only ~10% latency overhead while avoiding convergence instability.
  • Decoupled Perception and Grouping: Decoupling continuous geometric feature prediction from discrete instance grouping circumvents fixed-query bottlenecks, allowing GPU-parallel region growing to adapt dynamically to scenes of arbitrary geometric complexity.
  • Robust Cross-Domain Generalization: Models trained exclusively on indoor data maintain 60%–95% precision when transferred directly to outdoor datasets (Synthia, VKITTI2) and in-the-wild web images, suppressing spurious plane hallucinations in unseen environments.

Limitations & Future Work

  • Dependency on Monocular Geometry Quality: Segmentation precision relies on the backbone's depth and normal accuracy; artifacts or extreme noise in textureless walls and reflective surfaces propagate into region growing.
  • Fixed Geometric Thresholds: Global coplanarity thresholds (\(\theta = 5^\circ\), \(\tau_{\text{rel}} = 0.025\)) may not adapt optimally across mixed scenes containing both near-field micro-structures and far-field expanses.
  • Physical Sensor Limits of Real Ground Truth: Reconstructed ground truth remains bounded by the physical noise envelope of LiDAR scanners; integrating multi-view uncertainty estimation represents a promising avenue for further refinement.
  • vs PlaneRCNN / PlanarRecon: These Mask R-CNN style architectures rely on anchor proposals and parameter regression, suffering from bounding box artifacts and high false positive rates on curved surfaces. In contrast, this work employs dense planarity classification and parallel region growing to support arbitrary topological boundaries.
  • vs PlaneTR / Zeroplane: Transformer query-based methods enforce a fixed upper bound on plane instances and often merge separated coplanar surfaces. This method's parallel region growing handles unbounded instance counts and achieves nearly double the precision of Zeroplane at 0.1 cm.
  • vs MonoPlane / Sequential-RANSAC: Iterative RANSAC baseline methods are prohibitively slow (1.4 FPS) and prone to fitting phantom planes through noisy depth points. This paper leverages deep planarity gating and parallel GPU connected components to attain superior accuracy at 13.3 FPS.

Rating

  • Novelty: ⭐⭐⭐⭐ [Introduces dense planarity classification to monocular geometry and designs an efficient parallel GPU region growing operator]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive cross-benchmark evaluations spanning ScanNet++, Hypersim, Synthia, VKITTI2, and in-the-wild images with rigorous ground-truth revision]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, mathematically rigorous geometric formulations, and well-structured ablation investigations]
  • Value: ⭐⭐⭐⭐⭐ [Provides a modular, high-precision, and real-time planar prior extraction framework for 3D reconstruction, SLAM, and 3DGS pipelines]