Skip to content

SFD-Net: Sharp Feature Detection Network Based on Local Geometric Features

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/inyoungoh-cde/SFD-Net
Area: 3D Vision
Keywords: point cloud, sharp feature detection, local geometric descriptor, density variation, multi-scale normal statistics

TL;DR

SFD-Net proposes a plug-and-play multi-scale Local Geometric Descriptor (LGD) derived from second moments of normal differences, combining an enhanced PointNet++ backbone and an imbalance-aware Focal Tversky-Dice loss to achieve state-of-the-art sharp feature detection under joint density variation and noise.

Background & Motivation

Point clouds serve as the native discrete geometric representation for 3D perception across LiDAR, depth cameras, and structured-light systems, underpinning key downstream tasks such as 3D semantic segmentation, object detection, surface reconstruction, and wireframe extraction. However, unlike polygon meshes or regular voxels with explicit topological connectivity, unorganized point sets lack neighborhood connectivity priors. Consequently, local differential geometric properties, such as surface normals and curvatures, depend entirely on the regularity and quality of spatial neighborhoods. Among geometric structures, sharp features—loci where the surface-normal field exhibits discontinuity (\(G^1\) continuity fails and the tangent plane changes abruptly, such as creases at patch intersections and corners where multiple patches meet)—are fundamental for decomposing shapes into distinct faces, edges, and vertices, as well as enabling boundary-aware semantic segmentation and high-fidelity mesh reconstruction.

While learning-based detectors achieve commendable accuracy on clean and dense synthetic benchmarks, two coupled and pervasive perturbations severely degrade performance in real-world scanning pipelines. First, physical sensor measurement noise directly corrupts local covariance and normal estimations, destabilizing the differential cues required for sharp edge discrimination. Second, non-uniform spatial sampling and point thinning broaden the local \(k\text{-NN}\) spacing distribution, introducing severe density variation: fixed-scale neighborhoods either lack sufficient context on sparse regions or over-smooth sharp discontinuities over larger supports. Compounding these issues, sharp feature points constitute an extreme minority class (typically only 1%–10% of total points), causing standard cross-entropy objectives to bias predictions toward the dominant smooth background and collapse recall.

To address these joint challenges without resorting to fragile absolute coordinate modeling or orientation-sensitive global normals, this work revisits multi-scale differential geometric invariants. Core idea: develop a rigid-motion- and scale-invariant Local Geometric Descriptor (LGD) that aggregates weighted second-moment statistics of pairwise normal differences across nested multi-scale neighborhoods into an isotropy index, and integrate it into a dual-Transformer-enhanced PointNet++ backbone guided by an imbalance-aware composite loss.

Method

Overall Architecture

SFD-Net operates via a two-stage pipeline. In Stage 1, a compact, rigid-motion-invariant, scale-invariant, and normal-sign-flip-invariant 3D Local Geometric Descriptor (\(\boldsymbol{\phi}_i \in \mathbb{R}^3\)) is computed for every query point from multi-scale PCA normal statistics. In Stage 2, this geometric descriptor is concatenated with raw point coordinates into a 6D per-point input vector, which is fed into an enhanced PointNet++ encoder-decoder backbone equipped with two lightweight Transformer encoder blocks and optimized via a composite Focal Tversky and Dice objective.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Point Cloud P ⊂ ℝ³"] --> B["Nested Multi-Scale Neighborhoods<br/>Single k₂-NN Query via Prefix Slicing"]
    B --> C["Multi-Scale Normal-Difference Isotropy (LGD)<br/>Three-Scale PCA Normals + Weighted Second Moments + Isotropy Index"]
    C --> D["Feature Concatenation & Geometric Guidance<br/>Raw Coordinates || 3D LGD Descriptor (6 Channels)"]
    D --> E["Dual-Transformer Enhanced PointNet++ Backbone<br/>SA1 Cross-Patch Attention + FP1 Cross-Scale Attention"]
    E --> F["Class-Imbalance-Aware Composite Objective<br/>Focal Tversky Loss + Dice Overlap Loss"]
    F --> G["Point-Wise Sharp Feature Predictions ŷᵢ ∈ {0, 1}"]

Key Designs

1. Multi-scale normal-difference isotropy (LGD): robust geometry under density variation

Purely coordinate-based covariance metrics (such as surface variation) conflate absolute spatial scale with local point density, fluctuating erratically under sparse or irregular sampling. Furthermore, single-scale neighborhoods face an inherent dilemma between fine-scale sensitivity and coarse-scale noise suppression, while global normal orientation propagation introduces branch cuts and sign flips near complex creases. LGD circumvents these hurdles by querying a single \(k_2\text{-NN}\) neighborhood and extracting nested prefix subsets \(\mathcal{N}_0(i) \subset \mathcal{N}_1(i) \subset \mathcal{N}_2(i)\) corresponding to \(k_0 < k_1 < k_2\) (set to 20, 40, 80). At each scale \(s \in \{0, 1, 2\}\), unit PCA normals \(\mathbf{n}_i^{(s)}\) are derived from the eigenvector corresponding to the smallest eigenvalue of the local covariance matrix \(\mathbf{C}_i^{(s)}\) without orientation propagation. To capture directional divergence along surface creases, pairwise normal differences \(\Delta \mathbf{n}_{ij}^{(s)} = \mathbf{n}_j^{(s)} - \mathbf{n}_i^{(s)}\) are weighted by their squared magnitude to form a normalized second-moment tensor:

\[\boldsymbol{\Sigma}_i^{(s)} = \frac{1}{Z_i^{(s)}} \sum_{j \in \mathcal{N}_s(i)} w_{ij}^{(s)} \Delta \mathbf{n}_{ij}^{(s)} \Delta \mathbf{n}_{ij}^{(s)\top}, \quad w_{ij}^{(s)} = \|\Delta \mathbf{n}_{ij}^{(s)}\|_2^2 + \varepsilon_w\]

where \(Z_i^{(s)} = \sum_{j} w_{ij}^{(s)} + \varepsilon_z\) ensures valid normalization. On planar regions, normal differences remain nearly collinear, yielding a single dominant eigenvalue in \(\boldsymbol{\Sigma}_i^{(s)}\); near creases or corners, difference vectors span multiple directions, producing an isotropic tensor spectrum. With sorted eigenvalues \(\mu_{i,1}^{(s)} \le \mu_{i,2}^{(s)} \le \mu_{i,3}^{(s)}\), the dimensionless isotropy index is defined as:

\[\omega_i^{(s)} = 1 - \frac{\mu_{i,3}^{(s)}}{\mu_{i,1}^{(s)} + \mu_{i,2}^{(s)} + \mu_{i,3}^{(s)}} \in [0, 1)\]

Under similarity transforms \(T(\mathbf{x}) = c\mathbf{R}\mathbf{x} + \mathbf{t}\), translation is eliminated by mean centering, isotropic scaling leaves unit eigenvectors unchanged, and rotation transforms the tensor via \(\boldsymbol{\Sigma}_i'^{(s)} = \mathbf{R} \boldsymbol{\Sigma}_i^{(s)} \mathbf{R}^\top\). Crucially, normal sign ambiguity (\(\pm\)) cancels out in the outer product \(\Delta \mathbf{n} \Delta \mathbf{n}^\top\). Concatenating across three scales yields \(\boldsymbol{\phi}_i = [\omega_i^{(0)}, \omega_i^{(1)}, \omega_i^{(2)}] \in \mathbb{R}^3\), where \(k_0\) ensures spatial acuity for thin creases, \(k_1\) provides structural balance, and \(k_2\) dampens random noise and local undersampling variance.

2. Dual-Transformer enhanced PointNet++ backbone: bridging local geometry and non-local context

While PointNet++ hierarchically extracts metric-space features via Set Abstraction (SA) and Feature Propagation (FP), its strictly bounded spherical or \(k\text{-NN}\) receptive fields prevent information exchange across distant topological boundaries. SFD-Net augments this architecture by strategically integrating two lightweight Transformer encoder blocks: the first Transformer block is inserted immediately after the initial SA stage (\(d_{\text{model}}=96\), 6 attention heads, 2 layers, sequence length \(N_1=1024\)), enabling global cross-patch aggregation before aggressive subsampling; the second block is placed after the first FP upsampling stage (\(d_{\text{model}}=256\), 8 attention heads, 2 layers, resolution \(N_3=64\)), reinforcing cross-scale feature consistency and sharpening boundary localization during feature propagation. This design balances bounded local efficiency with global structural context.

3. Class-imbalance-aware composite objective: optimizing recall and boundary overlap

Because sharp feature points are rare (1%–10%), standard cross-entropy and basic focal losses fail to penalize false positives and false negatives flexibly. SFD-Net adopts a composite objective combining Focal Tversky and Dice losses:

\[\mathcal{L} = \lambda_1 \mathcal{L}_{\mathrm{FT}} + \lambda_2 \mathcal{L}_{\mathrm{Dice}}\]

Using continuous soft counts for true positives (\(\mathrm{TP} = \sum_i p_i y_i\)), false positives (\(\mathrm{FP} = \sum_i p_i (1-y_i)\)), and false negatives (\(\mathrm{FN} = \sum_i (1-p_i)y_i\)), the Tversky Index adjusts the trade-off via weighting hyperparameters \(\alpha\) and \(\beta\):

\[\mathrm{TI} = \frac{\mathrm{TP} + \epsilon}{\mathrm{TP} + \alpha \mathrm{FP} + \beta \mathrm{FN} + \epsilon}, \quad \mathcal{L}_{\mathrm{FT}} = (1 - \mathrm{TI})^\gamma\]

The exponent \(\gamma > 0\) focuses optimization on hard boundary examples, while the Dice term \(\mathcal{L}_{\mathrm{Dice}} = 1 - \frac{2\sum_i p_i y_i + \epsilon}{\sum_i p_i + \sum_i y_i + \epsilon}\) directly maximizes overall spatial overlap. Together, they simultaneously suppress false alarms and boost sharp feature recall.

Loss & Training

The framework is implemented in PyTorch on a single NVIDIA RTX 3090 GPU. For every 10K-point input cloud, a deliberate 500-point subsample is ingested per forward pass to stress-test detection under low-density regimes. Optimization runs for 200 epochs using Adam with an initial learning rate of \(10^{-3}\), batch size 16, and a cosine annealing warm restart schedule (\(T_0=10, T_{\text{mult}}=2, \eta_{\min}=10^{-5}\)).

Key Experimental Results

Main Results

On the ABC benchmark (1,000 CAD models split into 400 train, 90 val, 510 test), uniform random thinning is applied to generate density variation, followed by additive Gaussian noise at three relative standard deviations \(\sigma_{\text{rel}} \in \{0.00125, 0.006, 0.012\}\) (0.12%, 0.6%, 1.2% of the bounding-box diagonal). All baselines are retrained from scratch on identical splits:

Method ABCnone F1 (↑) ABCnone FPR (↓) ABC0.12% F1 (↑) ABC0.12% FPR (↓) ABC0.6% F1 (↑) ABC0.6% FPR (↓) ABC1.2% F1 (↑) ABC1.2% FPR (↓)
PointNet++ [26] 50.89 48.57 55.02 45.88 49.03 49.35 50.21 49.01
DGCNN [36] 49.89 48.90 50.65 48.51 48.44 49.58 47.53 49.99
RepSurf [27] 55.23 45.96 57.16 43.06 54.21 46.55 48.17 49.73
PointMLP [21] 54.59 46.30 47.53 50.00 47.52 50.01 47.65 49.94
PIE-Net [35] 52.55 38.50 52.44 37.35 47.61 40.54 49.56 39.33
BoundED [4] 48.12 49.88 47.98 49.94 47.97 49.90 47.70 49.98
SFC-Net [47] 47.92 49.84 48.64 49.53 48.27 49.69 48.03 49.80
MSL-Net [15] 59.41 29.83 59.11 30.79 56.11 36.23 52.74 47.53
EdgeFormer [38] 52.04 31.24 48.66 33.95 34.99 45.16 28.13 48.55
EDWG [41] 54.95 46.06 54.93 46.07 48.20 49.81 47.76 49.90
SFD-Net (Ours) 61.33 36.93 60.42 39.76 58.62 40.88 57.05 41.94

Furthermore, evaluating LGD as an architecture-agnostic plug-and-play frontend across various backbones demonstrates consistent gains:

Baseline Configuration Average F1 (↑) FPR (↓) \(\Delta\)F1 \(\Delta\)FPR
PointNet++ Baseline 50.89 48.57 — —
LGD + PointNet++ 57.89 43.70 +7.00 -4.87
RepSurf Baseline 55.23 45.96 — —
LGD + RepSurf 59.28 43.34 +4.05 -2.62
EDWG Baseline 54.95 46.06 — —
LGD + EDWG 60.80 36.98 +5.85 -9.08
SFD-Net (Full Model) 61.33 36.93 — —

Ablation Study

Ablation experiments on ABCnone isolate the effects of geometric descriptor configurations, Transformer placements, and loss objectives:

Category Module Variant Configuration Input Ch FPR (↓) F1 (↑)
Geom. Prior Descriptor (a) Coordinates only (no LGD) 3ch 41.08 58.55
Geom. Prior Descriptor (b) Smaller scale set \(k \in \{10, 20, 40\}\) 6ch 38.87 61.16
Geom. Prior Descriptor (c) Averaged descriptor across scales 4ch 38.87 60.05
Network Transformer (d) No Transformer blocks 6ch 37.60 61.01
Network Transformer (e) FP stage only 6ch 38.32 60.02
Network Transformer (f) SA stage only 6ch 39.91 60.18
Loss Term Objective (g) Focal Tversky only 6ch 38.14 59.93
Loss Term Objective (h) Dice loss only 6ch 36.80 59.16
Loss Term Objective (i) Standard Cross-Entropy 6ch 40.85 60.58
Full Model All Modules SFD-Net (LGD + Dual Transformer + Composite Loss) 6ch 36.93 61.33

On zero-shot transfer to real indoor scans (S3DIS Area 5 with semantic boundary pseudo-labels), SFD-Net achieves 53.20% F1 and 41.63% FPR, outperforming PointNet++ (50.48% / 49.59%), RepSurf (50.01% / 48.19%), and EDWG (50.88% / 45.40%). In computational efficiency, pairing LGD with the lightweight EDWG backbone achieves 60.80% F1 (within 0.53 pp of SFD-Net) at only 2.65 ms latency (over 60× faster than SFD-Net's 170.7 ms) and 1.20M parameters.

Key Findings

  • Crucial geometric prior: Removing LGD entirely (variant a) incurs the largest individual performance drop of \(-2.78\) pp in F1, demonstrating that neural networks cannot easily learn robust differential discontinuities from sparse, noisy point sets alone.
  • Resilience under severe noise: Under the highest noise level (ABC1.2%), MSL-Net's FPR degrades sharply to 47.53% and EdgeFormer collapses to 28.13% F1, whereas SFD-Net maintains 57.05% F1, outperforming the closest runner-up by 4.31 pp.
  • Concatenation vs. averaging: Averaging the multi-scale indices into a scalar (variant c) degrades F1 by 1.28 pp, showing that decoupling fine-scale crease sensitivity (\(k_0\)) from coarse-scale noise suppression (\(k_2\)) is vital for classification.

Highlights & Insights

  • Tensor outer-product formulation: By constructing second-moment tensors from unoriented normal differences, the method neatly sidesteps ill-posed global normal orientation propagation while guaranteeing invariance to rigid motion, scale, and sign flips.
  • Universal plug-and-play capability: LGD serves as a modular 3-channel input augmentor that consistently lifts the F1-score across PointNet++, RepSurf, and EDWG by +4.05 to +7.00 pp without modifying network architectures.
  • Ultra-fast deployment potential: When coupled with EDWG, LGD achieves near-SOTA accuracy at an inference speed of 2.65 ms per cloud, paving the way for real-time edge computing on robotic and autonomous driving platforms.

Limitations & Future Work

  • Reliance on local planar PCA assumption: In regions with severe non-uniform density or extreme noise, linear PCA normal estimation may drift, suggesting future exploration into lightweight learned normal estimators.
  • Binary classification limitation: The current framework detects sharp points as a single class without distinguishing creases, corners, and boundary silhouettes.
  • Lack of ground-truth real-world benchmarks: Zero-shot real-world evaluation relies on semantic boundary pseudo-labels as a proxy; developing accurately annotated real-scan sharp feature benchmarks remains an important direction.
  • vs MSL-Net [15]: MSL-Net computes normal angle histograms with a shallow MLP, but high measurement noise distorts raw angle distributions, causing its FPR to spike to 47.53%; SFD-Net's second-moment eigenvalue isotropy is inherently more stable, yielding a +4.31 pp F1 margin under heavy noise.
  • vs EdgeFormer [38]: EdgeFormer pairs handcrafted features with attention, but overfits to clean density distributions and collapses under noise; SFD-Net's nested multi-scale prefix structure ensures stable generalization.
  • vs EDWG [41]: EDWG is designed for efficient wireframing but suffers high false positives (FPR 46.06%); prepending LGD drops its FPR by 9.08 pp and boosts F1 to 60.80%, highlighting strong architectural synergy.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Elegant formulation of normal-difference second-moment tensors and isotropy index for point cloud edge detection.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 4 noise/density conditions, 3 backbone plug-and-play tests, extensive ablations, and zero-shot real-world transfer.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous geometric derivations, clean structure, and clear empirical insights.
  • Value: ⭐⭐⭐⭐⭐ Provides an effective, modular tool for CAD feature extraction, surface reconstruction, and edge-aware 3D perception.