Skip to content

Unified and Efficient Point-Line Local Features

Conference: ECCV 2026
Paper: ECCV Paper
Code: https://github.com/francois141/upal
Area: 3D Vision
Keywords: local feature extraction, joint point-line detection, knowledge distillation, visual localization, LSD acceleration

TL;DR

UPAL introduces a unified and lightweight local feature extractor that jointly extracts keypoints, line segments, and descriptors within a single ALIKED-based backbone using multi-teacher distillation and GPU-accelerated LSD, achieving a 4x speedup and 10x smaller parameter footprint while matching or surpassing SOTA accuracy.

Background & Motivation

Sparse keypoint detection and robust descriptor matching form the bedrock of classic multi-view computer vision, including structure-from-motion, visual localization, 3D reconstruction, and visual SLAM. However, in low-textured environments, repetitive architectural structures, or under aggressive lighting variations, point-only methods frequently degrade due to the loss of salient corners. Straight line segments, which naturally permeate human-made environments, offer strong geometric priors that remain robust under partial occlusions and textureless surfaces. Integrating line features alongside keypoints has therefore established state-of-the-art results across geometric estimation benchmarks.

Despite their complementary advantages, existing hybrid point-line systems face major computational bottlenecks that prevent real-time deployment on edge and embedded devices. Typical pipelines execute two completely separate heavy neural networks for points and lines respectively, incurring redundant feature extractions since both points and lines localize in high-gradient regions. Furthermore, state-of-the-art deep line detectors rely on deep U-Net architectures and depend on CPU-bound line segment detectors such as LSD for post-processing. Classical LSD entails quadratic computational complexity with respect to image dimensions due to approximate gradient sorting and sequential pixel-level seeding, creating a severe runtime bottleneck. Meanwhile, prior attempts at joint point-line learning suffered from negative multi-task interference and lagged significantly behind dedicated baselines.

Recognizing that points and lines share common underlying image gradients and mutual spatial configurations, this paper investigates how to achieve a unified, lightweight, and streamlined representation. Core idea: build upon a compact convolutional backbone, employ multi-teacher cooperative distillation to jointly learn keypoints, deformable descriptors, and line distance fields while discarding redundant angle field regression, and offload seed filtering to the GPU to realize high-throughput, memory-efficient point-line extraction without requiring dedicated line descriptors.

Method

Overall Architecture

The UPAL pipeline replaces disjoint point and line extractors with a single lightweight feed-forward network. Given an input image, a compact convolutional encoder equipped with deformable convolutions constructs a multi-scale bottleneck representation. From this shared bottleneck, three lightweight decoder branches simultaneously predict a keypoint score map, sparse deformable local descriptors, and a line distance field. For final line segment extraction, UPAL completely bypasses neural angle field prediction, computes gradient angles natively on the GPU, and performs aggressive seed filtering before handing over to an optimized LSD line-growing module.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image<br/>H x W x 3"] --> B["Shared Backbone and Multi-Branch Decoupling<br/>4-stage multi-scale CNN with DCNv2 bottleneck"]
    B --> C["Multi-Teacher Cooperative Distillation<br/>Ensemble soft-target supervision for joint training"]
    B --> D["Angle-Field-Free Efficient Distance Regression<br/>Only 3 conv layers predicting clean distance field"]
    D --> E["GPU-Accelerated Sparse Seed Line Growth<br/>Parallel GPU gradient angles with Top 20% seeds"]
    E --> F["Unified Geometric Feature Output<br/>Subpixel points + 128D descriptors + line segments"]
    B --> F

Key Designs

1. Shared Backbone and Multi-Branch Decoupling: eliminating redundant computation via joint multi-task representation

Instead of executing two separate networks for points and lines, UPAL builds upon the ALIKED-n16 architecture as its shared encoder. The backbone consists of four convolutional stages with progressively reduced spatial resolutions; the final two stages incorporate Deformable Convolutions (DCNv2) to capture non-rigid geometric variations and perspective distortions. Each stage output is projected via a \(1 \times 1\) convolution, bilinearly upsampled to the original input resolution, and channel-concatenated to form the shared bottleneck feature map \(F\). Three lightweight heads branch directly from \(F\): the Score Map Head (SMH) predicts keypoint probability heatmap \(S \in [0, 1]^{H \times W}\), accompanied by Differentiable Keypoint Detection (DKD) that performs non-maximum suppression (NMS) and \(3 \times 3\) softargmax subpixel refinement; the Sparse Deformable Descriptor Head (SDDH) extracts adaptive local spatial offsets around detected keypoints to sample and aggregate multi-scale features into compact 128-dimensional descriptors; simultaneously, the line branch reuses \(F\) directly, confirming that low-level gradient cues and mid-level geometric priors can be unified in an ultra-compact architecture.

2. Multi-Teacher Cooperative Distillation: combining heterogeneous SOTA teachers into an ultra-compact student

Direct multi-task training of a lightweight model from scratch often suffers from optimization conflicts and suboptimal localization in complex indoor scenes. To circumvent this, UPAL introduces an offline multi-teacher distillation strategy that draws on the distinct strengths of top-performing models. While ALIKED descriptors exhibit strong matching capability, its unsupervised keypoints underperform in indoor scenes. Conversely, SuperPoint and DaD provide sharp, highly repeatable corner detections in human-made environments. UPAL synthesizes a composite keypoint teacher heatmap by taking the element-wise maximum: $\(S^* = \max\left(S_{\text{SuperPoint}}, S_{\text{DaD}}\right)\)$ and supervises the keypoint score head via weighted binary cross-entropy loss to counteract foreground-background class imbalance. For descriptors, the top \(n = 1000\) keypoints are supervised via an \(L_1\) loss against ALIKED-n32 teacher descriptors: $\(\mathcal{L}_{\text{desc}} = \frac{1}{n} \sum_{i=1}^n \|d_i - \mathbf{d}_i^*\|_1\)$ The line branch is supervised by distance fields extracted from DeepLSD. This multi-teacher paradigm allows a 0.78M-parameter student model to inherit the complementary strengths of multiple heavyweight models.

3. Angle-Field-Free Efficient Distance Regression: shedding redundant parameters and avoiding angle drift

State-of-the-art line detectors such as DeepLSD and ScaleLSD typically regress both a distance field and an angle field, which significantly increases network depth and decoding latency. UPAL reveals that neural angle field prediction is actually detrimental: under challenging day-night illumination shifts and extreme viewpoint changes, deep angle fields exhibit high-frequency errors that misguide line tracking. UPAL completely eliminates the angle field prediction head, instead calculating image gradient angles via parallel GPU Sobel operators. The line decoder is reduced to just three convolutional layers with batch normalization and ReLU activations. The distance field \(D\) is supervised with an \(L_1\) loss against DeepLSD ground truth \(\mathbf{D}^*\), restricted to a 5-pixel radius \(r=5\) around lines: $\(\mathcal{L}_D = \frac{1}{|\Omega|} \sum_{p \in \Omega} \left| \frac{D(p)}{r} - \frac{\mathbf{D}^*(p)}{r} \right|\)$ This streamlined head saves millions of parameters while outperforming heavy U-Net decoders.

4. GPU-Accelerated Sparse Seed Line Growth: overcoming classical LSD quadratic complexity

While the classical LSD algorithm provides precise line localization with false detection control, its CPU-based approximate sorting and sequential pixel examination cause severe latency scaling with \(\mathcal{O}(H \times W)\) image dimensions. UPAL reconstructs the line post-processing into an efficient hybrid GPU-CPU pipeline. First, all image preprocessing, gradient magnitude, and orientation computations are executed entirely on the GPU. Second, leveraging the empirical insight that the vast majority of low-gradient pixels never contribute to valid lines, UPAL implements a two-fold seed pruning mechanism: downsampling the candidate seed grid with a spatial stride of 2, and retaining only the top 20% of pixels with the lowest distance field values (i.e. those lying closest to line backbones). Only this heavily filtered, sparse seed set is transferred to the CPU for line growing. This design reduces line post-processing latency by 3x without any loss in repeatability or geometric precision.

Loss & Training

UPAL employs a two-stage training scheme to ensure balanced multi-task convergence. Training is conducted on 10,000 distractor images from the Oxford-Paris retrieval dataset resized to \(800 \times 800\), providing extensive illumination and structural diversity. In stage one, only the 3-layer distance field branch is trained for 1 epoch (approx. 1 hour) with early stopping using the Adam optimizer at a learning rate of \(1 \times 10^{-4}\). In stage two, the distance field branch and encoder are frozen while the keypoint score head and deformable descriptor head are trained for 60 epochs (approx. 12 hours) with descriptor window \(K=3\) and keypoint loss weight \(\lambda=200\). Training runs across four NVIDIA TITAN GPUs (24GB each) with a per-GPU batch size of 4.

Key Experimental Results

Main Results

UPAL is extensively evaluated across point benchmarks (HPatches, MegaDepth, ScanNet), line detection benchmarks (HPatches, RDNIM), and hybrid point-line downstream tasks (7Scenes visual localization and ETH3D 3D reconstruction). The latency and parameter comparisons underscore its efficiency:

Method Parameters (M) โ†“ GPU Latency (ms) โ†“ CPU Latency (ms) โ†“ 7Scenes T / R Error (cm / deg) โ†“ 5cm / 5ยฐ Pose Acc. (%) โ†‘
SuperPoint + DeepLSD 9.8 293 2259 5.0 / 1.33 49.6
ALIKED + DeepLSD 9.2 286 2239 5.1 / 1.48 49.6
ALIKED + M-LSD 1.3 54 525 5.2 / 1.52 48.5
ALIKED + TP-LSD 25.0 56 1582 5.2 / 1.48 48.8
DaD + DeDoDe v2 + ScaleLSD 120.0 147 >200K - -
Wireframe (Joint) 5.8 266 2493 4.6 / 1.30 53.8
PLNet (Joint) 7.5 112 6233 5.0 / 1.36 50.1
UPAL (Points only) 0.78 - - 5.1 / 1.39 49.1
UPAL (Points + Lines) 0.78 70 976 4.4 / 1.23 54.6

On generic line detection across HPatches and the challenging day-night RDNIM benchmark, UPAL achieves top repeatability and low localization errors:

Benchmark Method Loc Err 50l โ†“ Loc Err 300l โ†“ H Estim AUC @3px โ†‘ Rep @3px โ†‘ Time (ms) โ†“
HPatches LSD (Classical) 0.50 1.42 88.3 53.6 150
HPatches DeepLSD (Teacher) 0.55 1.43 87.4 54.1 473
HPatches ScaleLSD (Heavyweight) 0.57 1.56 87.6 46.5 177
HPatches UPAL (Ours) 0.51 1.42 88.1 54.1 141
RDNIM (Day-Night) LSD (Classical) 1.78 1.85 44.8 30.0 85
RDNIM (Day-Night) DeepLSD (Teacher) 1.73 1.81 45.3 31.8 294
RDNIM (Day-Night) ScaleLSD (Heavyweight) 1.71 1.83 47.4 25.0 166
RDNIM (Day-Night) UPAL (Ours) 1.62 1.69 45.3 35.9 83

Ablation Study

The ablation study on HPatches verifies each progressive enhancement in the line detection pipeline, particularly highlighting the removal of neural angle fields and the incremental speedups from GPU offloading and seed subsampling:

Step Configuration & Stage Loc Err โ†“ H Estim AUC (%) โ†‘ Repeatability (%) โ†‘ Latency (ms) โ†“ Note
(1) DeepLSD Vanilla Baseline 1.432 90.6 27.2 233 U-Net with angle field + CPU LSD
(2) DeepLSD without Angle Field (no AF) 1.341 90.6 28.2 218 Removing angle field lowers localization error
(3) LSD Classical Baseline 1.414 91.1 26.3 36 Standard CPU LSD implementation
(4) (3) + Parallel Angle Computation 1.414 91.1 26.3 27 Identical accuracy, 25% lower latency
(5) (4) + Seed Grid Stride = 2 1.455 90.6 26.3 19 Filters dense redundant seeds, down to 19ms
(6) (5) + Top 20% Seed Subset 1.476 90.6 26.3 17 Discards uninformative low-gradient background
(7) (6) + Full GPU Preprocessing (UPAL) 1.343 90.7 28.9 11 Optimized GPU pipeline with lowest 11ms latency

Key Findings

  • Angle field elimination boosts robustness: Removing the learned angle field reduces localization error from 1.432 to 1.341 on HPatches. On RDNIM day-night transitions, UPAL's localization error (1.62) markedly surpasses its teacher DeepLSD (1.73), verifying that learned angle fields are vulnerable to domain shifts.
  • Line endpoint matching suffices without dedicated line descriptors: By simply matching line endpoints using the point descriptor head and resolving bipartite matching via Sinkhorn, UPAL eliminates the need for heavyweight graph neural network matchers while attaining superior localization accuracy (54.6% @ 5cm/5ยฐ on 7Scenes Stairs).
  • Point-line multi-task synergy: On indoor ScanNet pose estimation, standalone ALIKED scores 6.9/13.3/20.1 due to corner supervision deficit. In contrast, UPAL achieves 15.6/28.9/41.4, demonstrating that joint training with line distance fields anchors keypoints along structural edges.

Highlights & Insights

  • Minimalist line decoder design: Adding only three convolutional layers to a lightweight point extractor achieves state-of-the-art line detection, disproving the prevailing assumption that accurate line distance fields require massive U-Net architectures.
  • Hybrid edge-friendly algorithmic engineering: Offloading parallelizable gradient and candidate seed filtering to GPU while leaving sequential region-growing to CPU provides an optimal template for accelerating classical geometric vision algorithms.
  • Endpoint reuse simplifies downstream integration: By representing line correspondence entirely through point descriptors at endpoints, the downstream matching graph requires no specialized line descriptor representations.

Limitations & Future Work

  • Limitations acknowledged by authors: UPAL focuses strictly on feature detection and endpoint description; multi-view matching currently relies on external heuristic solvers and Sinkhorn iterations rather than a unified end-to-end differentiable point-line matcher.
  • Practical failure modes: In scenes undergoing heavy non-rigid deconstruction or extreme darkness where visual contours dissolve, distance field predictions can fragment, leading to broken line segments; endpoint occlusions also degrade bipartite line association.
  • Promising future directions: Integrating UPAL with a unified, lightweight transformer-based matcher (e.g. extending LightGlue to jointly handle point and line constraints) could deliver an end-to-end differentiable point-line pipeline.
  • vs DeepLSD: DeepLSD relies on a heavy U-Net to predict both distance and angle fields, executing line extraction entirely on CPU; UPAL reduces model parameters by 10x, discards the problematic angle field, and achieves a 4x overall speedup via GPU-accelerated LSD.
  • vs Wireframe / PLNet: Wireframe and PLNet jointly predict points and lines but employ larger networks (5.8M - 7.5M parameters) and suffer from multi-task performance degradation; UPAL leverages multi-teacher distillation on ALIKED, achieving superior point and line metrics with only 0.78M parameters.
  • vs ALIKED: ALIKED extracts point features only and struggles with indoor corner repeatabilities; UPAL inherits its efficient backbone and deformable descriptors while rectifying indoor weaknesses via SuperPoint/DaD distillation and adding SOTA line detection at negligible overhead.

Rating

  • Novelty: โญโญโญโญ [Elimination of angle fields, streamlined 3-conv line head, and GPU-hybrid LSD provide clear architectural innovations]
  • Experimental Thoroughness: โญโญโญโญโญ [Exhaustive evaluations covering point, line, visual localization, 3D reconstruction, and fine-grained latency ablations]
  • Writing Quality: โญโญโญโญโญ [Clear motivation, self-contained mathematical formulations, and compelling efficiency-accuracy Pareto comparisons]
  • Value: โญโญโญโญโญ [Highly practical for real-time SLAM, mobile AR, and robotics deployment on constrained edge hardware]