SemLight: Distilled SemanticโGeometric Fusion for Efficient Local Feature Matching¶
Conference: ECCV 2026
Paper: ECCV Original
Area: 3D Vision
Keywords: local feature matching, semantic-geometric fusion, knowledge distillation, lightweight network, relative pose estimation
TL;DR¶
Addressing matching ambiguities in low-texture, repetitive, and illumination-varying environments, SemLight distills high-level semantic priors from a vision foundation model into a compact CNN backbone and dynamically modulates appearance and surface normals via channel-wise reliability gating and 3-token cross-modal attention, delivering a 17.8% relative gain in AUC@5 over XFeat on MegaDepth while maintaining a real-time GPU runtime of 5.4 ms.
Background & Motivation¶
Local feature matching serves as the cornerstone of fundamental visual perception pipelines, including Structure-from-Motion (SfM), visual SLAM, visual localization, and relative camera pose estimation. Recent deep learning pipelinesโspanning joint detector-descriptor frameworks to learned matcher architecturesโhave substantially pushed matching performance in standard scenarios. Nonetheless, when deployed in environments dominated by extensive textureless walls, repetitive architectural patterns, or severe illumination shifts, conventional methods frequently stumble into matching geometrically similar but semantically distinct regions, resulting in pervasive semanticโgeometric ambiguity.
Existing attempts to mitigate this ambiguity with semantic cues generally follow two paradigms, both burdened by clear compromises. On the one hand, keypoint-level saliency selection or region-level correspondence grouping (e.g., SFD2, MESA) relies on discrete semantic segmentation masks or sparse keypoints; these approaches lack sub-pixel spatial granularity in low-texture zones and introduce cumbersome multi-stage overhead. On the other hand, methods directly infusing dense representations from vision foundation models (such as DINOv2 in DeDoDe, MATCHA, or Cadar et al.) provide rich semantic discrimination, but their heavy ViT backbones introduce hundreds of milliseconds of latency and massive memory consumption, rendering them impractical for resource-constrained platforms such as drones and mobile robots.
The core tension is that high-level semantic discrimination is desperately needed to resolve local geometric ambiguities, yet foundation backbones are computationally prohibitive for edge-scale deployment. The core idea is to distill task-relevant semantic priors from a vision foundation model into a shared compact CNN backbone, and locally integrate appearance, surface normals, and semantic cues via channel-wise reliability gating and 3-token cross-modal attention for real-time disambiguation.
Method¶
Overall Architecture¶
SemLight is designed as a single-pass, end-to-end local feature extraction and matching framework optimized for low deployment latency. Given an input image, a lightweight 5-block convolutional encoder extracts multi-scale representations, which are channel-normalized and upsampled to a shared 1/8-resolution grid. From this fused representation, four parallel heads predict keypoint heatmaps, base appearance descriptors, dense 3D surface normals, and distilled semantic embeddings. The Reliability Gated Fusion (RGFusion) module then evaluates local modality confidence and computes cross-modal token interactions for each keypoint, ultimately producing a 64-dimensional descriptor that combines appearance, geometry, and semantics.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Image I"] --> B["Lightweight 5-Block CNN Backbone<br/>Multi-scale Fused Feature F"]
B --> C1["Keypoint Detection Head<br/>Dense Heatmap Prediction"]
B --> C2["Appearance Descriptor Head<br/>Base Feature Projection d"]
B --> C3["3D Surface Normal Head<br/>Geometric Prior Supervision n"]
B --> C4["Distilled Semantic Embedding (DistilSem)<br/>Student Semantic Head Feature s"]
C2 & C3 & C4 --> D["Channel-wise Reliability Gating<br/>Bottleneck MLP Modality Weights"]
C2 & C3 & C4 --> E["Cross-Modal Token Attention<br/>3-Token Interaction Residual Update"]
D & E --> F["Feature Addition & Normalization<br/>Unified 64-d Robust Descriptor"]
Key Designs¶
1. Distilled Semantic Embedding (DistilSem): Compressing foundation semantics into a shared backbone To eliminate the severe latency of running foundation vision models at inference time, SemLight adopts a two-stage offline-to-online distillation workflow. In the first offline stage, a compact projection head \(P(\cdot)\) is trained on top of a frozen DINOv3 ViT-S backbone to compress high-dimensional token representations into a 64-dimensional matching-oriented map. This learned projection suppresses noise and background clutter while preserving boundary-aware semantic contours. In the second online training stage, the projection head is frozen as a teacher, and a lightweight student convolutional head \(\psi_\theta\) is attached directly to the shared CNN feature backbone \(F\). The student is trained via cosine distance loss against the teacher representation:
Because the foundation ViT teacher is entirely discarded during deployment, the student head introduces negligible parameter overhead, operates within the shared feature space, and eliminates external semantic inference latency.
2. Channel-wise Reliability Gating: Adaptive estimation of multimodal confidence In real-world scenes, the quality of individual modalities fluctuates dramatically: in low-texture planar regions, appearance channels exhibit poor signal-to-noise ratios and must defer to geometry and semantics; conversely, across specular reflections or depth discontinuities, surface normals are prone to distortion, requiring appearance or high-level semantics to dominate. SemLight concatenates the appearance descriptor \(\mathbf{d}\), surface normal \(\mathbf{n}\), and semantic feature \(\mathbf{s}\), passing them through a compact bottleneck MLP:
The gated representation is formed via element-wise multiplication: \(\mathbf{f}_{\mathrm{gate}} = \mathbf{g}_d \odot \mathbf{d} + \mathbf{g}_n \odot \mathbf{n} + \mathbf{g}_s \odot \mathbf{s}\). By computing channel-specific reliability scores on the fly, this module dynamically suppresses degraded modalities and amplifies reliable channels at minimal computational expense.
3. Cross-Modal Token Attention: Lightweight 3-token interaction and residual enhancement While channel gating handles "how much weight" each modality receives, it does not enable explicit semantic-to-appearance or geometric-to-appearance compensation. SemLight formulates the three modality vectors at each keypoint as a set of three tokens \(\mathbf{X} = [\mathbf{d}; \mathbf{n}; \mathbf{s}] \in \mathbb{R}^{3 \times D}\), and performs scaled dot-product attention across this compact token set:
The updated appearance token \(\mathbf{O}_{0,:}\) is then projected and added back to the original appearance feature via a residual connection: \(\mathbf{f}_{\mathrm{att}} = \mathbf{d} + \mathrm{Proj}(\mathbf{O}_{0,:})\). Restricting attention to just three tokens per keypoint completely avoids the quadratic overhead of spatial self-attention, allowing appearance representations to selectively absorb discriminative semantic context to resolve fine-grained geometric ambiguities.
Loss & Training¶
The framework is optimized end-to-end on image pairs sampled from MegaDepth and synthetic COCO, resized to \(800 \times 600\). The overall multi-task training objective is defined as:
Specifically: 1. Keypoint loss \(L_{\mathrm{keypoint}}\): Formulated as negative log-likelihood over a \(65\)-channel cell classification problem, using keypoints detected by ALIKE as pseudo-ground-truth targets. 2. Descriptor loss \(L_{\mathrm{desc}}\): Measures the negative log-likelihood of matching similarity across ground-truth correspondence pairs. 3. Normal loss \(L_{\mathrm{normal}}\): Supervised using cosine similarity \(1 - \frac{\mathbf{n}_{\mathrm{pred}} \cdot \mathbf{n}_{\mathrm{gt}}}{\|\mathbf{n}_{\mathrm{pred}}\| \|\mathbf{n}_{\mathrm{gt}}\|}\), where pseudo ground-truth normals are derived from Depth Anything v2 depth predictions. 4. Semantic loss \(L_{\mathrm{sem}}\): Minimizes cosine distance against the frozen DINOv3 teacher projection. The loss balancing hyper-parameters are set to \(\alpha_1=2, \alpha_2=1, \alpha_3=1\). The model is trained on a single NVIDIA RTX 4090D GPU for 30 hours using the Adam optimizer with an initial learning rate of \(4 \times 10^{-3}\) and batch size of 8.
Key Experimental Results¶
Main Results¶
Relative pose estimation is evaluated on MegaDepth-1500 (outdoor viewpoint and lighting variations) and ScanNet-1500 (indoor textureless walls and repetitive geometry). Pose recovery uses MAGSAC++ under mutual nearest neighbor (MNN) descriptor matching:
| Method | Desc. Format & Dim | ScanNet AUC@5 | ScanNet AUC@10 | ScanNet AUC@20 | MegaDepth AUC@5 | MegaDepth AUC@10 | MegaDepth AUC@20 |
|---|---|---|---|---|---|---|---|
| ORB (ICCV2011) | 256-b | 9.0 | 18.5 | 29.9 | 17.9 | 27.6 | 39.0 |
| SuperPoint (CVPR2018) | 256-f | 12.5 | 24.4 | 36.7 | 37.3 | 50.1 | 61.5 |
| DISK (NeurIPS2020) | 128-f | 9.6 | 19.3 | 30.4 | 35.0 | 51.4 | 64.9 |
| ALIKE (TMM2022) | 64-f | 8.0 | 16.4 | 25.9 | 49.4 | 61.8 | 71.4 |
| SiLK (ICCV2023) | 32-f | 15.9 | 30.1 | 44.5 | 39.9 | 55.1 | 66.9 |
| SFD2 (CVPR2023) | 128-f | 13.0 | 25.6 | 38.3 | 44.9 | 58.2 | 68.3 |
| Cadar et al. (ACCV2023) | 256-f | 20.3 | 36.8 | 52.4 | 35.7 | 51.2 | 64.2 |
| XFeat (CVPR2024) | 64-f | 16.7 | 32.6 | 47.8 | 42.6 | 56.4 | 67.7 |
| LiftFeat (ICRA2025) | 64-f | 18.5 | 34.9 | 51.2 | 44.7 | 59.5 | 70.3 |
| SemLight (Ours) | 64-f | 19.2 | 35.8 | 52.5 | 50.2 | 62.5 | 72.7 |
Model complexity and runtime are benchmarked at VGA resolution on both an ultra-low-power CPU (Intel i5-1135G7 @ 2.4 GHz) and a mobile GPU (NVIDIA GeForce MX450):
| Metric | SuperPoint | XFeat | LiftFeat | SemLight (Ours) |
|---|---|---|---|---|
| Params (M) | 1.30 | 0.66 | 0.85 | 0.92 |
| FLOPs (G) | 19.85 | 1.33 | 4.96 | 5.41 |
| Descriptor Dimension | 256-f | 64-f | 64-f | 64-f |
| CPU Runtime (ms) | 273 | 43 | 65 | 72 |
| Mobile GPU Runtime (ms) | 25.9 | 3.8 | 4.9 | 5.4 |
Ablation Study¶
An incremental ablation study on MegaDepth-1500 isolates the impact of distilled semantics, reliability modeling, and cross-modal attention (using LiftFeat as the baseline):
| Config | AUC@5 | AUC@10 | AUC@20 | GPU Runtime (ms) | Note |
|---|---|---|---|---|---|
| Baseline (LiftFeat) | 44.7 | 59.5 | 70.3 | 4.9 | Geometry + appearance only |
| BL + DistilSem + Naive Addition | 47.8 | 60.6 | 70.7 | 5.1 | Adding semantic features without gating |
| BL + DistilSem + Reliability Modeling | 49.4 | 61.7 | 71.9 | 5.3 | Adding channel-wise confidence gating |
| BL + DistilSem + Cross-Modal Reasoning | 48.9 | 61.3 | 71.4 | 5.2 | Adding 3-token attention interaction |
| Full Model (SemLight) | 50.2 | 62.5 | 72.7 | 5.4 | Combined gating and attention |
| Replace DistilSem with raw DINOv3 | 53.6 | 65.9 | 74.6 | 738.8 | High accuracy but 136x latency explosion |
Key Findings¶
- Knowledge distillation achieves high efficiency-accuracy balance: Naively injecting the distilled 64-d semantic embedding improves AUC@5 by +3.1% (44.7 to 47.8) with only 0.2 ms runtime overhead. In contrast, running full DINOv3 achieves 53.6 AUC@5 but incurs a prohibitive 738.8 ms latency.
- Complementarity of gating and attention: Channel-wise reliability gating provides adaptive coarse-grained weighting (+1.6% AUC@5), while token-level attention enables fine-grained cross-modal feature compensation (+1.1% AUC@5). Their combination delivers optimal discriminability.
- Robust cross-domain indoor generalization: Without indoor fine-tuning on ScanNet, SemLight reaches 52.5 AUC@20, surpassing Cadar et al. (52.4) which was trained directly on ScanNet.
Highlights & Insights¶
- Task-oriented semantic projection: Rather than distilling generic token embeddings from foundation models directly, SemLight trains a compact matching projection that filters background clutter before distillation, significantly lowering student learning difficulty.
- Ultra-compact 3-token cross-attention: Limiting token interaction to the three modalities of a single keypoint provides cross-modal reasoning without spatial quadratic expansion, keeping attention latency under a fraction of a millisecond.
- Practical blueprint for edge feature matching: SemLight demonstrates that compact semantic priors can be distilled into sub-1M parameter CNNs, setting a practical template for real-time robotic SLAM and mobile AR.
Limitations & Future Work¶
- Nocturnal domain degradation: Because distillation data originates predominantly from daytime outdoor sets (MegaDepth), semantic cues experience degradation in extreme low-light environments, leading to a minor drop in nighttime Aachen localization (81.7% vs 82.1% at 0.25m/2ยฐ).
- Dependency on pseudo-normal quality: Geometric supervision relies on monocular depth priors from Depth Anything v2, which can be noisy across transparent or dynamic surfaces.
- Future directions: Investigating unsupervised night-domain distillation adaptation and integrating thermal or event-camera modalities.
Related Work & Insights¶
- vs XFeat: XFeat delivers extreme speed using a lightweight CNN but lacks semantic and 3D geometric awareness, struggling in repetitive or textureless scenes; SemLight incurs only 1.6 ms additional latency while boosting MegaDepth AUC@5 by 17.8% relative.
- vs LiftFeat: LiftFeat introduces 3D surface normals to lightweight matching but suffers under symmetric or lighting-shifted structures; SemLight demonstrates that compact semantic cues provide essential orthogonal disambiguation.
- vs Cadar et al. & DeDoDe: Prior foundation-model-based matchers require heavy ViT inference; SemLight compresses semantics into a lightweight CNN via teacher-student distillation, preserving a compact 64-d descriptor and real-time efficiency.
Rating¶
- Novelty: โ โ โ โ โ Elegant combination of task-oriented semantic distillation and 3-token reliability-gated multimodal fusion.
- Experimental Thoroughness: โ โ โ โ โ Extensive benchmarking across relative pose, homography, visual localization, and multi-hardware runtime profiles.
- Writing Quality: โ โ โ โ โ Clear motivation, sound mathematical formulations, and thorough qualitative analysis.
- Value: โ โ โ โ โ High practical value for real-time mobile robotics, visual SLAM, and resource-constrained edge computing.