Skip to content

SHReg: Strictly Rotation-Equivariant Point Cloud Registration via Spherical Harmonics

Conference: ECCV 2026
Paper: ECCV 2026
Area: 3D Vision
Keywords: Point Cloud Registration, Rotation Equivariance, Spherical Harmonics, Irreducible Representations, Closed-form Pose Estimation

TL;DR

Addressing the vulnerability of learning-based point cloud registration to unseen rotations caused by heuristic data augmentation or fragile local reference frames, SHReg introduces a strictly SO(3)-equivariant backbone based on spherical harmonics representation theory, simultaneously decoupling strictly rotation-invariant descriptors for matching and leveraging higher-order equivariant tensors for closed-form rigid transformation estimation from single correspondences, outperforming state-of-the-art methods across 3DMatch, 3DLoMatch, and KITTI.

Background & Motivation

Point cloud registration is a fundamental problem in 3D computer vision and robotics, aiming to estimate an optimal rigid transformation that aligns two partially overlapping point clouds. In practical deployments, scans exhibit arbitrary 3D poses and orientations, whereas the extracted geometric descriptors are expected to remain strictly invariant to such global rotations. This enduring structural contradiction makes learning robust and discriminative local representations a primary bottleneck. Early patch-wise methods attempted to achieve rotation invariance either by extracting handcrafted geometric statistics (such as point-pair features, distances, and angles) or by constructing local reference frames (LRFs) to canonicalize local patches. However, handcrafted invariants aggressively discard fine-grained geometric information, while LRF estimation relies heavily on normals or principal curvatures that degrade rapidly under noise, point density variations, and low overlap, frequently suffering from severe axis flips.

With the advent of advanced point cloud backbones, scene-wise dense feature extractors (such as those based on KPConv or FCGF) have become dominant due to their computational scalability and global context modeling capabilities. Nevertheless, standard point convolutions are inherently rotation-sensitive. Prevailing approaches rely heavily on extensive random rotation data augmentation during training to force networks to "memorize" rotation invariance empirically. This heuristic practice entails critical drawbacks: the continuous SO(3) group cannot be sufficiently sampled through discrete augmentations, leading to sharp performance degradation under unseen orientations; furthermore, spending substantial model capacity on compensating for pose variations compromises the network's ability to focus on structural distinctiveness.

To resolve this dilemma, this work departs from empirical augmentation and fragile reference frame estimation by grounding the network in the algebraic representation theory of the special orthogonal group SO(3). Because 3D rigid rotations form an exact Lie group, intermediate features should transform strictly equivariantly under SO(3) by construction, preserving faithful local structural orientation cues while naturally enabling exact rotation-invariant readouts. Core idea: decompose point-wise features into irreducible representations of SO(3) via spherical-harmonic tensor-product convolutions, deriving strictly rotation-invariant descriptors for coarse-to-fine correspondence matching while directly exploiting higher-order equivariant tensors to hypothesize rigid transformations from individual correspondences in closed form.

Method

Overall Architecture

Given two partially overlapping point clouds \(P\) and \(Q\), SHReg executes a dual-branch decoupled equivariant-invariant pipeline. First, a strictly SO(3)-equivariant spherical-harmonic backbone extracts multi-resolution features, organizing all channels as direct sums of SO(3) irreducible representations (irreps). Along the invariant branch, orthogonal norm projection converts irrep blocks into strictly rotation-invariant descriptors, which are fed into a Geometric Transformer to establish coarse-to-fine superpoint and point-level correspondences. Along the equivariant branch, higher-order equivariant tensors from matched point pairs are extracted to construct local orthonormal frames, producing closed-form rigid transformation hypotheses from individual correspondences. Finally, a global geometric verification step selects the optimal transformation maximizing inlier support.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    P["Input Point Clouds P and Q"] --> A["Spherical Harmonic Irreps Tensor-Product Convolution<br/>Strict SO(3)-Equivariant Feature Extraction"]
    A --> B["Rotation-Invariant Descriptor Readout<br/>Orthogonal Norm Projection over Irrep Blocks"]
    A --> D["Feature-Based Single-Correspondence Hypothesis Proposer<br/>Local Orthonormal Frame via ℓ=1,2 Tensors"]
    B --> C["Coarse-to-Fine Correspondence Matching<br/>Superpoint and Dense Point Soft Assignment"]
    C --> E["Global Hypothesis Verification and Selection<br/>Inlier Support Maximization for Closed-form Pose"]
    D --> E
    E --> F["Final Rigid Transformation Alignment (R*, t*)"]

Key Designs

1. Spherical Harmonic Irreps Tensor-Product Convolution: strict SO(3) equivariance by algebraic construction To eliminate the rotation sensitivity of conventional convolutions without resorting to fragile local coordinate systems or empirical data augmentation, SHReg encodes local geometry into the irreducible representation (irrep) space of SO(3). For each point \(p_i\), its feature vector decomposes into a direct sum of type-\(\ell\) irrep blocks: $\(F_i = \bigoplus_{\ell=0}^{\ell_{\max}} F_i^{(\ell)},\qquad F_i^{(\ell)}\in\mathbb{R}^{C_\ell\times(2\ell+1)}\)$ where \(\ell=0\) corresponds to invariant scalar channels and \(\ell>0\) corresponds to higher-order geometric tensors. Under any rotation \(R \in \mathrm{SO}(3)\), each block transforms according to \(F_i^{(\ell)}(R \circ P) = F_i^{(\ell)}(P) D^{(\ell)}(R)^\top\), where \(D^{(\ell)}(R)\) is the Wigner-D matrix of order \(\ell\). In message passing, the unit direction \(\hat{r}_{ij} = (p_j - p_i) / \|p_j - p_i\|\) is lifted into the spherical harmonic basis \(Y^{(\ell_e)}(\hat{r}_{ij})\), while the radial distance \(r = \|p_j - p_i\|\) generates learnable weights via radial basis function expansion and a lightweight MLP \(a^{(\ell_e)}(r)\). Convolutional features are formed by the tensor product between input irreps and spherical harmonic kernels, projected onto valid output irreps via Clebsch-Gordan (CG) coefficients: $\(M_{ij}^{(\ell_o)} = \sum_{\ell_i, \ell_e} C_{\ell_i, \ell_e \to \ell_o} \Big[ F_j^{(\ell_i)} \otimes \big( a^{(\ell_e)}(r_{ij}) Y^{(\ell_e)}(\hat{r}_{ij}) \big) \Big]\)$ subject to triangle constraints \(|\ell_i - \ell_e| \le \ell_o \le \ell_i + \ell_e\). Aggregated neighborhood features are modulated by scalar-gated nonlinearities computed exclusively from invariant \(\ell=0\) activations, guaranteeing that every architectural layer strictly commutes with 3D rotations.

2. Rotation-Invariant Descriptor Readout: preserving structural discriminability under orthogonal contractions Feature matching requires evaluating metric distances across different poses; however, raw equivariant representations change orientation under rotation, precluding direct inner-product comparisons, while naive geometric projections discard critical details. Because Wigner-D matrices are orthogonal (\(D^{(\ell)}(R)^\top D^{(\ell)}(R) = I\)), the Euclidean norm across the \(2\ell+1\) components of any irrep block is strictly invariant under arbitrary 3D rotations. SHReg contracts each channel along its representation dimension: $\(X_{i,c}^{(\ell)} = \sqrt{\sum_{m=-\ell}^{\ell} \big| F_{i,c}^{(\ell,m)} \big|^2}\)$ Concatenating invariant invariants across all orders yields the point descriptor \(d_i = \mathrm{Concat}(\{X_i^{(\ell)}\}_{\ell=0}^{\ell_{\max}})\). These invariant descriptors are passed through a Geometric Transformer with self- and cross-attention, coupled with matchability and saliency heads, to perform coarse-to-fine matching across superpoints and dense points. This enables the matcher to focus exclusively on intrinsic geometric saliency without corruption from rotational variance.

3. Feature-Based Single-Correspondence Hypothesis Proposer: reducing hypothesis search space from cubic to linear Standard robust estimators (such as RANSAC or LGR) require sampling triplets of correspondences to solve for rigid transformations via SVD, resulting in an \(\mathcal{O}(N^3)\) hypothesis space that scales poorly under low overlap or low inlier ratios. In contrast, SHReg utilizes the directional geometric structure embedded within higher-order equivariant features. Specifically, \(\ell=2\) features are aggregated across channels into 5D vectors and mapped to symmetric trace-free tensors \(S_p, S_q\), whose dominant eigenvectors define unsigned principal axes \(a_p, a_q\). Vectorial features from \(\ell=1\) (\(v_p, v_q\)) then resolve the sign ambiguity: \(\tilde{a}_p = \mathrm{sign}(a_p^\top v_p) a_p\). Constructing right-handed orthonormal frames \(A_p, A_q \in \mathrm{SO}(3)\) via Gram-Schmidt orthogonalization allows the relative rotation and translation for each correspondence \((p \leftrightarrow q)\) to be recovered directly in closed form: $\(R = A_q A_p^\top, \qquad t = q - R p\)$ Each single correspondence deterministically produces one valid rigid transformation hypothesis, collapsing the search space to \(\mathcal{O}(N)\). In the global stage, counting inlier correspondences satisfying \(\|R_k \tilde{p}_{x_j} + t_k - \tilde{q}_{y_j}\|_2^2 < \tau\) efficiently pinpoints the optimal transformation without combinatorial sampling.

Loss & Training

The network is trained end-to-end with a multi-task objective \(\mathcal{L} = \mathcal{L}_c + \mathcal{L}_f + \mathcal{L}_r\): 1. Superpoint Matching Loss \(\mathcal{L}_c\): Applies an overlap-aware circle loss on superpoint invariant descriptors to maximize separation between positive and negative patches in overlapping regions; 2. Fine-level Point Matching Loss \(\mathcal{L}_f\): Supervises the soft assignment matrix \(Z\) and saliency heads via negative log-likelihood within ground-truth superpoint clusters; 3. Contrastive Rotation Loss \(\mathcal{L}_r\): Enforces directional consistency on equivariant features under the ground-truth rotation \(R_{gt}\) using a margin-based contrastive loss, ensuring reliable orientation estimation even in sparse overlap zones.

Key Experimental Results

Main Results

On the indoor 3DMatch benchmark (overlap \(>30\%\)) and the low-overlap 3DLoMatch benchmark (overlap \(10\% \sim 30\%\)), SHReg achieves state-of-the-art accuracy compared with scene-wise, patch-wise, and rotation-robust baselines:

Method Size (MB) 3DMatch FMR (%↑) 3DMatch IR (%↑) 3DMatch RR (%↑) Time (s↓) 3DLoMatch FMR (%↑) 3DLoMatch IR (%↑) 3DLoMatch RR (%↑) Time (s↓)
FCGF⋄ (ICCV 2019) 8.76 94.7 31.1 82.8 0.12 59.4 9.8 38.0 0.13
SpinNet⋄ (CVPR 2021) 1.41 97.6 47.5 88.6 9.85 75.3 20.5 59.8 9.03
Predator⋄ (CVPR 2021) 7.43 96.6 58.0 89.0 0.64 78.2 26.7 64.4 0.47
GeoTransformer (TPAMI 2023) 9.83 98.1 70.9 92.4 0.18 87.4 43.5 74.3 0.17
YOHO (ACM MM 2021) 12.38 98.2 64.4 90.8 2.81 78.9 25.9 66.0 2.62
RoReg (TPAMI 2023) 12.71 98.2 81.6 93.0 2.27 82.3 39.6 70.1 2.10
RoITr⋄ (CVPR 2023) 10.10 98.0 82.4 91.9 0.36 89.2 54.6 74.1 0.34
PEAL (CVPR 2023) 9.83 98.4 71.0 94.2 1.46 88.3 46.0 78.8 1.19
PARE-Net (ECCV 2024) 3.84 98.5 76.9 95.0 0.17 88.3 47.5 80.5 0.17
SHReg (Ours) 9.52 98.5 78.6 95.4 0.28 88.6 48.8 82.4 0.28

On Rotated 3DLoMatch, where full-range continuous 3D rotations are applied to test generalization under unseen pose distributions:

Method Size (MB) Standard 3DLoMatch RE (◦↓) Standard 3DLoMatch TE (cm↓) Standard 3DLoMatch TR (%↑) Rotated 3DLoMatch RE (◦↓) Rotated 3DLoMatch TE (cm↓) Rotated 3DLoMatch TR (%↑)
FCGF 8.76 4.84 12.87 39.6 4.74 13.39 24.5 (-15.1)
Predator 7.43 3.61 10.65 65.6 3.55 10.30 64.0 (-1.6)
GeoTransformer 9.83 2.91 8.71 75.4 2.94 8.85 72.6 (-2.8)
PEAL 9.83 2.84 8.64 81.2 2.86 8.53 78.7 (-2.5)
YOHO* 12.38 3.54 10.34 66.6 3.61 10.16 67.1 (+0.5)
RoReg* 12.71 3.01 9.26 71.3 3.03 9.28 71.0 (-0.3)
BUFFER* 0.92 3.03 9.86 74.4 3.02 9.99 74.7 (+0.3)
RoITr* 10.10 2.95 9.03 75.1 2.97 9.08 75.5 (+0.4)
PARE-Net* 3.84 2.87 8.83 81.3 2.84 8.71 81.8 (+0.5)
SHReg (Ours)* 9.52 2.92 8.96 83.0 2.88 8.94 83.3 (+0.3)

On the outdoor LiDAR KITTI Odometry dataset (zero-shot transfer using indoor 3DMatch weights):

Method Size (MB) RE (◦↓) TE (cm↓) TR (%↑) Inference Time (s↓)
FCGF 8.76 0.30 9.5 96.6 —
D3Feat 14.08 0.30 7.2 99.8 —
Predator 22.77 0.27 6.8 99.8 0.77
GeoTransformer 25.50 0.23 6.2 99.8 0.26
PARE-Net 2.08 0.23 4.9 99.8 0.21
SHReg (Ours) 24.82 0.23 4.7 99.8 0.24

Ablation Study

Ablation experiments across backbone components and pose estimation strategies on 3DMatch and 3DLoMatch:

Component Category Configuration / Variant 3DMatch FMR (%) 3DMatch IR (%) 3DMatch RR (%) 3DLoMatch FMR (%) 3DLoMatch IR (%) 3DLoMatch RR (%)
Backbone Rotation-sensitive baseline (no SH/irreps) 98.0 71.0 93.6 86.7 41.5 78.2
SH encoding w/o CG projection 98.2 73.5 94.2 87.5 43.5 79.3
SH + CG, w/o Gate nonlinearity 98.3 75.8 94.7 88.0 45.6 80.1
SH + CG + Gate, w/o invariant readout 97.6 68.0 92.8 84.0 36.5 75.5
Full SHReg backbone 98.5 78.6 95.4 88.6 48.8 82.4
Pose Estimator RANSAC (triplet sampling) 98.5 78.6 94.0 88.6 48.8 79.5
LGR / patch-based estimator 98.5 78.6 94.6 88.6 48.8 81.0
Ours (single-correspondence proposer) 98.5 78.6 95.4 88.6 48.8 82.4

Key Findings

  • Invariant readout is indispensable for correspondence matching: Removing the orthogonal norm readout layer causes 3DLoMatch RR to drop precipitously from 82.4% to 75.5% and IR to drop by 12.3 percentage points. This confirms that equivariant features couple structural shape with spatial orientation, requiring norm contraction to extract orientation-independent descriptors for metric matching.
  • Single-correspondence closed-form estimation overcomes low overlap: Holding correspondence sets constant, the proposed single-correspondence proposer improves 3DLoMatch RR by +2.9% over triplet-based RANSAC (82.4% vs. 79.5%). Under 10%~30% overlap, random triplet sampling frequently draws spurious combinations, whereas single-point tensor frames bypass multi-point combinatorial degradation and converge reliably with only 500 hypotheses.
  • Strict equivariance prevents orientation flips under severe rotations: While rotation-sensitive baselines (e.g. FCGF, GeoTransformer) degrade by 15.1% and 2.8% on Rotated 3DLoMatch, SHReg maintains stable performance (+0.3% TR). Equivariant tensor frames encode consistent spatial orientation, eliminating the 180° orientation ambiguity common in low-overlap symmetric geometry.

Highlights & Insights

  • Embedding group symmetries directly into network architecture: Instead of relying on approximate data augmentation or unstable local coordinates, SHReg leverages spherical harmonics and Clebsch-Gordan tensor products so that every layer transforms predictably under SO(3) by construction.
  • Decoupled equivariant-invariant synergy: By projecting irrep features into invariant scalars for matching while preserving high-order tensors for pose recovery, the pipeline successfully satisfies the conflicting requirements of invariant correspondence search and orientation-sensitive alignment.
  • Closed-form pose recovery from individual correspondences: Reconstructing symmetric trace-free tensors from \(\ell=2\) features and resolving axis orientation with \(\ell=1\) features yields deterministic local frames, turning pose hypothesis generation from a cubic sampling problem into a linear deterministic step.

Limitations & Future Work

  • Computational overhead of high-order tensor products: Computing Clebsch-Gordan tensor products and handling high-order irreps incur considerable memory and computational costs during dense point convolution, resulting in larger model footprints in memory.
  • Degeneracy on planar and spherical surfaces: On completely flat or rotationally symmetric local patches, higher-order tensor eigenvalues can degenerate, making principal axis estimation sensitive to noise.
  • Future directions: Exploring sparse, low-rank approximations of spherical harmonic tensor products, extending exact equivariance to Sim(3) similarity transformations, and applying the framework to real-time outdoor LiDAR SLAM.
  • vs GeoTransformer / Predator: Standard Transformer and sparse convolution methods depend on rotation augmentation and degrade under large, unseen rotations; SHReg provides exact mathematical equivariance, showing zero performance degradation under full-range rotations.
  • vs YOHO / RoReg: Previous rotation-robust methods rely on discrete rotation discretization or patch-based LRF estimation, incurring heavy computational costs and sensitivity to occlusion; SHReg achieves continuous closed-form alignment with higher efficiency and accuracy.
  • vs PARE-Net: While PARE-Net explores position-aware equivariance, its pose estimation still relies on multi-point sampling; SHReg integrates single-correspondence closed-form estimation, boosting registration recall especially in low-overlap regimes.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneering integration of SO(3) spherical harmonic irreps with single-correspondence closed-form pose estimation for point cloud registration]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations on 3DMatch, 3DLoMatch, Rotated 3DLoMatch, and outdoor KITTI, with thorough component ablations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear mathematical derivations, well-structured arguments, and self-contained explanations of group-theoretic concepts]
  • Value: ⭐⭐⭐⭐⭐ [Provides a rigorous, reliable framework for robust point cloud registration under extreme pose variations and low overlap]