Building Transformation Layers for Riemannian Neural Networks¶
Conference: NeurIPS2026
arXiv: 2609.35436
Code: https://github.com/GitZH-Chen/RieTrans
Area: Others (Riemannian neural networks)
Keywords: Riemannian geometry, multiple tangent spaces, manifold convolution, parameter trivialization, positive definite matrices
TL;DR¶
The paper reinterprets fully connected layers as signed point-to-hyperplane responses, constructs manifold-valued layers using multiple tangent spaces and a tractable pseudo-distance, and extends them to convolution through product manifolds, validating representational flexibility across ten geometric instantiations while performance and cost remain geometry- and data-dependent.
Background & Motivation¶
When intermediate representations are covariance matrices, linear subspaces, or hyperbolic embeddings, ordinary matrix multiplication can violate positive definiteness, orthogonality constraints, or curvature structure. Riemannian networks already have normalization, attention, residual, and classification layers, but their basic feature transformations often remain space-specific: SPDNet uses bilinear matrix maps, Grassmann networks use left multiplication followed by reorthogonalization, and hyperbolic networks transform features through the origin tangent space or Lorentz spacetime structure. Each approach has its uses, but changing the representation geometry within a consistent architecture is difficult.
A generic alternative applies the logarithmic map at the origin, performs a Euclidean fully connected transformation, and returns through the exponential map. However, all outputs then share the same expansion point, and feature transformation primarily occurs in one tangent space. Other geometry-specific layers construct outputs from true point-to-hyperplane distances, but those distances need not have closed forms on general manifolds; even when computable, multiple distance constraints need not admit a common output. Convolution based on weighted Fréchet means has a different limitation: its output remains on the original manifold, preventing the dimensional flexibility of ordinary convolution.
Rather than searching for matrix multiplication on every manifold, the paper identifies a more transferable geometric relationship behind fully connected layers: each output coordinate corresponds to the signed response of the input relative to a hyperplane. Core Idea: use tractable point-to-hyperplane pseudo-distances to turn output constraints into tangent coordinates, assign each output its own input expansion point, and recover manifold-valued fully connected and convolutional layers through exponential maps and product geometry.
Method¶
Overall Architecture¶
The input is a point on a known Riemannian manifold, and the output lies on another specified manifold; their intrinsic dimensions may differ. For each output direction, the layer learns an input base point and a tangent normal, computes the logarithmic map of the input at that base point and its normal response, assembles the scalar responses in the output origin's tangent space, and applies the exponential map to obtain a valid output.
Convolution does not introduce a separate manifold averaging operator. Instead, the points within a receptive field form one point on a product manifold, to which the same fully connected construction is applied. Parameter trivialization generates learnable base points and normals from Euclidean parameters in a fixed tangent space, while geometry-specific instantiations expand the unified formula into implementable vector or matrix operations.
Point-to-hyperplane pseudo-distance, multiple-tangent-space responses, parameter trivialization, and product-manifold convolution are therefore different levels of one construction, not four independently trained networks in a sequence. Because the contribution is primarily geometric, no flowchart is used to misrepresent theorems and instantiations as sequential modules.
Key Designs¶
1. Point-to-hyperplane pseudo-distance and multiple tangent spaces: making implicit output constraints explicit and tractable
Each weight row of an ordinary fully connected layer defines a normal direction, while the bias determines the corresponding hyperplane's location. The paper generalizes this relationship to manifolds: the input is expanded through a logarithmic map at a base point, and its inner product with a normal is evaluated using the local Riemannian metric. Points with zero response form the Riemannian hyperplane defined by that base point and normal; the sign distinguishes its two sides, while the magnitude also includes the normal's scale.
The important distinction is that this inner product is not unconditionally called a true geometric distance. The construction explicitly adopts the following surrogate. For a nonzero normal, the paper's point-to-hyperplane pseudo-distance is:
It avoids optimizing over the nearest point on a hyperplane. With a metric-orthonormal basis at the output origin, the signed pseudo-distances to output coordinate hyperplanes are precisely the coefficients of the output logarithmic map. The implicit constraints can therefore be solved by assembling a tangent vector, yielding the central transformation in Theorem 3.3:
Here, \(\mathcal N\) is the input manifold, \(\mathcal M\) is the output manifold, and \(m\) is the output intrinsic dimension; \(B_i\) is a metric-orthonormal basis of the output origin's tangent space. Each output direction has its own \(P_i\) and \(A_i\), so the same input is expanded at multiple learnable locations rather than every output reading a single logarithmic map at the origin. Output synthesis still uses one fixed origin, but the input responses use different local reference points.
This explains both the distinction from a single-tangent-space transformation and the ability to change input and output dimensions. The pseudo-distance agrees with the true geodesic point-to-hyperplane distance only under specific conditions, such as a manifold being isometric to Euclidean space; the paper explicitly includes Euclidean space and SPD geometries LEM and LCM. On a general curved space, “response” in this note should not be interpreted as an exact shortest distance.
2. Parameter trivialization: generating moving base points from a direction and displacement instead of directly optimizing tangent-bundle parameters
If base points and normals are learned independently, updating a base point changes the tangent space containing its normal, so an ordinary Euclidean optimizer cannot treat both as unconstrained vectors. In Euclidean space, many base points also represent the same hyperplane, making independent base-point learning potentially redundant. Following an existing compact parameterization, the paper learns a tangent vector and a scalar displacement at the input origin, then generates the geometric parameters through:
Here, \([Z_i]\) is a nonzero tangent vector normalized using the input origin's metric, \(\gamma_i\) controls displacement along that direction, and parallel transport carries the normal to the new base point. Training updates \(Z_i\) in a fixed tangent space and the real scalar \(\gamma_i\), rather than directly updating \(A_i\) in a moving tangent space. This permits back-propagation with a Euclidean optimizer while retaining manifold constraints on the generated parameters.
For input intrinsic dimension \(n\), each output direction replaces two \(n\)-dimensional geometric parameters with one \(n\)-dimensional tangent vector and one scalar, reducing \(2n\) to \(n+1\). This is the compact modeling choice adopted by the paper, not a proof that every independent base-point/normal pair on any curved manifold has a lossless equivalent in this parameterization. It also does not remove geometric computation costs: matrix logarithms under AIM and matrix square roots under BWM can remain expensive.
3. Product-manifold convolution: combining a receptive field before summing componentwise geometric responses
Ordinary convolution can be viewed as extracting a local window, concatenating its features, and applying a shared fully connected kernel. On a manifold, concatenating arrays of matrices or subspaces does not produce a valid point on the original manifold; the appropriate object is an ordered tuple of points, hence a product manifold. Its metric and logarithmic map decompose componentwise, so each output-direction response sums the base-point inner-product responses of the window's points before synthesis in the output tangent space and mapping back.
One kernel outputs one point on the target manifold, and multiple kernels provide multiple output channels; parameters are shared across windows. The target manifold may change matrix size or subspace dimension rather than matching the input. Unlike a weighted Fréchet mean, the operation is not averaging within the input manifold, so dimensional changes are not constrained by a mean having to remain in the original space. Appendix G.5 also qualifies the symmetry claim: these layers admit equivariance when parameters transform covariantly with the isometry, not ManifoldNet's fixed-parameter equivariance.
The main SPDNN experiments use only one global receptive field across channels and one kernel, making that convolution a fully connected layer on a product manifold. Additional SPDConvNet experiments in the appendix use sliding local windows and shared kernels to directly evaluate local convolution; the main SPDNN results should not substitute for spatial or temporal convolution validation.
4. Geometry-specific instantiations and batching: a unified interface does not imply identical numerical computations
The construction is instantiated on three hyperbolic models, five SPD metrics, and two Grassmann representations. HFC-P and HFC-K simplify responses using Möbius and Einstein addition, respectively, while HFC-H uses the Lorentz metric and parallel transport; these structures are computational tools for instantiation, not extra axioms required by the general theorem. Explicitly constructing all local vectors for every output produces large intermediate tensors. Theorem 4.3 reduces shared computations to batched inner products between inputs and unit directions, followed by elementwise curvature terms, avoiding materialization of the full batch-by-output-by-input tensor.
SPD instantiations assemble responses in matrix tangent spaces appropriate to the metric. LEM and AIM recover positive definite outputs through the matrix exponential at the identity; LCM constructs a Cholesky factor from strictly lower-triangular entries and diagonal logarithmic coordinates, then reconstructs the positive definite matrix. PEM and BWM have only locally defined exponential maps, requiring the positive-definite domain constraints in Appendix G.4 to be handled, for example through eigenvalue clipping. An output SPD matrix of size \(r\times r\) has intrinsic dimension \(r(r+1)/2\), not merely \(r\) scalar outputs.
Grassmann inputs represent subspaces through an orthonormal basis (ONB) or a projector perspective (PP). The target \(\operatorname{Gr}(q,r)\) has \(q(r-q)\) tangent degrees of freedom; assembled responses are mapped back using SVD and trigonometric functions, or an exponential of a skew-symmetric block matrix. This allows both ambient and subspace dimensions to change. PP logarithms use ONB logarithms to support back-propagation. Logarithmic maps near the cut locus require numerical treatment, so the formulas on this compact manifold are not unconditional global single-valued coordinates.
A Worked Example¶
Consider the paper's Radar SPDNN-LEM configuration: multiple \(20\times20\) covariance matrices form the input, the convolution uses a global receptive field, and the output is one \(8\times8\) SPD matrix. The following explains that configuration's mechanism, not an additional measurement.
First, treat the covariance matrices as a product-manifold input rather than averaging them. Each target tangent direction has learnable geometric responses for the input components; under LEM, these simplify to inner products and offsets in matrix-logarithm coordinates.
Second, sum component responses for each output direction. An \(8\times8\) symmetric matrix has \(36\) independent coordinates, requiring \(36\) output tangent coefficients; assemble them through the metric-orthonormal basis into a symmetric matrix and exponentiate it to obtain a positive definite output.
Third, an SPD multinomial logistic regression classifier reads this output and produces class scores. Supervision comes from Radar's three classes, not from directly treating the point-to-hyperplane pseudo-distance as a training loss; the classification objective back-propagates to the kernel's tangent parameters and displacements.
Loss & Training¶
The contribution is a transformation layer, not a new loss function. Hyperbolic experiments follow a link-prediction setup, with a two-layer HNN encoder mapping the input dimension to 16 and then 16 to 16, trained with Adam at learning rate \(10^{-2}\); weight decay and dropout are tuned separately for different methods. Cora uses no activation for any method, while the other three graphs use tangent-space ReLU. The appendix also specifies an existing bias translation after HFC, so the resource table does not measure an isolated layer.
SPDNN uses one convolutional layer followed by multinomial logistic regression. For LEM, PEM, and LCM, the classifier matches the convolution's metric; AIM and BWM use a LEM classifier for efficiency, so these models do not use one metric throughout convolution and classification. Training uses batch size 30, at most 150 epochs with early stopping, and primarily AMSGrad, with SGD for some configurations; matrix powers may also be applied before convolution to alter the latent geometry.
Grassmann networks likewise use one transformation followed by classification. The proposed models use AMSGrad at learning rate \(5\times10^{-3}\), batch size 30, and 150 training epochs. Baselines use SGD with a different learning rate, so these results are not a strict ablation that changes only the transformation while freezing all optimization settings.
Source notation requires caution: Section 3.1 first writes a Euclidean fully connected layer with a positive bias but then uses a negative bias in its coordinate expression; Section 5.1 writes an outer Log for LorentzTan, whereas the related tangent-transformation table writes Exp. This note does not silently harmonize them into an author-confirmed implementation formula; reproduction should check the code and final version.
Key Experimental Results¶
Main Results¶
The following selects two-layer HNN link-prediction results from Table 4. The metric is test AUC (%), reported as five-fold averages. This compares transformation layers within one backbone, not results across different tasks.
| Transformation | Disease | Airport | Pubmed | Cora |
|---|---|---|---|---|
| Möbius | 76.73 ± 4.86 | 93.26 ± 0.43 | 94.95 ± 0.06 | 90.75 ± 0.47 |
| Einstein | 77.34 ± 2.56 | 92.72 ± 0.07 | 94.99 ± 0.13 | 89.73 ± 0.21 |
| LFC | 78.00 ± 0.60 | 92.63 ± 0.27 | 94.22 ± 0.11 | 91.74 ± 0.12 |
| Poincaré FC | 79.19 ± 2.05 | 94.21 ± 0.43 | 94.30 ± 0.19 | 87.16 ± 0.95 |
| HFC-P | 80.66 ± 1.35 | 94.13 ± 0.47 | 94.77 ± 0.28 | 90.76 ± 0.57 |
| HFC-K | 80.58 ± 1.59 | 94.32 ± 0.27 | 94.61 ± 0.27 | 89.94 ± 0.49 |
| HFC-H | 82.93 ± 0.89 | 95.18 ± 0.18 | 93.99 ± 0.25 | 92.20 ± 0.25 |
HFC-H is best in this comparison on Disease, Airport, and Cora, but Einstein and LorentzTan both reach 94.99 on Pubmed, exceeding every HFC variant. The paper relates these outcomes to graph hyperbolicity; they suggest geometry-dependent suitability rather than proving that a richer Riemannian transformation always beats tangent-space models.
The following selects SPD classification accuracy (%) from Table 6, retaining the same-architecture GyroSPD++-LEM baseline and three proposed metrics. NTU60 evaluates only mutual actions, not the full standard NTU60 task.
| Model | Radar | HDM05 | FPHA | NTU60 mutual actions |
|---|---|---|---|---|
| GyroSPD++-LEM (reproduced in this paper) | 98.08 ± 0.26 | 77.63 ± 1.01 | 88.23 ± 0.62 | 85.48 ± 1.10 |
| SPDNN-LEM | 98.27 ± 0.48 | 81.16 ± 0.93 | 91.83 ± 0.41 | 86.72 ± 0.14 |
| SPDNN-PEM | 98.43 ± 0.44 | 78.77 ± 0.45 | 90.33 ± 0.37 | 82.61 ± 0.37 |
| SPDNN-BWM | 98.72 ± 0.14 | 72.49 ± 2.02 | 87.80 ± 0.56 | 82.64 ± 0.35 |
LEM is stronger on the last three datasets, while BWM is stronger on Radar. Reproduction boundaries matter: Appendix I.2.4 specifies 122 classes and a per-class 50/50 split for HDM05, and cross-view for NTU60, differing from some source papers. Even on FPHA's identical official split, reproduced GyroSPD++-LEM scores 88.23 versus 94.72 in its source paper. Winning the internal comparison therefore does not establish superiority over those published source results.
Ablation Study¶
The following uses Radar five-fold results from Table 8 to examine GrNN subspace and ambient dimensions. It is a dimensional analysis, not a module-removal ablation.
| Config | Subspace dimension | Ambient dimension | Accuracy (%) |
|---|---|---|---|
| GrNet | 4 | 20→16 | 90.48 ± 0.76 |
| GrNN-ONB | 4→4 | 20→16 | 93.92 ± 0.74 |
| GrNN-ONB | 4→4 | 20→20 | 92.83 ± 0.66 |
| GrNN-ONB | 4→6 | 20→16 | 95.23 ± 0.96 |
| GrNN-ONB | 4→8 | 20→16 | 94.77 ± 0.81 |
| GrNN-PP | 4→4 | 20→16 | 94.35 ± 0.42 |
| GrNN-PP | 4→6 | 20→16 | 94.51 ± 0.53 |
Changing ONB from retaining subspace dimension 4 to output dimension 6 increases accuracy from 93.92 to 95.23, a gain of 1.31 percentage points; increasing it further to 8 does not improve the result. Under PP, the same dimensional change increases 94.35 to only 94.51. Flexible subspace dimensions create a useful design space, but gains are not guaranteed to grow monotonically with dimension.
Key Findings¶
- HNN resources measure full-model registered parameter elements and peak allocated memory in MiB. On Cora, NestFC uses 2,076,931 parameter elements and 1234.25 MiB, versus HFC-H's 23,248 and 326.61 MiB; these should not be described as effective degrees of freedom or inference-latency measurements.
- Compact parameters do not ensure the lowest memory in every setting: HFC-H uses 89.82 MiB on Disease versus 77.20 MiB for Möbius. Geometric intermediates still incur costs.
- In the appendix's local-convolution Radar setting
[40,3,3], SPDConvNet-LEM reaches 97.68 ± 0.32 at 1.58 seconds/epoch, versus ManifoldNet's 77.89 ± 1.56 at 13.12 seconds/epoch. However, tangent ReLU was removed from MVC-Net because of numerical instability, so activations are not identical across methods. - SPD training speed depends strongly on the metric: SPDNN-LCM takes 0.65 seconds/epoch on Radar, while SPDNN-BWM takes 6.07 seconds/epoch. The appendix also states that some traditional SPD baselines run on CPU for HDM05 and FPHA while other configurations run on an A6000, preventing all cross-device timing differences from being attributed to the layer itself.
Highlights & Insights¶
- Replacing “matrix multiplication” with “hyperplane responses” as the essence of a fully connected layer allows input and output geometries to be defined separately. The transferable approach is to identify the geometric constraints behind Euclidean operations rather than inventing separate algebra for each space.
- The pseudo-distance is a solvability choice as well as an efficiency choice. It turns potentially incompatible true-distance constraints into output tangent coefficients, making it essential to the framework rather than an incidental approximation.
- Product manifolds provide a clear convolution interface: geometry is handled within components, and components combine through the summed metric. Channels, matrix sizes, and subspace dimensions consequently become distinct architectural variables.
Limitations & Future Work¶
- The authors restrict the framework to tractable Riemannian operators; it does not directly apply to unknown manifolds or spaces lacking manageable Exp and Log. Numerical approximations are a future direction, not an already validated universal solution.
- Equality between pseudo-distance and true geodesic distance is conditional, while Grassmann cut loci and the locally defined domains of PEM and BWM require treatment. Numerical regularization can alter the ideal geometric map, motivating reports of clipping frequency and gradient stability.
- Compact parameterization, geometry, classifier metric, matrix-power preprocessing, and optimizer choices can all affect results. Existing experiments do not fully isolate gains from multiple tangent spaces from these jointly changing factors.
- The main SPDNN is a shallow global-convolution model, and Grassmann evaluation is limited to Radar; the local-convolution extension also focuses on Radar. Stability in deeper networks, long time series, and broader tasks remains to be established.
- Future comparisons could align isometric representations, optimization budgets, and initializations across instantiations to distinguish representation effects, numerical conditioning, and parameterization, rather than attributing every empirical difference to a superior geometry.
Related Work & Insights¶
- vs Möbius / Einstein tangent-space layers: These primarily expand at a fixed origin before Euclidean transformation, whereas the paper learns a different input base point for each output. Pubmed shows that richer local responses are not always necessary.
- vs Poincaré FC / GyroSPD++: The general definition can include existing constructions through appropriate distance choices; HFC-P derived using the paper's pseudo-distance is not identical to the previous Poincaré FC. Compact trivialization is also a consequential change in the SPD comparison, so differences should not be reduced to “using Riemannian operators.”
- vs ManifoldNet: Fréchet-mean convolution retains fixed-parameter equivariance but cannot freely change the target manifold dimension. The proposed construction obtains dimensional flexibility and parameter-covariant geometric compatibility; these are not identical guarantees.
- vs Grassmann FRMap + ReOrth: Left multiplication and QR can change ambient dimension while keeping subspace dimension fixed; constructing the target tangent space opens both dimensions. Radar dimensional analysis directly supports the usefulness of this design space.
Rating¶
- Novelty: 4/5. The unified hyperplane interpretation and product-manifold extension are valuable, while pseudo-distances and compact parameterization build on prior ideas.
- Experimental Thoroughness: 4/5. Three manifold families, local-convolution tests, and reproduction boundaries are covered, but deep-network stability and strict factor isolation remain limited.
- Writing Quality: 3/5. The central theorem is clear and the appendix is detailed, but bias signs and LorentzTan's Log/Exp notation need checking.
- Value: 4/5. A useful unified interface for manifold transformation layers, not a default guaranteed to be stronger or faster on every task.