CrossFeat: Bridging Imaging Modalities in Feature Descriptor Space¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/paulschneider01/CrossFeat
Area: Medical Imaging
Keywords: multimodal matching, keypoint descriptor, geometry-appearance disentanglement, descriptor space crossing, multimodal image registration
TL;DR¶
CrossFeat introduces a lightweight framework to learn a non-linear crossing operator in feature descriptor space by disentangling monomodal descriptors into an invariant geometry code and a modality-specific variational appearance code, enabling off-the-shelf descriptors to perform cross-modal matching across medical, driving, and satellite domains without retraining.
Background & Motivation¶
Local keypoint detection and description form the foundational backbone for image registration, 3D reconstruction, and pose estimation. Over decades of refinement, descriptors such as SIFT, SuperPoint, DISK, and ALIKED have achieved remarkable discriminative power and robustness against geometric transformations, viewpoint changes, and illumination variations. However, nearly all classical and modern deep descriptors are built on the monomodal assumption, where correspondence matching relies on shared photometric properties and intensity distributions. In multimodal imaging scenarios—such as magnetic resonance imaging and ultrasound (MRI–Ultrasound) in medicine, visible-light and event cameras (RGB–Event) in robotics, and optical and synthetic aperture radar (RGB–SAR) in Earth observation—the underlying physical sensing principles differ drastically. Consequently, identical anatomical structures or physical surfaces exhibit completely disparate appearances, causing standard cosine similarity across monomodal descriptors to break down completely.
Prior efforts to bridge this multimodal gap have pursued two main paradigms: either training dedicated descriptors from scratch for specific modality pairs, or developing massive cross-modal foundation matchers (such as MatchAnything or MINIMA-RoMA). The former requires collecting paired data and retraining models whenever a new sensor modality is introduced, resulting in combinatorial overhead. The latter, while achieving broad generalization, relies on large correlation transformers with tens of millions of parameters, incurring multi-second inference latencies per image pair and preventing easy integration into established sparse pipelines such as surgical navigation or real-time visual odometry. This reveals a fundamental tension: established monomodal descriptors already capture rich, high-quality geometric structures through extensive pretraining and design; discarding them and retraining from scratch whenever new modality combinations arise is highly inefficient.
The core insight of CrossFeat is to reframe multimodal matching as a feature transformation problem entirely within descriptor space, rather than performing pixel-level image translation or end-to-end network retraining. The cross-modality gap is fundamentally driven by modality-specific appearance (intensity profiles, contrast inversions, speckle, and noise), whereas underlying geometric landmarks (edges, corners, and surface curvatures) remain physically congruent. Core idea: disentangle an off-the-shelf monomodal descriptor into a modality-invariant geometry representation and a modality-specific variational appearance representation, freeze the original descriptor extractor, and apply a lightweight modality-conditioned operator solely to the appearance code, achieving plug-and-play cross-modal matching with zero descriptor retraining.
Method¶
Overall Architecture¶
CrossFeat consists of four coordinated components: a shared descriptor encoder \(E_\phi\), a variational appearance reparameterization head, a modality-conditioned appearance crosser \(C_\psi\), and a modality-adaptive decoder \(G_\theta\). During inference, given a monomodal descriptor \(\mathbf{d}_a \in \mathbb{S}^{D-1}\) extracted from source modality \(m_a\), encoder \(E_\phi\) projects it into a deterministic geometry code \(\mathbf{z}_{geom}\) and a mean appearance code \(\boldsymbol{\mu}_{app}\). The geometry code \(\mathbf{z}_{geom}\) bypasses the crosser entirely to preserve structural properties intact. The appearance code is fed into crosser \(C_\psi\), where learned modality embeddings for \((m_a, m_b)\) guide affine modulation and residual refinement to synthesize target appearance code \(\hat{\mathbf{z}}_{app}\). Finally, decoder \(G_\theta\) recombines \(\mathbf{z}_{geom}\) and \(\hat{\mathbf{z}}_{app}\) under the target modality condition to reconstruct a crossed descriptor \(\hat{\mathbf{d}}\) projected back onto the unit hypersphere \(\mathbb{S}^{D-1}\), enabling standard nearest-neighbor (NN) matching against target descriptors \(\mathbf{d}_b\).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
InA["Source Descriptor d_a ∈ S^(D-1)<br/>Extracted from modality m_a"] --> Enc["Shared Backbone Encoder E_ϕ<br/>Projects to geometry & appearance heads"]
Enc --> GeomFlow["Geometry Code z_geom<br/>Deterministic bypass preserving spatial structure"]
Enc --> AppFlow["Appearance Code z_app<br/>Variational Gaussian distribution with prior alignment"]
AppFlow --> CrossMod["Modality-Conditioned Crosser C_ψ<br/>FiLM modulation & residual transformation"]
GeomFlow --> Dec["Modality-Adaptive Decoder G_θ<br/>Fuses geometry & crossed target appearance"]
CrossMod --> Dec
Dec --> NormOut["L2 Hypersphere Normalization<br/>Yields crossed descriptor d_hat ∈ S^(D-1)"]
NormOut --> MatchStage["Nearest-Neighbor Matching / Optional TTA<br/>Matches target descriptor d_b in modality m_b"]
Key Designs¶
1. Geometry–Appearance Disentangled Latent Space: Deterministic Structural Bypass and Variational Regularization Directly training a monolithic network to map the entire descriptor from one modality to another risks corrupting the delicate geometric coordinates or undertransforming the appearance statistics. CrossFeat addresses this by architecturally decoupling the latent representation via a shared encoder \(E_\phi: \mathbb{S}^{D-1} \to \mathbb{R}^{d_g} \times \mathbb{R}^{d_a}\) (with \(d_g=D, d_a=D/2\)). The geometry head outputs a deterministic point estimate \(\mathbf{z}_{geom}\) that physically bypasses the crosser, structurally guaranteeing that structural and geometric information remains untouched. In contrast, the appearance head outputs Gaussian distribution parameters \(\boldsymbol{\mu}_{app}\) and \(\log \boldsymbol{\sigma}_{app}^2\). During training, stochastic sampling via the reparameterization trick \(\mathbf{z}_{app} = \boldsymbol{\mu}_{app} + \boldsymbol{\sigma}_{app} \odot \boldsymbol{\varepsilon}\) (\(\boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\)) prevents the network from memorizing deterministic instance shortcuts, forcing it to distill modality-specific contrast and intensity statistics into the latent space.
2. Modality-Conditioned Feature Crosser: Identity-Centered FiLM Modulation and Residual Projection The appearance crosser \(C_\psi: \mathbb{R}^{d_a} \times \mathcal{M} \times \mathcal{M} \to \mathbb{R}^{d_a}\) maps appearance from source modality \(m_a\) to target modality \(m_b\). To guarantee gradient propagation to the encoder and maintain training stability from step zero, the crosser integrates Feature-wise Linear Modulation (FiLM) with a residual MLP. Concatenated learned embeddings of \(m_a\) and \(m_b\) are passed through a hypernetwork to produce per-dimension affine parameters \((\boldsymbol{\gamma}, \boldsymbol{\beta})\), with \(\boldsymbol{\gamma}\) initialized to zero for identity centering: \(\mathbf{z}_{\text{mod}} = (1 + \boldsymbol{\gamma}) \odot \mathbf{z}_{app}^{(a)} + \boldsymbol{\beta}\). A residual MLP refines this code:
This formulation ensures that the crosser initializes strictly as an identity mapping, preventing erratic divergence early in training and focusing capacity on learning non-linear cross-modal shifts.
3. Modality-Adaptive Decoding and Hyperspherical Normalization: Target Statistic Alignment Decoder \(G_\theta: \mathbb{R}^{d_g} \times \mathbb{R}^{d_a} \to \mathbb{S}^{D-1}\) recombines the unaltered source geometry code \(\mathbf{z}_{geom}^{(a)}\) with the crossed appearance code \(\hat{\mathbf{z}}_{app}\). Because different imaging modalities present distinct feature energy profiles (such as varying gradient magnitude distributions across MRI sequences), the decoder also receives FiLM conditioning based on target modality \(m_b\). The resulting vector is passed through an MLP and strictly normalized by its \(\ell_2\) norm, ensuring that the reconstructed crossed descriptor \(\hat{\mathbf{d}}\) resides on \(\mathbb{S}^{D-1}\) and seamlessly matches the distribution of genuine target descriptors.
4. Dual-Anchor Keypoint Sampling and Test-Time Adaptation (TTA): Eliminating Detection Bias and Iterative Refinement Extracting training keypoints exclusively from the source modality introduces an asymmetric sampling bias toward structures that are salient only in that sensor. CrossFeat employs dual-anchor sampling: half of the keypoints in each training batch are detected in the source modality and half in the target modality, exposing the encoder to salient geometry from both sensors. For inference-time deployment, an optional lightweight test-time adaptation (TTA) module is introduced. Following an AdaIN statistical warm-up, the system performs five iterations: mutual nearest-neighbor matches are filtered by rigid RANSAC to identify inlier correspondences as pseudo-labels; closed-form ridge regression then predicts a small residual correction on crossed descriptors, which is added residually and re-normalized before the next iteration.
Loss & Training¶
The overall training loss combines six complementary terms to guarantee fidelity, structural invariance, cross-modal alignment, and hard-negative discrimination:
- Reconstruction Loss \(\mathcal{L}_{recon}\): Minimizes cosine distance between reconstructed and input descriptors in each modality, preventing information loss in the latent codes: \(\mathcal{L}_{recon} = 1 - \cos(\mathbf{d}_a, G_\theta(\mathbf{z}_{geom}^{(a)}, \mathbf{z}_{app}^{(a)}))\).
- Geometry Alignment Loss \(\mathcal{L}_{geom}\): Enforces soft alignment between co-located keypoints across modalities: \(\mathcal{L}_{geom} = \|\mathbf{z}_{geom}^{(a)} - \mathbf{z}_{geom}^{(b)}\|_2^2\).
- Crossing Loss \(\mathcal{L}_{cross}\): Supervises the decoded crossed descriptor against the true target descriptor: \(\mathcal{L}_{cross} = 1 - \cos(\hat{\mathbf{d}}, \mathbf{d}_b)\).
- Adversarial Modality Loss \(\mathcal{L}_{adv}\): Connects a linear modality classifier to \(\mathbf{z}_{geom}\) trained via a gradient reversal layer (GRL), actively stripping modality information from the geometry representation.
- Variational Regularization \(\mathcal{L}_{kl}\): Minimizes the KL divergence between the appearance posterior and a learnable modality-conditioned prior \(p(\mathbf{z}_{app} \mid m) = \mathcal{N}(\boldsymbol{\mu}_m, \boldsymbol{\sigma}_m^2)\) with a free-bits mechanism.
- Contrastive Discrimination Loss \(\mathcal{L}_{nce}\): Imposes an InfoNCE loss across in-batch cross-target pairs to penalize non-matching keypoints and boost nearest-neighbor discriminability.
Hyperparameters are fixed across all tasks: \(\lambda_r=1.0, \lambda_g=0.5, \lambda_c=1.0, \lambda_{adv}=0.03, \lambda_{kl}=0.01, \lambda_{nce}=0.5\). The model contains only ~0.5–0.6M parameters, trained using AdamW (learning rate \(3 \times 10^{-4}\), batch size 2048, cosine annealing) for up to 500 epochs in 1–3 hours on a single NVIDIA A100 GPU.
Key Experimental Results¶
Main Results¶
CrossFeat was evaluated across three challenging multimodal domains: Medical imaging (ReMIND training; BRATS and RESECT testing), Autonomous Driving (EventScape training; DELIVER testing), and Satellite Earth Observation (WHU-OPT-SAR training; QXS-SAROPT testing).
The table below summarizes sparse matching performance (capped at 500 matches per image) across Precision (P), Recall (R), average match count (#M), Success Rate (SR@X pixels), and Area Under the Curve (AUC@X pixels):
| Task | Method | P | R | #M | SR@1 (%) | SR@3 (%) | AUC@1 | AUC@3 |
|---|---|---|---|---|---|---|---|---|
| Medical | SP + LG | 0.73 | 0.19 | 74 | 23.8 | 67.5 | 0.26 | 0.45 |
| Medical | DISK + LG | 0.28 | 0.05 | 43 | 2.5 | 8.8 | 0.05 | 0.12 |
| Medical | ALIKED + LG | 0.56 | 0.12 | 39 | 6.2 | 27.5 | 0.18 | 0.33 |
| Medical | SP + MINIMA-LG | 0.80 | 0.26 | 91 | 37.5 | 67.5 | 0.26 | 0.49 |
| Medical | SIFT + NN (Baseline) | 0.65 | 0.13 | 34 | 11.2 | 15.0 | 0.64 | 0.64 |
| Medical | Cross(SIFT) + NN (Ours) | 0.95 | 0.25 | 62 | 47.0 | 68.9 | 0.95 | 0.95 |
| Medical | Cross(SIFT) + NN + TTA (Ours) | 0.98 | 0.44 | 107 | 79.2 | 97.4 | 0.96 | 0.96 |
| Driving | SP + MINIMA-LG | 0.40 | 0.11 | 59 | 0.4 | 17.0 | 0.08 | 0.20 |
| Driving | SIFT + NN (Baseline) | 0.41 | 0.06 | 34 | 9.4 | 12.7 | 0.39 | 0.39 |
| Driving | Cross(SIFT) + NN (Ours) | 0.96 | 0.05 | 22 | 73.5 | 74.9 | 0.92 | 0.92 |
| Driving | Cross(SIFT) + NN + TTA (Ours) | 0.97 | 0.23 | 110 | 85.3 | 88.6 | 0.94 | 0.94 |
| Satellite | SP + MINIMA-LG | 0.01 | 0.00 | 9 | 0.0 | 0.0 | 0.00 | 0.00 |
| Satellite | SIFT + NN (Baseline) | 0.23 | 0.02 | 26 | 0.0 | 0.0 | 0.23 | 0.23 |
| Satellite | Cross(SIFT) + NN (Ours) | 0.87 | 0.02 | 9 | 34.0 | 38.0 | 0.87 | 0.87 |
| Satellite | Cross(SIFT) + NN + TTA (Ours) | 0.74 | 0.05 | 20 | 18.0 | 23.0 | 0.73 | 0.73 |
When compared against dense matching models (MatchAnything, MINIMA-LoFTR, MINIMA-RoMA), CrossFeat demonstrates strong competitiveness: in the Medical domain, CrossFeat + TTA achieves 79.2% SR@1 and 0.96 AUC@1 (outperforming MINIMA-RoMA's 66.2% SR@1 and 0.46 AUC@1); in Driving, CrossFeat achieves 73.5% (and 85.3% with TTA), whereas MINIMA-RoMA only reaches 3.3%; in the challenging optical-SAR Satellite task where dense matchers collapse (0.0% SR@1), CrossFeat achieves 34.0% SR@1 and 0.87 AUC@1. Furthermore, CrossFeat runs in 0.5–1.0s per image, roughly 5–10× faster than dense matchers (~5.0s).
Ablation Study¶
Ablation experiments on the held-out ReMIND medical dataset across six modality pairs and five random seeds evaluate each component's contribution:
| Configuration | P (%) | R (%) | #M | SR@1 (%) | SR@3 (%) | AUC@1 | AUC@3 |
|---|---|---|---|---|---|---|---|
| Full model | 91.7 ± 3.4 | 6.2 ± 0.3 | 91.2 ± 2.1 | 52.9 ± 4.8 | 65.1 ± 10.3 | 89.3 ± 6.0 | 89.3 ± 6.0 |
| Architecture: w/o disentanglement | 84.8 ± 7.6 | 5.9 ± 0.2 | 86.8 ± 2.1 | 39.6 ± 7.2 | 60.0 ± 7.0 | 80.8 ± 4.5 | 80.8 ± 4.5 |
| Architecture: w/o variational | 86.3 ± 5.8 | 6.4 ± 0.3 | 93.2 ± 4.7 | 48.3 ± 6.4 | 65.0 ± 4.4 | 84.2 ± 3.8 | 84.2 ± 3.8 |
| Architecture: w/o adversarial | 84.1 ± 7.9 | 5.8 ± 0.2 | 84.9 ± 4.6 | 40.6 ± 10.3 | 58.7 ± 11.7 | 84.0 ± 7.9 | 84.0 ± 7.9 |
| Architecture: w/o cond. decoder | 88.4 ± 9.8 | 6.2 ± 0.4 | 88.7 ± 5.2 | 40.9 ± 1.7 | 60.7 ± 6.1 | 82.4 ± 4.9 | 82.4 ± 4.9 |
| Losses: w/o \(\mathcal{L}_{nce}\) | 75.0 ± 5.5 | 3.2 ± 0.1 | 43.9 ± 1.6 | 15.1 ± 5.9 | 28.1 ± 7.2 | 73.4 ± 5.1 | 73.9 ± 5.5 |
| Losses: w/o \(\mathcal{L}_{geom}\) | 84.4 ± 8.8 | 6.0 ± 0.2 | 86.6 ± 2.8 | 48.9 ± 10.6 | 65.3 ± 6.2 | 82.2 ± 6.9 | 82.2 ± 6.9 |
| Losses: w/o \(\mathcal{L}_{recon}\) | 86.3 ± 6.2 | 5.5 ± 0.3 | 77.3 ± 3.2 | 44.4 ± 5.3 | 57.9 ± 5.7 | 84.3 ± 3.7 | 84.3 ± 3.7 |
| Losses: w/o latent reg. | 82.8 ± 8.2 | 6.2 ± 0.2 | 88.2 ± 4.4 | 41.0 ± 7.2 | 60.9 ± 5.9 | 78.7 ± 4.1 | 78.7 ± 4.1 |
| Training: w/o KP sampling | 69.5 ± 2.5 | 1.0 ± 0.1 | 11.2 ± 0.8 | 45.4 ± 6.6 | 60.3 ± 4.2 | 61.5 ± 2.1 | 61.7 ± 2.1 |
Key Findings¶
- Contrastive loss \(\mathcal{L}_{nce}\) is decisive: Removing InfoNCE collapses SR@1 from 52.9% down to 15.1% (-37.8 pp) and cuts match count in half. Minimizing distance to positive targets without pushing away hard negatives fails to create the margin required for nearest-neighbor retrieval.
- Disentanglement is essential: Treating the descriptor as an undivided monolithic vector causes a 13.3 pp drop in SR@1, confirming that isolating geometry from the crosser is critical for preserving spatial coherence.
- Dual-anchor sampling avoids detection blindness: Without balanced keypoint extraction across both source and target images, match count plummets from 91.2 to 11.2 and precision drops by 22.2 pp.
- Universal descriptor gains: CrossFeat substantially elevates classic SIFT and consistently enhances modern learned descriptors including SuperPoint, ALIKED, and DISK, demonstrating wide applicability across feature designs.
Highlights & Insights¶
- Architectural bypass avoids hallucination: Routing \(\mathbf{z}_{geom}\) completely around the crossing operator guarantees that structural coordinates are never distorted during appearance mapping, sidestepping mode collapse and false geometric hallucination.
- Extreme parameter and sample efficiency: Requiring only ~0.5M parameters and 1–3 hours of training per domain, CrossFeat revitalizes off-the-shelf descriptors for multimodal setups without requiring costly transformer pretraining.
- High inlier purity over raw density: CrossFeat prioritizes highly reliable matches over dense correspondence, producing a pronounced rightward shift in positive cosine similarity and higher sensitivity index \(d'\), facilitating clean RANSAC registration.
Limitations & Future Work¶
- Explicit modality conditioning required: The crossing operator currently requires explicit modality labels \((m_a, m_b)\) at test time, limiting applications in uncalibrated or continuously varying multispectral streams.
- TTA sensitivity in sparse regimes: In extreme noise settings (such as optical-SAR), inaccurate initial RANSAC inliers can lead closed-form ridge regression astray, degrading precision (SR@1 drops from 34.0% to 18.0%).
- Future directions: Developing unsupervised modality gap estimators from image statistics to dynamically guide crossing without manual labels, and expanding the formulation to 3D volumetric representations (such as CT-MRI).
Related Work & Insights¶
- vs MINIMA / MatchAnything: MINIMA and MatchAnything train massive dense matchers on synthetic cross-modal data, incurring high computational cost (~5s per pair). CrossFeat operates downstream in descriptor space with 100× fewer parameters and 5–10× faster inference, achieving higher accuracy in severe sensing discrepancies like optical-SAR.
- vs Steerers / Affine Steerers: Steerers apply Lie group operators to descriptors to enforce rotational and affine equivariance within a single modality. CrossFeat extends descriptor-space manipulation from monomodal geometric equivariance to cross-modal physical appearance translation.
- vs XoFTR: XoFTR tailors a coarse-to-fine transformer specifically for visible–infrared pairs. CrossFeat provides a modular, descriptor-agnostic crossing framework that generalizes across arbitrary modality combinations.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Reframing cross-modal matching into geometry-appearance disentanglement and descriptor-space crossing is elegant and highly effective]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive quantitative evaluation across medical, driving, and satellite domains with rigorous ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear mathematical formulation, disciplined narrative structure, and compelling empirical visualizations]
- Value: ⭐⭐⭐⭐⭐ [High practical impact for multimodal image registration, surgical guidance, and cross-sensor tracking with minimal compute overhead]