Skip to content

Two-Parameter Flow Map Learning for Continuous-Time Diffeomorphic Image Registration

Conference: ECCV 2026
Paper: ECCV Official
Project: https://mattkia.github.io/TPFMDIR/
Area: Medical Imaging
Keywords: Diffeomorphic Image Registration, Non-autonomous ODEs, Two-Parameter Flow Maps, Cocycle Consistency, Continuous-Time Dynamical Systems

TL;DR

Addressing the reliance of non-autonomous diffeomorphic registration on explicit numerical integration and heavy time discretization, TPFM-DIR proposes directly learning the continuous-time solution operator of non-autonomous ODEs as a two-parameter flow map; by regularizing with anchored cocycle consistency, it eliminates velocity integration during training and achieves high-accuracy, topology-preserving alignment via minimal progressive compositions at inference.

Background & Motivation

In medical image computing and clinical analysis, diffeomorphic image registration (DIR) serves as the foundational mechanism for establishing anatomically plausible correspondences across different subjects or longitudinal intra-subject temporal scans. In stark contrast to unconstrained deformable image registration pipelines that prioritize raw voxel intensity similarity at the expense of topological foldings, singularities, and unphysical tearing, diffeomorphic methods explicitly constrain the spatial mapping to be a smooth, invertible bijection with strictly positive Jacobian determinants. While unconstrained deformable architectures leveraging advanced Vision Transformers and multi-scale correlation volumes have pushed alignment accuracy to remarkable levels, diffeomorphic models have historically struggled with an inherent trade-off between strict topological regularity and empirical registration overlap.

A prominent class of modern neural diffeomorphic registration approaches formalizes spatial transformations as continuous flows generated by ordinary differential equations (ODEs). Autonomous ODE formulations assume a stationary velocity field, enabling computationally cheap deformation recovery via scaling-and-squaring algorithms; however, this time-invariance assumption severely restricts their expressiveness when modeling complex large deformations and highly dynamic physiological processes. Non-autonomous ODE formulations lift this constraint by introducing time-dependent velocity fields, yet state-of-the-art implementations (such as LDDMM, NODEO, and R2Net) must rely on explicit numerical integration solvers (e.g., Euler or Runge-Kutta schemes) during both forward training and backward gradient propagation. This coupling ties deformation fidelity to discrete solver step sizes and truncation errors, causes significant memory bottlenecks, and relies on heuristic surrogate regularizers that fail to guarantee continuous-time group flow structures.

The key insight of this paper is to shift the learning target from estimating the underlying instantaneous velocity fields to directly learning the continuous-time solution operator—the two-parameter flow map—in function space. Core idea: directly parameterize the solution operator of a non-autonomous ODE as a two-parameter flow map conditioned on start and end times, and enforce an anchored cocycle consistency regularization derived from the fundamental algebra of time-varying dynamical systems, thereby completely removing numerical integration during training while recovering topology-preserving continuous diffeomorphisms via minimal compositions at inference.

Method

Overall Architecture

TPFM-DIR bypasses the conventional paradigm of predicting velocity fields followed by discrete numerical integration, formulating the registration network as an end-to-end two-parameter mapping operator. Given a pair of moving image \(I_m\) and fixed image \(I_f\) along with sampled temporal coordinates \(s, t \in [0, 1]\) (\(s \le t\)), the model directly predicts the continuous deformation \(\phi_{s,t}\) that transports spatial coordinates from time \(s\) to time \(t\). During training, the framework combines temporal similarity supervision with an anchored cocycle consistency penalty to symmetrically guide continuous trajectories without invoking any numerical ODE solver. During inference, the continuous trajectory over the unit time interval is partitioned into \(N=4\) subintervals, whose local flow predictions are composed in sequence to construct global, folding-free diffeomorphic transformations.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Moving/Fixed Images and Time Scalars<br/>Im, If, s, t ~ Uni[0, 1]"] --> B["Two-Parameter Flow Map Network<br/>parameterized as phi_s,t = Id + (t-s) f_theta"]
    B --> C["Anchored Cocycle Consistency Regularization<br/>enforces algebraic cocycle flow law at endpoints"]
    C --> D["Time-Dependent Symmetric Similarity Loss<br/>alternates optimization along forward and backward paths"]
    D --> E["Inference-Time Progressive Compositions<br/>partitions [0, 1] into N=4 subintervals and composes"]
    E --> F["Output Diffeomorphic Deformation Field<br/>high Dice overlap with near-zero Jacobian folding"]

Key Designs

1. Direct Two-Parameter Flow Map Parameterization: Decoupling Flow Structure from Numerical Solvers To overcome the computational overhead and discretization sensitivity of classical non-autonomous integration schemes, TPFM-DIR models the solution operator of the non-autonomous ODE \(\dot{\phi}_t = v(t, \phi_t)\) directly as a two-parameter transformation family: $\(\phi_{s,t} = \text{Id} + (t - s) f_\theta(s, t; I_m, I_f)\)$ where \(f_\theta\) is a deep network operating in a shared coordinate space, conditioned on moving and fixed volumes as well as sinusoidal temporal positional embeddings of \(s\) and \(t\). This formulation analytically guarantees \(\phi_{s,s} = \text{Id}\) for all \(s\), ensuring that all spatial transformations are defined over a shared reference frame suitable for valid spatial composition. Moreover, an immediate theoretical benefit is that the underlying instantaneous velocity field can be extracted analytically in closed form without finite-difference approximations: \(v(t, \cdot) = f_\theta(t, t; I_m, I_f)\). This enables full mathematical transparency and physical interpretability of continuous-time velocities while completely bypassing ODE solver loops during training.

2. Anchored Cocycle Consistency Regularization: Replacing Handcrafted Penalties with Algebraic Group Structure Satisfying the identity condition alone does not ensure that the parameterized map constitutes a physically valid continuous-time flow. Any valid solution of a non-autonomous ODE must satisfy the cocycle property: \(\phi_{r,t} \circ \phi_{s,r} = \phi_{s,t}\) for all intermediate times \(s \le r \le t\). Enforcing this constraint across arbitrary triplets \((s, r, t)\) would require nested multi-step spatial warpings per iteration, which is computationally prohibitive. TPFM-DIR introduces an anchored cocycle regularization with fixed temporal boundaries (\(0\) for the moving image and \(1\) for the fixed image): $\(\phi_{s,t} \circ \phi_{0,s} = \phi_{0,t}, \quad \phi_{t,s} \circ \phi_{1,t} = \phi_{1,s}\)$ The authors formally prove via Theorem 1 that, under standard smoothness assumptions, satisfying the identity condition and these anchored cocycle constraints is both necessary and sufficient for \(\phi_{s,t}\) to be a two-parameter diffeomorphism solving the non-autonomous ODE. Minimizing the mean squared error of these compositions serves as an intrinsic geometric regularizer, eliminating the need for delicate, heuristic smoothness penalties (such as bending energy, displacement diffusion, or higher-order Jacobian penalties) typically required by prior learning-based DIR models.

3. Progressive Inference Compositions: Mitigating Uniform Sampling Bias and Folding During training, independent uniform sampling of \(s, t \sim \text{Uni}[0, 1]\) (with \(s \le t\)) naturally produces a triangular distribution over the interval length \(|t - s|\), heavily biasing the training distribution toward small time steps. Consequently, large transitions are relatively underrepresented during optimization, meaning that directly predicting the global mapping \(\phi_{0,1}\) in a single evaluation can exhibit mild deviations from ideal continuous flow behavior. To exploit the network's superior accuracy and topological stability over shorter intervals, TPFM-DIR partitions \([0, 1]\) into \(N\) uniform subintervals at inference time and sequentially composes the predicted local flow maps: $\(\phi_{0,1} = \phi_{t_N, 1} \circ \dots \circ \phi_{t_1, t_2} \circ \phi_{0, t_1}\)$ Because each local segment operates well within the well-conditioned, small-deformation regime of the trained network, this composition progressively accumulates large geometric displacements without triggering spatial singularities. Empirical evaluations verify that as few as \(N=4\) compositions saturate topological regularity, reducing the percentage of negative Jacobian voxels to near-zero.

Loss & Training

The entire framework is trained end-to-end under an alternating symmetric optimization protocol. In the forward-backward temporal trajectories, the image similarity is supervised using the local Normalized Cross-Correlation (NCC) loss with a window size of 7, aligning the forward-deformed moving image with the backward-deformed fixed image at time \(t\): $\(\mathcal{L}_{\text{sim}}^{s,t} = - \text{NCC}(I_m \circ \phi_{0,t}, I_f \circ \phi_{1,t})\)$ To avoid doubling GPU memory overhead from simultaneous forward and backward composite computations, the training loop alternates which branch carries the composite transformation at each epoch. The overall training objective is formulated as: $\(\mathcal{L}^{s,t} = \mathcal{L}_{\text{sim}}^{s,t} + \lambda \mathcal{L}_{\text{cocycle}}^{s,t}\)$ where the regularization weight is set to \(\lambda = 10\). The network is optimized using AdamW with a learning rate of \(10^{-4}\) and a batch size of 1 on a single NVIDIA RTX 3090 GPU (24GB VRAM), running for 200 epochs on most benchmarks (100 epochs on LungCT).

Key Experimental Results

Main Results

TPFM-DIR was extensively evaluated across nine diverse public benchmarks covering 2D and 3D modalities, including brain MRI (OASIS, IXI, Mindboggle101, LPBA40, CANDI), thoracic and abdominal CT (LungCT, AbdomenCT), and cardiac sequences (ACDC MRI, CAMUS ultrasound). Below is a consolidated quantitative comparison on representative 3D brain MRI (OASIS, IXI) and large-motion thoracic CT (LungCT):

Dataset Category Method Dice (%) ↑ |J|<0% (%) ↓ HD95 (mm) ↓ TRE (mm) / SDLogJ ↓
OASIS Classical Non-Auto DIR LDDMM 76.59 ± 2.42 0.0064 ± 0.0051 3.89 ± 0.93 0.012 ± 0.009 (SDLogJ)
OASIS Autonomous / Surrogate DIR CycleMorph 81.93 ± 2.14 0.0211 ± 0.0091 2.36 ± 0.81 0.040 ± 0.026 (SDLogJ)
OASIS Autonomous / Surrogate DIR GradICON 83.74 ± 1.42 0.0039 ± 0.0012 2.09 ± 0.36 0.011 ± 0.006 (SDLogJ)
OASIS Autonomous / Surrogate DIR TransMorph-diff 83.51 ± 1.52 0.0066 ± 0.0073 2.35 ± 0.76 0.071 ± 0.041 (SDLogJ)
OASIS Neural Non-Auto ODE DIR NODEO 81.97 ± 1.39 0.0024 ± 0.0010 2.18 ± 0.71 0.008 ± 0.005 (SDLogJ)
OASIS Unconstrained Deformable TransMorph 84.11 ± 1.30 1.0665 ± 0.5631 2.26 ± 0.68 0.821 ± 0.252 (SDLogJ)
OASIS Unconstrained Deformable HViT 85.07 ± 1.05 0.4812 ± 0.0114 1.92 ± 0.46 0.631 ± 0.412 (SDLogJ)
OASIS Ours TPFM-DIR 87.08 ± 1.07 0.0019 ± 0.0007 1.78 ± 0.44 0.006 ± 0.002 (SDLogJ)
IXI Autonomous / Surrogate DIR GradICON 76.43 ± 1.28 0.0018 ± 0.0022 3.19 ± 0.55 0.015 ± 0.006 (SDLogJ)
IXI Unconstrained Deformable CorrMLP 77.54 ± 1.68 0.3675 ± 0.2168 3.15 ± 0.39 0.547 ± 0.231 (SDLogJ)
IXI Unconstrained Deformable HViT 80.67 ± 1.67 0.5933 ± 0.1028 2.98 ± 0.46 0.679 ± 0.290 (SDLogJ)
IXI Ours TPFM-DIR 82.52 ± 0.98 0.0015 ± 0.0006 2.95 ± 0.68 0.007 ± 0.003 (SDLogJ)
LungCT Classical Non-Auto DIR LDDMM - 0.0033 ± 0.0015 - 3.09 ± 0.26 (TRE)
LungCT Autonomous / Surrogate DIR GradICON - 0.0009 ± 0.0004 - 2.64 ± 0.17 (TRE)
LungCT Unconstrained Deformable CorrMLP - 0.2673 ± 0.0711 - 2.48 ± 0.16 (TRE)
LungCT Ours TPFM-DIR - 0.0 ± 0.0 - 2.19 ± 0.14 (TRE)

Note: TRE denotes Target Registration Error in mm (lower is better); \(|J|<0\%\) measures the percentage of voxels with non-positive Jacobian determinants (lower is better); SDLogJ represents the standard deviation of the log-Jacobian determinant. On LungCT, TPFM-DIR also achieves 71.18% SSIM with strictly zero folding (\(0.0\%\)).

Ablation Study

To investigate the role of the cocycle consistency weight \(\lambda\), the authors conducted sensitivity ablations across three representative benchmarks (3D brain OASIS, large respiratory motion LungCT, and 2D cardiac ACDC):

Weight \(\lambda\) OASIS Dice (%) ↑ OASIS |J|<0% (%) ↓ LungCT TRE (mm) ↓ LungCT |J|<0% (%) ↓ ACDC Dice (%) ↑ ACDC |J|<0% (%) ↓
\(\lambda = 0\) (w/o cocycle reg) 79.48 3.1263 4.02 2.1821 80.19 3.0671
\(\lambda = 5\) 86.90 0.0415 3.36 0.0019 84.52 0.0039
\(\lambda = 10\) (default) 87.08 0.0019 2.19 0.0 87.42 0.0002
\(\lambda = 15\) 85.84 0.0001 2.41 0.0 87.08 0.0001
\(\lambda = 20\) 85.02 0.0 3.11 0.0 85.74 0.0

Furthermore, the framework demonstrated remarkable backbone-agnostic adaptability. When integrating the TPFM formulation into standard deformable architectures (CorrMLP and TransMorph), both models significantly outperformed their original deformable and conventional scaling-and-squaring (-diff) variants across all tested datasets:

Backbone & Setting OASIS Dice ↑ OASIS |J|<0% ↓ IXI Dice ↑ IXI |J|<0% ↓ AbdomenCT Dice ↑ AbdomenCT |J|<0% ↓
CorrMLP (Deformable) 84.66 0.4640 77.54 0.3675 50.28 0.1656
CorrMLP-diff (Scaling-and-Squaring) 83.12 0.0377 75.14 0.0782 48.73 0.0114
CorrMLP-TPFM (Ours) 86.32 0.0038 80.38 0.0016 51.47 0.0002
TransMorph (Deformable) 84.11 1.0665 77.72 1.2574 46.66 3.1310
TransMorph-diff (Scaling-and-Squaring) 83.51 0.0066 75.98 0.0018 41.41 0.0034
TransMorph-TPFM (Ours) 86.50 0.0122 78.19 0.0092 50.27 0.0083

Key Findings

  • Cocycle Regularization is Fundamental to Flow Validity: When \(\lambda = 0\), registration accuracy collapses (OASIS Dice drops from 87.08% to 79.48%) while folding rates surge by over three orders of magnitude (\(>3.1\%\)), confirming that similarity supervision alone cannot induce continuous-time group properties. Setting \(\lambda = 10\) provides an optimal balance between anatomical deformation flexibility and topological preservation.
  • Inference Composition Rapidly Saturates Topology Gains: Progressing from single-step direct evaluation (\(N=1\)) to \(N=4\) compositions reduces negative Jacobian percentages to nearly zero; increasing beyond \(N=4\) yields marginal topological gains, confirming that a small number of local transitions is sufficient.
  • Superior Runtime Efficiency over Prior Non-Autonomous Models: On 3D OASIS volumes, instance-optimization non-autonomous NODEO requires 214 seconds per inference pair, and learning-based R2Net takes 0.96 seconds (1.53 s/pair during training). In contrast, TPFM-DIR achieves 0.38 seconds per inference pair and 1.31 s/pair during training with only 2.7 GB GPU memory footprint, enabling practical clinical deployment.

Highlights & Insights

  • Operator-Level Flow Modeling: By shifting the paradigm from instantaneous velocity field integration to learning continuous solution operators directly, the method completely decouples flow expressiveness from numerical ODE integration errors and backward pass graphs.
  • Anchored Cocycle Reduction: The reduction of combinatorial full-trajectory cocycle constraints to boundary-anchored compositions provides a mathematically grounded yet computationally frugal regularization that guarantees non-autonomous ODE solutions.
  • Backbone-Agnostic Plug-and-Play: Standard registration architectures can be converted into topologically certified diffeomorphic flow operators simply by incorporating temporal embeddings into decoders, narrowing the long-standing performance gap between diffeomorphic and deformable models.

Limitations & Future Work

  • Cumulative Interpolation Artifacts: Although \(N=4\) linear interpolation warpings incur minimal overhead, repeated spatial resampling can theoretically induce slight smoothing on sub-voxel fine vascular structures or sharp anatomical boundaries. Future work could explore higher-order B-splines or analytical coordinate inversions.
  • Batch Size Scalability: Constrained by GPU memory for full 3D volumetric fields, training currently relies on batch size 1. Expanding to multi-GPU distributed groupwise registration and self-supervised foundation pretraining remains an open avenue.
  • vs LDDMM / NODEO / R2Net: Prior non-autonomous models parameterize time-varying velocity fields \(v(t)\) and rely on discrete forward solvers (Euler/RK4), incurring severe runtime bottlenecks and truncation errors. TPFM-DIR directly learns the two-parameter solution map without numerical integration, accelerating inference by orders of magnitude.
  • vs VoxelMorph / TransMorph-diff (Autonomous Scaling-and-Squaring): Autonomous approaches assume stationary velocity fields, which restrict deformation trajectories under large physiological motions (e.g., lung breathing and cardiac cycles). TPFM-DIR supports time-varying flows, outperforming stationary models by 2%~3% Dice across diverse anatomies.
  • vs GradICON / CycleMorph (Surrogate Regularization): Surrogate models apply ad-hoc inverse consistency or gradient penalties defined purely at discrete endpoints. TPFM-DIR enforces algebraic cocycle consistency, providing formal guarantees of continuous-time dynamical flow solutions.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates continuous-time non-autonomous diffeomorphic registration via direct two-parameter flow map learning with anchored cocycle constraints.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous validation across 9 multimodal 2D/3D datasets, comprehensive baselines, ablations on regularization and composition steps, and runtime analysis.
  • Writing Quality: ⭐⭐⭐⭐⭐ Mathematically sound, structurally coherent, with precise theorems, corollaries, and well-designed experimental verification.
  • Value: ⭐⭐⭐⭐⭐ Closes the historical performance gap between diffeomorphic and unconstrained deformable registration, offering a practical framework for flow-based medical image computing.