Short-to-Long Functional Connectivity Transfer via Structure-Aware Latent Diffusion¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Medical Imaging
Keywords: Functional Connectivity, Phenotype Prediction, Latent Diffusion Models, Reward-Guided Generation, Low-Rank Adaptation
TL;DR¶
Addressing the noise in short-scan fMRI and the attenuation of individual phenotypic variance in standard diffusion models, this paper introduces SALD, a structure-aware latent diffusion framework that combines structural connectivity priors with a reward-aligned LoRA adaptation steered by phenotype distillation and one-step denoising anchoring, significantly restoring downstream cognitive prediction while preserving network topology.
Background & Motivation¶
Functional connectivity (FC) derived from resting-state functional magnetic resonance imaging (rs-fMRI) serves as a principal representation for mapping brain connectome organization and predicting individual cognitive and behavioral phenotypes. However, empirical correlation matrices depend heavily on acquisition duration: in clinical and developmental cohorts such as the Adolescent Brain Cognitive Development (ABCD) study, participant compliance, head motion, and scanner costs restrict scans to brief intervals. These truncated acquisitions yield noisy, unstable estimates with severe variance, drastically constraining the statistical power of downstream brain-wide association studies (BWAS). While diffusion MRI structural connectivity (SC) provides an anatomical scaffold and demographic covariates offer population priors, reliably recovering high-fidelity, long-scan-equivalent (e.g., 20-minute) functional connectivity from short observations remains an open challenge.
Short-to-long FC transfer represents a conditional distribution transfer problem under partial observability. Crucially, standard conditional generative architectures (such as conditional diffusion models) trained with maximum-likelihood objectives encounter an inherent fidelity–utility tension. FC-to-phenotype prediction operates in a weakly explainable regime: even gold-standard 20-minute acquisitions explain only modest variance (predictive correlations typically plateau around 0.5, accounting for ~25% variance). When multiple plausible long-scan connectomes are consistent with a single short observation, likelihood-based objectives inherently favor conditional expectations. This averaging effect preserves population-level distributions and visual plausibility but attenuates subtle, subject-specific variations—the exact idiosyncratic variations that drive behavioral trait prediction.
The angle of attack in this paper is that unsupervised distribution matching alone cannot guarantee downstream scientific utility; one must ground generative paths in anatomical constraints and explicitly guide generation through task-oriented discriminative rewards. Core idea: develop a Structure-Aware Latent Diffusion (SALD) Transformer conditioned on structural connectivity priors and short-scan latents, regularized by one-step denoising anchoring and steered via LoRA fine-tuning on a differentiable phenotype distillation reward from a frozen long-scan predictor.
Method¶
Overall Architecture¶
SALD operates across two distinct stages: structure-conditioned latent diffusion base modeling, followed by reward-aligned LoRA steering with denoising anchoring. First, to handle the high dimensionality and positive semi-definite geometry of 100-parcel brain correlation matrices, a shared Variational Autoencoder (VAE) compresses both short- and long-scan FC into a unified latent manifold. The denoising backbone employs a Diffusion Transformer (DiT), concatenating the noisy target latent with the short-scan latent along the channel dimension. Concurrently, flattened log-transformed structural connectivity (SC) and demographic embeddings (age and sex) are projected through MLPs to yield auxiliary conditioning vectors that modulate DiT blocks via adaptive layer normalization (AdaLN-Zero). In the second stage, the diffusion backbone and VAE decoder are frozen, and Low-Rank Adaptation (LoRA) modules are inserted into DiT attention and projection layers. Gradients from a frozen long-scan phenotype predictor are backpropagated through the final DDIM sampling steps alongside a one-step noise anchoring regularization, optimizing the generator directly in connectivity space.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input short-scan fMRI and multimodal priors<br/>Fshort, SC structural connectivity, and covariates"] --> B["Shared manifold latent space encoding<br/>Unified VAE coordinate representation"]
B --> C["Structure-aware latent diffusion backbone<br/>DiT token concatenation and AdaLN-Zero modulation"]
C --> D["Phenotype distillation reward-guided adaptation<br/>Frozen long-scan predictor with LoRA gradients"]
E --> F["High-fidelity phenotype-sensitive long FC<br/>DDIM deterministic sampling and VAE decoding"]
D --> E["One-step denoising anchoring regularization<br/>Aligns u1 timestep noise to prevent drift"]
Key Designs¶
1. Shared manifold latent space encoding: Unifying short- and long-scan coordinates Parcellated functional connectomes typically scale to \(100 \times 100\) or higher, making data-space diffusion computationally prohibitive and prone to violating topological correlation constraints. More importantly, short-scan observations \(F_{\text{short}}\) and long-scan targets \(F_{\text{long}}\) represent the identical underlying neural connectome of a subject, differing primarily in finite-sample estimation noise rather than semantic identity. SALD trains a single VAE on paired short- and long-scan matrices, forcing the encoder to embed both noisy inputs and clean targets onto a unified biologically plausible manifold (\(z_{\text{short}} = E(F_{\text{short}}), z_{\text{long}} = E(F_{\text{long}}) \in \mathbb{R}^{4 \times 618}\)). This reduces cross-domain translation to a tractable trajectory denoising and refinement task within a shared coordinate system.
2. Structure-aware latent diffusion backbone: Conditioning on anatomy and short observations Functional signals alone cannot compensate for missing pathways during abbreviated acquisitions; however, white-matter tractography provides the physical constraints governing inter-regional communication. SALD utilizes a Diffusion Transformer (DiT) where the noisy latent \(z_t\) and condition latent \(z_{\text{short}}\) are channel-concatenated into tokens \([z_t, z_{\text{short}}] \in \mathbb{R}^{8 \times 618}\), allowing self-attention layers to directly inspect intact short-scan features during iterative denoising. Simultaneously, flattened streamline matrices \(\text{SC}_{\text{flat}}\) and 128-dimensional demographic embeddings \(e_{\text{cov}}\) (combining one-hot sex and sinusoidal age representations) are mapped to a conditioning vector \(h_{\text{aux}}\) via MLPs. This auxiliary representation modulates transformer blocks through AdaLN-Zero, providing macro-scale anatomical guidance while token concatenation preserves subject-specific topology.
3. Phenotype distillation reward-guided adaptation: Distilling signal without noise overfitting Under standard likelihood training, the model optimizes \(\mathcal{L}_{\text{diff}} = \mathbb{E}_{t,\epsilon}[\|\epsilon - \epsilon_\theta(z_t, t, \mathbf{h})\|^2]\), driving predictions toward conditional population averages under uncertainty and degrading individual-level discriminative features. Directly optimizing prediction errors against raw behavioral scores, however, causes the generator to overfit non-biological noise inherent to psychometric testing. SALD integrates Differential Reward Fine-Tuning (DRaFT) by embedding trainable LoRA parameters \(\theta_{\text{LoRA}}\) into DiT query, key, value, and projection layers while freezing the pretrained backbone. A frozen Kernel Ridge Regression predictor \(R(\cdot)\) trained on clean 20-minute FC evaluates decoded samples \(\hat{F}_{\text{LoRA}}\), supervising the model to match predictions derived from long-scan targets. This formulation extracts explainable trait-congruent variance and seamlessly utilizes subjects with missing behavioral labels.
4. One-step denoising anchoring regularization: Balancing fidelity and utility Unconstrained reward optimization risks steering generated samples away from the manifold of valid functional connectivity. Adapted from the KL regularization principle in DPOK and DRaFT, SALD incorporates a one-step noise prediction anchor evaluated at the earliest non-zero DDIM step \(u_1\). The complete steering loss is formulated as:
where \(\epsilon_{\text{base}}\) denotes the frozen base diffusion output, and \(\hat{F}_{\text{LoRA}} = D(z_{\theta_{\text{LoRA}}})\) is the connectome generated by backpropagating through the final \(K\) DDIM steps. The regularization weight \(\gamma\) controls the trade-off: strong anchoring preserves distributional fidelity, whereas flexible weighting permits LoRA adapters to adaptively recalibrate phenotype-relevant edge weights.
Loss & Training¶
The workflow proceeds in two phases. Phase one trains the full-parameter DiT backbone on 11,156 20-minute ABCD scans across durations \(\tau \in \{1, \dots, 10\}\) minutes using standard diffusion MSE objectives. Phase two freezes the VAE and diffusion backbone, updating only low-rank matrices \(\theta_{\text{LoRA}}\). Deterministic DDIM sampling accelerates inference, with gradients truncated to the final \(K\) steps under objective \(\mathcal{L}_{\text{steer}}\), ensuring rapid, stable convergence.
Key Experimental Results¶
Main Results¶
Evaluation was conducted on 613 independent test subjects from ABCD Release 4.0 with complete cognitive scores across input durations from 1 to 10 minutes, using 20-minute acquisitions (Schaefer 100 atlas) as the gold standard. Downstream evaluation tests cross-architecture generalization via Linear Ridge Regression (LRR) on the NIH Toolbox Cognition Total Composite Score (while KRR served as distillation reward). Metrics include Frobenius reconstruction distance, phenotype prediction MSE, and Differential Identifiability (\(I_{\text{diff}}\)) for subject specificity.
| Scan Duration \(\tau\) | Metric | Real Short FC | Cond. VAE | BicycleGAN | Cond. DDPM | Ours (SALD) | 20-min Reference |
|---|---|---|---|---|---|---|---|
| 1 min | Cognition MSE (↓) | 288.60 | 282.10 | 295.40 | 288.52 | 267.25 | 225.10 |
| 1 min | Frob. Distance (↓) | 16.40 | 14.85 | 16.20 | 15.13 | 14.27 | 0.00 |
| 1 min | Identifiability \(I_{\text{diff}}\) (↑) | 4.05 | 3.20 | 3.85 | 4.21 | 4.10 | 28.50 |
| 2 min | Cognition MSE (↓) | 275.40 | 271.80 | 284.10 | 274.23 | 252.82 | 225.10 |
| 2 min | Frob. Distance (↓) | 14.80 | 13.10 | 14.50 | 14.26 | 12.25 | 0.00 |
| 2 min | Identifiability \(I_{\text{diff}}\) (↑) | 8.10 | 6.50 | 7.90 | 8.49 | 10.57 | 28.50 |
| 3 min | Cognition MSE (↓) | 268.20 | 265.50 | 276.30 | 272.74 | 250.73 | 225.10 |
| 3 min | Frob. Distance (↓) | 13.90 | 12.45 | 13.70 | 13.55 | 11.77 | 0.00 |
| 3 min | Identifiability \(I_{\text{diff}}\) (↑) | 10.15 | 8.20 | 9.75 | 9.92 | 12.97 | 28.50 |
| 5 min | Cognition MSE (↓) | 258.10 | 256.40 | 268.20 | 261.30 | 245.80 | 225.10 |
| 5 min | Identifiability \(I_{\text{diff}}\) (↑) | 12.80 | 10.90 | 12.10 | 12.84 | 15.76 | 28.50 |
| 10 min | Identifiability \(I_{\text{diff}}\) (↑) | 18.05 | 15.60 | 17.20 | 18.17 | 20.83 | 28.50 |
Note: Baseline values sourced from Figures 3, 4, and Table 1 of the paper. SALD reduces cognition MSE by 15–22 points over Conditional DDPM in short acquisitions (1–3 min) and demonstrates widening identifiability advantages from 2 minutes onward.
Ablation Study¶
Component ablations across the noisy short-scan regime (\(\tau = 1 \sim 3\) min) isolate the contributions of multimodal conditioning \(-(\text{SC}, \text{cov})\), structural connectivity alone \(-\text{SC}\), short-scan inputs \(-\text{FC}_\tau\), one-step anchoring \(-\text{Anchoring}\), and direct label supervision \(\text{Direct-y}\).
| Duration \(\tau\) | Model Config | Full Model (Base) | w/o SC & Covariates | w/o SC | w/o Short FC Input | w/o Anchoring | Direct Label Superv. |
|---|---|---|---|---|---|---|---|
| 1 min | SALD (Frob. ↓ / MSE ↓) | 14.27 / 267.25 | 14.65 / 285.06 | 14.77 / 285.15 | 14.36 / 272.92 | 13.32 / 265.95 | 14.27 / 266.84 |
| 1 min | C-DDPM (Frob. ↓ / MSE ↓) | 15.13 / 288.52 | 15.83 / 304.55 | 16.62 / 307.50 | 17.78 / 311.95 | — | — |
| 2 min | SALD (Frob. ↓ / MSE ↓) | 12.25 / 252.82 | 14.26 / 257.55 | 13.58 / 256.90 | 14.36 / 272.92 | 12.44 / 250.56 | 15.43 / 257.97 |
| 2 min | C-DDPM (Frob. ↓ / MSE ↓) | 14.26 / 274.23 | 15.83 / 282.21 | 15.01 / 283.95 | 17.78 / 311.95 | — | — |
| 3 min | SALD (Frob. ↓ / MSE ↓) | 11.77 / 250.73 | 12.99 / 255.30 | 13.25 / 253.00 | 14.36 / 272.92 | 12.21 / 248.85 | 14.36 / 256.04 |
| 3 min | C-DDPM (Frob. ↓ / MSE ↓) | 13.55 / 272.74 | 15.53 / 290.48 | 13.85 / 276.23 | 17.78 / 311.95 | — | — |
Key Findings¶
- Direct supervision with raw behavioral scores (\(\text{Direct-y}\)) causes structural degradation at 2 minutes (Frobenius distance deteriorates from 12.25 to 15.43), confirming that noisy behavioral labels mislead generation, whereas distillation from a frozen long-scan model extracts stable predictive signal.
- Denoising anchoring highlights the fidelity–utility trade-off: omitting anchoring allows aggressive reward optimization with slight MSE drops at 1 min, but regularizing drift from the pretrained diffusion prior becomes vital for stability as duration increases.
- Structural connectivity provides indispensable anatomical guidance: excluding SC alone at 1 min worsens MSE from 267.25 to 285.15, demonstrating that tractography constraints stabilize functional topology during severe information deficits.
- In cross-phenotype evaluations, SALD models tuned exclusively on cognitive scores generalize effectively to unseen traits, recovering higher predictive variance in UPPS negative urgency, Child Behavior Checklist (CBCL), and Sleep Disturbance Scale for Children (SDS).
Highlights & Insights¶
- Formulates and empirically substantiates the fidelity–utility tension in brain network generation: likelihood-based diffusion produces visually plausible connectomes that attenuate the subtle inter-subject variations critical for phenotypic modeling.
- The phenotype distillation scheme elegantly bypasses behavioral noise by matching predictions against a frozen long-scan model rather than raw noisy labels, simultaneously unlocking unlabelled cohorts.
- Combines parameter-efficient LoRA tuning with truncated DRaFT backpropagation, delivering a lightweight, stable mechanism to steer complex diffusion models toward downstream neuroscience utility.
Limitations & Future Work¶
- Analysis is grounded in the single ABCD cohort using the SBCI pipeline and Schaefer 100-parcel parcellation; generalizability to finer atlases (e.g., 400 parcels) or adult clinical populations warrants validation.
- The framework targets group-level research applications; individual patient diagnosis requires rigorous prospective validation and uncertainty quantification.
- Future work could explore extending the paradigm to dynamic functional connectivity (dFC) and continuous flow matching formulations.
Related Work & Insights¶
- vs BrainNetDiff / DiffGAN-F2S: Prior connectome diffusion models focused primarily on unconditional synthesis or cross-modal translation (SC-to-FC or fMRI-to-SC); SALD specifically resolves duration-dependent super-resolution while mitigating discriminative feature loss.
- vs Palette / DRaFT: Adopts channel concatenation and reward-guided tuning from vision domains, but customizes them for symmetric positive-definite connectome manifolds with anatomical SC modulation and phenotype distillation.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates and solves the fundamental fidelity-utility conflict in neuroimaging generation with an elegant framework.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Evaluated across thousands of scans, multiple durations, diverse phenotypes, and rigorous ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear exposition, principled mathematical grounding, and comprehensive empirical insights.
- Value: ⭐⭐⭐⭐⭐ Offers a practical computational avenue to reduce acquisition burden while safeguarding behavioral predictive power.