Y-diff: Structure-Texture Decoupled Diffusion Distillation for H&E-to-pCLE Translation¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/Hdw2agon/Y_diff
Area: Medical Imaging
Keywords: Cross-Modality Translation / Probe-based Confocal Laser Endomicroscopy (pCLE) / Structure-Texture Decoupling / Flow Matching / Diffusion Distillation
TL;DR¶
Y-diff proposes a structure-texture decoupled teacher-student diffusion distillation framework that leverages DiffAE semantic flow matching to transfer pCLE optical textures and grayscale SPNet spatial priors to anchor local cellular topology, substantially eliminating unpaired translation artifacts and reducing downstream cell counting MAPE from 55.77% to 13.84%.
Background & Motivation¶
Probe-based Confocal Laser Endomicroscopy (pCLE) enables real-time in vivo cellular-level optical biopsies, substantially enhancing the immediacy and efficacy of clinical pathological diagnosis. However, constrained by physical optical imaging mechanisms, pCLE imagery universally suffers from degraded contrast and blurred cell boundaries. This deficiency not only makes diagnostic interpretation heavily reliant on senior clinical experts, but also renders the acquisition and annotation of large-scale, high-quality paired training data prohibitively expensive. Consequently, existing deep learning algorithms remain mostly confined to small-sample classification, falling short of supporting downstream computational pathology tasks such as cellular detection or instance segmentation that impose strict topological requirements.
Translating widely available ex vivo H&E-stained pathology images into synthetic pCLE images via unsupervised cross-modal translation represents a promising route to overcome this data bottleneck. Nonetheless, an enormous domain gap separates H&E and pCLE, and establishing pixel-aligned paired ground truth between these modalities is fundamentally impossible in clinical practice. Classical cycle-consistency models (e.g., CycleGAN) and advanced neural Schrödinger bridge models (e.g., UNSB) attempt to map tissue morphology and optical styles simultaneously in a highly entangled pixel space. As a result, they inevitably generate black-and-white speckle artifacts, streak noise, or catastrophic structural distortions in cellular geometry. Concurrently, virtual staining frameworks tailored for weakly paired histology slides tend to suffer topological collapse when subjected to such completely unpaired distribution shifts.
The fundamental tension underlying these failures is that direct pixel-space mapping convolves the target domain's optical textures with the source domain's cellular topology: when learning pCLE dark-field optical textures, the generative process irreversibly corrupts the inherent cellular architecture of the H&E image. Core idea: physically decouple unpaired cross-modal translation into independent "texture flow" and "structure flow," transferring optical textures via low-dimensional semantic flow matching while enforcing cellular topological preservation through time-conditioned spatial priors and teacher-student diffusion distillation.
Method¶
Overall Architecture¶
Y-diff departs from traditional black-box end-to-end pixel mapping by constructing an architecture composed of a shared spatial forward-noising branch, a flow-matching semantic path, and dual-branch (Teacher-Student) denoising networks. The framework physically segregates the translation into two independent information flows: - Texture Flow: A pre-trained Diffusion Autoencoder (DiffAE) compresses global color and optical attributes into low-dimensional semantic codes \(z_{\text{sem}}\). An independent Flow Matching network establishes an Ordinary Differential Equation (ODE) optimal transport trajectory between the H&E and pCLE semantic distributions. - Structure Flow: The input H&E image is converted to grayscale and fed into a Structure-Preserving Network (SPNet), which extracts spatial morphological features at each diffusion timestep and modulates the latent states to strictly anchor cellular geometry.
During training, an Oracle Teacher model is optimized on real pCLE semantics to approximate the target data manifold. Subsequently, the Teacher and Flow Matching modules are frozen, and the Student model is trained conditioned on the transport-mapped semantic codes, recombining target textures with source spatial topology under multi-scale distillation losses and adversarial discriminator supervision.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["H&E Pathology Input xA"] --> B["DiffAE Semantic Flow Matching<br/>Optimal Transport ODE Maps zA to zB"]
A --> C["SPNet Time-Conditioned Structure Constraint<br/>Grayscale Pixel-Wise MLP Spatial Modulation"]
A --> D["Diffusion Forward Noising Phase<br/>Truncated DDIM Obtains Intermediate Latent xt"]
B --> E["Oracle Teacher Knowledge Distillation<br/>Multi-Scale Feature & Perceptual Alignment"]
C --> E
D --> E
E --> F["Synthesized Pseudo-pCLE Image"]
G["Real pCLE Image xB"] --> H["Target Domain Adversarial Supervision<br/>Discriminator Distribution Alignment"]
F --> H
Key Designs¶
1. DiffAE Semantic Flow Matching: Bypassing Pixel Entanglement and Mode Collapse Directly optimizing cross-modal mappings within pixel space is heavily disrupted by high-dimensional noise and distribution discrepancies. DiffAE provides a semantic encoder that abstracts global color and optical characteristics into a low-dimensional code \(z_{\text{sem}}\). Freezing the intermediate latent \(x_t\) and simply swapping \(z_{\text{sem}}\) causes the generated image to exhibit target-style optical attributes. Y-diff deploys Flow Matching to train a velocity field network \(v_\phi\) between source H&E semantic distribution \(p(z_{\text{sem}}^A)\) and target pCLE semantic distribution \(p(z_{\text{sem}}^B)\). Given minibatch optimal-transport-coupled pairs \((z_{\text{sem}}^A, z_{\text{sem}}^B)\) and linear interpolation \(z_t = (1-t)z_{\text{sem}}^A + t z_{\text{sem}}^B\), the flow matching objective is: $$ \mathcal{L}{\text{FM}} = \mathbb{E}^A)|_2^2 \right] $$ During inference, starting from }}^A, z_{\text{sem}}^B} \left[ |v_\phi(z_t, t) - (z_{\text{sem}}^B - z_{\text{sem}\(z_{\text{sem}}^A\), a single numerical integration pass using the Dopri5 ODE solver deterministically yields the translated target semantic code \(z_{\text{sem}}^{A \to B}\), fundamentally circumventing the mode collapse observed in GAN baselines.
2. SPNet Time-Conditioned Grayscale Spatial Constraints: Microscopic Cellular Topology Anchoring Global semantic transfer alone cannot restrain local morphological drift, allowing reverse diffusion to distort cellular boundaries. To counter this, Y-diff introduces the Structure-Preserving Network (SPNet, denoted as \(S_\psi\)). SPNet employs a lightweight pixel-wise MLP that processes the H&E image converted to single-channel grayscale, eliminating staining hue interference. FiLM (Feature-wise Linear Modulation) layers inject temporal embeddings \(c_t\) into each hidden layer, enabling the network to adaptively modulate spatial constraint intensity across different DDIM denoising phases: $$ x_t^A \leftarrow x_t^A + \lambda_{\text{map}} \cdot \sigma(S_\psi(x_A, t)) $$ This design directly channels the microscopic geometric topology of original tissue sections into the latent trajectory, ensuring that cell nuclei do not shift, dissolve, or spuriously replicate during synthesis.
3. Oracle Teacher-Guided Multi-Scale Knowledge Distillation: Stable Recombination of Decoupled Priors Direct end-to-end joint optimization of decoupled features risks gradient conflicts and optimization instability. Y-diff resolves this with a two-stage training scheme: first, a frozen Oracle Teacher model is trained with real pCLE semantics \(z_{\text{sem}}^B\) to output stable Pseudo-Ground-Truth targets \(\hat{x}_{\text{Teacher}}^B\); second, the Student model receives the flow-transported semantics \(z_{\text{sem}}^{A \to B}\) alongside SPNet spatial constraints to synthesize \(\hat{x}_{\text{Student}}^B\). To transfer the Teacher's optical generative priors without compromising source topology, a comprehensive multi-scale Knowledge Distillation (KD) loss aligns pixel, perceptual, and semantic feature spaces: $$ \mathcal{L}{\text{KD}} = \lambda}} |\hat{x{\text{Student}}^B - \hat{x}^B|}1 + \lambda}} \text{LPIPS}(\hat{x{\text{Student}}^B, \hat{x}}}^B) + \lambda_{\text{CLIP}} (1 - \text{CosSim}(\text{CLIP}(\hat{x{\text{Student}}^B), \text{CLIP}(\hat{x}^B))) $$ This structured teacher supervision enables the Student model to reliably render pCLE textures over source histological boundaries without hallucinating nonexistent cellular structures.}
4. Target-Domain Adversarial Distribution Supervision: High-Frequency Perceptual Sharpening While distillation effectively bridges feature representations, relying solely on L1 and perceptual regression can oversmooth generated images, losing the granular optical speckles and cytoplasmic textures unique to pCLE. Y-diff introduces a discriminator \(D\) to minimize divergence against the real pCLE data manifold, formulating the adversarial objective as: $$ \mathcal{L}G = \mathbb{E}}[\log D(x_B)] + \mathbb{E{x_A \sim p_A}[\log(1 - D(\hat{x}^B))] $$ The composite training objective for the Student model is }\(\mathcal{L}_{\text{Student}} = \lambda_{\text{KD}} \mathcal{L}_{\text{KD}} + \lambda_G \mathcal{L}_G\). The discriminator's adversarial feedback compels the Student to recover authentic high-frequency optical textures while strictly adhering to structural priors.
Loss & Training¶
The training pipeline executes in two sequential stages: 1. Oracle Prior Preparation: The Oracle Teacher model is trained independently for 30 epochs on real pCLE images guided by DiffAE semantics and SPNet constraints; in parallel, the Flow Matching MLP \(v_\phi\) is trained on minibatch OT-coupled semantic pairs. 2. Cross-Modality Distillation: Freezing both the Teacher model and the Flow Matching ODE module, the Student network and discriminator are trained for 10 epochs using the combined distillation and adversarial losses. Optimization uses Adam on a single RTX PRO6000 GPU. 3. Truncated DDIM Sampling: To curtail computational overhead, diffusion adopts a truncated schedule (1,000 total timesteps, 0.5 truncation ratio), compressing forward encoding and reverse sampling to 25 steps each from intermediate state \(x_t\), drastically accelerating inference while retaining crucial structural latents.
Key Experimental Results¶
Main Results¶
Evaluation is conducted on the Clinical pCLE Dataset (CCD), comprising 296 ex vivo colon H&E slides and 296 in vivo pCLE frames (split into 240 training and 56 testing images per modality), and the weakly paired MIST immunohistochemistry benchmark (ER-004 subset, 4,000 training / 1,000 testing). Quantitative metrics include Structure Similarity (SS), Luminance and Contrast (LC), Histogram Correlation (Hist), and LPIPS perceptual distance.
The quantitative comparison on the CCD benchmark is summarized below:
| Method | Paradigm | SS ↑ | LC ↑ | Hist ↑ | LPIPS ↓ |
|---|---|---|---|---|---|
| CycleGAN | Cycle-Consistent GAN | 0.614 | 0.908 | 0.731 | 0.511 |
| PyramidPix2Pix | Virtual Staining Pix2Pix | 0.114 | 0.926 | 0.823 | 0.974 |
| PPT | Patch Contrastive Learning | 0.217 | 0.925 | 0.854 | 0.907 |
| SynDiff | Adversarial Diffusion Translation | 0.226 | 0.927 | 0.815 | 0.529 |
| D-VST | Staining Decoupled DiT | 0.103 | 0.953 | 0.853 | 0.811 |
| UNSB | Neural Schrödinger Bridge | 0.488 | 0.920 | 0.798 | 0.655 |
| CycleGAN-Turbo | One-Step Accelerated Diffusion | 0.316 | 0.887 | 0.672 | 0.777 |
| Y-diff (Ours) | Structure-Texture Decoupled Distillation | 0.689 | 0.939 | 0.868 | 0.495 |
Methods tailored for paired virtual staining (e.g., D-VST, PyramidPix2Pix) suffer topological collapse in completely unpaired regimes (SS drop to ~0.10). In contrast, Y-diff achieves top performance across SS (0.689), Hist (0.868), and LPIPS (0.495), while maintaining competitive LC (0.939). On MIST, Y-diff similarly achieves the lowest FID (41.40) and highest structural fidelity.
Ablation Study¶
A systematic component ablation on CCD demonstrates the individual contributions of each module (w/o Student denotes the Oracle upper bound using ground-truth target semantics):
| Config | Mechanism / Setting | SS ↑ | LC ↑ | Hist ↑ | LPIPS ↓ |
|---|---|---|---|---|---|
| w/o Flow matching | Removing flow matching; using source H&E semantics | 0.5397 | 0.9339 | 0.8646 | 0.4810 |
| w/o SPNet | Removing time-conditioned spatial SPNet | 0.6090 | 0.9092 | 0.7814 | 0.4915 |
| w/o Discriminator | Removing adversarial discriminator supervision | 0.6039 | 0.9199 | 0.7802 | 0.4992 |
| Ours (Full) | Full Y-diff Student Framework | 0.6892 | 0.9387 | 0.8684 | 0.4948 |
| w/o Student (Oracle) | Directly utilizing ground-truth pCLE semantics | 0.7015 | 0.9664 | 0.8807 | 0.4460 |
Key Findings¶
- Flow Matching Governs Semantic Alignment: Removing Flow Matching drops SS by 0.1495, highlighting that severe cross-modal distortion largely stems from semantic and color misalignment corrupting the generative prior.
- SPNet Anchors Cellular Topology: Omitting SPNet degrades the Student into unconstrained imitation of the Teacher, causing substantial drops in SS and Hist and verifying that spatial MLP guidance is vital against geometric drift.
- Downstream Clinical Validation: In a downstream cell counting task, pseudo-pCLE images generated from NuInsSeg H&E slides were used to augment a U-Net trained on 35 manually annotated pCLE images and tested on 30 real pCLE images:
| Data Augmentation Strategy | MAE ↓ | RMSE ↓ | MAPE (%) ↓ |
|---|---|---|---|
| Manual Data Only (w/o Virtual pCLE) | 83.29 | 103.29 | 55.77% |
| + CycleGAN Augmentation | 63.71 | 73.17 | 30.86% |
| + Y-diff Augmentation (Ours) | 28.43 | 34.30 | 13.84% |
Y-diff augmentation slashes cell counting MAPE from 55.77% down to 13.84%, substantially surpassing CycleGAN (30.86%). Blind evaluations by three pathologists further confirm that Y-diff achieves statistically superior Mean Opinion Scores (MOS) in both morphological fidelity and texture realism (Welch's t-test, \(p < 0.05\)).
Highlights & Insights¶
- Decoupled-and-Recombined Paradigm: Rather than performing monolithic pixel-space transformation, isolating global optical style into an ODE flow path and preserving local histology via a grayscale MLP successfully breaks the trade-off between texture transfer and topological stability.
- Asymmetric Distillation Evades Training Instability: Establishing a static, high-capacity Oracle Teacher prior and transferring knowledge to the Student via multi-scale distillation bypasses the severe gradient conflicts typical of joint end-to-end unpaired translation.
- Closing the Loop to Clinical Downstream Tasks: Unlike studies that focus solely on visual metrics, Y-diff demonstrates that synthesized pseudo-pCLE data directly bolsters cell counting accuracy, providing practical utility for low-resource computational pathology.
Limitations & Future Work¶
- Sampling Latency and Fixed Resolution: Relying on a 25-step DDIM sampling process remains computationally heavier than one-step GAN baselines, and the current implementation is restricted to 256×256 resolution, limiting direct application to whole-slide imaging (WSI).
- Lack of Physical Optical Modeling: Texture synthesis relies purely on learned statistical distributions rather than explicit formulations of fluorophore decay, scattering, or objective focal depth.
- Future Directions: The authors suggest incorporating Latent Diffusion Models (LDM) to enable real-time megapixel-scale translation, alongside embedding physical confocal imaging equations into the texture pipeline.
Related Work & Insights¶
- vs CycleGAN [50] & CycleGAN-Turbo [28]: CycleGAN relies on pixel-space cycle consistency that induces severe ghosting under large domain gaps, while one-step CycleGAN-Turbo collapses on complex pathology distributions. Y-diff's structural decoupling prevents these artifacts.
- vs Neural Schrödinger Bridge (UNSB) [15]: UNSB attempts optimal transport in a unified pixel-noise space, which struggles to converge on small medical datasets and produces speckled artifacts. Y-diff restricts optimal transport to a low-dimensional semantic latent space.
- Takeaway for Medical Cross-Modal Translation: When tackling extreme modality gaps lacking pixel alignment (e.g., OCT-to-histology, MRI-to-CT), decoupling invariant structural scaffolding (e.g., grayscale topology) from modality-specific texture representations offers a far more robust generative framework than monolithic pixel translation.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ [Combines semantic flow matching with time-conditioned spatial MLP modulation in a diffusion distillation framework]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive comparisons, component ablations, downstream U-Net cell counting, and blind expert evaluations provide comprehensive validation]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, structured methodology, rigorous formulation, and high-quality figures]
- Value: ⭐⭐⭐⭐⭐ [Offers a robust, high-fidelity data synthesis pipeline for rare in vivo pCLE optical biopsy analysis]