Skip to content

HASSL: Hierarchy-Aware Self-Supervised Learning Framework for Single Cell Microscopy

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/tum-ai/HASSL
Area: Medical Imaging
Keywords: Self-supervised Learning, Single Cell Microscopy, Hierarchical Representation, Double-Teacher Distillation, HDBSCAN

TL;DR

To prevent coarse imaging modalities from suppressing fine-grained morphological hierarchies in single-cell self-supervised learning, HASSL introduces a double-teacher distillation framework guided by segmentation weak priors and a stability-weighted HDBSCAN hierarchical contrastive loss, achieving robust cross-modality alignment and morphology-driven subcluster separation across 2.3 million cells.

Background & Motivation

High-throughput microscopy imaging is indispensable in modern cell biology, drug discovery, and functional genomics. However, the scarcity of large, versatile, and accurately annotated single-cell datasets has prompted widespread adoption of self-supervised representation learning (SSL). While self-distillation frameworks (such as the DINO family) scale effectively without manual annotations or negative-pair overhead, deploying them directly to cellular microscopy reveals a critical domain limitation. In practice, experimental setups, imaging modalities (e.g., brightfield, fluorescence, phase contrast), and plate batch effects dominate the low-level statistical distribution of microscopy images. Consequently, standard SSL models inevitably fall into shortcuts, clustering cells into broad, modality-driven "superclusters" that obscure subtle intracellular and cross-modal morphological features.

This representation bias creates a fundamental tension: cellular analysis inherently relies on subtle, multi-level hierarchical morphology. Identical cell types (such as neurons or oligodendrocyte progenitor cells [OPCs]) exhibit drastically different visual appearances under different modalities, while entirely distinct cell types often display deceptively similar coarse silhouettes within the same modality. When the latent space is governed by imaging modalities, models fail to separate fine-grained subtypes within a modality and cannot align identical biological phenotypes across modalities. This deficiency impairs downstream tasks such as phenotype classification, mechanism-of-action (MOA) identification, and drug perturbation profiling. Supervised hierarchical learning suffers from the severe lack of multi-level labels, while existing hierarchical SSL methods (e.g., hard-mining or recursive prototype mining) struggle with noisy boundary gradients and unstable optimization on continuous morphological transitions.

This paper tackles the challenge by decoupling morphological structure from modality cues and actively structuring the latent space via the intrinsic geometry of unlabeled data. Core idea: integrate a segmentation teacher using zero-shot masks as weak geometric priors into double-teacher distillation to break modality superclusters, and mine multi-resolution prototypes from in-batch HDBSCAN condensed trees with persistence-stability weighting to pull ancestor clusters while repelling sibling subtypes.

Method

Overall Architecture

HASSL operates on single-modality crops of single cells without requiring any hierarchical or class supervision. The end-to-end framework consists of four stages: generating weak geometric segmentation priors, conducting double-teacher self-distillation to initiate morphology-driven breakout, constructing in-batch condensed cluster trees via HDBSCAN, and enforcing stability-weighted hierarchical prototype contrastive learning.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Single-Cell Image Crop Input<br/>Heterogeneous modalities and subtypes"] --> B["Stage 1: Weak Prior Segmentation Generation<br/>Zero-shot CellposeSAM geometric contours"]
    B --> C["Stage 2: Double-Teacher DINO Distillation<br/>Global image teacher + segmentation teacher alignment"]
    C --> D["Stage 3: HDBSCAN Hierarchical Tree Construction<br/>In-batch MST & ancestor-sibling prototype mining"]
    D --> E["Stage 4: Stability-Weighted Prototype Contrast<br/>Hinge contrastive loss weighted by persistence ฮป"]
    E --> F["Output: Hierarchy-Aware Latent Space<br/>Cross-modality alignment & compact subclusters"]

Key Designs

1. Double-Teacher DINO Distillation: Guiding Morphology Awareness with Geometric Weak Priors
To prevent the student backbone from shortcutting onto low-level modality textures, HASSL introduces binary segmentation masks pre-extracted by a generalist zero-shot model (CellposeSAM) as weak structural priors. Alongside the standard EMA image teacher that provides global semantic targets, a second segmentation teacher is introduced to ingest both images and segmentation maps. Both teachers produce sharpened, centered target distributions (\(q_{\mathrm{img}}\) and \(q_{\mathrm{seg}}\)) via the Sinkhorn-Knopp algorithm. The student ViT-S/16 network is trained to simultaneously minimize the cross-entropy against global image views and cross-modal distillation objectives between image crops and mask crops. The resulting joint loss encourages cells sharing the same underlying morphology across different imaging modalities to converge toward common geometric targets, breaking modality-dominated superclusters.

2. Multi-Resolution Prototype Mining: Extracting Ancestor Positives and Sibling Negatives from Condensed Trees
Flat clustering enforces a rigid resolution that either aggregates distinct subtypes or fragments representations into micro-clusters. HASSL builds a minimum spanning tree (MST) and condensed cluster tree over in-batch student embeddings using HDBSCAN (with min_cluster_size=2 to explore deep hierarchies). For any non-noise anchor sample \(i\), the centroids of all clusters encountered along its path from leaf to root define the multi-resolution positive prototype set \(\mathcal{P}_i = \{\mu_{c_{ik}}\}\). Conversely, the negative prototype set \(\mathcal{S}(i)\) is restricted to direct sibling clusters that split off from the anchor at each hierarchical level. This mining strategy removes redundant negative prototypes from deeper descendant sub-branches, maintaining a minimal efficacious set of negative constraints that preserve parent-level coherence while sharpening inter-subtype boundaries.

3. Stability-Weighted Hinge Contrastive Loss: Adaptive Hierarchical Pull and Repulsion via Cluster Persistence
In an unsupervised hierarchy, clusters at varying tree depths exhibit disparate confidences; treating reliable deep subclusters identically to coarse, unstable groupings distorts the representation space. HASSL leverages the density persistence score \(\lambda\) of HDBSCAN nodes to dynamically parameterize positive and negative prototype weights (\(\alpha_{ik}\) and \(\beta_{ic}\)). For ancestor positives, higher weights are assigned to deeper, more persistent clusters to tighten intra-subtype compactness. For sibling negatives, inverted normalized persistence values are applied, ensuring distinct cell types are decisively separated while closely related biological subtypes avoid excessive repulsion. The stability-weighted barycenters are aggregated via cosine similarity within a hinge margin objective:

\[\mathcal{L}_i = \max\left(0,\, m + \sum_{c\in\mathcal{S}(i)}\beta_{ic}\,s(a_i,\mu_{c}) - \sum_{k=0}^{K}\alpha_{ik}\,s(a_i,\mu_{c_{ik}})\right)\]

Optimized jointly with the distillation objective via an annealed schedule, this loss forces distinct cell types with subtle morphological variations (such as astrocytes and neurons in brightfield) to form well-separated, compact subclusters.

Loss & Training

The backbone is a ViT-S/16 trained for 100 epochs on 2ร— NVIDIA RTX A6000 GPUs. The segmentation teacher weight \(\gamma\) is linearly ramped from 0 to 0.2 across pre-training. The HDBSCAN hierarchical contrastive loss is introduced during the final 20 epochs, with its effective weight linearly scaled up to 0.1 (\(\lambda=1\)). Multi-crop augmentation (2 global + 8 local crops) and optimization hyperparameters strictly adhere to DINOv3 standards.

Key Experimental Results

Main Results

Evaluations were conducted on a curated multi-modality corpus encompassing 2,390,832 single cells across 20 instance segmentation benchmarks with 208 annotated classes. Performance was evaluated via k-NN retrieval (\(K \in \{1, 3, 5, 9\}\)) in the global embedding space, alongside Normalized and Adjusted Mutual Information (NMI/AMI).

Method Acc@1 (%) Acc@3 (%) Acc@5 (%) Acc@9 (%) mAP (%) NMI (%) AMI (%)
Cellpaint-DINO 44.8 55.7 61.7 68.9 49.7 47.0 46.4
ChadaViT 46.3 58.1 71.7 79.4 52.9 45.2 44.6
OpenPhenom 45.3 57.1 63.9 71.9 50.8 40.3 39.5
scDINO 50.4 61.9 68.0 75.3 55.4 43.9 43.3
HCSC* 46.1 61.6 64.5 79.7 53.9 45.3 44.6
Baseline DINOv3* 49.4 64.4 72.1 80.9 56.5 46.8 46.2
HASSL (Ours) 50.9 68.0 75.5 83.5 57.8 47.9 47.3

Note: Models marked with an asterisk were retrained on the unified benchmark for fair comparison.

Ablation Study

The ablation isolates the impact of flat clustering (DBSCAN), unweighted HDBSCAN, individual double-teacher distillation, and the standalone hierarchical contrastive loss.

Config Acc@1 (%) Acc@9 (%) mAP (%) NMI (%) AMI (%) Note
HASSL (Full model) 50.9 83.5 57.8 47.9 47.3 Synergistic combination of double-teacher and weighted HDBSCAN
w/o Double Teacher 50.9 82.8 58.1 47.6 46.9 Removes mask guidance; deeper neighborhood metrics decline
w/o HDBSCAN 50.6 82.9 57.4 47.2 46.2 Lacks hierarchical contrast; subtype boundary sharpness degrades
Baseline DINOv3 + DBSCAN 49.1 75.6 55.2 47.2 46.6 Flat single-scale clustering causes micro-cluster over-fragmentation
Baseline DINOv3 + Unweighted HDBSCAN 48.4 74.5 54.3 47.0 46.3 Equal weighting across depths distorts latent optimization

Key Findings

  • Substantial Gains on Deep Hierarchies: On multi-level hierarchy datasets (tree depth > 1), HASSL surpasses baseline DINOv3 by +6.3% in top-9 retrieval accuracy, while maintaining superior performance on flat single-level datasets without regression.
  • Superior Biological Drug Perturbation Profiling: On the unseen Allen Institute Cell Science perturbation benchmark (classifying paclitaxel vs. brefeldin perturbation across 7 cell lines), an MLP trained on frozen HASSL embeddings achieved an F1weighted of 92.0%, outperforming Cellpaint-DINO (81.7%) by +7.8% and DINOv3 (81.8%) by +10.2%.
  • Zero-Shot Out-of-Distribution Generalization: Evaluated on the unseen Human Protein Atlas (HPA) dataset, HASSL obtained 54.2% F1weighted, substantially exceeding DINOv3 (51.5%) and approaching Cellpaint-DINO (54.8%), which had seen closely matching fluorescent dye profiles during pre-training.
  • Minimal Compute Overhead: HASSL incurs only +43.6 minutes (+2.3%) in training wall-clock time and +4.44 GB (+10.2%) in peak VRAM relative to DINOv3 over 100 epochs, making it practically deployable.

Highlights & Insights

  • Segmentation Prior as a Modality Orthogonalizer: Leveraging zero-shot segmentation masks as weak, label-free geometric priors successfully nudges self-distillation representations away from dominant imaging modality artifacts and toward structural morphology.
  • Persistence-Weighted Topological Contrast: Converting HDBSCAN stability values (\(\lambda\)) into continuous, confidence-aware contrastive weights mitigates the instability of discrete hard negative mining across biological morphology gradients.
  • Cross-Modality Latent Alignment: In global t-SNE space, HASSL shifts the fluorescent cluster of the same cell type to become the 2nd nearest neighbor of the brightfield cluster (compared to the 8th for DINOv3 and 22nd for scDINO), demonstrating genuine biological semantic alignment.

Limitations & Future Work

  • Dependence on Baseline Mask Quality: While the weak prior is tolerant to moderate noise, catastrophic failure of the zero-shot segmenter on severely degraded or out-of-focus fields could introduce spurious shape priors.
  • In-Batch Hierarchy Resolution: The granularity of the condensed tree depends on in-batch sample diversity; very small batch sizes or highly homogeneous sampling could limit the depth of discovered biological hierarchies.
  • Framework Generality: While instantiated on DINOv3, the stability-weighted hierarchical prototype mining loss can be naturally extended to InfoNCE-based SSL, Barlow Twins, or Masked Image Modeling paradigms.
  • vs DINOv3 / scDINO: Conventional self-distillation models are heavily biased toward global texture and channel statistics; HASSL introduces structural segmentation guidance and hierarchical tree contrast to actively dismantle modality clusters.
  • vs HCSC (Hierarchical Contrastive Selective Coding): HCSC uses recursive k-NN with discrete hard positives and negatives, which causes gradient oscillations on continuous biological phenotypes; HASSL provides smoother, more stable convergence through condensed tree stability weighting.
  • vs Cellpaint-DINO / OpenPhenom: Existing foundation models in cell biology often rely on specialized multi-channel tokenization for specific assays (like Cell Painting); HASSL offers an assay-agnostic formulation that generalizes across 8 distinct microscopy modalities.

Rating

  • Novelty: โญโญโญโญโ˜† [Combines geometric segmentation weak priors with HDBSCAN persistence-weighted hierarchical prototypes in an elegant and effective pipeline]
  • Experimental Thoroughness: โญโญโญโญโญ [Evaluated on 2.3M cells, 20 datasets, k-NN retrieval, clustering metrics, and biologically critical downstream drug perturbation tasks]
  • Writing Quality: โญโญโญโญโญ [Clear motivation, rigorous mathematical formulation, and solid empirical closure]
  • Value: โญโญโญโญโญ [Sets a strong design paradigm for overcoming modality shortcuts in single-cell microscopy and biological foundation models]