Skip to content

Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity

Conference: NeurIPS2026
arXiv: 2609.32876
Paper: NeurIPS2026 Poster
Area: Medical Imaging
Keywords: histological similarity, cross-institution generalization, batch effects, relative similarity, multimodal large language models

TL;DR

Using three-image similarity choices across six histopathology cohorts, MOSAIC directly contrasts same-class cross-domain candidates with different-class same-domain candidates and finds that zero-shot multimodal large language models usually resist acquisition-context shortcuts better than Euclidean distances between pathology foundation model embeddings, without establishing clinical diagnostic competence or reliable cross-domain performance.

Background & Motivation

Computational pathology foundation models use large-scale histology tile pretraining to produce reusable visual representations; models such as UNI2H and Virchow2 are intended to support cross-hospital retrieval, few-shot transfer, and multi-center data integration, not just downstream classifiers. These applications require proximity in the representation space to primarily reflect tissue morphology rather than stain color, scanners, or tissue preparation. However, within-cohort linear probes or nearest-neighbor classification can succeed through correlations between acquisition source and class, so strong performance does not rule out reliance on hospital signatures rather than biological features.

Studies such as the Robustness Index and PathoROB already show that pathology representations contain substantial source information, but their analyses generally require continuous embeddings and cannot directly evaluate generative multimodal models that return textual choices. Rather than proposing another pathology encoder, this paper therefore constructs a comparison task both model families can answer: given a reference image, choose between same-class tissue from another source and different-class tissue from the same source. This creates explicit competition between morphological consistency and acquisition-style consistency, exposing cross-domain shortcuts more directly than randomly chosen candidates.

General-purpose multimodal models can receive inference-time instructions to compare cells and tissue architecture while ignoring non-biological differences, providing a reference distinct from fixed embedding geometry. Their success would at least show that cross-source morphological comparison is not inherently impossible, but would not by itself establish a particular training objective as the cause of encoder failures. Core idea: use triplets in which class and acquisition source suggest conflicting answers to jointly measure whether encoders and generative models prioritize histological similarity, rather than merely checking their within-domain classification performance.

Method

Overall Architecture

Each evaluation contains three H&E histology tiles: a reference, a same-class positive candidate, and a different-class negative candidate. In cross-domain configurations, the positive must come from another slide or institution, while the negative shares the reference's source at the corresponding level. The model chooses one candidate; existing tissue-class labels determine correctness, not the model's explanation or newly collected expert similarity annotations.

Encoders independently embed the three images and select the candidate with the smaller Euclidean distance to the reference; multimodal large language models (LLMs) receive all three images and a text prompt and directly return a choice. Randomized candidate order gives both routes the same two-choice task and a 50% chance baseline. This is an evaluation protocol, not a trained network, so the datasets, distance calculation, and prompts are not forced into a model architecture diagram.

MOSAIC covers cross-slide comparisons on HER2ST, CCTGS, and TIGER, and cross-institution comparisons on CAMELYON16, TCGA, and BEETLE. The six cohorts span breast, colorectal, and lung tissue, with 2–6 classes; the cross-institution component involves 9 institutions or tissue source sites. All configurations together contain 13,800 comparison triplets, including 8,400 cross-domain triplets: 6,200 cross-slide and 2,200 cross-institution. The 13,800 comparisons are not all cross-institution tests, and the millions of source tiles are not the number of actual comparisons.

Key Designs

1. Source-conflicting triplets: make technical shortcuts and tissue classes imply opposite answers

In ordinary nearest-neighbor evaluation, same-class images may also share a hospital, so a correct choice does not establish morphological reasoning. Here the positive preserves the reference class but changes its source, while the negative preserves its source but changes its class, explicitly separating these cues. Choosing the negative is not merely a missed class: in this controlled comparison, the model ranks different-class same-source tissue above same-class cross-source tissue. Below-chance results are therefore especially informative and consistent with systematic source preference, although not every error can be attributed to color alone.

The three comparison levels serve different purposes. Within-slide comparisons take all three images from one slide and primarily assess class discrimination without source conflict at that level; cross-slide comparisons take the reference and negative from one slide and the positive from another. Cross-institution comparisons elevate the constraint to the institution: the reference and negative share an institution and the positive comes from another, but the first two need not share a slide. The appendix also provides within-institution controls whose images may come from different slides inside an institution, so within-institution must not be confused with within-slide.

Cross-slide sampling is stratified by ordered positive–negative class pairs, with 100 comparisons per pair: HER2ST's 5 classes give 20 pairs and 2,000 comparisons, CCTGS's 6 classes give 30 pairs and 3,000 comparisons, and TIGER's 4 classes give 12 pairs and 1,200 comparisons. Cross-institution sampling uses 100 comparisons per reference-institution and positive-class combination, producing 400 on CAMELYON16, 600 on TCGA, and 1,200 on BEETLE; the negative class is sampled uniformly from available different classes. This stratification prevents large classes from fully dominating the task, but overall accuracy still reflects a designed sampling distribution, not a hospital's real case distribution.

2. A shared choice task through two routes: compare fixed embedding geometry with instruction-guided visual judgment

Relative comparison reduces different model outputs to one shared question: does the same-class candidate rank above the different-class candidate? The paper expresses the correct ranking event through similarity:

\[ \mathbb{1}\bigl[s(t_r,t_p)>s(t_r,t_n)\bigr]. \]

The reference and positive share a class, while the negative belongs to another class; relative similarity accuracy is the fraction of valid comparisons with a correct choice. For encoders, the actual rule is Euclidean distance, not cosine similarity or a score calibrated by a downstream classifier:

\[ d_k=\|\mathbf{e}_{\mathrm{ref}}-\mathbf{e}_{c_k}\|_2,\quad k\in\{1,2\}. \]

The model selects the candidate with the smaller distance. Thus, this tests cross-domain ranking with existing representations and a specified distance rule, not whether those representations would fail under every possible adaptation. For LLMs, the biology-focused prompt directs attention to cell size, shape, and nuclear features, as well as glandular organization, stromal patterns, and cell density, while asking the model to ignore non-biological differences such as stain intensity and imaging artifacts. Both prompts require a JSON response with a more_similar_candidate field, with positive and negative candidates randomly assigned to positions 1 and 2.

The evaluation includes 17 models: 5 pathology foundation models, 6 general foundation models or extracted visual encoders, and 6 generative multimodal models. Qwen3.5, GLM-4.6V, and Gemma 3 each appear as both full models and visual encoders, offering useful contrasts between access to the same images and different judgment mechanisms. However, LLMs also receive explicit task instructions and perform joint three-image reasoning, whereas encoders use a fixed distance rule; this is a system-level comparison, not a controlled ablation changing only the training objective.

3. Stratified statistics and valid responses: prevent averages from hiding difficult tissue pairs or nonresponses

Overall accuracy aggregates valid comparisons, with additional analyses by reference class and positive–negative class pair to distinguish uniform gains from improvements concentrated on difficult pairs. For example, on HER2ST, UNI2H has only 25% accuracy on the connective-versus-immune-infiltrate pair and 32% on cancer versus immune infiltrate, suggesting that source competition and morphological resemblance can compound errors. Within-domain controls address a different question: is low cross-domain accuracy simply caused by an inability to discriminate classes at all? Strong within-slide but substantially weaker cross-slide performance for pathology models on HER2ST supports the interpretation that domain changes expose representation problems.

Comparisons without valid model outputs are excluded rather than counted as errors. LLM accuracy is therefore conditional on successfully returning a valid choice; if difficult comparisons are more likely to produce invalid outputs, exclusion can introduce selection bias. In Appendix F, Qwen3.5 has 354/400 valid responses on CAMELYON16, GLM-4.6V has 370/400, and GLM-4.6V has 1,995/2,000 on HER2ST, so these should not be treated as full-coverage two-choice systems. More complete reporting should jointly include response coverage, valid-response accuracy, and sensitivity analysis counting invalid outputs as errors.

Statistical testing covers only HER2ST cross-slide and CAMELYON16 cross-institution comparisons. The authors use 10,000 comparison-level bootstrap resamples for 95% confidence intervals, while McNemar's tests use paired comparisons valid for both models; fewer than 25 discordant pairs trigger an exact binomial test, otherwise a chi-squared test with continuity correction is used. The appendix explicitly states that category-best models were selected post hoc and the comparisons are exploratory. These results do not give every dataset and tissue pair the same statistical support, and comparison-level resampling does not directly replace slide- or institution-level independence assessment.

Loss & Training

The paper does not train or fine-tune models and introduces no loss function, optimizer, or supervised learning stage. Tissue-class labels construct and score the triplets rather than providing labeled examples during LLM inference; zero-shot means no additional task training or labeled demonstrations, not guaranteed absence of pathology data during pretraining.

WSIs are filtered using an Otsu tissue mask and tiled with 512×512-pixel windows and a 256-pixel stride, retaining tiles with at least 1% foreground. Local embedding extraction and model inference use an NVIDIA RTX 4090; input-token prices are the authors' cost proxy, not a substitute for end-to-end latency, output-token charges, or total retrieval-system cost.

Key Experimental Results

Main Results

The following table selects overall accuracies from Tables 1 and 2 and preserves their two-decimal reporting. The first three columns are cross-slide and the final three are cross-institution; these are not six repetitions at identical difficulty or class granularity.

Model HER2ST CCTGS TIGER CAMELYON16 TCGA BEETLE
UNI2H 0.46 0.59 0.19 0.61 0.31 0.56
Midnight 0.46 0.72 0.20 0.64 0.36 0.48
Virchow2 0.46 0.71 0.23 0.63 0.44 0.57
DINOv2 0.55 0.60 0.36 0.52 0.44 0.51
Gemini 3 Flash 0.55 0.75 0.35 0.71 0.55 0.54
Gemini 3.1 Flash Lite 0.53 0.72 0.37 0.73 0.56 0.58
Qwen3.5 0.55 0.74 0.40 0.68 0.59 0.57
GPT-5 Nano 0.57 0.64 0.46 0.59 0.52 0.57

At higher precision, GPT-5 Nano scores 0.570 on HER2ST versus 0.464 for Midnight, the best pathology model; their 95% CIs are [0.547, 0.591] and [0.442, 0.485], respectively. On CAMELYON16, Gemini 3.1 Flash Lite scores 0.730 versus Midnight's 0.640, with corresponding 95% CIs of [0.685, 0.775] and [0.593, 0.685]. The paired tests report p<0.0001 and p<0.001 for these two comparisons; GPT-5 Nano versus DINOv2 on HER2ST instead has p=0.153, so a significant advantage over the best general encoder is not established.

Within-domain controls more directly show that within-domain leadership does not guarantee cross-domain leadership. The table below uses within-slide values from Appendix Table 4 and cross-slide values from Appendix Table 13, avoiding precise drop calculations from two-decimal values.

HER2ST model Within-slide accuracy Cross-slide accuracy Drop (percentage points, calculated from table values)
UNI2H 0.728 0.459 26.9
Midnight 0.697 0.464 23.3
Virchow2 0.735 0.459 27.6
Gemini 3 Flash 0.706 0.549 15.7
GPT-5 Nano 0.616 0.570 4.6

Ablation Study

The prompt ablation covers only two models and two datasets, not all six cohorts. The following values come from Table 3; gains are percentage-point differences between biology-focused and minimal prompts.

Model and configuration Biology-focused prompt Minimal prompt Gain (percentage points)
Gemini 3 Flash, HER2ST cross-slide 0.549 0.510 3.9
GPT-5 Nano, HER2ST cross-slide 0.570 0.560 1.0
Gemini 3 Flash, CAMELYON16 cross-institution 0.705 0.668 3.7
GPT-5 Nano, CAMELYON16 cross-institution 0.590 0.555 3.5

Biology-focused guidance improves all four configurations, but minimal prompts do not universally outperform the best pathology model. On CAMELYON16, that claim applies to Gemini 3 Flash's 0.668 versus Midnight's 0.640; GPT-5 Nano's minimal-prompt accuracy is 0.555 and does not satisfy that comparison.

Key Findings

  • On TIGER, the best LLM, GPT-5 Nano, still scores 0.46, below the 0.50 chance baseline: relative improvement does not mean the task is solved.
  • On TCGA, Qwen3.5's precise accuracy is 0.589 versus 0.503 for Gemini Emb. 2, the best general encoder, but many cross-institution comparisons remain incorrect.
  • On BEETLE, the best LLM, Gemini 3.1 Flash Lite, scores 0.579 versus Virchow2's 0.565, a small gap; not all institution shifts support equally strong advantages.
  • Training-scale analysis is observational across different models, with nonuniform units including WSIs, tiles, and image–text pairs; it supports the claim that reported scale did not guarantee robustness among evaluated models, not that increasing data must fail.

Highlights & Insights

  • Expose shortcuts through comparison relationships: source consistency and class consistency compete without requiring a trained classifier head. This evaluation design can audit whether other medical imaging representations depend on acquisition devices or data centers.
  • Specify similarity criteria at inference time: the prompt ablation shows that generative models can shift feature emphasis, while fixed embedding distances have no equivalent immediate interface. This motivates relational supervision for representation learning, although the paper does not test such a training strategy.
  • Retain within-domain controls: strong within-slide but weak cross-slide performance is more diagnostic than a single low score. A favorable within-domain average should not replace out-of-domain auditing.

Limitations & Future Work

  • Limited scope: evaluation covers only H&E tiles and six cohorts, not immunohistochemistry, multiplexed imaging, whole-WSI reasoning, or prospective clinical use; zero-shot similarity choices are not clinical diagnostic validation.
  • Classes substitute for fine-grained similarity: same-class tiles need not always be morphologically closer than different-class tiles, especially with heterogeneous normal or stromal categories. Independent expert relationship annotations and class-granularity sensitivity analyses would help.
  • Valid-response selection bias: excluding invalid outputs can change the comparison sets for LLMs versus encoders. Coverage should be reported alongside checks for concentration of invalid responses in difficult classes or particular institutions.
  • Insufficient causal evidence: prompts, architecture, training data, objectives, and compute vary together, so language reasoning or a specific pretraining objective cannot yet causally explain the full advantage.
  • Source generalizations conflict with tabulated results: the text and Figure 5 description say all models lose accuracy across institutions on CAMELYON16 and TCGA, but Appendix within-institution values for GLM-4.6V-enc. are 0.365 and 0.315, versus cross-institution values of 0.580 and 0.403. This note preserves those numbers rather than treating universal degradation as a verified conclusion.
  • Distinguish release promises from available resources: the paper uses “release MOSAIC” language, while its abstract and checklist state that code and data will be released upon acceptance. The cache provides no verifiable repository link, so it does not establish current public availability.
  • vs Robustness Index / PathoROB: these studies analyze representation neighborhoods or embedding robustness, whereas MOSAIC uses discrete comparisons to include generative models. The approaches are complementary; discrete choices do not replace analysis of the entire embedding space.
  • vs stain normalization, RandStainNA, and STRAP: prior methods try to suppress color or style variation in the input, while this paper directly tests whether models can compare across such variation. Controlled before-and-after normalization comparisons would be needed to establish staining as the primary cause of source confusion.
  • vs single-image pathology VLM classification: instead of asking a model to guess a class name, this task supplies reference and candidate images for relational judgment. Stronger visual reference reduces ambiguity, but this task advantage is not equivalent to unrestricted diagnostic competence.

Rating

  • Novelty: 4/5 — Source-conflicting triplets place pathology representation robustness and generative visual judgment in a shared evaluation.
  • Experimental Thoroughness: 4/5 — Six cohorts, 17 models, and within-domain controls provide broad coverage, but response validity and statistical independence need further analysis.
  • Writing Quality: 4/5 — The task is clearly defined, while some universal claims and resource-release statements need greater precision.
  • Value: 4/5 — Useful for multi-center representation and retrieval auditing, but insufficient to establish clinical deployment readiness.