Skip to content

FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs

Conference: NeurIPS2026
arXiv: 2609.33158
Area: Medical Imaging
Keywords: Fundus images, cross-dataset generalization, probability calibration, subgroup fairness, low-rank adaptation

TL;DR

FOCUS connects harmonized disease labels from ten fundus datasets and the prediction interfaces of three foundation-model families to a shared evaluation layer, finding that ranking, calibration, and subgroup behavior do not align, while MLLM supervised fine-tuning primarily improves binary decisions and calibration with little average AUROC gain.

Background & Motivation

Color fundus photography research has expanded from disease-specific classifiers to vision encoders such as DINO and RETFound, vision-language models such as CLIP and FLAIR, and generative multimodal large language models such as Gemma and MedGemma. Although all consume fundus images, their outputs differ: image features, image-text similarities, or generated answers with token probabilities. Consequently, the same reported classification accuracy can arise from different training supervision, prompt templates, answer parsers, and probability transformations, rather than directly measuring differences in medical competence.

A second blind spot concerns differences between datasets. Acquisition devices, regions, populations, image quality, and annotation protocols can change together; strong performance within one dataset does not establish reliability elsewhere. Existing fundus datasets supply valuable disease labels, while FunBench and LMOD emphasize generative question answering or failure modes. These resources do not jointly cover same-task cross-dataset comparisons of vision encoders, dual encoders, and generative models while systematically examining ranking, calibration, and subgroup disparities.

FOCUS does not introduce a new diagnostic network. It establishes extensible dataset, task, and model-interface registries, reporting base-model results separately from post-adaptation external transfer. Core idea: establish what labels and prediction interfaces mean, then use a shared multidimensional protocol to examine which capabilities models retain and which failure modes change across data sources, rather than treating one leaderboard score as reliability.

Method

Overall Architecture

FOCUS takes fundus images, valid disease labels, and the age, sex, and quality metadata available in some of its ten datasets. It performs “Task and Metadata Harmonization,” obtains scores and hard predictions through “Interface-Separated Evaluation,” and conducts “Multidimensional Reliability Diagnostics.” A separate “Matched-Base Adaptation Comparison” examines whether LoRA adapters trained on BRSET or mBRSET transfer to other datasets supporting the same task.

The shared element is the analysis protocol, not an ensemble of the three model families or equal supervision for every model. Linear probes in the base benchmark already use the target dataset's training labels, whereas zero-shot image-text matching and MLLM prompting require no task-specific fitting. Explicit cross-dataset adaptation transfer is evaluated in the separately reported MLLM SFT analysis.

The diagram describes only the evaluation process. Solid edges indicate evaluation flow, and dashed edges identify training supervision. The three model families produce results independently.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Ten fundus datasets"] --> B["Task and Metadata<br/>Harmonization"]
    B --> C["Interface-Separated Evaluation<br/>VM: frozen-feature probe<br/>VLM: probe or image-text matching<br/>MLLM: binary prompting"]
    T["Each dataset's training labels"] -.->|Probe training only| C
    C --> D["Multidimensional<br/>Reliability Diagnostics"]
    D --> E["Matched-Base Adaptation Comparison<br/>Separate adapter inference<br/>Same-task ID and OOD"]
    S["BRSET or mBRSET<br/>Task labels"] -.->|LoRA SFT supervision| E
    E --> F["Task- and interface-specific reports"]

Key Designs

1. Task and Metadata Harmonization: identical disease names do not ensure identical labels

The registry contains BRSET, mBRSET, PAPILA, RFMiD, RFMiD 2.0, IDRiD, Messidor-2, REFUGE, G1020, and JSIEC1000. All three targets are binary, but each dataset participates only in tasks with valid labels after harmonization. The ten datasets should not be interpreted as uniformly supporting all three tasks.

“Any diabetic retinopathy (DR)” corresponds to ICDR grade at least 1; “referable DR” corresponds to grade at least 2 and/or suspicious macular edema. The glaucoma-related task is more precisely glaucomatous optic neuropathy (GON): labels use a vertical cup-to-disc ratio of at least 0.6 or specialist annotations of glaucomatous structural damage, rather than claiming that a single fundus photograph supplies all evidence required for a definitive glaucoma diagnosis.

BRSET and mBRSET support all three tasks; PAPILA, REFUGE, and G1020 support GON; the remaining five datasets support the two DR tasks. This mapping reduces naming discrepancies but cannot eliminate differences in annotation noise, diagnostic standards, or sample composition between sources.

Metadata coverage is also uneven. Age and sex analyses primarily depend on BRSET, mBRSET, and PAPILA; quality analyses depend on BRSET and mBRSET. Availability is a condition for calculation, not an invitation to infer missing demographic attributes or quality labels.

2. Interface-Separated Evaluation: preserve the families' different supervision and score sources

Vision models (VMs) freeze the image encoder, fit a class-balanced logistic-regression head on each dataset's training split, and evaluate on that dataset's test split. General encoders include ViT-L/16, DINOv2-L, and DINOv3-ViT-L; ophthalmic encoders include RETFound variants and VisionFM-Fundus. This evaluates pretrained features plus a lightweight classifier trained with task labels, not unsupervised zero-shot diagnosis by the vision model.

Dual-encoder vision-language models (VLMs) expose two interfaces. Linear probing uses only the frozen image tower with the same head protocol as VMs. Zero-shot evaluation matches the image to negative and positive text prompts, multiplies wrapper-processed similarities by 100, and applies a two-class softmax. This fixed scaling changes score sharpness, so VLM calibration depends on the probability-conversion convention as well as the representation.

The VLM group includes CLIP, SigLIP2, MedSigLIP, EyeCLIP, RET-CLIP, and FLAIR. The source's Table 2 places MedSigLIP in a Generalist row, so group names should not be interpreted as a strictly controlled distinction between medical and entirely nonmedical pretraining. Cross-group averages are not a causal experiment on medical pretraining.

Multimodal large language models (MLLMs) receive task-specific binary prompts requiring yes or no. Continuous scores and hard predictions are generated separately: the score is the maximum probability of affirmative token variants in the full-vocabulary softmax at the first generated position, while the hard prediction comes from parsing the first standalone yes/no in the generated text.

\[ s_i=p_i^+=\max_{v\in\mathcal{Y}_{\mathrm{single}}}p_\theta(v\mid x_i,q_i). \]

Affirmative variants cover capitalization, leading spaces, and leading newlines, retaining only variants encoded as one token. The implementation also extracts the maximum negative-variant probability \(p_i^-\), but AUROC, AUPRC, and ECE use raw \(p_i^+\), not \(p_i^+/(p_i^++p_i^-)\) or the sum of all affirmative-variant probabilities.

This distinction matters: a model may prefer yes over no while allocating most vocabulary probability to other answer forms, producing a low raw affirmative score. The hard prediction comes from generated text, so it need not equal thresholding that score at 0.5. Accuracy and ECE should not be treated as equivalent descriptions of one decision process.

Rows without a valid yes/no parse or affirmative score are excluded. A common interface makes results analyzable but does not ensure that every model retains the same sample denominator; parsing failures and missing-score rates should be interpreted alongside metrics.

3. Multidimensional Reliability Diagnostics: distinguish ranking, probabilities, and between-group gaps

Ranking metrics are AUROC and AUPRC, with AUPRC implemented as scikit-learn average precision; accuracy and F1 use hard predictions. AUROC describes positive-negative ordering without guaranteeing performance at a fixed threshold, while AUPRC also depends on each dataset's positive prevalence.

ECE measures positive-class probability calibration, not confidence that the predicted class is correct. Scores are divided into 10 equal-width bins, comparing each bin's mean positive-class score with its observed positive frequency:

\[ \operatorname{ECE}=\sum_{m=1}^{10}\frac{|B_m|}{n}\left|\overline{s}_{B_m}-\overline{y}_{B_m}\right|. \]

Age groups use the median valid age within the current analyzed result rows, with ages above that median assigned to one group; a single fixed age threshold is not imposed across datasets. Sex is normalized to male/female using dataset fields, and unrecognized or missing values are excluded from sex analyses.

Fairness diagnostics calculate absolute between-group differences in predicted-positive rate, accuracy, true-positive rate, false-positive rate, and AUROC. The equal opportunity gap is the true-positive-rate difference; the equalized odds gap is the maximum of the true-positive-rate and false-positive-rate differences, not their average. These are descriptive differences, not direct estimates of causal discrimination.

Quality robustness compares two quality strata within the entire model–dataset–task result, rather than further partitioning every demographic subgroup. BRSET uses adequate/inadequate quality and mBRSET uses the final accepted/not-accepted flag. Absolute differences between quality groups are calculated separately for accuracy, AUROC, AUPRC, F1, and ECE.

The default minimum demographic subgroup size is 10, and the minimum quality-group size is 5. Metrics are skipped if fewer than two groups remain or the smaller group fails the threshold. Diagnostic coverage can therefore differ between models. Two equally poor groups can have a small gap; a small disparity does not establish safety or high absolute performance.

Table 3 labels Fairness, Quality, and Shift as higher-is-better, but the appendices define the underlying fairness and quality gaps without adequately specifying these high-score summaries' transformations, weights, or the construction of Shift. This note preserves the reported values and directions without inventing formulas or interpreting the summaries as simple complements of a particular defined gap.

4. Matched-Base Adaptation Comparison: compare before and after fine-tuning on the same model and test task

Adaptation trains separately for each of the three tasks on BRSET and mBRSET, because both provide task labels and relevant metadata. Each adapter is evaluated on the training source's held-out test split and on other datasets supporting the same task. OOD means that training and test datasets differ, not transfer between disease tasks.

There are 228 adapter evaluation configurations: 36 in-domain and 192 same-task out-of-domain. Each result is paired with the same base MLLM on the identical test dataset and task. Deltas are fine-tuned minus base metrics; increases are desirable for AUROC, AUPRC, accuracy, and F1, while decreases are desirable for ECE.

Adapters and base models use the same prompts, answer parser, and score extraction, preventing evaluation-code changes from being counted as adaptation gains. Statistical analysis aligns paired predictions by image identifier and resamples images with replacement within each comparison. Domain-level summaries additionally bootstrap the mean paired deltas across adapter–test comparisons, applying Benjamini–Hochberg correction within each metric family.

This is more interpretable than comparing two unmatched leaderboards, but the paired unit remains the image, not an explicitly defined patient cluster. Both eyes or multiple images from one patient may be correlated; patient-level resampling and patient-disjoint splits not reported by the paper should not be assumed to have been performed.

Loss & Training

The six adapted models are Gemma-3-27B, MedGemma-4B, MedGemma-1.5-4B, LLaVA-NeXT-8B, Qwen3-VL-8B, and MedGemma-27B. SFT optimizes cross-entropy on assistant binary-answer tokens, rather than directly optimizing AUROC, fairness gaps, or calibration error.

LoRA targets all linear modules with rank 16, alpha 16, dropout 0.05, and no trained bias. Training uses bf16, gradient checkpointing, and the configured 16-bit mode, for 1 epoch at a learning rate of \(2\times10^{-5}\).

The 27B-class models use batch size 1 with gradient accumulation 8; smaller models use batch size 4 with gradient accumulation 4. Training jobs request 4 H100 GPUs with 80GB each; base evaluation and adapter inference request 1 H100.

The 18h training and 20h evaluation values in Appendix D are Slurm job time limits, not measured runtimes, and cannot support comparisons of inference efficiency.

Key Experimental Results

Main Results

The base bundle covers 532 model–dataset–task–method configurations: 190 VM, 228 VLM, and 114 MLLM. The table below reproduces the Overall rows of source Table 3. These are group summaries, not a controlled competition with identical supervision and metadata coverage for every model.

Model group Method Accuracy AUROC AUPRC ECE ↓ Fairness ↑ Quality ↑ Shift ↑
General VM Linear probing 0.803 0.828 0.617 0.209 0.899 0.885 0.654
Ophthalmic VM Linear probing 0.733 0.667 0.414 0.231 0.917 0.899 0.762
General VLM-encoders Linear probing 0.802 0.829 0.614 0.201 0.923 0.863 0.661
General VLM-encoders Zero-shot 0.552 0.549 0.303 0.275 0.934 0.891 0.800
Ophthalmic VLM-encoders Linear probing 0.763 0.773 0.592 0.182 0.915 0.911 0.732
Ophthalmic VLM-encoders Zero-shot 0.636 0.647 0.395 0.240 0.948 0.918 0.729
General MLLMs Zero-shot prompting 0.575 0.658 0.423 0.388 0.950 0.950 0.603
Medical MLLMs Zero-shot prompting 0.791 0.803 0.642 0.194 0.939 0.966 0.567

General vision encoders substantially outperform ophthalmic VMs, but the general VLM linear-probe AUROC of 0.829 is also slightly above the general VM value of 0.828. “VMs are strongest” is therefore not a strict conclusion across all methods. Ophthalmic VLMs outperform general VLMs in zero-shot evaluation but trail them under linear probing, showing that interfaces can reverse group rankings.

Medical MLLMs have better average AUROC and ECE than general MLLMs, but their Shift summary of 0.567 is below the general MLLM value of 0.603. Even following the table's higher-is-better direction, no group wins every axis. Because the summary is underdefined, this column supports comparison only at the level of the original report.

Source Table 4 reports general/ophthalmic VM mean AUROC of 0.754/0.592 on mBRSET, a difference of +0.162. Although mBRSET is the only mobile-phone fundus dataset, device, population, and label composition change together. This indicates a gap under compound domain shift, not isolated evidence that mobile acquisition causes degradation.

Ablation Study

The core adaptation analysis compares matched base models with SFT adapters rather than removing network modules. The table combines Table 5 and appendix Table 11. Every delta is SFT minus base, and intervals are 95% confidence intervals.

Test domain Metric Mean delta 95% CI q-value: Table 5 / Table 11
In-domain AUROC +0.004 [-0.002, 0.010] 0.477 / 0.521
In-domain Accuracy +0.033 [0.003, 0.076] 0.004 / 0.004
In-domain ECE -0.040 [-0.085, -0.008] 0.004 / 0.004
OOD AUROC +0.002 [-0.002, 0.006] 0.521 / 0.521
OOD Accuracy +0.012 [0.005, 0.021] 0.005 / 0.005
OOD ECE -0.018 [-0.026, -0.009] 0.004 / 0.004

After SFT, mean in-domain AUROC is 0.776 and ECE is 0.164; corresponding OOD values are 0.724 and 0.290 (appendix Table 10). Average accuracy and ECE improvements pass corrected significance tests, whereas AUROC intervals include zero. This supports the interpretation that adaptation primarily changes answer and score mappings, but does not directly establish an internal mechanism or guarantee benefit for every adapter.

Key Findings

  • In Table 3, medical MLLM AUROC is 0.868 for any DR and 0.656 for GON; general VM values are 0.867 and 0.742. GON is substantially harder, although heterogeneous labels and imaging protocols may contribute to this gap.
  • The source reports that VLM linear probing generally improves AUROC while increasing fairness gaps for all six compared models. Better ranking and smaller subgroup error-rate gaps cannot substitute for one another.
  • At the same 27B scale, the main text reports that MedGemma increases mean AUROC from Gemma-3's 0.648 to 0.797 and reduces ECE from 0.36 to 0.19. Scaling the medical family from 4B to 27B does not reliably improve mean AUROC. These are descriptive comparisons of the tested models, not causal evidence controlling all training differences.
  • The source contains conflicting numerical summaries: the text around Figure 4 gives GON means of 0.701 for general VMs, 0.714 for general VLM probes, and 0.642 for medical MLLMs, versus 0.742, 0.745, and 0.656 in Table 3. This note uses method-explicit Table 3 and does not mix the two sets of averages.
  • Table 4's JSIEC1000 row reports 0.958, 0.851, and a delta of +0.106; subtracting the displayed values yields 0.107. Its RFMiD row reports 0.914, 0.849, and +0.066, while displayed-value subtraction yields 0.065. The author-reported deltas are preserved; undisplayed precision may explain them, but the source does not provide enough detail to verify this.

Highlights & Insights

  • The evaluation interface is itself an experimental variable. Switching a VLM from image-text matching to linear probing changes ranking and subgroup gaps as well as accuracy. One pretrained model does not have a unique medical classification performance.
  • Raw affirmative-token probability deserves a separate audit. MLLM ECE reflects both disease signal and probability mass allocated to answer forms. Future controlled work could compare raw scores, binary-normalized scores, and independent calibration, but these alternatives must not be presented as experiments performed here.
  • Matched deltas are more informative for adaptation than absolute leaderboards. Pairing the same base model, task, and test source reveals whether adaptation improves ranking or primarily shifts decision behavior. The valid sample intersection and failure coverage should still be reported.

Limitations & Future Work

  • This is not clinical validation. The study performs retrospective classification evaluation and does not establish diagnosis, screening triage, patient-decision, or deployment readiness claims. Section 6 and Appendices A and S explicitly define this intended-use boundary.
  • Label harmonization cannot create a common gold standard. Structural GON labels differ from definitive glaucoma diagnoses, and DR annotation protocols remain noisy. Mapping sensitivity and annotation agreement should be reported in future work.
  • Missing metadata and outputs change comparison denominators. Minimum-group filtering, missing sex values, and invalid MLLM answers affect coverage. Each metric needs its effective sample size, not only an aggregate score.
  • Patient dependence remains insufficiently specified. The paper describes image-paired bootstrap but does not adequately establish patient-clustered statistics or patient-disjoint splits for every dataset. Both eyes or repeated captures may affect uncertainty and split independence; this is a risk, not proven leakage.
  • Summary definitions and reporting consistency need improvement. Fairness, Quality, and Shift formulas are insufficiently specified, and in-domain AUROC q-values differ between the main text and appendix. A fixed result version, aggregation weights, and complete computation protocol would improve reproducibility.
  • Image redistribution is constrained by licenses. Appendix S recommends releasing metrics, metadata, manifests, and code only where permitted. Public accessibility should not be equated with unrestricted redistribution.
  • vs RETFound / VisionFM: These works emphasize ophthalmic foundation representations; FOCUS offers an external, cross-dataset and interface-aware evaluation perspective. Its results motivate retaining general vision baselines such as DINO rather than assuming domain pretraining is uniformly superior.
  • vs FLAIR / RET-CLIP / EyeCLIP: These models incorporate ophthalmic image-text supervision into dual encoders. FOCUS evaluates both zero-shot image-text matching and frozen-image-feature probing, demonstrating why the two interfaces should be reported separately.
  • vs FunBench / LMOD: Fundus question-answering skills and generative failure modes are not equivalent to same-task disease classification transfer. FOCUS brings generative models into a shared analysis layer with nongenerative models, but does not evaluate the medical correctness of free-text explanations.
  • Future research direction: Controlling device, population, and labeling standards separately, then testing calibration methods with patient-level splits and fixed valid sample sets, would more precisely explain the observed compound domain shifts.

Rating

  • Novelty: 4/5 — Joint evaluation across interfaces, data sources, and reliability axes, rather than a new network architecture.
  • Experimental Thoroughness: 4/5 — Broad base and adaptation coverage, with remaining limits in subgroup metadata, patient-level statistics, and controlled domain shifts.
  • Writing Quality: 3/5 — The main protocol is clear, but summary definitions and main-text/appendix numerical consistency need improvement.
  • Value: 4/5 — A reusable framework for retrospective medical foundation-model evaluation, not evidence of clinical safety.