Skip to content

๐Ÿฅ Medical Imaging

๐Ÿง  NeurIPS2026 ยท 8 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (110) ยท ๐Ÿ“ท CVPR2026 (173) ยท ๐Ÿ”ฌ ICLR2026 (87) ยท ๐Ÿงช ICML2026 (28) ยท ๐Ÿค– AAAI2026 (76) ยท ๐Ÿง  NeurIPS2025 (77)

๐Ÿ”ฅ Top topics: Medical Imaging ร—3 ยท Multimodal/VLM ร—2

Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning

Aegis adds a task-relevant synthetic-image gradient update after each client-local epoch to increase sample mixing in linear leakage; it reduces the measured reconstruction rate to 9.38%โ€“11.63% across three MedMNIST modalities, without providing zero-leakage or differential privacy guarantees.

DeepArrhythmia: Segment-Contextualized ECG Arrhythmia Classification via Selective Evidence Acquisition

DeepArrhythmia produces R-peak-aligned beat labels within 10-second ECG segments and uses initial segment confidence to decide whether to acquire numerical and morphology evidence, demonstrating the value of context and physiological grounding across four datasets without outperforming always-rich inference on every metric.

FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs

FOCUS connects harmonized disease labels from ten fundus datasets and the prediction interfaces of three foundation-model families to a shared evaluation layer, finding that ranking, calibration, and subgroup behavior do not align, while MLLM supervised fine-tuning primarily improves binary decisions and calibration with little average AUROC gain.

Information Bottleneck-Guided Adaptive Hypergraph Transformer for Brain Disease Diagnosis

IBAHGT combines adaptive message passing on a fixed nearest-neighbor hypergraph with a parallel global Transformer, then uses node-level fusion and information-bottleneck auxiliary training to improve fMRI brain-network classification, achieving 77.6% / 86.3% accuracy on randomly split ABIDE / ADNI, although its mutual-information derivations contain important conditional and sign issues.

Modeling Whole-Slide Images as Dynamic Tumor Microenvironment Fields

TMEvolve models pathology slides as repeatedly updated latent region fields, learning slide representations through intra-region diffusion and concept-guided directed boundary interactions; it achieves the highest mean C-index on three of four survival cohorts, but its โ€œevolutionโ€ denotes representation refinement rather than actual tumor chronology.

Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity

Using three-image similarity choices across six histopathology cohorts, MOSAIC directly contrasts same-class cross-domain candidates with different-class same-domain candidates and finds that zero-shot multimodal large language models usually resist acquisition-context shortcuts better than Euclidean distances between pathology foundation model embeddings, without establishing clinical diagnostic competence or reliable cross-domain performance.

SheafStain: Sheaf-Theoretic Schrรถdinger Bridge for Spatially and Biologically Coherent Virtual Staining

SheafStain adds neighborhood VFM spatial conditioning, overlap-consistency losses, and pathology supervision to an unpaired Schrรถdinger bridge, then uses adaptive extra patches and weighted stitching to generate IHC from H&E, substantially reducing seams on 1024ร—1024 assembled regions from two breast pathology datasets without guaranteeing diagnostic correctness.

Tokenizer-Generator Coupling in Medical Image Generation

This controlled factorial study treats the tokenizerโ€“generatorโ€“sampler triple as its experimental unit: across 70 generation cells on ChestMNIST-64, reconstruction quality does not reliably predict generation quality, quantizer rankings depend on the generator, and validation-selected reduced-step sampling lowers LFQ-1024 + D3PM FID-192 from 0.44 to 0.09.