Skip to content

๐Ÿงฌ Computational Biology

๐Ÿง  NeurIPS2026 ยท 11 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (6) ยท ๐Ÿ“ท CVPR2026 (21) ยท ๐Ÿ”ฌ ICLR2026 (155) ยท ๐Ÿ’ฌ ACL2026 (5) ยท ๐Ÿงช ICML2026 (52) ยท ๐Ÿค– AAAI2026 (20)

๐Ÿ”ฅ Top topics: Biomolecules ร—3

At FullTilt: Real-Time Open-Set 3D Macromolecule Detection Directly from Tilted 2D Projections

FullTilt feeds aligned 2D tilt-series directly into a visually prompted multiclass 3D detector, replacing volumetric sliding-window detection with cross-tilt row attention, near-zero-tilt query initialization, and training-time geometric augmentation to achieve subsecond zero-shot detection on three real-world cryo-ET datasets, while still requiring simulated-data pretraining and input alignment.

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

NMO tests molecular optimization beyond pharmaceutical priors through three nanophysics tasks constrained by electrode binding, and combines explicit-anchor GGS representations, random-molecule pretraining, and genetic-guided GFNs to discover scientifically promising candidates, although the full model does not maximize AUC on every task.

CellMSA: Context Modeling for Single-Cell Representation Learning

CellMSA aligns cross-batch and related-type cells by gene identity, extracts gene-pair dependencies from low-dimensional context, and uses them to guide target-cell encoding, achieving strong results in label-informed integration, classification without test-label retrieval, and perturbation prediction combined with STATE-ST.

Data-Driven Soft Labeling Scales DNA Read Classification to Whole-Body Cell-Type Deconvolution

Syto replaces single-origin hard labels with the cell-type distribution associated with methylation patterns in reference data, then separately models read classification, sample deconvolution, and proportion calibration; its best reported configuration reduces pseudobulk MSE across 39 cell types from CelFiE's \(3.28\times10^{-4}\) to \(0.88\times10^{-4}\), without establishing clinical cancer-detection performance.

Evolutionary foraging in grids: Intermittent search dynamics emerge in finite, depletable landscapes

On finite two-dimensional grids with non-renewable resources, the empirical distributions of step length, velocity, and turning angle evolve through a genetic algorithm rather than a prescribed power law; the resulting second and fourth displacement moments favor an intermittent-search description, without proving explicit state switching or globally optimal foraging performance.

FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases

FlyAOC evaluates fruit-fly knowledge base curation end to end, from large-scale full-text retrieval to ontology-grounded candidate outputs, showing that multi-agent context partitioning improves recall of known recoverable annotations while semantic scores do not establish exact biological validity.

LEMON-ZEST: Evolution-Informed Tokenization for Efficient Protein Language Modeling

LEMON-ZEST clusters evolutionarily conserved fragments into shared tokens and combines stochastic segmentation with a dual-head encoder for global retrieval and residue information, allowing a 200M-parameter model to outperform large baselines on several fold-level retrieval metrics without being best at every classification level or downstream task.

PHOEBI: An Open-World Benchmark for Multi-Label Bacterial Identification in Phase-Contrast Microscopy

PHOEBI tests multi-label identification of known species in unseen mixtures using 40 combinations of six bacteria and approximately 120,000 phase-contrast images, revealing severe failures of per-image trained classifiers and greater stability of geometric prototype decoders on frozen features, although their gains must be interpreted against the all-present baseline and combination-level confidence intervals.

PocketVE: Stable and Property-Guided Structure-Based Drug Design with Variance-Exploding Diffusion

PocketVE combines VE/EDM denoising that preserves the shared proteinโ€“ligand coordinate scale, training-time pocket perturbation, and inference-time multi-property CFG, raising 3D validity from 58.6% to 80.6% and reducing strain energy from 457.4 to 127.9 relative to guided TAGMol on CrossDocked2020, without attributing these gains to VE alone or claiming superiority on every docking metric.

Preserving DEG Rankings for Gene Discovery in Histology-Based Spatial Gene Expression Prediction

The paper shifts histology-based spatial expression prediction from matching each gene's spatial profile to preserving which genes should be prioritized for a tissue contrast, using morphology-derived proxy groups, differentiable U statistics, and an across-gene correlation loss to improve DEG ranking and pathway overlap without guaranteeing better conventional spatial PCC.

Towards Scalable Context-Aware Single-Cell Spatial Transcriptomics Prediction from Histology Images

CELLO shares one pathology foundation model forward pass per histology tile and predicts cell-level gene expression through continuous location querying and distance-decay cross-attention, improving average PCC on 52 paired H&Eโ€“Xenium samples while achieving a mean 14.0ร— speed-up over DeepSpot2Cell in timing that excludes upstream segmentation.