LEMON-ZEST: Evolution-Informed Tokenization for Efficient Protein Language Modeling¶
Conference: NeurIPS2026 (task-list assignment; the version read is an arXiv preprint)
arXiv: 2609.37675v1
Area: Computational Biology
Keywords: protein language models, evolution-informed tokenization, remote homology retrieval, hierarchical contrastive learning, stochastic segmentation
TL;DR¶
LEMON-ZEST clusters evolutionarily conserved fragments into shared tokens and combines stochastic segmentation with a dual-head encoder for global retrieval and residue information, allowing a 200M-parameter model to outperform large baselines on several fold-level retrieval metrics without being best at every classification level or downstream task.
Background & Motivation¶
Protein language models typically treat each amino acid as a token and learn sequence regularities through masked language modeling. This preserves fine-grained information but feeds long sequences directly into attention and requires the model to rediscover which local fragments have stable biological significance. Remote homology retrieval is also not a matter of literal similarity: proteins with substantially different sequences can retain related folds. Recovering masked residues well does not automatically make cosine similarity between global embeddings suitable for retrieval.
Standard BPE shortens inputs, but merges adjacent strings according to frequency rather than evolutionary conservation. A statistically common fragment need not group substitutions that preserve structural significance into the same unit. Conversely, directly using structural tokens introduces stronger information but may require predicted or experimental structures in the evaluated configuration. This paper absorbs conservation information during vocabulary construction so that inference still requires only sequences; that is different from claiming that training contains no structure-derived information.
Rather than only increasing model size, the paper changes the basic units presented to the model: conserved fragments and related variants share compressed tokens, contrastive learning shapes the global retrieval space, and residue recovery constrains information loss. Core idea: let evolutionary conservation define equivalence classes in the input vocabulary, then use multigranular segmentation and dual-head training to align compressed sequence representations with remote homology relationships.
Method¶
Overall Architecture¶
ZEST is the vocabulary and tokenization mechanism, while LEMON is the encoding model that uses it. Offline construction extracts and clusters conserved zones from TEDLH HMM consensus sequences; online processing segments an input sequence into variable-length tokens and records the residue span of each token.
A shared Transformer contextualizes the shorter token sequence. Its global branch uses attention pooling to produce a vector for cosine retrieval, while its local branch expands token representations to residue positions and predicts amino acids; a separate token-level MLM output is tied to the input embeddings.
Solid arrows below show data flow, and dashed arrows show training-only supervision. Homology retrieval at inference does not require structures, MSAs, or CATH labels for the query, but training uses structural classification levels.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400, 'subGraphTitleMargin': {'top': 8, 'bottom': 16}}}%%
flowchart TD
P["TEDLH consensus sequences"] --> V["Evolutionary equivalence vocabulary"]
V --> T["Stochastic trie segmentation"]
X["Input sequence"] --> T
subgraph Y["Dual-head representation"]
E["Shared token encoder"] --> R["Attention pooling<br/>global retrieval vector"]
E --> S["Residue broadcasting<br/>global and local positions"]
S --> A["20-way residue prediction"]
E --> M["Embedding-tied<br/>token prediction"]
end
T -->|tokens and spans| E
C["Training: CATH hierarchy"] -.->|contrastive supervision| R
G["Training: masked targets"] -.->|MLM supervision| A
G -.->|MLM supervision| M
R --> O["Inference: cosine retrieval"]
Key Designs¶
1. Evolutionary equivalence vocabulary: map related conserved-fragment variants to one input unit
The authors use HHsuite consensus sequences for 765,248 TEDLH HMM profiles to extract conserved zones containing at least 3 residues and excluding homopolymers. These HMM profiles originate from multiple sequence alignments, so the zones represent patterns retained across related sequences rather than frequent strings selected directly from the pretraining corpus. MMseqs2 linclust then clusters them at 70% sequence identity, and a score considering cross-profile frequency and zone length ranks the clusters. The text does not provide the full ranking formula, so it should not be rewritten as a precisely specified mathematical objective.
The vocabulary retains the top 31,975 clusters, adds 20 standard amino acids as fallback tokens and 5 special tokens, and contains 32,000 IDs. The crucial step is not merely merging multiple residues: different zone members of the same cluster share one ID, making distinct surface forms an evolutionarily related input unit. This is a many-to-one mapping rather than a separate token for every fragment. It can suppress some sequence variation but also removes genuine residue differences, requiring a subsequent constraint on information loss.
The mean ZEST token span is approximately 4 residues, so a fixed token window typically covers a longer original sequence. The paper also reports 4–6-fold compression and illustrates a 1024-token window covering approximately 4,000 residues; these are average or empirical descriptions, not a guarantee that every 4,000-residue sequence fits. Sequences with fewer matching fragments rely more heavily on single-residue fallback, and stochastic segmentation can also increase token counts.
2. Stochastic trie segmentation: expose valid shorter fragments nested inside longer ones
Greedy longest matching always consumes the longest zone first, potentially hiding shorter vocabulary entries behind longer prefix matches. The coverage analysis reports approximately 139 unused tokens after visiting 5 million sequences. The authors organize the vocabulary into a prefix trie that returns valid matches from longest to shortest at each starting position. With a specified probability, tokenization rejects the longest match and samples uniformly among the shorter valid matches. The remaining suffix is tokenized subsequently, allowing the same sequence to have different boundaries across iterations.
The expression below organizes the selection probabilities stated in the paper. It describes local match sampling, not an additional loss function proposed by the authors.
When only one valid match exists, there is no shorter alternative and the available match should be retained; single-residue tokens provide fallback for unmatched zones. This dropout changes segmentation boundaries rather than masking token embeddings or uniformly sampling the entire vocabulary. It increases exposure to short fragments but does not guarantee equal occurrence counts during finite training, or even that every vocabulary entry appears.
The paper states that maximal dropout falls back to character-level segmentation, but the stated rule can still select multi-residue fragments when intermediate-length matches exist. The safer conclusion is that increasing dropout favors finer segmentation; a complete character-level limit requires additional selection conditions or implementation evidence. That limit is not a necessary consequence of the sampling rule in the text.
3. Dual-head representation: constrain global relationships and local content separately while compressing computation
The shared encoder uses a pre-norm Transformer, RoPE, SwiGLU, and Flash Attention through PyTorch scaled_dot_product_attention. Its main attention computation operates on the compressed token sequence, rather than first expanding to all residues and performing equally large attention. Input token embeddings are tied to the token MLM output projection, so token recovery directly constrains input representations; this output has a different granularity from residue prediction.
The global representation head uses a learned query vector and 4-head attention to pool all token representations. Pooling is followed by LayerNorm, an alignment projection initialized to identity, bottleneck projection blocks, a final contrastive embedding projection, and L2 normalization. Attention pooling can emphasize zones relevant to global structural relationships instead of assuming every position contributes equally to retrieval. Hierarchical contrastive learning then makes those relationships visible in vector distances; pooling alone does not automatically create this geometry.
The local branch broadcasts each token vector over its recorded residue span. Broadcasting alone would give all positions within a token the same representation and could not distinguish their residues. Two positional signals are therefore added: the global position along the chain and the local offset within the token, the latter described as ALiBi-style in the text. A two-block bottleneck MLP outputs 20-way amino acid logits, allowing training gradients to penalize local information lost by coarse encoding. This is a predictive constraint on the shared encoder, not a mathematical guarantee of reversible decoding.
Because the vocabulary is many-to-one, different fragments can begin with the same ID, and the recovery branch must still rely on context, spans, and positions. It mitigates compression ambiguity without proving lossless reconstruction of arbitrary sequences. The global retrieval and local recovery heads have distinct responsibilities, explaining why strong remote homology performance does not imply strong performance on every residue-level task.
A Worked Example¶
Consider an abstract fragment used only to illustrate segmentation: valid matches at one starting position have spans of 6, 4, and 1. With stochastic segmentation disabled, the span of 6 is selected. When dropout is triggered, each shorter candidate has half the conditional probability of selection, rather than forcing the span of 1. If the span of 4 is selected, subsequent matching handles the remaining two positions. This is not a real protein sequence or a reported experimental example.
Each resulting token occupies one attention position in the shared encoder. The global head aggregates contextual vectors into a retrieval embedding; the local head broadcasts the span-4 token vector to 4 residue positions and distinguishes them through local offsets. Training compares masked positions with their target residues and supervises the global vector using CATH relationships. Retrieval inference follows only the sequence-to-global-embedding path and needs no structural labels for the query.
Multiple stochastic segmentations of the same sequence also enable test-time augmentation. The following is the embedding-averaging mechanism given in the paper; each Tokenize call represents a fresh stochastic segmentation.
Weights remain frozen, and embeddings from different segmentations are averaged rather than training a new ensemble. The experiment uses 5 forward passes. Sequential processing can control peak GPU memory, but computation grows linearly with the number of passes, and varying segmentation lengths affect each pass's cost. This is not a free improvement.
Loss & Training¶
Pretraining uses 50% of UniRef90, approximately 90M sequences, with token-level MLM and residue-level MLM after expansion. Fine-tuning uses 59.7M protein pairs constructed from TEDLH domain annotations and contrastive learning at the CATH Architecture, Topology, and Superfamily levels. The text does not sufficiently specify the complete loss, coefficients, temperature, optimizer, or training schedule; this note does not invent an exact combined objective.
The main LEMON model contains 200M parameters, and the authors report one week of training on a single H100. CD-HIT separates training and validation pairs to reduce sequence leakage within the fine-tuning data. “Sequence-only” describes model inputs and inference modality: CATH is a structural classification, and TEDLH/TED construction also has structural origins. The conclusion's broad claim of no explicit structural supervision is therefore in tension with the structural hierarchy labels used in the method.
An audit of 726,922 fine-tuning domain sequences found no exact string matches against the three benchmarks. However, approximately 10% of benchmark domains have at least 70% global identity with a training sequence, and approximately 2.5% reach at least 90%. In CATH S20, 80.9% of fold types appear in training. These findings support discussing generalization across sequence divergence within known folds, not claiming complete absence of close-homology overlap or demonstrated generalization to unseen fold classes.
Key Experimental Results¶
Main Results¶
Remote homology retrieval maps each sequence to one vector, ranks candidates by cosine similarity, and uses structural classification levels to define positive and negative relationships. AUROC assesses overall ranking, while mAP complements it with top-of-list retrieval quality. The table below extracts LEMON and the key contrastive baseline ProtTucker from Table 1, without extending a particular level's advantage into superiority across the entire table.
| Benchmark / Level | LEMON AUROC | ProtTucker AUROC | LEMON mAP | ProtTucker mAP |
|---|---|---|---|---|
| CATH S20 / Architecture | 0.825 | 0.783 | 0.356 | 0.250 |
| CATH S20 / Topology | 0.898 | 0.874 | 0.320 | 0.293 |
| SCOPe / Fold | 0.903 | 0.876 | 0.347 | 0.263 |
| SCOPe / Superfamily | 0.959 | 0.971 | 0.555 | 0.693 |
| SCOP / Fold | 0.904 | 0.855 | 0.287 | 0.212 |
| SCOP / Superfamily | 0.948 | 0.961 | 0.421 | 0.541 |
LEMON has clear fold-level advantages but trails ProtTucker on both SCOPe and SCOP Superfamily metrics. Figure 3 separately reports SCOP fold AUROC of 0.875 at the strictest th10 identity threshold and 0.904 at th95; this threshold-specific setting should not be interchanged with the overall columns in Table 1.
The text claims mean AUROC of 0.86 and mAP of 0.31 across the three benchmarks without clearly specifying the averaged levels and weights in the supplied source. Direct arithmetic averaging of the six Table 1 columns instead gives approximately 0.906 and 0.381 for LEMON and does not reproduce those numbers. They are retained as author-reported values, not presented as a verified six-column mean.
Table 1 lists 5m for LEMON and 9m for ProtTucker, while the caption only describes total GPU wall-clock on one H100. In the benchmark context, these should be treated as reported embedding/evaluation times rather than training times; whether complete retrieval evaluation is included is insufficiently clear. They must not be conflated with the main model's one-week training cost.
Ablation Study¶
Table 2 compares vocabularies with the same 50M encoder and 8M head. Effective batch size is 1024; pretraining uses a stratified 20M UniRef90 subset, followed by fine-tuning on the full TEDLH data. Evaluation uses SCOPe at 10% identity and 3 random seeds. ROC-AUC and fold ROC are separate labels in the original table, and the cache does not sufficiently explain their aggregation difference; they are not merged here.
| Tokenizer | ROC-AUC | mAP | fold ROC | Pretraining time (mins) | Fine-tuning time (mins) | Total time (mins, summed) |
|---|---|---|---|---|---|---|
| ZEST (32K) | 0.739 ± 0.017 | 0.068 ± 0.002 | 0.749 ± 0.014 | 304.7 | 985.5 | 1290.2 |
| BPE (32K) | 0.729 ± 0.008 | 0.036 ± 0.002 | 0.743 ± 0.008 | 278.0 | 1030.0 | 1308.0 |
| Char (25) | 0.708 ± 0.011 | 0.059 ± 0.002 | 0.721 ± 0.010 | 650.7 | 1043.1 | 1693.8 |
ZEST and BPE each have approximately 16.4M embedding parameters, whereas Char has approximately 0.01M. The encoder and heads are matched, not the entire model parameter count. Total time is this note's sum of the two stages: ZEST has the lowest total, but BPE pretraining alone is faster. The result supports this restricted training configuration, not universal training-time superiority across data scales, and does not fully isolate the contributions of vocabulary priors, compression, and stochastic segmentation.
Key Findings¶
“Zero-shot” in Table 3 means frozen embeddings with no task-specific fine-tuning. Functional tasks still transfer annotations from labeled training neighbors through cosine kNN with K=15, so this is not annotation-free evaluation. Domain boundary detection uses sliding-window representation clustering, and circular permutation uses pairwise comparison; the tasks should not all be reduced to one classifier.
| Task | Metric | LEMON (200M) | ESM-2 (650M) | ProtTucker (3B) |
|---|---|---|---|---|
| Circular Permutation | AUROC | 0.779 | 0.476 | 0.720 |
| Circular Permutation | AUPRC | 0.285 | 0.081 | 0.238 |
| EC Number | F-max | 0.821 | 0.800 | 0.878 |
| GO Mol. Function | F-max | 0.639 | 0.603 | 0.635 |
| GO Mol. Function | AUPR | 0.625 | 0.584 | 0.616 |
| Domain Boundary | NDO | 0.758 | 0.771 | 0.750 |
| Domain Boundary | NRes | 0.637 | 0.658 | 0.674 |
- Circular permutation and some functional metrics are strong, but EC Number trails ProtTucker and domain-boundary NRes trails both baselines. The results do not support superiority over SOTA on every task.
- NDO measures overlap between predicted segments and true domains, while NRes measures residue assignment accuracy. Better domain-level coverage does not imply equally accurate residue-level boundaries.
- TTA uses 3 seeds and 5 segmentations for SCOP fold retrieval, with the text describing an effective moderate-dropout range of approximately 0.05–0.5. The cache does not provide all plotted values, so no precise improvement table is invented.
Highlights & Insights¶
- Input equivalence classes can carry inductive bias. The method does not merely compress strings: evolutionarily related fragments share IDs, changing which distinctions the model must preserve and which it can suppress.
- Segmentation boundaries can provide representation augmentation. Sampling valid shorter fragments gives multiple views of the same sequence and remains usable with frozen weights, but extra inference costs must be counted.
- Retrieval geometry and local fidelity benefit from separate supervision. Attention pooling supports global comparison and residue recovery constrains compression loss; neither metric is a substitute for all downstream quality.
Limitations & Future Work¶
- Sequence-only inference does not imply an absence of structure-derived training supervision. Fairer comparisons should separately report pretraining corpora, vocabulary sources, CATH labels, and inference modalities.
- The audit excludes identical training strings but retains some high-identity neighbors and many seen folds. Strict homology isolation and unseen-fold evaluation should be assessed separately rather than describing these results as completely leakage-free.
- The character baseline has far fewer embedding parameters, and the restricted ablation does not fully separate all mechanisms. Matching total parameters and reporting stochastic segmentation coverage distributions and multiseed confidence intervals would clarify the source of gains.
- Many-to-one tokens lose detail, and the recovery head is a learning constraint rather than a losslessness guarantee. The paper also acknowledges weaker fine-grained sequence-property performance than models trained at larger scales.
- Missing loss coefficients, full hyperparameters, main-table timing definitions, and some averaging definitions limit strict reproducibility. The supplied text provides no verifiable code repository URL, so no guessed Code metadata is included.
Related Work & Insights¶
- vs BPE / BPE-dropout: BPE merges by frequency and BPE-dropout randomly omits merges. ZEST constructs a vocabulary from conserved-fragment clusters and randomly selects shorter valid trie matches. Both alter granularity, but their equivalence classes have different origins.
- vs ProtTucker: Both use CATH contrastive supervision to shape retrieval geometry. LEMON additionally changes input units and achieves better fold-level results with a smaller model, while ProtTucker remains stronger at some finer levels and on EC metrics.
- vs SaProt / ProstT5: The structure-conditioned configurations compared here use 3Di information. LEMON puts domain priors into offline vocabulary construction and training supervision, allowing sequence-only query inference. This does not mean every use of those other methods necessarily requires structures.
- vs HHblits: Profile alignment explicitly uses MSA statistics, whereas LEMON compresses some conservation information into discrete inputs and learns global distances through contextual encoding and contrastive learning. Preprocessing costs and supervision conditions should not be ignored.
- The text calls reference [33] EvoBPE, whereas the bibliography title is PUMA: Discovery of Protein Units via Mutation-Aware Merging. This note preserves the naming discrepancy without asserting that the two are identical.
Rating¶
- Novelty: 4/5. Shared tokens from conserved-fragment clusters and stochastic valid-match segmentation form a clearly domain-specific combination.
- Experimental Thoroughness: 3/5. Multiple benchmarks and multiseed ablations are valuable, but parameter matching, data isolation, and averaging definitions remain incomplete.
- Writing Quality: 3/5. The main narrative is clear, but supervision claims, limiting segmentation behavior, and timing definitions need greater precision.
- Value: 4/5. Input-unit design can improve remote homology retrieval in smaller models, but this is not proof of a comprehensive replacement for large protein models.