Skip to content

FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases

Conference: NeurIPS2026
arXiv: 2602.09163
Code: https://github.com/xingjian-zhang/flyaoc
Area: Computational Biology / Knowledge Base Curation
Keywords: Ontology curation, literature retrieval, agent evaluation, high-recall candidate generation, evidence recoverability

TL;DR

FlyAOC evaluates fruit-fly knowledge base curation end to end, from large-scale full-text retrieval to ontology-grounded candidate outputs, showing that multi-agent context partitioning improves recall of known recoverable annotations while semantic scores do not establish exact biological validity.

Background & Motivation

A biological knowledge base such as FlyBase is not a collection of paper summaries: it organizes gene functions, expression locations, and historical names into queryable structured records. Experts must locate literature, interpret experiments, distinguish the target gene from similarly named entities, and map authors' language to controlled vocabularies. Traditional named entity recognition or relation extraction benchmarks often supply relevant articles, bypassing the expensive questions of which papers to read and whether newly encountered names should change subsequent searches.

Extracting information from selected passages does not establish that a large language model can maintain a knowledge base. Evidence for a gene spans papers from different periods and naming conventions; retrieval failures, excessive context, and ontology-granularity errors can all create omissions in the final record. Evaluating against every FlyBase annotation would also be unfair, because the released corpus may lack the original evidence or contain only text referring to figures or supplements.

The paper uses existing expert annotations and accessible full text to construct a reproducible candidate-curation task with explicit evidence boundaries. It is not a biological discovery, wet-lab, or clinical decision system: its objective is to recover structured candidates for expert verification within a fixed review capacity. Core idea: first identify which existing expert labels are recoverable from the released text, then jointly evaluate retrieval, cross-paper extraction, ontology resolution, and candidate ranking so that model–harness failures appear in a common final recall metric.

Method

Overall Architecture

FlyAOC is a benchmark and evaluation protocol, not a network architecture that sequentially executes every baseline. Dataset construction combines FlyBase expert labels, linked literature, and ontology files into an evidence-filtered evaluation set; at inference time, a system receives a gene symbol, a FlyBase Gene Snapshot, the literature corpus, and ontology resources, and returns confidence-ranked structured annotations. Known answers used for label filtering belong to evaluation-set construction, not the agent's inference inputs; the paper does not train a new biological model.

Outputs cover three complementary tasks: Task 1 returns Gene Ontology (GO) identifiers, Task 2 returns anatomy–developmental-stage tuples, and Task 3 returns gene synonyms. The canonical primary metric for Task 2 evaluates anatomy only, not the exact correctness of complete tuples; stage identifiers remain in the output. Consequently, expression recall denotes anatomy-ontology coverage, not successful reconstruction of complete spatiotemporal expression records.

Retrieval-based systems share BM25 literature search, full-text access, GO/anatomy/stage ontology lookup, and output-schema validation tools. The four harnesses are parallel comparators: Memorization without corpus retrieval, a fixed Pipeline, an adaptive Single-Agent, and a Multi-Agent that delegates reading. They are therefore not drawn as a serial flowchart, which would incorrectly turn experimental comparators into components of one system.

Key Designs

1. Evidence-recoverability filtering: separate known facts from agent-visible evidence

FlyBase release FB2025_04 supplies expert target labels, but the corpus contains only 16,898 PMC open-access papers totaling over 140 million words. Coverage of post-2000 FlyBase references is 24%, or 16,665 / 69,404; accessible text does not constitute all the evidence available to experts. Dataset construction therefore first checks whether an annotation's curator-linked publication maps to a corpus article, then checks whether the released structured text actually supports the label. A reference match alone is insufficient: the paper may be present while a supporting microscopy image or essential supplement is absent from the text.

GPT-5 performs this second check for GO and expression labels. The verifier receives the target gene, one known expert label with ontology identifiers/evidence metadata, and a curator-linked article's title, abstract, and section text; it judges recoverability and returns a rationale and supporting span. For labels with multiple sources, OR aggregation retains a label if any cited article receives a supported verdict. Paraphrased or implicit evidence counts when a trained reader could reasonably derive the annotation, whereas tangential mentions and evidence outside the supplied text do not. GPT-5 does not search for new facts or create new biological annotations, and qualitative expert spot checks do not constitute an independent formal validation set.

Synonym support is lexical and is established by text matching rather than this label verifier. The abstract's 7,397 expert annotations are not the primary evaluation denominators: filtering and gene-specific deduplication yield 770 GO items, 252 anatomy items, and 457 synonym strings. This separates absent evidence from failure to locate evidence, but model-based filtering remains fallible and should not be described as noise-free ground truth.

Gene selection also defines the evaluation population: 3,999 candidates with Gene Snapshots become 3,446 after requiring at least 10 corpus full-text papers, then 314 after requiring recoverable function and expression labels, with 100 finally selected. The final ranking favors balanced annotation coverage across tasks rather than genes richly annotated in only one task. This gives every test gene evaluable evidence while favoring well-studied genes, so it does not represent open-world curation difficulty across all fruit-fly genes.

2. Bounded candidates and semantic recall: measure coverage without implying factual accuracy

Systems must rank candidates because expert review capacity is limited. Canonical cutoffs are GO@30, anatomy@10, and synonyms@20, leaving a relatively broad candidate space based on average recoverable-fact counts per gene. Function and anatomy use Wang ontology semantic similarity: exact matches receive full credit, while parents or related terms receive partial credit based on ontology distance; synonyms use case-insensitive exact string matching.

For each target label, the evaluator takes its maximum similarity to the gene's top-ranked candidates and normalizes by the number of ground-truth labels. The following combines the per-gene definition with canonical micro-aggregation; it is equivalent to the paper's metric, not an additional training loss:

\[ R_{\mathrm{micro}}@k= \frac{\sum_i\sum_{g\in\mathrm{GT}_i}\max_{p\in\mathrm{Pred}_{i,1:k}}\operatorname{sim}(g,p)} {\sum_i|\mathrm{GT}_i|}. \]

Here \(\operatorname{sim}(g,p)\in[0,1]\); for synonyms, it is replaced by an indicator of equality after case normalization. Failed or empty-output runs contribute zero numerator while retaining their denominators; failed genes must not be removed before reporting successful runs. Micro-aggregation weights richly annotated genes more heavily, whereas diagnostic macro-aggregation gives each gene equal weight.

A prediction can receive relatedness credit against multiple targets, so high semantic recall does not mean that every target has been recovered by a distinct, exact candidate. The appendix also reports semantic precision/F1, but the reference knowledge base does not exhaust all correct annotations; candidates absent from it still require human inspection of original evidence. These metrics suit high-recall candidate generation, not direct automatic updates to a biological database.

3. Harness comparison: contrast feedback capability with context fidelity

Memorization does not access the literature corpus but retains ontology lookup tools and feedback mechanisms similar to Single-Agent. It measures label recovery from parametric knowledge rather than serving as a tool-free closed-book baseline; evidence records use a null PMCID with explanatory text. Gains from retrieval-based systems therefore more closely represent the added value of literature access, not simply tools versus no tools.

Pipeline follows five fixed stages: search, retrieve top-ranked papers, extract natural-language descriptions per paper, batch-resolve ontology identifiers, and compile outputs. Paper processing can run in parallel with partitioned reading contexts at low cost, but newly discovered historical names do not trigger renewed retrieval, and incorrect mappings lack an adaptive correction loop. Extracting descriptions before uniform resolution supports mechanical execution but can accumulate many weakly relevant or incorrectly granular candidates.

Single-Agent uses a ReAct-style loop to select tools, reformulate queries, revisit ontology lookups, and submit results. It can feed aliases, interaction partners, or ortholog names encountered during reading back into search, but all raw article text accumulates in a single context. This preserves details and enables direct cross-paper reasoning at the cost of context growth, attention dilution, and higher API expenditure.

Multi-Agent delegates reading through an orchestrator; subagents return extracted annotations with resolved ontology identifiers rather than all raw article text. The orchestrator deduplicates, resolves conflicts, filters noise, and combines evidence across papers while retaining the ability to arrange further tool calls. Partitioning reduces the main context burden but may discard qualifiers or local evidence, so higher aggregate recall can coexist with losses on individual annotations. The paper does not isolate context dilution, retrieval mix, and tool failures through single-factor interventions; these explanations are diagnoses consistent with observations, not independently established causal findings.

4. Budgets and missing ontology terms: control retrieval opportunities and probe normalization limits

Standard experiments permit 1 / 2 / 4 / 8 / 16 papers per gene, with the budget stated in prompts and enforced through tool calls. Reference coverage measures the fraction of unique curator-cited sources for retained labels that the system reads; it does not measure all relevant literature or complete biological knowledge. This diagnostic can diverge from final scores when a system retrieves a source but misses a label or substitutes a nearby ontology term for an exact one.

The missing-term experiment hides rare leaf terms occurring in only one corpus paper among recoverable GO gene–term pairs. Ontology search filters these identifiers, and agents cannot submit hidden IDs: they must select visible alternatives or produce natural-language descriptions; this is a proxy task for novel concepts, not evidence that the underlying facts are genuinely new. A post-hoc resolver can access the full ontology to assess proximity to hidden targets, but that access must not be confused with the agent's inference-time capabilities.

Loss & Training

The paper introduces no optimization loss or fine-tuning of its baselines; performance comes from existing models combined with prompts, tools, orchestration, and context management. Temperature is set to 1.0, execution is capped at 50 turns per gene, and model-dependent cost limits are 1–10 USD; other settings follow provider defaults. Exact reproduction of paper tables relies on frozen normalized predictions and fixed evaluation files; rerunning APIs reproduces the procedure, not necessarily historical outputs under current provider behavior.

Key Experimental Results

Main Results

The following reports GPT-5-mini micro-averaged recall over recoverable labels, in percentages; retrieval-based harnesses use the same 16-paper budget per gene, whereas Memorization uses 0 papers. Values come from appendix Table 7: GO and anatomy award semantic partial credit, Syn uses exact matching, and the last column explicitly retains GO precision to avoid interpreting candidate recall as factual accuracy.

Harness GO R@30 Anatomy R@10 Syn R@20 GO precision@30
Memorization 49.6 39.7 37.4 48.5
Pipeline 48.9 41.9 40.5 26.4
Single-Agent 47.6 48.9 49.2 47.4
Multi-Agent 53.6 58.5 57.3 44.0

Multi-Agent's recall lead does not extend to every dimension: its GO precision is below Memorization and Single-Agent. Pipeline approaches other harnesses' GO recall under the larger cutoff but produces more weak candidates, so similar recall does not establish equivalent curation quality.

The following model comparison fixes Multi-Agent and a 16-paper budget per gene, using main-text Table 4. Recall is in percentages, and ± denotes the half-width of a 95% gene-bootstrap interval, not a standard deviation or uncertainty from repeated API calls; cost, tokens, and candidate counts are per-gene averages.

Model Cost USD Tokens Candidates GO Anatomy Syn Average
GPT-5-mini 0.40 1.3M 23 53.6±3.1 58.5±4.4 57.3±6.6 56.5±3.4
GPT-5 5.89 1.4M 29 57.1±4.6 60.9±7.4 53.8±8.0 57.3±5.8
Claude Sonnet 4.6 15.81 5.0M 30 59.7±5.7 58.4±5.7 46.4±6.0 54.8±4.2

Close estimates with overlapping intervals do not establish one unequivocal model winner. GPT-5-mini stands out for cost-adjusted average performance and its synonym point estimate, not superiority over larger models on every task. Some interval half-widths differ between main-text Table 4 and appendix Table 13; this table preserves the main-text version rather than reconciling the source's numbers.

Ablation Study

The following summarizes appendix setting analyses without combining different settings into one causal ablation. Standard budget extension uses canonical primary metrics; the hidden-term diagnostic retains its separate GO@20 and must not be treated as directly equivalent to standard GO@30.

Analysis setting GO metric Anatomy R@10 Syn R@20 Boundary
Multi-Agent, budget 16 R@30 = 53.6% 58.5% 57.3% Extended-budget comparison
Multi-Agent, budget 32 R@30 = 45.0% 50.7% 48.1% More reading does not guarantee better recall
Available GO terms, 646 pairs R@20 = 55.4% Not evaluated Not evaluated Separate hidden-term setting
Hidden GO terms, 124 pairs R@20 = 48.4% Not evaluated Not evaluated Rare-leaf proxy task

The hidden-term experiment generates 49 natural-language descriptions, of which 44 resolve to some GO term; mean similarity to hidden targets is only 0.19, with just 6/44 exact matches. Mapping to an ontology is not equivalent to recovering the hidden target, much less obtaining expert acceptance of a new ontology term. This experiment uses a separate hidden-label configuration and is not part of the public standard no-API table reproduction path.

Key Findings

  • At 16 papers per gene, Multi-Agent, Single-Agent, and Pipeline cost approximately 0.4, 0.85, and 0.10 USD. Partitioned reading improves cost effectiveness in this setting, not free scaling at arbitrary budgets.
  • Function, expression, and synonym tasks depend differently on literature, and historical names can change the retrieval space. Evaluating a retriever's ranking without final extraction misses this feedback value.
  • Main-text Table 3 uses diagnostic denominators different from canonical ones: GO has 184 + 372 + 389 = 945, and expression has 86 + 165 + 83 = 334, matching appendix retained row-level counts rather than deduplicated 770 / 252.
  • The synonym win/draw/lose row totals 37 + 108 + 745 = 890, also differing from the canonical deduplicated denominator of 457; the table does not fully explain this change of counting unit, so these values must not be used to reconstruct primary metrics or silently corrected.
  • Tool reliability affects scores. DeepSeek V3.2 has failed or empty outputs on 36/100, 94/100, and 40/100 genes for Memorization, Single-Agent, and Multi-Agent, respectively; denominators remain and failures receive zero credit, so low scores cannot be attributed solely to insufficient biological knowledge.

Highlights & Insights

  • Recoverability filtering separates unavailable evidence from retrieval/extraction failure. Transferring this design to another specialized knowledge base requires auditing source accessibility and evidence modalities before using all database labels as full-text task answers.
  • Historical names serve simultaneously as output targets and retrieval clues, genuinely coupling retrieval with extraction. Systems should feed reliable extracted names back into queries instead of summarizing once after a fixed search.
  • Agent compression does more than reduce tokens: it changes how much literature the main agent can track. Better intermediate representations should retain evidence spans, qualifiers, and provenance to support selective rereading rather than returning identifiers alone.
  • Separating frozen predictions from API reruns is a transferable benchmark-release practice. The former audits evaluation and numerical results; the latter studies current model behavior, addressing a different reproducibility question.

Limitations & Future Work

  • Both corpus and evaluation are text-limited, excluding evidence from images, layout, supplements, and inaccessible articles. High scores do not establish complete coverage of fruit-fly biology or experts' real evidence sets.
  • Test genes favor substantial prior research and balanced annotations, while FlyBase itself has curation lag. Existing denominators cannot fully evaluate discovery for low-resource genes or correct candidates absent from the database.
  • GPT-5 recoverability judgments receive qualitative expert spot checks but lack a formal validation set, error rate, or quantified agreement. Independent annotation of recoverability and evidence support would prevent filtering errors from becoming hidden ground-truth noise.
  • Semantic partial credit can conceal granularity errors, and anatomy-only primary evaluation ignores stage correctness. Exact matching, complete-tuple evaluation, and candidate–evidence fidelity should supplement recall rather than merely increasing candidate volume.
  • Hiding rare leaf terms simulates missing ontology support; it does not show discovery of new biological facts or creation of valid ontology terms. Genuine ontology extension requires expert review, concept-boundary checks, and evidence verification.
  • Gene-bootstrap intervals do not cover provider stochasticity, version drift, or between-run variation. Close point estimates and descriptive appendix analyses do not support strong causal or significance claims.
  • PMC-OA licenses are heterogeneous, not uniformly CC-BY; downstream use must inspect article-level licenses and provenance. Open downloading does not imply uniform rights for commercial redistribution.
  • vs BC4GO / CRAFT / GOTA: These resources evaluate entities, ontology terms, or annotation in supplied literature; FlyAOC starts with a gene query and evaluates document selection itself. It better captures end-to-end workflows but makes error sources harder to isolate.
  • vs PaSa: PaSa emphasizes academic literature retrieval, whereas FlyAOC passes retrieved articles onward to ontology-grounded extraction and aggregation. A system that finds the right article but maps evidence at the wrong granularity cannot be considered successful merely for retrieval.
  • vs SPIRES / CurateGPT / STRUCTSENSE: These works supply structured extraction or assisted-curation mechanisms; FlyAOC supplies a comparison platform spanning a full-text corpus, common outputs, and expert target labels. It neither replaces human curation systems nor requires new methods to adopt the baseline harnesses.
  • vs deep-research question answering and report benchmarks: Single-answer tasks permit stopping after finding an answer, while report completeness is difficult to quantify. FlyAOC tests continued evidence collection through bounded structured candidate coverage, but expert verification remains necessary to establish factual validity.

Rating

  • Novelty: High; jointly evaluates retrieval, extraction, and ontology grounding in a specialized knowledge base workflow.
  • Experimental Thoroughness: High within scope; covers harnesses, models, budgets, and hidden terms, but lacks formal recoverability validation and repeated runs.
  • Writing Quality: Good; tasks and candidate use are clear, although some appendix counting units and interval half-widths require clarification.
  • Value: High; useful for candidate generation before human review, not a justification for direct automatic updates to biological databases.