AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering¶
Conference: NeurIPS 2026 (author-reported acceptance to the Evaluation and Datasets track; not independently verified for this note)
arXiv: 2609.34428
Code: https://github.com/JohnnyNLP/agenthop
Data: https://huggingface.co/datasets/nlpai-lab/agenthop
Area: LLM Agent
Keywords: scientific multi-hop question answering, diagnostic evaluation, citation-graph navigation, evidence synthesis, resource constraints
TL;DR¶
AgentHop places 1,011 scientific four-option questions in a seven-tool sandbox with three resource caps and diagnoses failures through paper recall, conditional conversion, tool-call patterns, and resource failures; Gemini-3 Pro achieves 89.1% accuracy in the paper's evaluation, while similar overall scores can conceal different retrieval and synthesis bottlenecks.
Background & Motivation¶
Scientific question answering involves more than finding and restating a text passage: researchers often start from a paper, follow citations to related methods, read experimental sections, and connect findings across papers. A language agent must navigate, read content, synthesize evidence, and decide when to answer within one trajectory. Final correctness alone cannot distinguish failure to find the target paper, reading irrelevant sections, misunderstanding evidence already obtained, and exploring until resources run out.
Existing evaluations cover different aspects of this process. HotpotQA and 2WikiMultiHopQA emphasize multi-hop question answering, OpenScholar emphasizes scientific synthesis, BFCL focuses on function-call correctness, and AgentBoard provides trajectory-progress diagnostics; their settings and reporting do not jointly offer the four-axis view sought here. Scientific papers' section structure and citation graph enable that view: questions can specify gold papers and evidence locations, tool logs can identify the sections actually delivered, and resource status can be disclosed at every step.
AgentHop is therefore an evaluation instrument rather than a new agent algorithm: it measures finding evidence, using evidence, organizing actions, and controlling expenditure together. Its controlled sandbox sacrifices open-web realism but reduces environmental variation, giving models the same questions and budgets. Core Idea: combine scientific questions with paper- and section-level evidence labels and complete tool trajectories to turn accuracy into inspectable conditional diagnostics, rather than treating every failure as an undifferentiated capability deficit.
Method¶
Overall Architecture¶
Each item provides a system prompt, one seed-paper ID, a research question, and four options AโD; the gold papers are not supplied. The agent searches, follows citations, inspects sections, and reads content before submitting an option through submit_answer. The evaluation retains the answer, paper text actually received, tool sequence, and resource consumption, allowing four diagnostic views from the same run.
The dataset contains 7,205 papers and over 450K citation edges. Seeds come from nine computer-science venues during 2022โ2025: NeurIPS, ICML, ICLR, ACL, EMNLP, NAACL, CVPR, ECCV, and SIGIR. This does not establish that every target in the citation neighborhoods belongs to those years and venues; the appendix example includes earlier work such as WebGPT. Two-hop neighborhoods contain 311 papers on average, with median 315 and range 82โ667; the median number with readable section text is 137, with range 8โ274.
Items cross two distinctions. Single-target (ST) items require one gold paper, whereas multi-target (MT) items require two; depth d1 means a direct citation from the seed, and d2 introduces one bridge paper along the path. Counts are ST/d1 90, ST/d2 478, MT/d1 396, and MT/d2 47, totaling ST 568 and MT 443. These are citation-path depth and target-multiplicity designations, not actual tool-call turn counts; ST/d2 does not itself require synthesizing two gold papers.
There is no trainable network to diagram. The following explanation instead covers question construction, constrained interaction, metric definitions, and trajectory analysis.
Key Designs¶
1. Evidence-constrained question construction: distractors correspond to different information sources
The eight-stage pipeline expands 953 seed papers into 17,983 candidate citation chains, produces 1,651 questionโanswer pairs, and adds three distractors to each. Chain screening requires adjacent papers' shared-reference Jaccard score to be at least 0.08, substantive methodology- or result-related citation contexts of at least 80 characters, and a GPT-5.4 rating of at least 4 on a 1โ5 scale for whether the chain motivates further inquiry. The generator sees seed and target full text plus citation metadata but must not explicitly name the target title or method in the question, preventing evidence discovery from collapsing into keyword matching. Distractors come from answering with only the seed, with a non-gold neighboring paper, or with parametric knowledge and no paper context. The generator attempts to answer sincerely from the wrong or absent source rather than writing an obviously false statement. This gives incorrect options some diagnostic meaning, but does not uniquely identify the mechanism behind every wrong choice.
An ensemble of GPT-4.1, Claude Sonnet 4.6, and deepseek-chat subsequently removes seed-answerable shortcuts and cases lacking answer consensus even with full evidence, eliminating 502 items. Structural checks remove another 71 and ensure parsed body text for every retained seed, bridge, and gold target. Seven graduate students from the research group provide single-auditor coverage of the remaining 1,078 items, checking questions, answer support, and evidence-section labels; 116 are flagged, and lead-author triage removes 35. Suspicious evidence labels are then repaired and five irrecoverable items removed; finally, a p99.9 single-section length threshold of 70,015 characters removes 27 parser-collapse cases, leaving 1,011 items. Human review prioritizes single-pass coverage rather than designed independent double coding, and GPT-5.4's audit guide may create anchoring. A substantial curation pipeline does not substitute for measured inter-rater agreement.
2. Seven tools and three resource caps: selective reading and timely submission become part of the task
The literature tools are get_paper_info for metadata such as titles and abstracts, get_references for citations and their contexts, list_sections for section names, alphabet aliases, and character counts, and search_papers for up to 10 matching papers' titles and abstracts; each costs one point. read_section returns a full section and costs five points, accepting a full header, alphabet alias, or unique case-insensitive substring. Only this tool delivers body-text evidence credited by the evaluation; abstracts can assist navigation but do not satisfy paper recall. think preserves a deliberation note in the trajectory without changing the environment, and submit_answer terminates with an answer; both cost zero tool points, but still consume turns or tokens. The main text calls the keyword tool keyword_search, whereas the catalogue and ablations use search_papers; this note follows the latter and preserves the naming discrepancy.
Each item allows at most 20 assistant turns, 200K cumulative input-plus-output tokens, and 30 tool-budget points. Cumulative tokens sum usage across turns rather than measuring the context-window occupancy at one instant; history repeatedly included in inputs also contributes. Reading costs a fixed five points irrespective of text length, so points and tokens constrain different expenditures. Even reading alone could support at most six calls, and navigation further reduces that allowance. Metadata-only neighborhood nodes can support citation traversal but lack readable body text; inspecting sections identifies that state before spending scarce reading opportunities. Every tool result discloses the current turn, cumulative tokens, and budget used, so the test measures resource-aware planning under disclosed constraints.
A call rejected for insufficient budget costs nothing and delivers no content. After the first rejection, the prompt instructs the model to submit on the next turn using available evidence; a second consecutive budget rejection terminates the trajectory as budget failure, while a successful call resets the rejection counter. Thus the main text's one-turn retry does not replenish the budget or grant a free read; zero remaining points still permit free deliberation or submission. Cumulative tokens are checked after every assistant response, with immediate termination on overflow; failure to submit within 20 turns triggers max_turns. Unknown tools, missing arguments, and malformed JSON return tool errors rather than restarting the sample. Hard API errors permit at most two sample retries, while transient errors retry in place without consuming turns.
3. Four-axis metrics: reaching a paper is different from receiving answer evidence
The search axis uses binary per-item paper recall: at least one section of every gold paper must have been successfully delivered through read_section. Reading only one MT target, receiving metadata alone, rejected calls, or malformed arguments do not satisfy recall. Section recall is stricter: each gold paper must have a delivered section containing a direct-labelled evidence location; supporting labels provide interpretive context but do not replace direct evidence. Human labels may refer to native subsections while tools return top-level sections, so evaluation maps subsection labels to containing readable sections. This avoids withholding credit merely because a delivered parent contains the labelled subsection under a different header. Across 11,072 recall-credited trajectories pooled over 19 models, the appendix reports that 85.7% received direct-labelled evidence and 14.3% did not. Paper recall is therefore not synonymous with complete evidence recall.
The synthesis axis, conversion, is the fraction correct among items satisfying paper recall. The tool-use axis, calls/turn, is the mean number of tool calls emitted per assistant turn: values above one imply some same-turn batching, but higher values are not inherently better and do not prove parallel server execution. The resource axis reports mean tokens, turns, and points plus rates of hitting each cap before submission; accuracy counts non-submissions as incorrect. With \(r_i\) indicating paper recall and \(c_i\) indicating correctness, the central conditional diagnostic is:
This conditional quantity does not causally isolate reasoning. Models recall different item subsets, and differing proportions of easy items can alter conversion; moreover, recall requires only any section from a gold paper. Without recall, a model can still answer using prior knowledge, abstracts, option elimination, or guessing, so accuracy cannot be written as recall multiplied by conversion. The event decomposition is instead:
To compare synthesis performance, the paper applies exact McNemar tests on the same items recalled by both models, described as cap-corrected in the appendix. For example, Gemini-3 Pro versus Sonnet on MT uses 226 common items with discordant pairs 20/3 and two-sided \(p=0.0005\); Sonnet versus GPT-5.4 uses 193 common items and yields \(p=0.27\), so their separate conditional proportions do not establish a reliable ordering. Common-item pairing reduces item-difficulty confounding, but does not remove all recall-selection, interface, and reasoning-configuration confounds or extend conclusions to unrecalled items.
4. Conditional trajectory analysis: stopping policies and tool transitions explain bottlenecks
Trajectory analysis examines actions after the first gold-paper read, after reading both MT gold papers, and after a wrong-paper read. Failed MT recall is subdivided into reaching no gold paper, stopping prematurely after one, continuing exploration without finding the second, and staying within the already-read gold paper. This distinguishes navigation from stopping-policy problems more clearly than an aggregate MT recall drop. Actions after ST recall or reading both MT targets are categorized as immediate submission, deliberation, or further reading, with a small navigation-only residual. Bigram analysis compares accuracy between trajectories containing and not containing a consecutive tool pair; trigram analysis records frequent three-call patterns and their share of all trigrams, describing action habits rather than introducing another success metric.
Interpretation must remain conservative. Better outcomes after think can reflect stronger models using it more often or easier trajectories having resources left; they do not prove that inserting the tool improves accuracy. The appendix applies Bonferroni correction over 49 candidate bigrams, but the associations remain observational. Claimed family fingerprints describe a few tested checkpoints through their native serving interfaces: same-turn tool support, parsers, hidden reasoning, and defaults differ. Two checkpoints' habits cannot establish a permanent training mechanism for an entire provider, much less identify causal architectural deficits.
Loss & Training¶
The paper does not train models and introduces no loss function. Question generation and ensemble filtering are dataset-construction procedures, not fine-tuning of the 19 evaluated models.
Task prompts, tools, and budgets are standardized, but reasoning configurations are not fully matched: Opus uses extended thinking with an 8,192-token budget plus external think, Sonnet uses external think, and GPT models use medium reasoning effort; service defaults and local parsers also vary.
Model names denote the checkpoints reported in the paper, not current official availability or live rankings. Appendix Table 7 states that deepseek-chat and deepseek-reasoner both resolved to DeepSeek-V3.2 during evaluation, invoking non-thinking and thinking modes respectively. Older labels in Tables 8 and 15 must not be read as evidence that these were separate V3 and R1 generations.
Key Experimental Results¶
Main Results¶
The following selection comes from main-text Table 2. Recall and conversion are proportions, tokens are in K, and accuracy covers all 1,011 items. ST and MT conversion use different conditional denominators and should not be averaged unconditionally into a single synthesis-capability score.
| Checkpoint tested in the paper | Paper recall ST / MT | Conversion ST / MT | Calls/turn | Mean tokens (K) | Accuracy |
|---|---|---|---|---|---|
| Gemini-3 Pro | 0.868 / 0.752 | 0.957 / 0.970 | 1.05 | 92.4 | 89.1% |
| GPT-5.3 codex | 0.812 / 0.623 | 0.896 / 0.902 | 1.44 | 69.5 | 82.5% |
| Claude Opus 4.6 | 0.873 / 0.777 | 0.871 / 0.855 | 1.25 | 110.5 | 77.5% |
| Claude Sonnet 4.6 | 0.798 / 0.600 | 0.936 / 0.891 | 1.32 | 95.8 | 76.6% |
| GLM-5.1 | 0.768 / 0.673 | 0.830 / 0.768 | 1.20 | 102.0 | 69.6% |
| Gemma-4-31B | 0.621 / 0.521 | 0.788 / 0.861 | 1.01 | 87.6 | 64.8% |
| Kimi-K2.5 | 0.764 / 0.569 | 0.613 / 0.504 | 1.08 | 115.5 | 44.1% |
Opus and Sonnet differ by just 0.9 accuracy percentage points while showing higher recall and higher conditional conversion, respectively. Gemma-4-31B's MT conversion is 0.861 but MT recall is only 0.521, motivating a test of navigation improvements; this does not establish frontier-level synthesis over all items.
Ablation Study¶
Table 4 evaluates three models under the same budgets. Removing keyword search retains citation navigation, section reading, deliberation, and submission; it does not supply gold papers directly. Closed-book evaluation instead uses no tools and a single-turn answer from the question and four options.
| Checkpoint | Closed-book | No keyword search | Full seven tools | Gain from citation navigation and other tools | Gain from restoring keyword search |
|---|---|---|---|---|---|
| DeepSeek-chat | 18.3% | 53.3% | 56.4% | +35.0 pp | +3.1 pp |
| Gemma-4-31B | 31.7% | 54.8% | 64.8% | +23.1 pp | +10.0 pp |
| GPT-5.3 codex | 48.6% | 79.3% | 82.5% | +30.7 pp | +3.2 pp |
Citation following and content access provide most of the tool-conditioned gain on these models, but keyword search still adds 10.0 pp for Gemma. Calling keyword search dispensable cannot be generalized to all models or tasks. Table 18 also reports Gemini-3 Flash dropping from 57.3% closed-book to 54.7% with full tools, showing that agent-loop behaviors such as non-submission can reduce performance.
Key Findings¶
- Resource caps discriminate between behaviors: Table 2 gives token-overflow rates of 27.7% for Gemini-3 Flash and 37.0% for Gemma-4-26B-A4B. Low token usage can instead indicate early abandonment, not efficiency.
- Across three full runs, accuracy is 0.564/0.553/0.542 for DeepSeek-chat, 0.825/0.818/0.810 for GPT-5.3 codex, and 0.648/0.585/0.581 for Gemma-4-31B. Table 17 gives standard deviations 0.009/0.006/0.031 respectively; not all models vary within one percentage point.
- MT/d2 contains only 47 items: Table 10 reports DeepSeek-chat at 48.9% with a Wilson 95% confidence-interval half-width of 13.8 pp. Small cells do not support stable fine-grained rankings.
- The reported Cronbach's \(\alpha=0.997\) uses only 19 models as respondents and 1,011 questions as items. High internal consistency does not independently establish four-axis construct validity or cross-domain generalization.
Highlights & Insights¶
- Recording both what content arrived and what happened next localizes failures to concrete trajectories. Distinguishing paper recall from section-level evidence contact is especially useful for literature RAG evaluation.
- Separate tool points and cumulative tokens reveal that fixed call cost and text-length cost are different constraints. Disclosed resource state makes budget awareness an evaluated behavior.
- Similar accuracy need not imply similar improvement directions. Controlled interventions on navigation, evidence synthesis, and stopping policies can test diagnostic usefulness more directly than simply selecting the higher-scoring model.
Limitations & Future Work¶
- Closed-book success can reflect literature exposure during pretraining, four-option cues, or guessing. Table 19's absence of monotonic citation-count or year gradients does not rule out contamination; correct answers without recall do not establish successful evidence synthesis.
- GPT-5.4 generates questions, distractors, and audit guides. On a 94-item subset rebuilt using Inkling, GPT-5.4 scores 78.7% with a 95% interval of 69.4โ85.8%, versus 79.0% on the original set. Similar small-sample estimates neither prove absence of generator bias nor support four-axis comparison.
- The corpus focuses on computer science, answers are four-option selections, and gold paths have guaranteed body text. Incomplete evidence, open-ended argumentation, erroneous citations, and other disciplines are not covered; single-auditor review lacks independent double-coded agreement estimates.
- The paper contains unresolved reporting discrepancies: bigram effects are +9.8/+10.6/โ6.8 pp in the main text but +13.6/+12.6/โ11.5 pp in Table 13; some termination rates differ between Table 2 and cap-corrected Table 15, including Flash token failures at 27.7% versus 30.4%. These values are not mixed or silently corrected here.
- Appendix G describes Kimi's SayCan read as budget-rejected but credits paper recall as 1, apparently conflicting with the requirement that a section from every gold paper be successfully delivered. The cached narrative does not establish whether another successful read occurred; this example is not treated as a verified recall case.
- Future work should randomize tool policies on paired items, standardize or stratify serving interfaces, add cross-domain open-ended questions, and directly test whether prescribed changes improve the intended axis rather than converting observational associations into prescriptions.
Related Work & Insights¶
- vs AgenticRAGTracer: Chain-geometry diagnostics identify which hop fails, while AgentHop adds conditional answering after paper recall and budget termination. The views are complementary, but conditional conversion remains a confounded reasoning diagnostic.
- vs OpenScholar / HotpotQA: All involve evidence retrieval or multi-document reasoning. AgentHop emphasizes seed-based navigation, section-delivery logs, and constrained tool protocols; its contribution does not require claiming that prior question-answering benchmarks never evaluate retrieval.
- vs BFCL / AgentBoard / ReAct: Function-call correctness, progress diagnostics, and actionโobservationโdeliberation offer complementary perspectives. AgentHop supplies behavioral statistics relevant to those views, not a new ReAct training method.
Rating¶
- Novelty: 4/5 โ Four-axis trajectory diagnostics combined with scientific citation navigation offer more explanation than a single score.
- Experimental Thoroughness: 3/5 โ Nineteen models, paired tests, and tool ablations provide breadth, but interface confounds, audit design, and small subsets remain limiting.
- Writing Quality: 3/5 โ Metrics and appendices are detailed, with unresolved reporting, naming, and example inconsistencies.
- Value: 4/5 โ Useful for controlled retrieval-agent diagnosis, not a cross-domain capability measure or provider ranking.