Skip to content

๐Ÿฆพ LLM Agent

๐Ÿง  NeurIPS2026 ยท 4 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (30) ยท ๐Ÿ“ท CVPR2026 (42) ยท ๐Ÿ”ฌ ICLR2026 (162) ยท ๐Ÿ’ฌ ACL2026 (82) ยท ๐Ÿงช ICML2026 (59) ยท ๐Ÿค– AAAI2026 (33)

๐Ÿ”ฅ Top topics: LLM ร—2

AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering

AgentHop places 1,011 scientific four-option questions in a seven-tool sandbox with three resource caps and diagnoses failures through paper recall, conditional conversion, tool-call patterns, and resource failures; Gemini-3 Pro achieves 89.1% accuracy in the paper's evaluation, while similar overall scores can conceal different retrieval and synthesis bottlenecks.

ContractBench: Can LLM Agents Preserve Observation Contracts?

ContractBench uses a virtual clock, byte checks, and HTTP-trace validation to evaluate whether agents preserve cross-step contracts attached to tool artifacts; in the paper's experimental snapshot of 38 model variants, the highest success rate is 77.8%, and greater scale or newer versions do not guarantee improved reliability.

MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens

Without updating model weights, MoMHa uses a proposer to rewrite Python harnesses from execution traces, jointly optimizing accuracy, behavioral safety, and token cost; the authors report leading joint means on synthetic and real benchmarks, but source conflicts concerning metric scaling, the two-phase definition, and cost accounting require clarification.

ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

ToolSearcher trains a multi-turn tool selector using a category-constrained curriculum, event-level advantages for first discoveries of target tools, and progress-dependent credit assignment, improving Qwen2.5-7B-Instruct's overall StableToolBench F1 from GDPO's 0.496 to 0.513 and improving AppWorld task outcomes with a separate execution agent.