๐ฆพ LLM Agent¶
๐ง NeurIPS2026 ยท 4 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (30) ยท ๐ท CVPR2026 (42) ยท ๐ฌ ICLR2026 (162) ยท ๐ฌ ACL2026 (82) ยท ๐งช ICML2026 (59) ยท ๐ค AAAI2026 (33)
๐ฅ Top topics: LLM ร2
- AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering
-
AgentHop places 1,011 scientific four-option questions in a seven-tool sandbox with three resource caps and diagnoses failures through paper recall, conditional conversion, tool-call patterns, and resource failures; Gemini-3 Pro achieves 89.1% accuracy in the paper's evaluation, while similar overall scores can conceal different retrieval and synthesis bottlenecks.
- ContractBench: Can LLM Agents Preserve Observation Contracts?
-
ContractBench uses a virtual clock, byte checks, and HTTP-trace validation to evaluate whether agents preserve cross-step contracts attached to tool artifacts; in the paper's experimental snapshot of 38 model variants, the highest success rate is 77.8%, and greater scale or newer versions do not guarantee improved reliability.
- MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens
-
Without updating model weights, MoMHa uses a proposer to rewrite Python harnesses from execution traces, jointly optimizing accuracy, behavioral safety, and token cost; the authors report leading joint means on synthetic and real benchmarks, but source conflicts concerning metric scaling, the two-phase definition, and cost accounting require clarification.
- ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning
-
ToolSearcher trains a multi-turn tool selector using a category-constrained curriculum, event-level advantages for first discoveries of target tools, and progress-dependent credit assignment, improving Qwen2.5-7B-Instruct's overall StableToolBench F1 from GDPO's 0.496 to 0.513 and improving AppWorld task outcomes with a separate execution agent.