ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning¶
Conference: NeurIPS2026
arXiv: 2609.30906
Code: https://github.com/zhenlongDai/ToolSearcher
Area: LLM Agent
Keywords: large-scale tool selection, multi-turn retrieval, reinforcement learning, event-level advantage, trajectory credit assignment
TL;DR¶
ToolSearcher trains a multi-turn tool selector using a category-constrained curriculum, event-level advantages for first discoveries of target tools, and progress-dependent credit assignment, improving Qwen2.5-7B-Instruct's overall StableToolBench F1 from GDPO's 0.496 to 0.513 and improving AppWorld task outcomes with a separate execution agent.
Background & Motivation¶
Tool use involves more than supplying arguments to a known function. With approximately 16k tools across 49 categories, a model must first find a set that meets the request, while the full documentation cannot fit into its context. Retrieved candidates may have similar names and functions but different input constraints or output structures; semantic relevance alone does not ensure that one tool can consume another's output. Tool selection therefore requires maintaining a plan, reading interfaces, and revising candidate combinations through multiple retrieval rounds.
RAG retrieves documents in a single pass, while methods such as Search-R1 learn to interleave reasoning and retrieval. However, finding evidence for a knowledge question differs from finding and distinguishing composable interfaces. A single final reward does not tell the model whether it failed to retrieve a necessary tool or retrieved everything but selected incorrectly. Repeatedly finding a familiar tool can also receive the same signal as first discovering a missing one. High search coverage does not imply good final selection: the untrained 7B multi-turn model achieves SRecall of 0.724 under global search but final F1 of only 0.098.
The paper does not train a general-purpose agent to execute all business APIs; it optimizes upstream search and selection as a distinct task. Core Idea: first train fine-grained discrimination among same-category candidates, then reward only search events that first discover target tools, and enable final-selection optimization only after a trajectory has retrieved every target tool.
Method¶
Overall Architecture¶
The inputs are a user request, a tool repository, and the search engine's calling schema; the output is a set of tool identifiers, not an executed program. Each tool document describes functionality, input constraints, and output structure. The policy generates structured search calls, receives documents from the search engine, continues planning and retrieval, and finally submits a tool set.
Three designs work together during training: category-constrained tool discrimination changes the initial search environment; event-level search modeling assigns advantages to steps that add target coverage; trajectory-aligned credit allocation disables search or selection rewards that are inappropriate for the current progress. The first two are not additional external models, and the third is not a judge required at deployment.
Target tool IDs come from training labels and are used only to identify retrieval hits, first discoveries, and exact final-set matches. At inference time, the model has no target IDs and must rely on the request and retrieved documents; the supervision branch in the diagram must not be interpreted as an inference input.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Request + search schema"] -->|Early training| B["Category-constrained<br/>tool discrimination"]
B --> C["Event-level<br/>search modeling"]
A -->|Inference: global search| C
C -->|Generate final selection| D["Tool set"]
C -.->|Training trajectories| E["Trajectory-aligned<br/>credit allocation"]
D -.->|Training outcomes| E
T["Target tool IDs<br/>Training supervision only"] -.-> C
T -.-> E
E -.->|Training updates| C
Key Designs¶
1. Category-constrained tool discrimination: learn to read interfaces among similar candidates
If early training searches only the full repository, coarse topical queries can avoid many irrelevant tools without teaching the model to distinguish close neighbors within a functional domain. Category-Constrained Tool Discrimination (CCTD) restricts early retrieval to a specified category. The search schema adds a category argument, making returned candidates more similar and forcing the model to compare functionality and interfaces rather than accept topical relevance alone.
This is a curriculum, not permanent category isolation. The paper uses category-specific search for the first 30% of training data and then switches to global search to learn fine-grained queries across the repository. The appendix specifies 56 training steps, with category constraints in the first 17. The category-specific environment is itself harder: untrained Qwen3-4B-Instruct achieves F1 of 0.265 there, compared with 0.408 under global search. The benefit comes from learning discrimination before global search, not from an assumption that narrowing the candidate space always improves inference.
2. Event-level search modeling: reward first discoveries of target tools, not all retrieval activity
Event-level Search Modeling (ESM) treats a search call and its returned results as one event. Within a trajectory, the tools worth optimizing at search step \(j\) are those returned now, included in the labeled target set, and absent from previous results. Using the paper's notation, \(\mathcal{M}\) denotes retrieved tools and \(\mathcal{T}_s\) is the target tool set; the first-discovery set is:
Only events with a nonempty first-discovery set enter the effective-event set. Here, novelty means additional target-tool coverage within the current trajectory, not a differently worded query, any previously unseen document, or a newly added tool in the repository. Repeated retrieval of an already seen target and retrieval of only nontarget tools receive no search advantage from this mechanism.
For the same request, the model samples \(G\) trajectories and computes a binary retrieval reward separately for each target tool: 1 if it appears anywhere in the complete search trajectory, otherwise 0. An event may first discover several targets; the authors take the maximum of their group-relative advantages, not their mean:
An event discovering a target found by relatively few group trajectories thus receives a stronger signal. If it also discovers an easy target, the maximum does not dilute the important discovery as a mean would. Tool-level rewards compare whether the complete trajectory retrieved the tool, while credit is assigned to the event that first discovered it in that trajectory. This is not an independently computed hit-rate reward at every round.
Retrieved-document tokens remain context but are excluded from policy-gradient optimization through retrieved-token loss masking. The optimization targets model-generated tokens rather than increasing the generation probability of external documents. The model can still read the documents for subsequent decisions. Having no search advantage also does not mean that every gradient is zero, because the objective includes a KL regularization term.
3. Trajectory-aligned credit allocation: train search before full coverage, then train selection
Trajectory-Aligned Credit Allocation (TCA) first checks whether each trajectory has retrieved all labeled target tools. If not, its final-selection advantage is set to zero, avoiding training selection on guesses made with missing information. Once coverage is complete, a binary reward for exact equality between the selected and target tool sets is used to construct the group-relative selection advantage. This is an exact-match reward, not partial F1; complete retrieval enables selection optimization but does not guarantee correct selection.
On the search side, advantages are set to zero for events associated with target tools already mastered within the group, rather than repeatedly rewarding a retrieval capability shared by every trajectory. Different trajectories can therefore focus on different weaknesses: those missing tools continue learning discovery, while those with complete information but incorrect sets learn final decision-making. This is not simply the sum of SRecall and final Match rewards broadcast across the trajectory.
The source equations contain ambiguities that must be retained. Equation (5) does not specify a numerical stabilizer in the standard-deviation denominator. Identical group rewards produce zero variance; the text explicitly zeros mastered search capabilities but does not fully describe implementation handling for every zero-variance case. Equation (7) uses \(j\) as the collection index for group means and standard deviations but repeats the reward of trajectory \(i\) inside the collection, which would form a constant collection if read literally. This note follows the prose for the selection gate and does not silently rewrite the equation into an author-confirmed implementation; exact reproduction requires checking the code.
A Worked Example¶
The following ordinary data-conversion example illustrates the mechanism and is not an experimental trajectory reported in the paper. The request is to convert CSV records into JSON. The target set contains a CSV parser returning a list of records and a JSON encoder accepting that list; a similarly named parser returning plain text does not directly satisfy the planned interface connection.
The first round returns 5 documents and first discovers the correct parser, but not the encoder. This event can receive a search advantage. If the model submits a final set immediately, its selection advantage is zero because target coverage remains incomplete. The second round returns the same parser and other nontarget tools, so it adds no target coverage and earns no additional first-discovery credit.
The third round retrieves the correct encoder. The model reads its documentation, checks that its input structure matches the parser's output, and submits the two-tool set; only now is it eligible for a selection advantage. If all 5 group trajectories found the parser but only 2 found the encoder, training should focus on the latter discovery events rather than continue rewarding the mastered parser search.
The example concerns reward placement and information progress. Matching a labeled tool set does not prove actual execution compatibility; interpreting documented interfaces and successfully running an external program are different levels of evidence.
Loss & Training¶
The search component uses a GRPO-style clipped policy objective, assigns event advantages to corresponding model-generated tokens, and adds KL regularization. The selection component separately optimizes the final answer with the selection-event advantage. Both update the same policy model; no additional value function is trained to annotate each step.
Training uses 14,418 filtered StableToolBench samples, excluding non-English or unnatural instructions and instructions leaking API names. I1/I2/I3 contain 11,329/2,253/836 samples. The repository contains 16,464 REST APIs, 49 coarse categories, and 500+ collections. Collections can span categories, so I3 is not merely another name for same-category tasks.
The selector backbone is Qwen2.5-7B-Instruct or Qwen3-4B-Instruct, with Qwen3-Embedding-0.6B as the retriever. Each round returns 5 documents, each request has 5 sampled rollouts, and interaction is capped at 8 turns. The learning rate is \(10^{-6}\), the KL coefficient is 0.001, and the clip ratio is 0.2.
The appendix reports a single node with 8 A100 80GB GPUs, total batch size 256, and 1 training epoch. Training responses are capped at 5,000 tokens, with 768 retrieved-content tokens per round. At inference time, these limits rise to 30,000 response tokens and 4,096 retrieved-content tokens per round, while the limits of 8 turns and 5 documents per round remain unchanged. Single-pass RAG baselines use top-100 documents, so the comparison does not establish performance under identical context-budget arrangements.
Key Experimental Results¶
Main Results¶
StableToolBench has 765 test samples. F1, Recall, and Precision compare the final selected tool set with the labeled set; Match requires exact set equality. SRecall measures whether target tools were retrieved during search, without requiring correct final selection or testing calling arguments and execution compatibility.
The table below selects results from source Table 1. I3-Inst. is F1 for unseen-instruction generalization in collection-based multi-tool tasks, not an execution success rate.
| Selector | Method | Overall F1 | Recall | Precision | Match | I3-Inst. F1 |
|---|---|---|---|---|---|---|
| Qwen2.5-7B-Instruct | Multi-turns | 0.098 | 0.110 | 0.100 | 0.046 | 0.054 |
| Qwen2.5-7B-Instruct | Search-R1 | 0.327 | 0.311 | 0.361 | 0.152 | 0.194 |
| Qwen2.5-7B-Instruct | GDPO | 0.496 | 0.487 | 0.524 | 0.255 | 0.228 |
| Qwen2.5-7B-Instruct | ToolSearcher | 0.513 | 0.511 | 0.534 | 0.278 | 0.294 |
| Qwen3-4B-Instruct | Multi-turns | 0.408 | 0.439 | 0.410 | 0.169 | 0.272 |
| Qwen3-4B-Instruct | Search-R1 | 0.505 | 0.504 | 0.516 | 0.307 | 0.215 |
| Qwen3-4B-Instruct | GDPO | 0.518 | 0.513 | 0.537 | 0.305 | 0.238 |
| Qwen3-4B-Instruct | ToolSearcher | 0.531 | 0.527 | 0.546 | 0.316 | 0.284 |
Relative to GDPO, ToolSearcher's overall F1 increases by 0.017 and 0.013, or 1.7 and 1.3 percentage points, respectively; 7B I3-Inst. improves by 6.6 percentage points. It is not best on every subset: 4B I2-Cate. F1 is 0.491, below GDPO's 0.505, and 7B I1-Tool F1 is 0.561, below GDPO's 0.567.
AppWorld tool-selection evaluation uses 147 tasks across 49 scenarios from its original training split, solely as test data and not for ToolSearcher training. Downstream execution uses the D-1/D-2 subset of Test-N, with 105 tasks across 34 scenarios; D-3 is excluded. Selection and execution are evaluated on different samples.
The executor is fixed to FullCodeRefl + gpt-5-mini: it generates complete code using the selector's tools, reflects on error messages after execution failures, and retries. The TGC/SGC values below are therefore results of the system combining a Qwen selector with a separate GPT executor, not execution success rates of Qwen alone.
| AppWorld Selector | Method | Selection F1 | SRecall | Average TGC | Average SGC |
|---|---|---|---|---|---|
| Qwen2.5-7B-Instruct | GDPO | 0.483 | 0.693 | 0.277 | 0.171 |
| Qwen2.5-7B-Instruct | ToolSearcher | 0.514 | 0.697 | 0.334 | 0.228 |
| Qwen3-4B-Instruct | Search-R1 | 0.555 | 0.635 | 0.314 | 0.228 |
| Qwen3-4B-Instruct | ToolSearcher | 0.562 | 0.664 | 0.372 | 0.257 |
TGC measures the proportion of tasks passing all tests, while SGC measures the proportion of scenarios whose tasks all pass. For 7B, average TGC and SGC both improve by 5.7 percentage points over GDPO. For 4B, they improve by 5.8/2.9 percentage points over Search-R1. This does not mean strict superiority at every difficulty or on every metric: 7B D-2 TGC/SGC tie GDPO at 0.188/0.062, and 4B selection Precision is 0.612, below Search-R1's 0.617.
Ablation Study¶
The following results come from source Table 3 and all use Qwen2.5-7B-Instruct. Removing CCTD means global search throughout training; removing ESM replaces event modeling with final-outcome-based trajectory supervision; removing TCA disables progress-dependent credit corrections.
| Dataset | Config | F1 | Recall | Precision | SRecall |
|---|---|---|---|---|---|
| StableToolBench | Full model | 0.513 | 0.511 | 0.534 | 0.680 |
| StableToolBench | w/o CCTD | 0.477 | 0.466 | 0.510 | 0.583 |
| StableToolBench | w/o ESM | 0.394 | 0.377 | 0.433 | 0.456 |
| StableToolBench | w/o TCA | 0.461 | 0.444 | 0.501 | 0.538 |
| AppWorld | Full model | 0.514 | 0.444 | 0.669 | 0.697 |
| AppWorld | w/o CCTD | 0.433 | 0.375 | 0.576 | 0.452 |
| AppWorld | w/o ESM | 0.328 | 0.280 | 0.447 | 0.303 |
| AppWorld | w/o TCA | 0.469 | 0.431 | 0.567 | 0.565 |
The source has minor numerical and cross-reference inconsistencies: AppWorld 7B full-model Precision is 0.670 in main Table 2 but 0.669 in ablation Table 3. Each value is retained in its respective context rather than silently reconciled. The ablation prose repeatedly cites โTable 5,โ although numerical ablations are in Table 3, search-setting comparisons in Table 4, and Table 5 is captioned as training curves.
Key Findings¶
- Removing ESM has the largest effect: StableToolBench F1 drops by 11.9 percentage points and SRecall by 22.4; AppWorld drops are 18.6 and 39.4 percentage points. This supports directly supervising productive search events rather than waiting only for final-set correctness.
- TCA benefits more than search alone: StableToolBench F1 increases from 0.461 to 0.513, supporting the separation of search and selection learning according to information completeness. Single-component removals do not establish independent, additive contributions.
- The paper reports that removing ESM reduces average search rounds from approximately 4 to approximately 2. More rounds are not the goal: first-discovery credit makes searching for missing tools useful to learning, while repeated retrieval receives no equivalent reward.
Highlights & Insights¶
- Combining tool-level group statistics with event-level credit is more informative than a single SRecall. Two trajectories with identical coverage can miss different tools, and credit at first-discovery events distinguishes their contributions.
- Gating selection rewards separates missing information from incorrect decisions. The transferable principle is to check whether necessary evidence was actually observed before training a decision, rather than reward an unsupported correct guess.
- Category constraints shape training difficulty rather than impose a permanent deployment restriction. They can train fine-grained API discrimination, but their effectiveness depends on category definitions and sample coverage; a smaller search space is not automatically easier.
Limitations & Future Work¶
- The authors acknowledge that synthetic StableToolBench training combinations are mostly parallel, with limited strong sequential dependencies and state changes. AppWorld APIs modify databases or user data and better resemble real workflows; current improvements do not establish mastery of complex state management.
- Tool-set F1 and Match do not validate interface connections, arguments, or actual execution. AppWorld provides limited downstream evidence, but its executor is a separate model and only D-1/D-2 are evaluated, so results cannot be extended to all AppWorld difficulties.
- Training is limited to 8 turns and 5,000 response tokens, while AppWorld tasks involve 9.5 APIs on average and up to 26. Turns and API counts do not correspond one-to-one, but long documentation and complex state dependencies still constrain context budgets; context summarization learned during training is a possible direction.
- The experiments lack error bars and multiple-seed statistics, as explicitly acknowledged in the checklist. In particular, the 1.3โ1.7 percentage-point overall F1 advantages need repeated experiments to assess stability.
- Zero-standard-deviation handling, Equation (7) indexing, and some table references require code-level verification. Executable reward definitions and boundary tests, alongside tasks with interface constraints and state transitions, would improve reproducibility and complement set-level evaluation.
Related Work & Insights¶
- vs RAG / RAGSFT: These methods select after retrieving top-100 documents once and cannot adaptively retrieve missing tools based on observed interfaces. ToolSearcher learns interactive search, but differing retrieval-budget arrangements mean the comparison changes both training objectives and retrieval style.
- vs Search-R1 / GSPO: These baselines primarily apply final-selection rewards across the trajectory. ToolSearcher assigns different advantages to first-discovery events and final-selection events; its central change is credit placement, not a replacement clipping formula.
- vs GDPO: In these experiments, GDPO separately normalizes search SRecall and final exact-match rewards for joint optimization. ToolSearcher further differentiates signals by target tool and trajectory progress, reducing inappropriate incentives for mastered capabilities or decisions made without complete information.
- vs MARAG-R1: The authors adapt it to the same search engine while retaining answer, coverage, and exploration rewards. These results compare adapted tool-selection methods, not every capability of the original multi-retriever system.
Rating¶
- Novelty: 4/5 โ First discovery, tool-level group advantages, and progress gates form an explicit credit-assignment mechanism for tool selection.
- Experimental Thoroughness: 4/5 โ Two backbones, cross-domain execution, and three component ablations provide useful coverage, but repeated runs and the hardest execution setting are missing.
- Writing Quality: 3/5 โ Motivation and component roles are clear; equation indexing, zero-variance handling, and table references reduce reproduction precision.
- Value: 4/5 โ The framework offers reusable upstream search-and-selection designs for large tool repositories, but does not replace interface and state validation.