AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems¶
Conference: NeurIPS 2026 (acceptance information from the author list; local full text is arXiv v2)
arXiv: 2609.08572
Area: Multi-Agent Systems
Keywords: prompt optimization, sequential intervention, agent-level supervision, textual gradients, semantic abstraction
TL;DR¶
AgentGrad identifies repairable prompts through single-agent interventions, extracts textual gradients from differences between original and intervention-adjusted intermediate outputs, and clusters and abstracts shared corrective patterns; with GPT-5-mini, it improves the five-task average over unoptimized prompts by 11.76 percentage points, although some performance and cost summaries require the exceptions in the tables.
Background & Motivation¶
A multi-agent system (MAS) typically produces its final answer through several prompt-controlled steps. An earlier agent may retrieve information or rewrite a request, while a later agent integrates evidence, solves the problem, or produces the answer. A system failure can therefore arise within a step or through information passed between steps. Knowing that the final answer is wrong does not directly identify which prompt needs revision. TextGrad-style methods propagate natural-language feedback across components, GEPA searches prompts through trajectory reflection, and MIPROv2 jointly searches instructions and demonstrations. These approaches supply optimization directions, but do not necessarily first verify whether changing one particular agent can repair the system.
This uncertainty accumulates at two stages. During feedback extraction, suggesting changes to every prompt can spend expensive candidate evaluations on components that cannot repair the current failure. Even after selecting a prompt, final-answer correctness alone provides little direct evidence about the intermediate output that its agent should produce instead. During aggregation, randomly selected failures can demand different or interfering corrections, such as identifier leakage in privacy rewriting and an unrelated output-format problem. Concatenating those signals does not necessarily provide a coherent, testable direction for revision.
AgentGrad connects these problems: it temporarily supplies a training hint to one agent and checks whether the entire task succeeds. Only successful interventions provide candidate supervision for that agent, after which related local corrections are abstracted into reusable prompt rules. The objective is to find an effective repair target in the current trajectory, not to establish a unique causal root cause. Core Idea: use single-agent interventions to obtain local evidence about where and how to repair behavior, abstract similar evidence into general prompt updates, and check them with minibatch and validation gates.
Method¶
Overall Architecture¶
The inputs are an existing MAS, its current agent prompts, training and validation sets, and a task reward function. The output remains a set of prompts for the same system: model weights are not updated, and no intervention module is required at deployment. Each round runs the current system on training examples, collects trajectories whose reward is below the maximum, and then performs repair-target identification, local feedback extraction, semantic aggregation, and candidate-prompt validation.
Two different orders matter. Target identification tests agents in reverse execution order, with each test intervening on only one agent. What carries forward is the set of unresolved examples, not a system retaining the previous hint injection. Prompt updating instead considers semantic clusters from largest to smallest, and replaces prompts only when candidates pass both evaluations. Thus, an intervention-adjusted output is not a deployed prompt, and a candidate prompt is not an accepted update.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Training failures<br/>Training labels form hints"] --> B["Reverse Single-Agent<br/>Intervention"]
B --> C["Output-Contrast<br/>Supervision"]
C --> D["Semantic Clustering<br/>and Abstraction"]
D --> E["Two-Gate<br/>Update Validation"]
E -->|Next round within budget| A
E -->|Keep accepted prompts| F["Deployment inference<br/>New input, no hints"]
Labels, interventions, and output contrasts in the diagram belong to offline prompt optimization. At deployment, a new input passes through the original agents using their optimized prompts; training ground truth is not injected alongside it. The validation set filters revisions for generalizability, while the test set measures performance after optimization. Their roles are not interchangeable.
Key Designs¶
1. Reverse Single-Agent Intervention: verify whether a local change can repair the system
For a failure, AgentGrad starts with the last agent, temporarily appends a hint to its current prompt, and checks whether the complete task reaches the maximum reward. If it succeeds, the failure is assigned to that agent and is not tested at earlier agents. Otherwise, the procedure tries an earlier agent. The rationale for reverse order is the authors' observation that failures often concentrate in later agents, making this order potentially cheaper in expected intervention count; it is not a structural theorem for every system.
The repaired-subset definition in the paper expresses how unresolved failures are processed:
The intervened system changes only agent \(n\)'s behavior; other agents retain their current prompts, and downstream execution continues from the changed intermediate output. Repaired examples are removed from the unresolved set, so each example contributes to at most one selected target in that round. This shrinking set must not be interpreted as first modifying the last agent, retaining that modification, and then cumulatively modifying earlier agents until the whole system is repaired.
Hints are constructed from the example's ground-truth answer or constraints that the final output must satisfy, together with dataset descriptions, system structure, and agent roles. They supply additional training information, not a universal repair mechanism for unlabeled new tasks. If no single-agent intervention repairs a failure, it produces no gradient in that round. It can reappear when the training set is executed again in the next round; difficult examples are not permanently removed.
This target definition asks where a local change sufficient to repair the result exists under the current input, context, hint, and testing order. An earlier agent may also be faulty, and a different agent may also be capable of repairing the result. A later agent might even compensate for an upstream error. Identifying a repairable target is therefore more precise than identifying a uniquely responsible agent, and it constrains how the resulting supervision should be interpreted.
2. Output-Contrast Supervision: turn successful intervention into evidence about intermediate behavior
A task label specifies the final output but may not specify what retrieval, rewriting, or evidence integration should produce. AgentGrad retains both the selected agent's original output from the failed execution and its corrected output after hint injection under the same input context. The latter is model-generated, but its corresponding trajectory has passed the system-level success check. It therefore serves as an agent-level pseudo-label that is more directly relevant to the prompt being revised than a final-error description alone.
The gradient extractor receives the current prompt, the agent's input, the original output, and the corrected output:
It returns a natural-language correction describing which rules the prompt should add or revise to encourage the corrected behavior rather than the failed behavior. A textual gradient is an analogy for an optimization direction, not a true derivative of a discrete prompt and not back-propagation into LLM parameters. Matching the input context focuses the contrast on behavioral change, but does not guarantee that the pseudo-label is the unique correct or simplest intermediate solution.
The paper describes this extraction step as requiring no explicit loss because the original-output/pseudo-label contrast directly provides local supervision, without a separately written system-level loss description for the extractor. The overall framework still uses ground truth or task constraints to construct hints and rewards to identify failures, successful repairs, and beneficial updates. Consequently, loss-free does not mean label-free, supervision-free, or evaluator-free.
3. Semantic Clustering and Abstraction: convert instance-specific suggestions into a shared correction
For each agent, the aggregator collects sample-level textual gradients assigned to it in the current round and performs semantic grouping and within-group abstraction in a single LLM call. Rather than randomly constructing minibatches and concatenating all suggestions, it groups suggestions that share a corrective pattern and produces one generalized textual gradient for each cluster. Each cluster also identifies the training failures that generated its suggestions, forming the semantic minibatch used to evaluate the proposed revision.
Abstraction should preserve a genuinely shared behavioral rule instead of listing instance-specific entities. In privacy rewriting, a missed organization name, an unredacted location, and a fictional-looking personal identifier treated as safe can jointly suggest a policy for redacting identifying entities. Such an entity-level rule has a better opportunity to transfer than adding three particular strings to the prompt. Feedback about an unrelated problem should remain in a different cluster.
The aggregator determines the number of clusters, while a soft lower bound on cluster size guides abstraction granularity. The recommended lower bound cycles through \(5\to3\to1\to5\to\cdots\): larger values encourage cross-example patterns, and smaller values accommodate finer or rarer corrections. This does not mean forming exactly 5, 3, or 1 clusters, nor strictly requiring every cluster to reach that size. Smaller clusters are allowed when gradients are too dissimilar, preventing unrelated failures from being merged merely to meet a numerical constraint.
The value of this stage is not simply shorter feedback. It converts several instance-specific demands into a coherent account of what the prompt should learn from their common pattern. However, clustering and abstraction remain LLM judgments and can overmerge signals or omit conditions. Their outputs therefore require subsequent evaluation rather than being accepted as correct updates immediately.
4. Two-Gate Update Validation: repair the target pattern and then verify generalization
Semantic minibatches are considered in decreasing size order, trying broadly applicable rules before narrower repairs. Given an agent's current prompt and a generalized gradient, the prompt optimizer proposes a candidate. Evaluation substitutes only that prompt, keeping other agents' prompts at their current values. Unlike the temporary hint used for target identification, this candidate attempts to encode a correction strategy in reusable instructions rather than provide an answer to a particular example.
The first gate compares the old and new systems' average reward on the semantic minibatch that generated the gradient. Only a strict improvement triggers evaluation on the validation set. The second gate also requires a strict improvement in average reward. The candidate replaces the current prompt only when both gates pass; otherwise it is discarded and optimization proceeds to other gradients. The first gate checks whether the intended correction was implemented, and the second rejects revisions that memorize current training examples or damage other inputs.
The two-gate rule is:
Here, \(R\) denotes mean task reward on the corresponding dataset. This is an acceptance criterion, not a differentiable loss; strict improvement also does not constitute a significance test with confidence intervals. The loop repeatedly reconstructs the failure set within a rollout budget. Accepted revisions change later trajectories and repairable examples, so targets and feedback must be obtained again under the current system rather than fixed permanently after one identification pass.
A Worked Example¶
Figure 5 provides a qualitative privacy-rewriting example. An agent should replace sensitive identifiers with clear placeholders, but its original outputs retain organization names, geographical information, or personal names that look fictional. After system failure, candidate agents are tested in reverse order. The relevant intervention checks whether supplying a training hint to the rewriting agent alone produces a redacted request and makes the full task succeed.
Successful intervention produces a corrected output. The extractor compares it with the original output and converts a particular missed entity into a suggestion for revising the prompt. Three related cases in the figure enter one cluster, whose common identifier-redaction rule is abstracted by the aggregator. A case with a different corrective signal is excluded rather than merged merely to enlarge the cluster.
The optimizer then tries incorporating that rule into the rewriting prompt and checks whether the new prompt works without hints on the corresponding failed examples. A minibatch improvement triggers validation, and both improvements are required to retain the prompt. This example illustrates the mechanism; Figure 5 does not report the candidate's exact reward change or acceptance count, so it should not be presented as a numerical training log.
Loss & Training¶
The learned objects are prompt sets. Optimization seeks higher validation reward within the permitted system-rollout budget, followed by held-out test evaluation. A failure is defined by reward below the maximum, rather than solely by a manually written error description. In continuous-reward tasks, examples below maximum reward can also enter the failure set.
Task execution and optimizer components use the same backbone: GPT-5-mini experiments use GPT-5-mini for the task, gradient extraction, aggregation, and prompt revision; Qwen3-8B experiments use Qwen3-8B throughout. This same-backbone setup applies across optimization algorithms and avoids giving one method a stronger optimizer model. It does not imply equal costs across backbones or equal call counts for every optimization step.
The local cache is v2 and contains substantive method and experiment sections, but ends after the references without the appendix described in the task package. The main text states the rollout-budget constraint without specifying concrete budget values, all split sizes, or complete task-scoring details. These hyperparameters are therefore not supplied here, and the timing, cost, and transfer experiments below are not treated as verified equal-budget comparisons.
Key Experimental Results¶
Main Results¶
The table compares AgentGrad with the strongest reported competitor for each task under the same backbone. Original Tables 1 and 2 also include MIPROv2, TextGrad, GEPA, and unoptimized prompts. Scores use a percentage scale and are means ยฑ standard errors over 3 random seeds, not standard deviations. Different tasks use different rewards, so score magnitudes do not directly rank task difficulty.
| Backbone | Dataset | AgentGrad | Strongest competitor and score | Difference (percentage points) |
|---|---|---|---|---|
| GPT-5-mini | HotpotQA | 73.89 ยฑ 1.09 | GEPA: 68.33 ยฑ 1.55 | +5.56 |
| GPT-5-mini | HoVer | 64.78 ยฑ 1.44 | TextGrad: 63.22 ยฑ 1.47 | +1.56 |
| GPT-5-mini | PUPA | 95.17 ยฑ 0.49 | GEPA: 91.87 ยฑ 1.55 | +3.30 |
| GPT-5-mini | IFBench | 76.08 ยฑ 0.45 | GEPA: 75.23 ยฑ 0.32 | +0.85 |
| GPT-5-mini | MATH | 87.62 ยฑ 0.09 | GEPA: 86.37 ยฑ 0.50 | +1.25 |
| Qwen3-8B | HotpotQA | 60.45 ยฑ 1.68 | MIPROv2: 58.33 ยฑ 2.37 | +2.12 |
| Qwen3-8B | HoVer | 52.11 ยฑ 1.66 | TextGrad: 51.44 ยฑ 0.78 | +0.67 |
| Qwen3-8B | PUPA | 91.51 ยฑ 0.72 | GEPA: 91.03 ยฑ 1.71 | +0.48 |
| Qwen3-8B | IFBench | 41.42 ยฑ 0.99 | TextGrad: 42.52 ยฑ 0.45 | โ1.10 |
| Qwen3-8B | MATH | 85.81 ยฑ 0.25 | GEPA: 85.05 ยฑ 0.48 | +0.76 |
With GPT-5-mini, AgentGrad has the highest reported mean on all five tasks and gains an average of +11.76 percentage points over unoptimized prompts. The corresponding Qwen3-8B average gain is +9.67. Although that average gain is the strongest, Qwen3-8B IFBench performance falls below TextGrad. The paper's broad statement of winning every task under both backbones conflicts with Table 2 and should not be repeated as an exception-free conclusion.
Ablation Study¶
Table 3 progressively adds intervention-guided target identification (TI), agent-level supervision (AS), and semantic textual gradient abstraction (STGA) on GPT-5-mini HotpotQA and PUPA. Vanilla is the starting point of this ablation and should not be assumed identical to the main-table TextGrad configuration.
| Config | HotpotQA | PUPA | Note |
|---|---|---|---|
| Vanilla | 67.89 ยฑ 0.80 | 85.74 ยฑ 1.65 | None of the three added components |
| TI | 69.33 ยฑ 0.38 | 89.58 ยฑ 1.89 | Intervention-guided target identification only |
| TI + AS | 70.89 ยฑ 0.62 | 92.23 ยฑ 0.68 | Corrected intermediate-output supervision |
| TI + STGA | 71.89 ยฑ 0.29 | 93.13 ยฑ 0.53 | Semantic abstraction without AS |
| TI + AS + STGA | 73.89 ยฑ 1.09 | 95.17 ยฑ 0.49 | Full AgentGrad |
Relative to TI, AS adds 1.56 and 2.65 percentage points on the two tasks, while STGA adds 2.56 and 3.55. The full method exceeds these partial combinations. This supports complementary contributions but is not a complete factorial experiment containing every individual component and combination.
Key Findings¶
Timing and cost in Tables 4 and 5 concern GPT-5-mini only. GEPA is retained below as a fixed reference, with cheaper task-specific exceptions identified separately; it is not simultaneously the fastest and cheapest baseline on every task. Timing/cost measurements, main performance, and additional transfer testing should not be merged into one equal-budget comparison. The cache lacks the budget details required to verify those boundaries further.
| Dataset | Time: AgentGrad / GEPA (minutes) | Cost: AgentGrad / GEPA (USD) | Cost or reference boundary |
|---|---|---|---|
| HotpotQA | 109 / 346 | 30.43 / 35.31 | 13.8% cheaper than GEPA |
| HoVer | 244 / 390 | 52.18 / 58.64 | 11.0% cheaper than GEPA |
| PUPA | 151 / 319 | 20.53 / 29.29 | MIPROv2 is cheaper at 19.81 |
| IFBench | 90 / 269 | 19.45 / 23.28 | 16.5% cheaper than GEPA |
| MATH | 88 / 360 | 13.13 / 27.12 | TextGrad at 20.87 is the second-cheapest method |
| Reported average | 136 / 337 | 27.14 / 34.73 | About 2.5ร faster and 21.8% cheaper than average GEPA |
- Figure 4 defines the minibatch improvement ratio as the fraction of all candidates improving their minibatch and triggering validation; the validation improvement ratio is the fraction of validation calls that improve validation reward. Averaged over HotpotQA/PUPA, these are 0.72 and 0.27 for AgentGrad, 0.44 and 0.21 for TextGrad, and 0.28 and 0.14 for GEPA. Without exact denominators, these ratios cannot determine the total number of accepted updates.
- Table 6 is an additional test of optimized prompts on unseen benchmarks within the same domain, without further optimization, not retraining on the target task. HotpotQA โ 2WikiMultiHopQA yields 51.22 ยฑ 1.63 versus GEPA's 44.89 ยฑ 4.96; MATH โ OlympiadBench yields 68.33 ยฑ 1.21 versus GEPA's 66.00 ยฑ 1.43.
- Table 5's 21.8% reduction compares average costs; it does not mean a 21.8% saving on every task. On PUPA, AgentGrad is more expensive than the cheapest method, MIPROv2. The MATH bottom row reports 37.1%, which matches the reduction calculated against second-cheapest TextGrad, whereas the reduction against GEPA is about 51.6%; these references must remain distinct.
Highlights & Insights¶
- Intervention does more than attach responsibility labels to failures: it generates intermediate-output supervision suitable for comparison. This converts system-level feedback into local prompt rules, although its reliability still depends on hints and task evaluators.
- A single aggregator call groups and abstracts feedback instead of mixing unrelated cases into one revision request. The soft cluster-size cycle provides granularities ranging from common patterns to rare repairs without strictly discarding small clusters.
- The two gates separately check implementation of the correction and resistance to instance overfitting. Figure 4's minibatch and validation improvement ratios support this interpretation, but do not independently establish that one component explains the entire speed advantage.
Limitations & Future Work¶
- Only failures repairable through a single-agent hint produce supervision in the current round. Errors requiring simultaneous changes to multiple agents may remain excluded, and revisiting difficult examples does not solve the joint-intervention problem.
- Reverse-order, first-success selection introduces a targeting preference. Comparing intervention orders, multiple repairable targets, and repair stability would be useful without treating the current target as a unique causal explanation.
- Hints depend on ground truth or verifiable constraints, and pseudo-labels may contain example-specific information. Unlabeled tasks, weak evaluators, and label-leakage risks require separate evaluation.
- Standard errors from only 3 seeds and strict-improvement gates do not establish universal statistical significance. Small performance margins particularly warrant more repetitions and explicit tests.
- The local text lacks an appendix and concrete budgets, preventing full reproduction of timing, cost, and transfer configurations. The IFBench exception in Table 2 also limits claims of being best on every task.
Related Work & Insights¶
- vs TextGrad / ProTeGi: All use natural-language feedback as a prompt-revision direction. AgentGrad first localizes a target and generates an intermediate pseudo-label through successful intervention, then applies semantic abstraction; it does not obtain true numerical gradients.
- vs GEPA: GEPA uses trajectory reflection and evolutionary prompt search; AgentGrad emphasizes single-agent repair evidence and shared corrective patterns. Main-table means and additional transfer tests support its effectiveness, but do not erase the IFBench exception or missing budget information.
- vs MIPROv2: MIPROv2 jointly searches instructions and demonstrations, whereas AgentGrad centers on local update supervision derived from failed trajectories. The cost advantage is not absolute for every task: MIPROv2 is cheaper on PUPA.
- Implications for failure attribution: Counterfactual corrected outputs from diagnosis can become optimization data, but a position sufficient to repair a failure must remain distinct from a unique error source. Recording intervention conditions would help assess supervision reliability.
Rating¶
- Novelty: 4/5 โ Combines intervention-based targeting, local pseudo-labels, and semantic abstraction in a prompt-optimization loop.
- Experimental Thoroughness: 4/5 โ Two backbones, five tasks, and component ablations offer broad coverage, but budget details are absent from this cache.
- Writing Quality: 3/5 โ The method is reconstructable, but the performance summary conflicts with a table, and some efficiency references require careful distinction.
- Value: 4/5 โ Offers an interpretable improvement route for existing multi-agent pipelines without updating model weights.