MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens¶
Conference: NeurIPS2026 (task-queue assignment; this note analyzes arXiv v1)
arXiv: 2609.30967v1
Area: LLM Agent
Keywords: Harness search, multi-objective optimization, execution traces, behavioral safety, token cost
TL;DR¶
Without updating model weights, MoMHa uses a proposer to rewrite Python harnesses from execution traces, jointly optimizing accuracy, behavioral safety, and token cost; the authors report leading joint means on synthetic and real benchmarks, but source conflicts concerning metric scaling, the two-phase definition, and cost accounting require clarification.
Background & Motivation¶
LLM task performance depends not only on the model, but also on how external code assembles inputs, calls the model, decides whether to retry, and extracts results. The paper calls this Python program a harness. The same model can be wrapped in a single direct call or placed in a pipeline with retrieval, conditional verification, and output normalization. In the baseline implementations used here, APE, OPRO, DSPy, and TextGrad primarily change prompts or demonstrations within a fixed skeleton; MH searches over the entire Python harness, but primarily targets accuracy.
Accuracy-only search ignores two distinct costs: repeated verification may improve reliability while increasing token consumption, and satisfying an output format or static code rules does not establish behavioral safety. Conversely, refusing every request can reduce inappropriate responses while sacrificing helpfulness on ordinary requests. Rather than adding a filter after finding an accuracy-optimal harness, the authors expose accuracy, safety scores, and call costs while candidate structures are being proposed. The proposer needs to identify whether a prompt, parser, redundant call, or routing decision caused a failure, rather than seeing only a final aggregate score.
The central comparison is single-phase joint search versus a two-phase procedure that optimizes accuracy before cutting cost. If the first phase fixes an expensive two-call structure and the second only permits pruning, the search may not reach a different, cheaper call topology. This is a mutation-neighborhood restriction of the evaluated two-phase implementation, not a claim that all staged optimization is structurally incapable of change. Core idea: make the harness itself editable, diagnose structural failures from per-example traces, and account for accuracy, behavioral safety, and token cost together from the first search iteration.
Method¶
Overall Architecture¶
Inputs are task examples, an LLM client, an initial harness, and domain skill files; the output is a search-selected Python harness that is subsequently frozen. During search, the proposer reads historical source, per-example scores, model-call traces, and metadata to propose a new candidate. The evaluator executes it on the visible search set, the logger records feedback, and the next iteration diagnoses that feedback. Final evaluation runs only the frozen harness, without further search or target-model fine-tuning.
The pipeline contains three search designsโTrace-Driven Rewriting, Constrained Evaluation, and Joint-Reward Selectionโfollowed by Frozen-Harness Transfer. Domain skills supply both strategy priors and safety boundaries. Static checks and sandboxing protect execution, whereas behavioral safety is measured by a separate evaluation protocol; neither substitutes for the other.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
I["Initial harness and domain skills"] --> A["Trace-Driven Rewriting"]
A -->|proposer submits candidate| B["Constrained Evaluation"]
S["Visible search set"] --> B
B -->|evaluator scores and logger traces| C["Joint-Reward Selection"]
C -->|historical feedback: continue search| A
C -->|search ends: select source| D["Frozen-Harness Transfer"]
T["Hidden test set and target models"] --> D
D --> O["Predictions and evaluation metrics"]
Key Designs¶
1. Trace-Driven Rewriting: localize feedback to individual calls rather than reporting only better or worse scores
A candidate harness can change prompts, few-shot context, retrieval, output parsing, verification or retry loops, and routing between prompt variants or models. It does not generate the answer to the current task; it generates the program that will answer future examples. A valid program must expose initialization and per-example execution interfaces and return a dictionary containing at least prediction; the appendix also describes tokens_used. Thus, the optimization surface is program structure and textual configuration, not model parameters.
The proposer state includes historical candidate source, parent relationships, per-axis scores, common errors, and call traces from the current best harness. Traces record call inputs, outputs, latency, and tokens, allowing the proposer to distinguish an incorrect generation from a parsing mistake or a verifier that corrupts a correct answer. History resides in the filesystem and is read on demand rather than inserted into one large prompt; the reported median is 82 files read per iteration. This workflow increases diagnostic context, but also adds proposer cost that is not fully accounted for.
Skills combine generic search instructions, domain strategies, and domain safety specifications. Domain files provide a candidate strategy library, such as conditional verification, mathematical subject routing, or context pruning. Search discovers combinations, ordering, prompts, and thresholds rather than inventing every strategy from scratch. Appendix O describes a single structural transformation of the current best harness and a preference for additive changes. Consequently, rewriting full Python does not mean unconstrained program search without priors.
2. Constrained Evaluation: distinguish executable-code validity from behavioral safety of responses
Candidates first undergo AST-Guard static checks and interface validation, followed by inexpensive execution with a mock client. Only surviving candidates invoke the real evaluator on the search set. Appendix O adds resource-bounded screening and a reward floor before a candidate may compete to replace the current best. Approximately 30 candidates out of roughly 100 proposer calls per domain pass all gates and reach acceptance; this does not mean all 30 are accepted. Logged rejections also become search feedback.
At a high level, the safety architecture combines static analysis, domain-calibrated safety skills, and an execution sandbox. Static analysis checks program structure and disallowed operations; domain skills adjust constraints to task risk; sandboxing limits execution time, tokens, and call counts. These mechanisms address generated-code execution risks, but do not prove that protections are impossible to bypass or that deployment environments are absolutely safe. This note does not detail bypass techniques or reproduce harmful prompts.
Capability tasks use task-specific evaluators: classification and multiple-choice tasks use label matching, mathematics uses normalized answers, coding uses syntax and test execution, SQL uses execution results, and NER uses entity F1 with type matching. Safety tasks measure both refusal on matched risk-bearing requests and helpfulness on ordinary requests, preventing blanket refusal from becoming optimal. The source defines the safety reward as:
The primary judge is Claude Sonnet 4.6, which labels responses as refusal, compliance, or partial compliance, with half credit for partial cases. Here, safety means an average behavioral score under a particular user-context benchmark protocol, not deployment certification. The code-structural safety table's 1.00 instead refers to static checks. A variant without a safety skill can still retain AST checks, so its behavioral safety cannot be inferred to match the full system.
3. Joint-Reward Selection: price safety and cost while choosing the call structure
Search selects candidates with a linear scalar reward, while the proposer retains access to all three raw axes and per-example traces. When the aggregate improves, it can inspect whether the gain came from accuracy, safety, or cost rather than compressing different failure types into an undiagnosable number. Equation (3) defines:
The main text sets \(\lambda_s=1.0\) and values a 1k-token reduction as one percentage point of accuracy; one implementation in Appendix O gives \(\lambda_t=10^{-5}\). A screened candidate replaces the current best only with a strict reward improvement, with lower tokens breaking ties. The paper also discusses a Pareto candidate pool and post-hoc hypervolume, but this scalarized local-rewrite search should not be equated with directly maximizing global hypervolume.
Joint search can decide whether a step should exist, rather than merely shortening its prompt. For easy examples, a verification call may add cost and even introduce errors; for complex examples, extra verification may remain worthwhile. Trace-driven diagnosis can reserve computation for branches that need it. The two-phase baseline first finds a high-accuracy structure and then cuts cost within an accuracy tolerance, with its second phase restricted to the existing structure's neighborhood. This restriction is the authors' explanation for joint search's advantage.
Leaderboard tables use a different metric, not the search reward above. Equation (4) in the main text is:
Accuracy and tokens are the variant's cross-model means on a domain, whereas \(\mathrm{Safe}_{v}\) is its mean score on the three U-SafeBench safety domains. This variant-level constant multiplies every capability cell; it is not a safety measurement specific to that capability domain. A leading capability-domain joint score can therefore partly reflect a global safety advantage measured elsewhere, rather than leading raw task accuracy.
Both metrics prefer greater accuracy, greater safety, and lower cost, but linear and logarithmic penalties need not rank non-dominated candidates identically. More importantly, the written metric and reported score scale do not directly agree: if both scores lie in the unit interval, the upper bound at 672 tokens is approximately 0.154, yet reported joint means are around 0.48. Appendix N.3 instead writes \(J=6\cdot\mathrm{acc}\cdot\mathrm{Safe}/\ln(\mathrm{tokens})\). This is textual evidence of a scaling conflict, but does not establish which tables actually use that version; it does not justify reconstructing or correcting every table.
4. Frozen-Harness Transfer: separate discovered structures from cross-model testing
Each synthetic capability domain contains 100 LLM-generated examples verified by another model, split into 50 visible search examples and 50 hidden test examples. The seven capability domains total 724 harness evaluations, with a per-domain range of 98โ106; the three safety domains use roughly 25 each. The search evaluator is Claude Haiku 4.5 and the proposer is Claude Code. Appendix N.2 labels the primary proposer Claude Sonnet 4.6, so proposer and evaluator should not be conflated.
After search, the same program runs without modification across a 12-model fleet spanning four model families. Transfer here means transfer of harness structure and prompts, not trained weights. MoMHa wins the joint score on 8/12 models, not every model. The seven real benchmarks are described in the main text as a generalization check excluded from search, but real-domain skills, second-phase descriptions, and baseline provenance in the appendices leave the implementation of direct transfer in need of verification.
A Worked Example¶
Consider ordinary fact verification. An initial harness might perform evidence decomposition followed by verification for every claim. After search-set evaluation, the proposer observes in traces that a second call does not improve some clear examples but repeatedly consumes tokens. It proposes a candidate that returns a verdict in one pass for clear cases and invokes further checking only for uncertain cases. Accuracy, safety, and cost are evaluated again, results are logged, and the joint reward determines whether the candidate replaces the current best.
This illustrates how feedback changes the call graph, not why all verification should be removed. Appendix K describes fact-verification structures at approximately 536 versus 1483 tokens, but another passage there gives accuracy as 0.457, differing from Table 8's 0.508; these values cannot be combined into one unified test result. Once search ends, the selected program is retained. Hidden test examples do not trigger further proposer calls or test-answer-driven structural changes.
Loss & Training¶
There is no model-weight training, gradient descent, or back-propagation. Optimization consists of generating, screening, evaluating, and selecting Python harnesses. In the main text and Appendices D/K/O, the two-phase setup first optimizes accuracy subject to safety, then reduces tokens within a two-percentage-point accuracy tolerance. Appendix F instead describes accuracy/tokens followed by safety, creating a conflict in the experimental definition.
Appendix O also lists another implementation, \(R=\mathrm{acc}-\lambda(\mathrm{tokens}/\tau)\), with \(\lambda=0.1\) and \(\tau=1000\). It lacks an explicit safety term and has a different token coefficient from the preceding implementation. This note explains the joint objective as stated in the main text while retaining reward-path ambiguity as a reproducibility issue, rather than assuming complete equivalence.
Key Experimental Results¶
Main Results¶
The following excerpt from Table 2 contains author-reported joint scores, not accuracy. Original SEM values are retained, but are not treated as validated independent-sample uncertainty estimates. Real benchmarks use approximately 11 models; Appendix E states that GPT-5.2 participates only in synthetic and safety evaluation.
| Variant | LawBench | NuminaMath | FEVER | Spider | HumanEval | MBPP | MMLU-Pro | Mean |
|---|---|---|---|---|---|---|---|---|
| MH | 0.140ยฑ.004 | 0.360ยฑ.005 | 0.421ยฑ.006 | 0.222ยฑ.005 | 0.338ยฑ.007 | 0.066ยฑ.002 | 0.389ยฑ.006 | 0.277 |
| DSPy | 0.218ยฑ.003 | 0.453ยฑ.001 | 0.487ยฑ.001 | 0.296ยฑ.001 | 0.438ยฑ.007 | 0.299ยฑ.003 | 0.448ยฑ.002 | 0.377 |
| TextGrad | 0.220ยฑ.003 | 0.392ยฑ.007 | 0.606ยฑ.002 | 0.211ยฑ.007 | 0.373ยฑ.008 | 0.068ยฑ.002 | 0.473ยฑ.004 | 0.335 |
| MoMHa | 0.205ยฑ.005 | 0.516ยฑ.007 | 0.641ยฑ.008 | 0.331ยฑ.005 | 0.498ยฑ.009 | 0.588ยฑ.009 | 0.446ยฑ.004 | 0.461 |
MoMHa wins 5/7 real-benchmark columns; TextGrad retains LawBench and MMLU-Pro. Synthetic-track Table 3 reports a joint mean of 0.482, versus TextGrad's 0.422 and MH's 0.305. MoMHa wins 7/10 columns, but loses Math to DSPy and FV and Safeill to TextGrad. The difference 0.461โ0.377=0.084 conflicts with the introduction's โ+7.9 percentage points.โ This note reports the tabulated numbers rather than silently correcting the introduction.
Raw capability performance is not uniformly superior: MoMHa averages 0.539 and DSPy 0.542. In Table 8, compared with MH, MoMHa's NER accuracy falls from 68.6% to 56.2%, MCQ from 69.8% to 57.8%, synthetic mathematics from 51.2% to 49.0%, and LawBench from 35.1% to 34.5%. Joint-metric wins do not erase these non-wins.
Ablation Study¶
The following ten-domain aggregate comes from Appendix K. Overall averages the seven capability and three safety domains and is not the joint metric. Its Avg Tokens value of 572 also differs from the seven-capability-domain value of 672 and must not be mixed with it.
| Variant | Capability | Safety | Overall | Avg Tokens |
|---|---|---|---|---|
| MoMHa | 0.539 | 0.781 | 0.611 | 572 |
| 2-phase | 0.520 | 0.735 | 0.584 | 667 |
| MoMHa-ns | 0.528 | 0.716 | 0.584 | 620 |
| MoMHa-noTok | 0.512 | 0.748 | 0.583 | 641 |
| MoMHa-scalar | 0.539 | 0.754 | 0.604 | 640 |
Single-phase search improves Overall by 0.027 over two-phase search while using 95 fewer tokens per example. Removing the safety skill lowers behavioral safety by 0.065. Scalar-only feedback leaves mean capability unchanged but lowers safety by 0.027. This supports the diagnostic value of traces under this protocol, not equal benefits on every task. In Table 9, two-phase HumanEval reaches 85.5%, exceeding the full system's 80.3%; MoMHa-noTok reaches 69.1% on NER versus 56.2%.
The cost excerpt below comes from Table 12, in millions of tokens. It counts search evaluators and final evaluation only, not complete proposer API costs.
| Variant | Search-eval (M) | Final-eval (M) | Total (M) |
|---|---|---|---|
| APE | 9.55 | 5.21 | 14.76 |
| MoMHa (v3joint) | 9.20 | 6.41 | 15.61 |
| 2-phase | 11.02 | 7.31 | 18.33 |
| DSPy | 7.69 | 29.35 | 37.04 |
Key Findings¶
- Table 4 reports cross-model joint means of 0.434 for MoMHa and 0.384 for DSPy; MoMHa wins 8/12. GEPA wins on GPT-5.4, while DSPy wins on GPT-5-mini, o4-mini, and DeepSeek-R1. This is not dominance on all models.
- Appendix K reports hypervolume of 0.481 for MoMHa, 0.362 for MH, and 0.425 for two-phase search. This is a post-hoc metric with normalized axes, not the search reward; reference scales and aggregation settings affect comparisons.
- Appendix N.3 ranks MoMHa first under all tested alternative scalarizations on the synthetic track, but second under weighted sums and third under raw Pareto-dominance count on real benchmarks. The robustness conclusion is continued competitiveness, not first place under every metric.
- Appendix N.4 empties skill files on only two domains: SQL joint score falls from 0.263 to 0.247, whereas the physical-safety domain rises from 0.712 to 0.855. Domain priors can guide search or restrict exploration; two domains do not establish universal independence from them.
- The cached full text contains captions for several plots without readable curve values; missing plot readings are not supplied. Appendix N.5 adds a second-judge check, but its sample counts of 60, 30, and 30 do not replace cross-environment safety validation.
Highlights & Insights¶
- The most reusable design is to search the call structure. Cost reduction need not mean prompt compression alone: search can assess whether a step should exist, which is more informative than a joint-score gain by itself.
- Selection scalars and diagnostic multi-axis traces can be separated. Scalarization simplifies candidate ordering, while raw traces preserve failure locations and help distinguish safety, parsing, and cost problems.
- Measuring both refusal and helpfulness on ordinary requests is more meaningful than refusal rate alone. The benefit remains conditional on paired data and judge rules, rather than establishing deployment safety.
Limitations & Future Work¶
- Metric and reward-path conflicts: the main joint metric lacks a scaling factor, Appendix N.3 includes a factor of 6, and another reward expression in Appendix O omits safety. Each table needs its actual implementation, units, and raw per-example records, rather than a formula inferred from reported values.
- Unclear generalization and domain-file boundaries: the main text claims zero search cost on real benchmarks, yet Appendix D describes second-phase skills spanning all domains. Appendix C says MH uses curated baselines on real and safety domains rather than uniformly search-discovered harnesses. Real mathematics is also named NuminaMath, MATH, and MATH-500 in different places. The text alone does not establish mappings, adaptation details, or fully matched baseline procedures.
- Potential pseudoreplication in error bars: Table 2 writes \(\mathrm{SEM}=\sigma_{\mathrm{cross-model}}/\sqrt{11\times50}\). If the numerator is a standard deviation of model-level means but the denominator treats model-example products as independent trials, uncertainty may be understated. Predictions on the same examples across models have clustering dependencies. Per-example definitions, model/example-stratified bootstrap, and independent search repetitions are needed; small SEM values alone do not prove significance.
- Cost is not complete API cost: the main text defers the full proposer API comparison, while Appendix N.1 calls proposer cost negligible without supporting token breakdowns or pricing. The comparison of 15.61M with 18.33M and 37.04M holds only within Table 12's accounting scope. Table 7 actually reports all-domain mean harness inference tokens of 629 for MoMHa versus 522 for MH, so uniform savings cannot be claimed.
- Unmatched search spaces and budgets: fixed-skeleton prompt optimization and full Python rewriting differ in expressiveness, skill priors, and proposer workload, not just objectives. Objective ablations under the same structural search space and total budget would more cleanly attribute gains to joint optimization.
- Additional source non-wins and version drift: Appendix F's two-phase definition differs from K; O gives a ten-domain joint mean of 0.472 whereas Table 3 gives 0.482; some domain numbers in K differ from Tables 8/9. This note retains their separate scopes rather than merging them into a conflict-free result set.
- Limited safety scope: 0.781 is an aggregate benchmark score over three specific safety domains, and passing static checks does not establish complete sandbox protection. Larger evaluations, multiple judges, out-of-distribution user contexts, helpfulness breakdowns, and independent code audits remain necessary for assessing deployment risks.
Related Work & Insights¶
- vs MH / AutoHarness: all permit full-harness generation, but this paper explicitly includes safety and cost in search. MH is the closer structural-search comparison, although differences in baseline provenance described in the appendices weaken the strong attribution that all gains arise solely from changing objectives.
- vs DSPy / MIPROv2 / TextGrad / GEPA: the baseline implementations here retain fixed skeletons, while MoMHa can change the call graph; GEPA's instance-level Pareto diversity also differs from objective-axis analysis here. Conclusions should be limited to the reported implementations, not generalized into an inability of these frameworks to express complex programs.
- vs Guardrails, soft prompts, and activation steering: external filters mainly control input/output boundaries, whereas soft prompts and activation steering generally require model-internal access. MoMHa searches inspectable API-level programs. These approaches can be complementary, but the paper does not evaluate their joint deployment.
- Research direction: compare joint, staged, and safety-constrained search within a matched structural search space, publishing raw accuracy, per-domain behavioral safety, full API accounting, and clustered uncertainty. This is an evaluation and reproducibility recommendation, not a newly validated method introduced by this note.
Rating¶
- Novelty: 4/5 โ Combines structural search with multi-axis per-example feedback, while inheriting harness search and strategy priors.
- Experimental Thoroughness: 3/5 โ Extensive domains, models, and ablations, but search-space matching, statistical independence, and complete cost accounting remain insufficient.
- Writing Quality: 2/5 โ The main narrative is clear, but metric scaling, the two-phase definition, mathematical dataset names, and local numerical results conflict.
- Value: 4/5 โ Offers a reusable harness-design perspective; joint scores and safety conclusions should be used only with implementation ambiguities acknowledged.