ContractBench: Can LLM Agents Preserve Observation Contracts?¶
Conference: NeurIPS 2026 Evaluations & Datasets
arXiv: 2605.17281
Area: LLM Agent
Keywords: observation contracts, temporal validity, byte-level integrity, tool use, programmatic evaluation
TL;DR¶
ContractBench uses a virtual clock, byte checks, and HTTP-trace validation to evaluate whether agents preserve cross-step contracts attached to tool artifacts; in the paper's experimental snapshot of 38 model variants, the highest success rate is 77.8%, and greater scale or newer versions do not guarantee improved reliability.
Background & Motivation¶
Correct tool use does not guarantee a reliable workflow. An agent may understand the request, select the right API, and produce a syntactically valid call, yet the downstream service can reject it if the agent reorganizes a presigned URL as ordinary text or submits a temporary token after expiration. The issue is not merely whether an answer is factual: an external system has issued an artifact with conditions that subsequent actions must preserve. OAuth state, session tokens, ETags, and signed requests all create such cross-step dependencies.
Existing agent benchmarks primarily assess final task outcomes, tool selection, or compliance with domain policies, often without independently diagnosing intermediate-artifact handling. TicToc studies tool re-invocation when cached information becomes stale, but it cannot replace byte-level transmission checks: refreshing an expired URL repairs validity without preventing an interface from truncating the fresh URL. Conversely, preserving a token exactly does not automatically produce deadline-aware planning. Real tool pipelines can also alter artifacts through input limits, line wrapping, URL re-encoding, or display-text conversion, making failures a joint concern of model decisions and execution infrastructure.
ContractBench therefore turns these conditions into executable evaluation targets rather than asking another language model whether an agent's explanation sounds credible. Parameterized tasks, controlled time, and server logs let the authors apply temporal and integrity pressure separately and diagnose the violated constraint through failure labels. Core idea: treat intermediate tool artifacts as observation contracts requiring both temporal validity and byte-level integrity, and validate subsequent HTTP actions directly rather than evaluating tool-call competence alone.
Method¶
Overall Architecture¶
ContractBench is a benchmark, not a new neural network, and it introduces no training loss to optimize. Its inputs are a task instruction and an interactive simulated API. The agent uses tools to interact with the server, receives temporary URLs, tokens, or protocol states, and reuses those artifacts in later requests. A hidden programmatic validator produces the outputs from HTTP request logs: a success flag, a primary failure label, failure details, and trace metadata.
The design consists of observation-contract formalization, dual-axis tasks and controlled mutations, programmatic validation and failure attribution, and repeated evaluation and label feedback. The first two define what an agent must preserve and how evaluation applies pressure; the latter two establish whether it actually preserves the constraints and whether diagnostic feedback helps a subsequent attempt. No network diagram is included because the evaluation components should not be mistaken for a model architecture.
Each task contains TOML metadata, an agent-visible Markdown instruction, a FastAPI server, and a pytest validator. The metadata fixes difficulty, task category, and budgets; the server returns interaction outcomes and records events; the validator computes rewards from the logs. The agent sees only the instruction and server HTTP responses, not the hidden metadata or validation rules. This visibility boundary makes scoring depend on executed behavior rather than a description of the test script.
Key Designs¶
1. Observation-contract formalization: check timely use and exact transmission separately
The authors represent an observation contract as \(C=(o,t_{\text{issue}},\tau,\pi)\): \(o\) is the issued artifact, \(t_{\text{issue}}\) its issue time, \(\tau\) its time-to-live, and \(\pi\) an integrity predicate. A subsequent submission must arrive within the allowed time and satisfy the artifact's integrity condition; semantic equivalence or visual similarity cannot substitute for that condition. Definition 2.1 uses the validity window \(W(C)=[t_{\text{issue}},t_{\text{issue}}+\tau)\), while the implementation's default integrity check compares the SHA-256 digests of the submitted and original content. The following expression combines the two checks while preserving their independent roles.
Here \(h(o)\) is the original artifact's digest, while \(o'\) and \(t'\) are the submitted content and submission time. Trace compliance requires every required contract to hold; successful steps cannot offset a violation. The orthogonality proposition states that timely-and-intact, expired-but-intact, timely-but-altered, and expired-and-altered submissions are possible in principle. It does not assert statistical independence of failure probabilities in the dataset. The abstraction also admits deterministic predicates such as HMAC or ETag checks; using digest comparison to emphasize transmission fidelity does not make every real API's protocol validation equivalent to a byte hash. The source has an unresolved endpoint discrepancy: Definition 2.1 and Appendix F require submission strictly before expiration, whereas Figure 2 and Section 3.2 use the endpoint-inclusive condition \(t_{\text{fetch}}\leq t_{\text{issue}}+\tau\).
2. Dual-axis tasks and controlled mutations: expose temporal planning and transmission fidelity separately
The 33 tasks are generated from parameterized templates with difficulty controls such as TTLs, rate-limit windows, and token lengths. One seed controls resource ordering, artifact content, and timing jitter. The task organization distinguishes low-pressure Q1, validity-dominant Q2, integrity-dominant Q3, and joint-pressure Q4. The main text states that 24/33 tasks belong to Q4, aiming to reflect artifacts that both expire and are signature-bound. Tasks cover OAuth/auth, signed requests, state chains, resource management, and multi-service workflows rather than only copying a short string. However, counting the entries in Appendix D yields 5 Q2 tasks, 5 Q3 tasks, and 23 Q4 tasks, with no separately listed Q1 task. The main text's count of 24 should therefore not be treated as verified by the catalog.
To distinguish forgotten constraints from pipeline-induced artifact changes, the benchmark reproduces five deterministic transmission mutations: truncation, inserted line breaks, URL re-encoding, query-parameter reordering, and mismatches between displayed text and the actual link. These are evaluation pressures in the execution path, not security mechanisms that an agent should bypass. For example, an interface accepting only 200 characters can transmit incomplete content when the original signed query string is 256 characters long, even if the model intends to relay it directly. Appendix H illustrates a framework-level remedy: store the complete artifact server-side and resolve a short handle at execution time. Temporal pressure instead requires agents to recognize maintenance windows, rate-limit backoff, and resource deadlines; repeatedly resending the same request is not a universal recovery method. Evaluating both pressures prevents a temporal repair from being mistaken for an integrity repair.
3. Programmatic validation and failure attribution: recognize compliance only in HTTP traces
The FastAPI server records chronologically ordered request events under a virtual clock, and the pytest validator derives a binary outcome from them. Validity checks concern simulated time, versions, or service windows; integrity checks concern the content and protocol conditions that actually reach the server. An agent's final claim of success does not participate in scoring. The virtual clock separates environment time from network latency and model response speed, making the checks for a given task instance reproducible. Experiments also impose a 600-second wall-clock timeout per episode, alongside task-specific step and virtual-time budgets. The 600-second limit is neither an artifact TTL nor the rule for advancing simulated time.
The taxonomy contains 15 labels: 4 validity failures, 9 integrity failures, and 2 meta labels, SUCCESS and OTHER. It is therefore not a set of 15 failure types exclusively. Validity labels cover expiration, rate limits, scheduled unavailability, and version conflicts. The integrity group includes MUTATED_TOKEN, SIGNATURE_MISMATCH, and WRONG_HASH, but also protocol outcomes such as missing constraints, incorrect values, and compensation failure. It is a practical, broad attribution scheme rather than a taxonomy of byte changes alone. An episode can emit multiple events, but the most-severe rule retains one primary label; state events from successful workflows are excluded from failure distributions. This aids model comparison but conceals co-occurring failures. Moreover, the displayed severity table in Appendix K lists only some labels, so it does not justify reconstructing a complete set of weights.
4. Repeated evaluation and label feedback: distinguish tool use, compliance, and recovery
Every task ships with a reference solution running under the same virtual clock. It must pass validation before the task enters the benchmark. All 33 reference solutions pass a grid of 10 seeds per task, totaling 330 oracle episodes. These establish intended-path solvability and validator regression coverage, not an estimated ceiling on model capability. The main model leaderboard explicitly uses 3 rollouts per task over 33 tasks, producing 99 episodes per model at temperature 0 with recorded pinned model identifiers. Section 3.1 separately states that every model is evaluated on 10 seeds per task, conflicting with the leaderboard protocol. The 330 oracle episodes and 99 model episodes are reported separately here rather than merged into a nonexistent uniform setting.
The main metric, success rate SR, is the fraction of episodes whose primary label is SUCCESS. Per-task pass rate is the mean reward over a model's 3 runs on that task, while the failure-label distribution counts primary labels only among failed episodes. Finally, the authors select 42 GPT-5.1 failures for paired retries with no hint, the original correct label, or an incorrect label from a different axis. This separates the effect of retrying from that of receiving a correct diagnosis. The intervention is inference-time contextual feedback, with no weight updates; using labels for reinforcement learning post-training remains future work. The 42 paired failures are the retry sample, not GPT-5.1's total leaderboard failures: 48/99 successes imply 51 failures.
The protocol deliberately separates levels of reliability. Emitting a valid tool call does not guarantee entering the protocol correctly; entering the protocol does not guarantee preserving contracts throughout a trajectory; receiving a failure label does not make an invalidated artifact recoverable. The contribution is primarily an observable, diagnostic evaluation of these levels, not an agent algorithm that guarantees success.
Key Experimental Results¶
Main Results¶
The table below selects results from main-text Table 3 and Appendix Table 10, all using 99 episodes per model. It is a snapshot from the cached paper version, not a current online leaderboard. Results across providers or model versions are not generalized into universal capability claims.
| Model variant | Successful episodes / 99 | SR (%) | Main observation |
|---|---|---|---|
| Claude Opus 4.6 | 77 | 77.8 | Highest in this paper, still below 80% |
| GPT-5.2 | 74 | 74.8 | Main Table 3; Appendix Table 10 gives 74.7 |
| GPT-5 | 70 | 70.7 | Outperforms the subsequent GPT-5.1 |
| GPT-5.1 | 48 | 48.5 | Version upgrade exhibits regression |
| GPT-4o | 23 | 23.2 | Lower starting point of the GPT version comparison |
| Qwen3.5-397B-A17B | 70 | 70.7 | Matches GPT-5 |
| Qwen3.5-27B | 64 | 64.6 | Leaderboard value; other passages conflict |
| Qwen3.5-9B Instruct | 56 | 56.6 | Capability jump from 4B to 9B |
| Qwen3.5-4B Instruct | 0 | 0.0 | Does not reach a successful contract trajectory |
| Qwen3.5-9B Base | 0 | 0.0 | Base/Instruct gap also involves protocol entry |
Qwen3.5-27B's 64.6% comes from the leaderboard, but Section 5 gives 76.5%, and Appendix Table 12 gives a mean of 0.62. These values cannot be directly reconciled. This note identifies its leaderboard convention through the explicit success count and preserves the discrepancies.
Ablation Study¶
The paper does not introduce a new model with module-removal ablations. Its closest mechanism-level control is the label-feedback retry experiment. The following table concerns the same 42 GPT-5.1 failed episodes and must not be conflated with leaderboard SR.
| Retry condition | Successful episodes / 42 | Retry success rate (%) | Relative to no hint |
|---|---|---|---|
| No hint | 6 | 14.3 | Baseline |
| Wrong-label hint | 5 | 11.9 | -2.4 percentage points |
| Correct-label hint | 8 | 19.0 | +4.8 percentage points |
Correct-label coaching recovers 3 more episodes than wrong-label coaching, a +7.1-percentage-point gap. This is neither the improvement over no hint nor a 7.1-percentage-point increase on the full 99-episode leaderboard. The sample is small, and the table provides no significance test or confidence interval.
Appendix S further limits which errors appear recoverable instead of supporting a claim that labels help every failure.
| Original failure label | Paired episodes | No-hint successes | Correct-label successes | Wrong-label successes |
|---|---|---|---|---|
| WRONG_VALUE | 19 | 3 | 4 | 2 |
| MISSING_CONSTRAINT | 1 | 0 | 1 | 0 |
| EXPIRED_BEFORE_USE | 9 | 2 | 2 | 1 |
| RATE_LIMITED | 1 | 0 | 0 | 1 |
| Other rare labels | 12 | 1 | 1 | 1 |
| Total | 42 | 6 | 8 | 5 |
Key Findings¶
- Validity and integrity cannot substitute for each other. Appendix P reports that integrity mean reward falls from 0.80 to 0.47 between GPT-5 and GPT-5.1, a decrease of 0.33; validity and hybrid tasks decline by 0.25 and 0.20 respectively. Regression is concentrated more strongly on integrity, not absent from the other axes.
- The Qwen 3.5 jump is not merely better tool-call formatting. The authors observe that 4B already emits valid calls but stops after early failures, whereas models at 9B and above begin waiting, backing off, and adapting strategies. Mid-trajectory restraint is therefore an important diagnostic clue.
- Oracle solvability does not imply model stability. In Appendix L's subset of 25 models, 2,259 episodes, and 733 repeated cells, 86.6% of modelโtask cells are perfectly consistent and 13.4% vary. These percentages concern that subset, not the entire 38-model evaluation.
- Every model in the paper's cohort fails multi-turn-recall, which requires preserving an 8,192-byte URL across turns. This motivates artifact storage outside conversational context but does not prove that every framework or future model must fail.
- Correct labels do not guarantee recovery of invalidated resources. Among 9 paired EXPIRED_BEFORE_USE episodes, both correct hints and no hints recover only 2. The single RATE_LIMITED episode succeeds only under the wrong-label condition; one observation cannot establish a general strategy.
Highlights & Insights¶
- Treating an intermediate output as a contract rather than information better matches reliable execution. It distinguishes immutable artifacts from explanations that can be summarized, making the concept useful for long-running API agents.
- Direct HTTP-log validation prevents natural-language claims from substituting for executed compliance. A fixed simulated environment also turns reliability regression after a model upgrade into a testable concern rather than relying only on aggregate task scores.
- Paired correct- and wrong-label controls are more informative than showing retry improvements alone. They reveal directional diagnostic value while also showing that expiration and backoff may require execution-framework protections instead of longer self-reflection.
Limitations & Future Work¶
- The virtual clock improves reproducibility but omits real network latency, jitter, and API state changes. Production reliability still requires testing under real-time and fault conditions.
- Base models lack chat templates or tool-call formats, so their 0% scores mix protocol-entry failures with contract deficiencies. Base/Instruct comparisons alone cannot identify the causal contribution of post-training to contract preservation itself.
- Model scale, post-training, and provider execution stacks are not fully controlled. GPT version regression is observed under this setup, but attributing it further to a specific training objective or sycophantic behavior requires more direct intervention evidence.
- The source contains inconsistencies in endpoint handling, seeds versus rollouts, quadrant counts, and local results versus the leaderboard. Some prose also reports a per-task reward of 0.80, inconsistent with the granularity of a mean over 3 binary rollouts; it is not used here to construct a precise inverse-scaling table.
- The 33 templates and Q4-heavy distribution cannot represent every production API, and there is no human baseline. Oracles establish intended-path solvability, not coverage or human difficulty.
- Testable next steps include machine-readable TTLs, versions, and backoff windows, followed by controlled evaluations of artifact-handle storage, temporal scheduling, and backoff middleware. Such studies should compare changes on both failure axes rather than reporting aggregate success alone.
Related Work & Insights¶
- vs TicToc: TicToc concerns stale tool information and re-invocation; ContractBench adds proactive deadline planning and an independent integrity axis. The two are complementary, and byte errors should not be collapsed into staleness.
- vs SWE-bench / Terminal-Bench: These benchmarks assess code or terminal task completion, whereas ContractBench focuses on subsequent use of external API artifacts. Its diagnosis is narrower and more specific, not a measure of general software-engineering ability.
- vs ฯ-bench / ToolBench: Domain-policy and tool-use evaluations are not equivalent to intermediate-artifact time and byte constraints. Contract tests are suitable as an additional regression suite rather than a replacement for existing benchmarks.
- vs Reflexion / Self-Refine: The paper uses deterministic server-generated failure labels instead of model self-critique and adds a wrong-label control. The practical lesson is to distinguish recoverable errors from framework-level problems before choosing feedback or execution protections.
Rating¶
- Novelty: 4/5, separates cross-step artifact constraints into two programmatically verifiable axes.
- Experimental Thoroughness: 3/5, broad model coverage but few repetitions, a small retry sample, and conflicting reporting conventions.
- Writing Quality: 3/5, clear problem definition but protocol and appendix inconsistencies complicate reproducibility judgments.
- Value: 4/5, useful for agent-version regression diagnosis and reliable execution-framework design.