Skip to content

CTE-Bench: Counterfactual Trace Evaluation for Stateful Software Simulators

Conference: NeurIPS2026 (note archive; the read version is arXiv v1, which does not establish acceptance)
arXiv: 2609.36647
Area: Code Intelligence / LLM Evaluation
Keywords: stateful software, counterfactual traces, executable oracle, effect steps, free rollout

TL;DR

CTE-Bench executes software with and without an intervention to evaluate sustained response prediction on fixed future calls: effect-step accuracy reaches 54.3%โ€“61.5% with correct earlier answers, falls to 24.8%โ€“33.2% in free rollout, and whole-trace exact match peaks at just 1.2%.

Background & Motivation

Coding agents often modify a service that has already been running, rather than a function starting from an empty state. Bank ledgers, expiring tokens, shopping carts, file systems, and rate limiters retain history. Editing a condition can first change whether a request is accepted and then affect later calls through a state update that never occurred. Recognizing a code difference, or predicting whether one request succeeds, therefore does not establish an ability to anticipate the patched service's responses.

Existing evaluations cover neighboring parts of this problem. CRUXEval and REval examine function or program execution, Executable Counterfactuals introduces executable interventions, and ToolSandbox evaluates agent behavior in stateful environments. When an agent chooses its own actions, however, failure can originate in planning, argument selection, or environment prediction, making software-dynamics understanding difficult to isolate. This paper removes action selection: future calls are fixed, the model only predicts responses, and deterministic Python services determine correctness by execution.

The setting also exposes an evaluation shortcut: most future calls may be unaffected by the intervention, so ignoring the edit and replaying the original service can achieve high overall accuracy. The authors retain a no-intervention trace as a control and vary whether the model receives correct answers to earlier future calls. Core idea: use paired executable traces to identify calls genuinely affected by an intervention, and report one-step prediction with correct-answer assistance separately from continuous simulation based on the model's own predictions.

Method

Overall Architecture

Each scenario provides service source code, observed pre-intervention calls and responses, one source edit or state overwrite, and 40 predetermined future calls. The model receives one call to predict per request and outputs its complete JSON response. It neither executes code nor chooses the next action or changes the environment. Source-edit scenarios show the complete edited source along with the changed line and tokens.

On the evaluation side, a fresh service instance executes the observed history to recover the real state at the intervention point, applies the intervention, and executes the future calls to produce the target trace. A no-intervention execution branch starts from the same historical state and executes the same future calls. These traces support grading and effect-step identification; they are not tools accessible to the model. Only the answers-revealed protocol inserts earlier target responses into later prediction contexts.

This is a benchmark and evaluation protocol, not a new neural network. Its core designs are post-history intervention, paired-trace grading, three memory protocols, and auditable scenario and statistical boundaries. They need not be recast as training modules or an invented prediction network. Three roles must remain distinct: the model makes text predictions, the executor maintains real service state, and the grader compares their JSON responses.

Key Designs

1. Post-history intervention: change runtime rules without rewriting the past

A source intervention does not mean running the edited program from the beginning. The observed history executes under the original program, and its resulting state is retained. Only then is one code token replaced, for example by inverting a comparison or changing a numeric threshold. An edit is admitted only if the service can load the existing state and replay deterministically. This constraint represents patching a running service, rather than reinterpreting its entire history.

A state intervention instead retains the source and directly modifies stored state at the same point. Core-v1 uses service-clock jumps and rate-limit counter overwrites; the authors do not claim these are the most common developer changes. Moving the clock can expire existing tokens, and overwriting a counter can alter subsequent rate-limit decisions. Knowing the edit alone is insufficient: the model must infer the intervention-time state from the observed history.

The prediction target is the future visible response, not an explicit internal-state report at every step. A candidate is excluded from Core-v1 if no response changes anywhere in its stored trace, even when hidden state changes. This selection ensures observable intervention effects in the main collection, but does not represent a natural distribution over all possible edits.

2. Paired-trace grading: measure intervention understanding where responses actually change

When an edit affects only a few calls, overall accuracy rewards behaving as if it did not exist. The authors execute the same future calls twice to obtain intervention and no-intervention traces. Positions where their responses differ are effect steps. These need not be calls that directly execute the edited line: if a rejected request leaves a counter unchanged, a later counter read may also be an effect step.

The following notation expresses the paper's metric definition rather than introducing a training objective. Let \(y_{s,t}^{\xi}\) and \(y_{s,t}^{0}\) denote the intervention and no-intervention responses in scenario \(s\) at future call \(t\), and let \(\hat y_{s,t}\) be the parsed model prediction. The main score is:

\[ \mathcal E=\{(s,t):y_{s,t}^{\xi}\ne y_{s,t}^{0}\},\qquad \mathrm{VM}_{\mathrm{effect}}=\frac{1}{|\mathcal E|}\sum_{(s,t)\in\mathcal E}\mathbf 1[\hat y_{s,t}=y_{s,t}^{\xi}]. \]

The denominator contains all effect steps within the scored window: 2,476 in Core-v1. An unparseable JSON response is not correct. Value match (VM) requires equality of the complete parsed response, not character-by-character equality of output formatting. All-call VM instead uses all 10,200 predictions, so it answers a different question from effect-step VM.

Two diagnostic metrics retain other granularities. Status match (SM) compares success versus error only; field match flattens each oracle response into scalar fields and awards partial credit. SM can conceal wrong balances, counts, paths, or expiry times, while field match reveals near misses with only some incorrect fields. Both help explain VM but cannot replace the intervention-effect metric.

3. Three memory protocols: separate external correct feedback from self-generated predictions

Answers revealed includes the most recent 20 calls, whether observed or future. Earlier future calls carry the executor's correct responses. An incorrect prediction cannot contaminate the next prompt, so the score measures one-step prediction conditioned on a correct recent history, not independent simulation of the entire service.

No feedback includes all observed calls but omits earlier future calls. Each response must be predicted from the observed history, intervention, and current query alone. The protocol removes not only correct responses but also earlier future-call context. It is therefore not merely answers revealed with the answers masked, nor equivalent to supplying the full call prefix while requiring the model to maintain state independently.

Free rollout uses the same 20-call window as answers revealed, but replaces earlier future responses with the model's own predictions. The real executor still generates target responses along the fixed call sequence; incorrect predictions do not change grading ground truth. They change the history available to later model requests. This reveals how mistaken expectations persist, but differs from a real agent that selects new actions in response to outcomes.

Every protocol makes one request per future call rather than generating the entire trace in one request. The authors recommend attaching a protocol name to every score. Free rollout is closer to an agent relying on its own expectations, while answers revealed compares next-response prediction with correct recent feedback. Differences also involve memory content and feedback source, so score gaps cannot directly identify a single cognitive component.

4. Auditable scenario and statistical boundaries: retain ground truth and clarify scored coverage

The six services implement authentication, bank ledgers, shopping carts, an in-memory file system, rate limiting, and reservation quotas. Of 255 scenarios, 210 use source interventions: 150 branch inversions and 60 threshold changes drawn from 49 distinct patches. Another 45 use state interventions: 30 clock jumps and 15 counter overwrites. Each service contributes 30 scenarios except ratelimiter, which contributes 105 across four intervention cells. The overall average is not a uniformly weighted service average.

Observed histories contain 5โ€“56 calls, and only the first 40 future calls are scored. Stored executable traces are longer; scenario selection requires at least one changed response in that longer trace. Changes fall inside the scored window in 247 scenarios, while the first change occurs later in the remaining 8. Those 8 contribute to all-call VM but not effect-step VM; they are not interventions that never have any effect.

The release retains both traces and final service states, enabling offline regrading without an LLM judge. Confidence intervals use 2,000 bootstrap resamples of scenarios instead of treating the 40 responses within a scenario as independent samples. Service and intervention breakdowns are descriptive failure profiles, not module ablations; they do not separately identify causal contributions from state inference, rule application, execution simulation, or output formatting.

Key Experimental Results

Main Results

All four configurations are evaluated through hosted APIs. DeepSeek V4-Flash uses the deepseek-chat compatibility alias in non-thinking mode; Kimi K2.5 disables thinking; Qwen3.6-35B-A3B requests reasoning off and audits usage; Claude Sonnet 4.6 uses Bedrock defaults without explicitly requesting thinking. The following values come from Table 2. Units are percentages, and \(\pm\) denotes the half-width of a scenario-bootstrap 95% interval, not a standard deviation.

Model Answers revealed: effect-step VM No feedback: effect-step VM Free rollout: effect-step VM Answers revealed: all-call VM
DeepSeek V4-Flash 60.7 ยฑ 3.2 24.6 ยฑ 3.5 27.1 ยฑ 4.0 74.1 ยฑ 1.8
Kimi K2.5 61.5 ยฑ 3.1 28.9 ยฑ 3.8 33.2 ยฑ 3.7 71.8 ยฑ 1.4
Qwen3.6-35B-A3B 54.3 ยฑ 2.8 23.2 ยฑ 3.2 24.8 ยฑ 3.4 65.6 ยฑ 1.4
Claude Sonnet 4.6 58.4 ยฑ 3.7 25.0 ยฑ 3.5 27.3 ยฑ 4.0 61.8 ยฑ 2.3

The highest and lowest effect-step VM point estimates under answers revealed differ by 7.2 percentage points, whereas removing correct feedback produces drops of approximately 28โ€“36 points. The answers-revealed intervals overlap for DeepSeek, Kimi, and Sonnet; Qwen's also overlaps Sonnet's. Point estimates therefore do not establish a definitive model ranking.

Ablation Study

The paper contains no training-module removal ablations; protocol comparisons change evaluation conditions. The next two tables select the non-model controls from Table 1 and rollout analysis from Table 3 to examine whether scoring detects interventions and when errors begin accumulating.

Non-model predictor All-call VM (%) Effect-step VM (%) All-call SM (%) Interpretation
Replay without intervention 75.7 0.0 91.6 High overall score without predicting any changed response
Copy the last observed response 1.5 1.1 72.3 Correct status does not imply a correct payload
Predict the most common observed status 0.0 0.0 76.8 Success/error prediction alone can achieve relatively high SM
Model Exact 40-step whole trace (%) Median first wrong call All-call VM, calls 1โ€“10 (%) All-call VM, calls 21โ€“40 (%)
DeepSeek V4-Flash 1.2 3 66.9 53.6
Kimi K2.5 0.0 2 61.3 47.5
Qwen3.6-35B-A3B 0.0 2 53.3 39.2
Claude Sonnet 4.6 0.0 2 40.4 24.4

The rollout table applies only to free rollout. First-error positions are counted from future call 1 and include parse failures. DeepSeek's 3 fully correct scenarios are excluded only from the mean and median first-error positions, not from whole-trace exact match.

Key Findings

  • Replay without intervention has higher all-call VM than all four models under answers revealed, but zero effect-step VM. This does not demonstrate patch understanding: only 2,476 of the 10,200 calls are affected by the intervention.
  • Conditioning on self-generated predictions can outperform having no earlier future-call context, but remains far below correct feedback. The first free-rollout error typically occurs at call 2โ€“3, and later accuracy continues to decline. Whole-trace exact match alone provides almost no leaderboard separation.
  • Formatting cannot explain the entire protocol gap. The first three models have at most 1.4% parse failures under any protocol; Sonnet has 14.3%, 13.5%, and 16.7% respectively. Formatting substantially lowers Sonnet's absolute scores, so not every failure should be interpreted as a state-reasoning error.
  • The source contains a statistical inconsistency: the main text reports 33/10,200 unparseable DeepSeek responses under no feedback, but Appendix H Table 12 reports 0.0% parse failures and 100.0% parseability. The ratio is about 0.32%, which one-decimal rounding cannot explain. This note preserves the conflict rather than correcting it without evidence.

Highlights & Insights

  • Use an intervention/no-intervention control to expose scoring shortcuts. Effect steps are positions where executed responses differ, not merely where the edited line executes. This captures indirect effects visible in later reads after a rejected update.
  • Treat correct feedback as an evaluation condition, not an assumed model capability. The same model performs very differently with correct recent history and with its own expectations. Tool-use evaluations should likewise distinguish externally returned state information from internally maintained state.
  • Pair strict response equality with diagnostic metrics. The main score does not overlook wrong balances or counts, while SM and field match clarify different failure levels. This avoids both overly permissive success/error grading and treating all non-exact responses as identical failures.

Limitations & Future Work

  • Coverage is narrow and uneven. Six deterministic Python services do not represent concurrency, randomness, or distributed-system complexity, and ratelimiter supplies 105/255 scenarios. Equal-service weighting is a sensitivity check, not a substitute for broader sampling.
  • Protocols are not pure single-variable ablations. No feedback also omits earlier future calls, while free rollout and answers revealed use a fixed 20-call window. Window-size and prompt-wording sensitivity are unmeasured, so the entire gap cannot be attributed solely to feedback correctness.
  • Effect-step VM cannot penalize every invented change. On the separate 106-scenario no-effect probe, answers-revealed VM is 81.6% for DeepSeek, 83.0% for Kimi, and 82.3% for Sonnet, while replay should achieve 100%. Its coverage does not match Core-v1, and Qwen is not reported; these results cannot be merged into the main leaderboard.
  • Error sources remain entangled. Revealing the true intervention-time state, allowing controlled code execution, and recording the first state divergence could separate historical-state inference from execution simulation. These are proposed follow-ups, not mechanisms already established by the paper.
  • Results do not certify production deployment. The main evaluation covers four API configurations and one prompt format; the 24-scenario reasoning-mode probe cannot establish general superiority. High scores do not imply safe agent edits or generalization to other languages and real action policies.
  • vs CRUXEval / REval: These study execution-related program predictions; CTE-Bench emphasizes persistent state across calls and interventions applied after history. Extensions should preserve runtime state rather than reduce services to independent function problems.
  • vs Executable Counterfactuals: Both use executable programs to define counterfactual semantics. The key extension here is retaining existing state, applying edits only to the future, and measuring response changes across multiple calls.
  • vs ToolSandbox / Agent-Diff / TRAJECT-Bench: These evaluate actions, tool calls, or final states. CTE-Bench fixes future actions to isolate response prediction; it is a complementary diagnostic, not a replacement for end-to-end agent success.
  • Research direction: Build coverage-matched effect and no-effect scenario pairs to measure both missed changes and invented changes. Cross them with true-state revelation to test whether models lack state recovery or rule execution. This follows directly from the paper's limitations and is not a validated method.
  • Resources: CTE-Bench-Core-v1 dataset. The paper reports releasing scenarios, executable services, graders, reference outputs, and an evaluation card; their availability was not checked online for this note.

Rating

  • Novelty: 4/5 โ€” Post-history interventions, fixed future calls, and effect-step scoring define a clear combined task.
  • Experimental Thoroughness: 3/5 โ€” Three protocols and non-model controls are informative, but service coverage, prompt sensitivity, and error attribution remain limited.
  • Writing Quality: 4/5 โ€” The evaluation card separates supported and unsupported claims; the parse statistics require clarification.
  • Value: 4/5 โ€” Useful for diagnosing software-state prediction in coding agents, not for deployment safety certification.