Skip to content

CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models

Conference: NeurIPS2026
arXiv: 2602.17684
Area: Code Intelligence
Keywords: code reward models, verifiable rewards, preference data, reward shaping, test-time compute

TL;DR

CodeScaler distills high-quality execution feedback into a code reward model, uses strict code extraction and reward shaping to support reinforcement learning and candidate reranking without online execution, and exceeds RLVR by 1.55/4.23 percentage points in average Avg@8 across four benchmarks for 8B/14B policies trained on DeepCoder.

Background & Motivation

Code generation offers executable feedback: running a program in a sandbox and checking its tests provides supervision more directly grounded than language preferences. Methods such as DeepCoder use reinforcement learning from verifiable rewards (RLVR), but their bottleneck is not the supply of problem descriptions; it is the supply of trustworthy, high-coverage tests. Generating problems is easier than generating tests that distinguish correct algorithms from corner-case failures. Even agreement between generated tests and several model solutions can reflect shared omissions. In Appendix K, a problem permits inputs up to 30, but its tests cover only inputs up to 9, allowing exponential-time enumeration to be mislabeled as correct.

A reward model (RM) can read the problem and candidate code and produce a continuous score, potentially transferring a limited amount of reliable execution supervision to new problems without tests. However, replacing the reward is not merely an interface change: an executor rejects broken code, whereas an RM may reward syntactically invalid, irrelevant, or superficially sophisticated programs. Both the general-purpose SkyworkRM and code-specific AceCodeRM underperform RLVR in the paper's RL experiments. At inference time, they avoid test generation and sandbox execution but may still select the wrong solution.

The paper addresses data quality and reward integration together: it first learns that correct solutions should outrank incorrect solutions from actual policy-training trajectories, then assigns unparseable outputs an explicit minimum reward so that the policy cannot exploit malformed formatting. Core Idea: train a reusable proxy for code correctness on execution-verified on-policy preference data, then use syntax gating and positive-valued reward shaping to deploy the same RM for policy optimization and Best-of-N selection without online execution.

Method

Overall Architecture

The inputs are programming problems and policy-generated responses, and the system produces two outputs: a code policy trained with RM rewards and a scorer for reranking multiple candidates. The pipeline starts with Verified-Trajectory Preference Learning to obtain CodeScaler. The subsequent RL branch applies Syntax-Aware Code Extraction and Validity-Preserving Reward Shaping before using the reward for GRPO optimization through Dual Use in Training and Inference. The inference branch instead uses the trained RM to score candidates and select the highest-scoring solution, without generating or executing tests for those candidates.

The dependencies of these stages must be distinguished: RM training labels come from actual execution on DeepCoder tests, whereas deployed RM scoring does not require tests. Execution-free refers to subsequent scoring and policy optimization, not to a system that never used tests, nor to evaluation that dispenses with execution verification. Solid edges below represent data or reward flow; dashed edges represent reuse of the trained RM.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["DeepCoder tests<br/>and RLVR trajectories"] --> B["Verified-Trajectory<br/>Preference Learning"]
    C["Problem and policy response"] --> D["Syntax-Aware<br/>Code Extraction"]
    B -.->|trained RM| E["Validity-Preserving<br/>Reward Shaping"]
    D --> E
    E -->|RL reward| F["Dual Use in Training<br/>and Inference"]
    B -.->|inference candidate scoring| F
    F -->|training branch| G["GRPO policy update"]
    F -->|inference branch| H["BoN selects highest-scoring code"]

Key Designs

1. Verified-Trajectory Preference Learning: establish trustworthy correctness labels before expanding the scorer's use

The authors use DeepCoder's 24K problems as the supervision seed and collect on-policy training trajectories from Qwen3-8B-Base optimized with GRPO/RLVR, rather than simply sampling solutions independently from several fixed models. As the policy changes, the trajectories cover solutions and mistakes from different capability stages, exposing the RM to code distributions that subsequent optimization may actually produce. For each problem, solutions passing every test are positives, while any test failure makes a solution negative. These labels still establish correctness relative to the available test suite, not a mathematical proof of program correctness.

Appendix B specifies the sampling procedure: train the policy for 250 steps, sample 8 responses per problem, and discard trajectories from the first 100 steps, yielding 52,574 preference pairs. To prevent the RM from judging only whether a program resembles good code, approximately 30% additional misaligned negatives pair a problem with a solution to another problem. A well-structured program solving the wrong task must therefore rank below a correct solution to the actual task, adding semantic correspondence between specifications and implementations to the supervision target.

CodeScaler-8B continues training from Skywork-Reward-V2-Qwen3-8B rather than starting from random parameters. Its Bradleyโ€“Terry objective requires a positive solution to score above a negative solution for the same problem, with a larger margin indicating greater confidence in that ordering. This learns relative rankings, not calibrated probabilities of passing tests. Scores may be negative, and the loss itself does not prescribe a minimum score for invalid code.

2. Syntax-Aware Code Extraction: restrict scoring to a single parseable program

Execution in RLVR naturally exposes parsing failures, whereas an RM lacks that constraint. The authors therefore require a response to contain a single well-defined code block; multiple fragmented blocks are rejected rather than concatenated. The extracted code then undergoes a static abstract syntax tree (AST) check. Failure at either stage replaces the extraction with an empty string instead of allowing the RM to score problematic fragments freely.

This gate constrains output format and syntax, not algorithmic semantics. A program with incorrect boundary conditions, reversed traversal, or excessive complexity can still pass the AST check. The gate blocks one inexpensive shortcut: obtaining abnormal rewards through broken formatting or incomplete syntax. Its cost is rejecting multi-block responses that a human might otherwise combine, making the method better suited to coding evaluations expecting a single complete solution.

3. Validity-Preserving Reward Shaping: rank syntactically invalid outputs below every valid output

Simply assigning invalid outputs a reward of 0 is insufficient because valid code can receive negative Bradleyโ€“Terry scores. If valid code scores below 0 while an empty output receives 0, the policy can prefer the empty output. The authors apply softplus to the raw score of valid code and assign 0 to extraction or AST failures. In the equation below, โ€œvalidโ€ means only that the code passed the extraction and syntax checks described above.

\[ R'(q,c)=\begin{cases} \ln(1+e^{r_\phi(q,c)}), & c\text{ is valid},\\ 0, & \text{otherwise}. \end{cases} \]

Valid code consequently receives strictly positive rewards, whereas invalid code always receives 0. Softplus is monotonic, preserving rankings among valid candidates. It does not automatically give an incorrect algorithm a low score; it establishes a stable reward floor for syntactic validity. The transformed scale also affects within-group reward differences in GRPO, so preserving rankings does not mean leaving RL dynamics unchanged.

4. Dual Use in Training and Inference: reuse a correctness proxy on new problems rather than reuse their tests

In the training branch, the policy generates a group of responses for a problem, extracts their code, and obtains the shaped rewards. GRPO forms advantages from relative rewards within the group and updates the policy. The authors retain KL regularization against a reference policy to limit excessive pursuit of RM preferences when synthetic problem quality is not strictly controlled. Once the RM is trained, subsequent training problems need only descriptions and policy responses, not execution environments or test labels for every problem.

This enables data expansion: the authors extract concepts from TACO problems, connect co-occurring concepts in a graph, sample combinations through random walks of up to six steps, and ask GPT-4o to generate 20K new problems. Combining them with 24K DeepCoder problems yields a 44K training set. The concept graph generates new problems; it is not an internal reasoning structure of the RM, and the RM does not need the graph or test cases during RL.

In the inference branch, the policy first generates multiple candidates, and CodeScaler reads the problem and each candidate's code to select the highest-scoring one directly. There is no policy update or GRPO group-relative advantage computation: the RM is a reranker, not a candidate generator. Candidate generation still takes time, and the highest RM score does not guarantee correctness. The branches share a scorer, but the RL branch's syntax gating and shaping should not be assumed to be additional processing in every BoN experiment without explicit evidence.

A Worked Example

Consider a problem asking for the minimum number of swaps required to arrange two kinds of balls in a specified order. In Appendix I, the correct program scans left to right and counts black balls to the left of each white ball. The incorrect program traverses in the opposite direction and effectively solves the reversed goal. Both form complete programs, so syntax gating does not reject the incorrect one; the RM must recognize the relationship between the requested arrangement and the counting direction.

In the reported case, the correct program contains 278 characters and receives an RM score of 3.969. The incorrect program contains 661 characters but receives 7.531. BoN would select the wrong program, and monotonic softplus would preserve this incorrect ordering if the candidates were used for RL rewards. The example separates the components' responsibilities: AST checks address parsing, shaping establishes a reward floor for invalid outputs, and preference learning handles semantic judgment, which can still be misled by length and keyword density.

Loss & Training

The RM training loss is shown below; \(q\) denotes a problem, \(c^+\) and \(c^-\) are its positive and negative candidates, and \(\sigma\) is sigmoid.

\[ \mathcal{L}_{\mathrm{RM}}=-\mathbb{E}_{(q,c^+,c^-)\sim\mathcal{D}}\left[\log\sigma\left(r_\phi(q,c^+)-r_\phi(q,c^-)\right)\right]. \]

The RM uses AdamW with a learning rate of \(10^{-6}\). The cached RM epoch field contains an ambiguous repeated character, so no definite epoch count is inferred here. RL also uses a learning rate of \(10^{-6}\), with KL coefficient \(\beta=0.005\), batch size 128, and mini-batch size 64. Runs normally last 250 steps, sampling 8 responses per problem with a maximum response length of 16,384 tokens. Training temperature is 0.6, top-p for RM data collection is 0.95, and experiments use 8 A100 80GB GPUs.

Group-relative advantages subtract the group mean from shaped rewards and normalize by the group standard deviation; the corrupted GRPO advantage equation in the cache is not reconstructed here. Rising mean rewards alone cannot establish improving program correctness, which requires an independent execution check. Appendix G extends 8B training on DeepCoder to 650 steps and reports RM scores and pass@1 increasing together. This supports the absence of obvious reward hacking in that configuration, not a guarantee for arbitrary training distributions.

Key Experimental Results

Main Results

Trained policies are evaluated with Avg@8: execution correctness averaged over 8 independent generations per problem, not BoN@8 selection of one solution from 8 candidates. Evaluation temperature is 0.6 throughout. LiveCodeBench uses problems dated 2024-08-01 through 2025-02-01; CodeContests uses 239 sampled problems with difficulty no greater than 2; CodeForces uses 467 sampled problems; MBPP uses its standard test set.

The table retains the DeepCoder training results from Table 1, in percent. The average is the mean across four benchmarks, not an accuracy weighted by pooling all their problems.

Policy and reward LiveCodeBench CodeContests MBPP CodeForces Average Avg@8
Qwen3-8B-Base 13.75 19.03 61.70 5.35 24.96
8B + RLVR 23.60 32.00 76.01 16.43 37.01
8B + CodeScaler 24.80 33.94 75.90 19.61 38.56
Qwen3-14B-Base 21.37 26.25 71.09 8.45 31.79
14B + RLVR 25.94 34.30 75.45 20.79 39.12
14B + CodeScaler 27.55 39.33 81.61 24.89 43.35

The 8B average improves by 1.55 percentage points, but MBPP decreases from 76.01 to 75.90, so not every cell improves. With 44K training problems, 8B + CodeScaler reaches an average of 39.60: a 14.64-point improvement over the base model's 24.96, but only a 1.04-point increase over the DeepCoder-only result of 38.56. These reference points answer different questions.

Ablation Study

The following table uses Section 5.3 and the text accompanying Figure 4, reporting Avg@8 after training Qwen3-8B-Base on DeepCoder. โ€œExtraction + shapingโ€ is a joint intervention; it does not isolate the contribution of AST checks or softplus.

Reward configuration LiveCodeBench CodeContests MBPP CodeForces
SkyworkRM 18.50 23.22 67.59 8.00
SkyworkRM + extraction + shaping 20.74 27.45 69.79 10.00
CodeScaler + extraction + shaping 24.80 33.94 75.90 19.61

Extraction and shaping improve the general-purpose RM but do not close the gap to CodeScaler, indicating that both reward integration and preference data quality matter. Another ablation labels solutions with a pass ratio greater than 0.7 as positives even when they are not fully correct. Its average RL Avg@8 is 37.94, below strict binary labeling's 38.56, but its CodeContests and CodeForces scores of 35.04 and 21.44 exceed the binary variant's 33.94 and 19.61. Strict labels therefore win on the reported average, not on every benchmark.

Inference latency comes from Table 4: the same 467 CodeForces problems, 8 candidates per problem, and 8 additional tests per problem for CURE, using one A100 with vLLM. The table measures selection-stage overhead and explicitly excludes shared candidate-generation time.

Selection-stage time CURE CodeScaler
Total test generation (seconds) 979.3 Not used
Total execution (seconds) 516.7 Not used
Total RM scoring (seconds) Not used 146.1
Average per problem (seconds) 3.20 0.31

The text initially defines latency as end-to-end wall-clock time including candidate generation, but the actual comparison subsequently excludes the shared generation term. The approximately 10-fold claim therefore applies only to the measured selection stage, not the complete generation pipeline. As shared generation costs grow, the full-pipeline speedup approaches 1.

Key Findings

  • On actual DeepCoder trajectories, correct-versus-incorrect pairwise ranking accuracy is 87.9%, and the highest-scoring candidate in a group is correct 79.8% of the time. The RM discriminates useful differences but is far from a reliable verifier.
  • Training the RM on DeepCoder rather than KodCode trajectories, then using it for RL on rStarCoder, increases CodeForces Avg@8 from 11.58 to 15.95. This supports prioritizing trustworthy supervision over merely expanding trajectory sources.
  • Increasing misalignment augmentation from 0% to 30% raises AUROC from 0.911 to 0.990 in Appendix H's separate held-out analysis; 50% still yields 0.990. These results should not be directly conflated with the actual-trajectory AUROC of 0.879 above.
  • BoN tables contain an unexplained reporting difference: Table 9 gives average BoN@8 of 42.88 for the Binary variant, whereas Table 10 gives 44.15 for the 8B variant. The paper does not explicitly reconcile them, so they are not merged into a single result.

Highlights & Insights

  • The value of tests shifts from execution at every optimization step to establishing a trustworthy scorer first. This does not eliminate dependence on tests; it reuses previous verification effort on subsequent problems without tests.
  • Positive-valued shaping targets a concrete optimization loophole: an invalid output's reward of 0 can outrank a valid program's negative score. Clarifying the numerical meaning of rewards before integrating RL is more direct than merely asking the model to follow an output format.
  • Training rewards and inference ranking impose different pressures. Discriminating preferences in a static dataset does not establish reliability when a changing policy actively pursues high scores, making joint evaluation of RL stability and BoN more informative.

Limitations & Future Work

  • RM supervision still depends on test coverage. If a program passing all tests misses a specification boundary, strict labeling still makes it positive. Passing an AST check must not be interpreted as proof of semantic correctness or appropriate complexity.
  • Appendix I identifies surface-level biases: incorrectly high-scored programs are 28% longer and contain 37% more algorithmic keywords than correctly high-scored programs; concise, modular, or recursive correct solutions can be undervalued. Reward shaping does not repair these ranking mistakes.
  • Reward-hacking analysis mainly covers extended 8B training on DeepCoder, not all new synthetic distributions, languages, or longer runs. Execution audits excluded from RM training could continuously monitor divergence between rising scores and declining correctness.
  • Multilingual and non-Qwen BoN results support limited transfer, not cross-language RL stability, repository-level repair, or reliability for safety-critical code. Latency benefits also require remeasurement in a deployment pipeline that includes candidate generation.
  • vs DeepCoder / RLVR: These methods reward policies directly through test execution; CodeScaler first learns a proxy from such execution trajectories and then uses it to optimize on new problems. Reduced online testing requirements come at the cost of judgment errors that a policy may exploit.
  • vs AceCodeRM: Both learn code preferences, but this paper emphasizes passing every test for positive labels and adds syntax gating and shaping for RL. Ablations support an average advantage for strict labels without attributing every difference to the labeling threshold.
  • vs CURE / CodeT: Test-based selection depends on generated tests and execution feedback, whereas CodeScaler depends on learned problemโ€“code correspondence. A possible extension is to execute only close high-scoring candidates or candidates with uncertain RM judgments, then measure the budget curve between latency and selection errors.

Rating

  • Novelty: 4/5, combines verified-trajectory preferences with RL-compatible reward integration rather than introducing a new preference loss.
  • Experimental Thoroughness: 4/5, covers training, BoN, scale, and failure analysis, but some reporting differences and generalization boundaries remain unresolved.
  • Writing Quality: 4/5, clearly presents the mechanism, but end-to-end latency wording must be distinguished from the actual timing scope.
  • Value: 4/5, offers a reusable approach to code post-training with limited testing resources, without replacing final functional verification.