Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization¶
Conference: NeurIPS2026 (acceptance reported by the authors)
arXiv: 2609.39402
Area: LLM Reasoning
Keywords: reinforcement learning with verifiable rewards, credit assignment, proximal entropy, advantage weighting, mathematical reasoning
TL;DR¶
PEPO weights token updates in successful reasoning trajectories by uncertainty relative to their neighborhoods while preserving each trajectory's total weight, outperforming the compared GRPO and global-entropy baselines in the main mathematical reasoning experiments, with explicit limits on trend robustness and cross-task generalization.
Background & Motivation¶
Reinforcement learning with verifiable rewards (RLVR) can supervise a large language model (LLM) using final-answer correctness, but does not identify which reasoning steps determined the outcome. Group Relative Policy Optimization (GRPO) estimates advantages by comparing rewards across multiple responses to the same prompt, avoiding a value network. However, every token in one response inherits the same advantage. Connective words, mechanical derivations, and branching decisions within a successful response are therefore reinforced together. Value models, process reward models, or additional rollout branches can provide finer signals, but increase training, memory, or sampling costs. This paper instead uses token entropy from the existing policy distribution as a lightweight proxy: positions with several plausible continuations may deserve more attention than deterministic grammatical continuations.
High entropy, however, does not imply importance alone. The 80/20 method selects the highest-entropy 20% of tokens across the training batch, potentially placing routine tokens from difficult prompts ahead of crucial decisions from easier prompts. Entropy also tends to be higher near the beginning of reasoning than near the end. On Llama-3.2-3B-Instruct, the paper groups prompts by the number of successes among 8 generations: hard prompts have 0–1 successes, moderate prompts 2–6, and easy prompts 7–8. Their mean raw entropies are 0.315, 0.270, and 0.221, respectively. This difference shows that global ranking mixes in a trajectory-level entropy baseline; it does not establish that every token in a difficult prompt is more important. The positional analysis likewise finds that early raw entropy can be about 80% above the sequence mean.
Rather than training another token-value estimator, the paper changes the reference frame: a token should first be compared with nearby generation positions before receiving a share of the update weight within its successful response. This requires no additional correctness labels, but still produces an importance proxy rather than each token's true causal contribution. Core Idea: construct neighborhood-relative weights using an entropy softmax over a centered sliding window, renormalize them across the full sequence, and apply them only to successful trajectories to reduce interference from absolute entropy levels in token-level credit assignment.
Method¶
Overall Architecture¶
Proximal Entropy Policy Optimization (PEPO) retains GRPO's sampling, answer verification, and clipped policy objective. For each prompt, the policy generates a group of responses; a verifier supplies binary rewards, and group normalization produces a response-level advantage. Shannon entropy at each position is obtained from the policy distribution during generation. The additional processing is confined to training-weight computation: local entropy comparison, sequence-level renormalization, and advantage modulation on successful trajectories.
Two normalization levels must be distinguished. GRPO compares rewards across responses to the same prompt to decide which responses should be reinforced or suppressed. PEPO compares relative uncertainty across tokens within a successful response to decide how its existing advantage should be allocated. The latter cannot turn an incorrect answer into a correct one and does not generate a new process-correctness score.
The input is a completed trajectory together with per-position entropy, and the output is token-level advantages for the policy update. Because a centered window includes entropy after the current token, this is a training-side computation performed after a complete rollout, not an online decoding rule restricted to the available prefix. At deployment, the trained policy still generates autoregressively without additional verifier calls or rollout branches.
The three key designs are, in order, “Local Relative Entropy,” “Sequence-Level Renormalization,” and “Successful-Trajectory-Only Weighting.” This is an advantage-weighting/objective modification rather than a multi-stage network architecture. The formulas below explain the weight flow without presenting sampling and loss computation as new modules.
Key Designs¶
1. Local Relative Entropy: compare tokens with their neighbors instead of ranking them against the whole batch
The method first obtains policy entropy at each generation position. This entropy comes from the next-token probability distribution over the entire vocabulary, not the negative log-probability of the sampled token. The former measures overall uncertainty among possible continuations; the latter describes only the outcome already sampled. Let \(H_{i,t}\) denote entropy at position \(t\) of response \(i\), and let \(\mathcal{W}(t)\) be its centered window. The default window size is \(W=101\), covering the current token and 50 positions on either side at interior positions.
The core definition of proximal entropy is:
This quantity is not a newly defined Shannon entropy. It is the central position's relative share in a softmax over neighborhood entropies. When neighboring entropies are nearly identical, the center receives approximately the uniform window share; a position that is more uncertain than its neighbors receives a larger share. Thus, a token with modest absolute entropy can receive relatively high weight when it forms a local spike in an otherwise deterministic region. Conversely, an ordinary token in a uniformly high-entropy region does not automatically dominate merely because the entire region is uncertain.
The appendix's temperature analysis divides entropy in the exponent by a temperature, with a default of 1. Lower temperature amplifies local differences, whereas higher temperature flattens the weights. This is not the generation sampling temperature. The main method uses neither token-level correctness labels nor a separate correctness reward for high-entropy positions; it uses uncertainty as a proxy on top of the existing success signal.
One property holds exactly: adding the same constant to every entropy in a sequence leaves the local share unchanged because the exponential factor cancels between numerator and denominator. This removes a particular confounder—a constant shift in the entropy level of the whole trajectory—but does not remove every effect of prompt difficulty. Difficulty can still change local entropy variation and therefore the relative weights.
The positional-trend result is only approximate. The appendix decomposes raw entropy into a smooth trend and a token-specific residual, then applies a local Taylor expansion. The trend's constant value at the center cancels, but its slope remains in the denominator. Its effect can be treated as a small correction only when variation inside the window is sufficiently small, for example \(|f'(t)|W/2\ll1\), and the curvature remainder is controlled.
In particular, a “symmetric window” does not imply exact removal of arbitrary linear trends. Although offsets are symmetric, the denominator weights them by unequal residual exponentials, so their weighted sum generally does not vanish. The appendix gives corrections of approximately \(\mathcal{O}(f'(t)W)\) plus the curvature remainder, not unconditional invariance. Smooth monotonicity alone also does not guarantee the appendix's assumption that curvature decreases with sequence length.
Windows at sequence boundaries must be filled. The implementation uses reflect padding, mirroring nearby entropy values from the same trajectory rather than inserting artificially low-entropy neighbors through zero padding. Reflection preserves constant-shift invariance, but folds the original positional trend at the boundary, so the interior symmetric-neighborhood argument does not directly apply. The affected fraction of tokens is approximately \(W/T_i\) for long trajectories, estimated by the authors at about 5%; the fraction is larger for short responses.
2. Sequence-Level Renormalization: redistribute credit without directly increasing a response's advantage budget
The local softmax denominator changes with its center, so different tokens are evaluated in different windows. The sequence of central shares is therefore not a probability distribution summing to 1. Multiplying advantages by these shares directly would alter the trajectory's overall scale and make update strength depend on window size. PEPO renormalizes these shares across the complete response and multiplies them by response length:
The average token weight remains 1, while credit shifts within the response from relatively less uncertain positions to relatively more uncertain ones. More uniform local shares yield weights closer to standard GRPO; pronounced local spikes produce larger differences. Continuous weighting also differs from 80/20's hard selection: the latter excludes many tokens from updating, whereas PEPO preserves a nonzero signal at every position while giving selected positions a larger share.
Conservation of total weight is an algebraic property, not a guarantee that training outcomes are conserved. It preserves the sum of scalar advantage weights within a sequence, but does not preserve gradient norm, actual policy change, or the clipped objective's value. Probability ratios, clipping activation, and parameter gradients differ across positions, so moving advantage weight between positions still changes the optimization direction.
The appendix's sequential ablation helps distinguish these designs. Replacing global-entropy binary selection with proximal-entropy binary selection first tests the reference frame. Replacing binary selection with continuous weighting then tests whether the remaining positions should retain a learning signal. It does not change all factors simultaneously and attribute all gains to locality.
3. Successful-Trajectory-Only Weighting: avoid excessively penalizing uncertain exploration within failures
Answer verification still supplies the original binary reward. PEPO multiplies the trajectory advantage by the preceding weights for reward-1 responses, while retaining the original uniform advantage for reward-0 responses. The implemented rule can be expressed as:
This asymmetry matters. A high-entropy position in a failed response may represent an attempt at an immature solution strategy; concentrating negative advantage there could suppress exploration more strongly. A successful response at least has final-correctness support for reinforcing local decisions. Nevertheless, it can still contain redundant steps, accidental correctness, or steps unnecessary for the outcome.
“Weight successful trajectories only” does not mean failures are excluded from training. Failed tokens remain in the original GRPO objective without entropy modulation. If every reward in a sampled group is identical, there is no reward distinction for the group advantage to exploit; local entropy weights do not create correctness differences. The paper does not present this design as a solution for such uniform-reward groups.
In the appendix's sequential ablation, continuous weighting on all trajectories achieves a mean of 57.53, versus 58.01 for successful trajectories only. This supports retaining asymmetric treatment in this setting. However, the intermediate configuration lacks run-level uncertainty, so the 0.48-percentage-point difference is descriptive rather than evidence of statistical significance or general superiority.
Loss & Training¶
The optimizer still maximizes the standard clipped surrogate, replacing \(A_i\) with \(\hat A_{i,t}\). The probability ratio compares the current and old rollout policies for the same token under the same prefix. Lower and upper clipping thresholds limit excessive individual policy changes. The objective is normalized by the group's total token count, without training an additional value model or process reward model.
The main experiments use ROLL and DeepMath-103K, with 8 responses per prompt and 500 training steps. AdamW uses a learning rate of \(1\times10^{-6}\), with lower/upper clipping thresholds of 0.2/0.28. Sampling temperature and top-p are both 0.6, and maximum response length is 4096. Typical responses contain about 2000–3000 tokens, so the default 101-token window covers approximately 3–5%.
The Single-stream Policy Optimization (SPO) extension is called PESPO. It retains local weighting but uses one rollout per prompt, with advantages supplied by SPO rather than a same-prompt GRPO group. It uses VeRL and DAPO-Math-17K. PESPO can therefore be compared with its baselines within the SPO setting, but differences from the main PEPO experiment cannot be attributed solely to single-stream operation because the framework and data also change.
Local window computation adds no rollouts, and the authors argue that entropy can reuse existing generation-engine computation. Under the same 2-H200, 500-step setup, the appendix reports Qwen3-4B total training times of 1 day 19 hours 25 minutes for GRPO and 1 day 19 hours 45 minutes for PEPO. This supports low added overhead empirically, rather than proving zero overhead at the kernel level.
Key Experimental Results¶
Main Results¶
The following selection from Table 2 reports accuracy percentages as mean ± sample standard deviation across 3 independent runs, not confidence intervals. MATH500 uses avg@1; AIME2024, AIME2025, and AMC use avg@16, the average correctness across sampled responses—not pass@16, which measures whether at least one of 16 samples succeeds. Mean is an equal-weight average of the four benchmark scores, not a problem-count-weighted average.
| Model | Method | AIME2024 | AIME2025 | AMC | MATH500 | Mean |
|---|---|---|---|---|---|---|
| Qwen3-1.7B | GRPO | 16.58 ± 0.44 | 18.65 ± 0.62 | 50.28 ± 0.99 | 80.27 ± 1.01 | 41.45 ± 0.50 |
| Qwen3-1.7B | 80/20 | 20.97 ± 1.48 | 20.83 ± 0.75 | 52.13 ± 0.44 | 80.73 ± 1.50 | 43.67 ± 0.30 |
| Qwen3-1.7B | PEPO | 21.74 ± 1.06 | 22.43 ± 0.67 | 55.26 ± 1.11 | 82.67 ± 0.31 | 45.53 ± 0.46 |
| Qwen3-4B | GRPO | 32.75 ± 0.54 | 23.54 ± 0.76 | 59.96 ± 0.68 | 84.33 ± 1.22 | 50.15 ± 0.43 |
| Qwen3-4B | 80/20 | 35.35 ± 2.47 | 30.62 ± 0.21 | 67.25 ± 1.53 | 89.73 ± 1.15 | 55.74 ± 0.86 |
| Qwen3-4B | PEPO | 39.24 ± 1.27 | 31.81 ± 0.13 | 69.58 ± 0.59 | 91.40 ± 1.83 | 58.01 ± 0.14 |
| Llama-3.2-3B-Instruct | GRPO | 7.22 ± 0.84 | 0.90 ± 0.52 | 21.43 ± 0.57 | 47.67 ± 1.33 | 19.31 ± 0.20 |
| Llama-3.2-3B-Instruct | 80/20 | 8.60 ± 0.61 | 0.49 ± 0.12 | 22.57 ± 0.91 | 47.67 ± 0.50 | 19.83 ± 0.03 |
| Llama-3.2-3B-Instruct | PEPO | 9.10 ± 0.52 | 1.46 ± 0.55 | 24.12 ± 0.36 | 49.33 ± 0.12 | 21.00 ± 0.15 |
The original table also includes Entropy Adv., whose Mean scores for the three models are 43.71 ± 0.52, 54.91 ± 0.23, and 20.02 ± 0.44. Within Table 2, PEPO exceeds the listed trained baselines on all four benchmark means for each model. For Qwen3-4B, Mean increases by 7.86 percentage points over GRPO and 2.27 over 80/20. However, Llama's absolute AIME2025 accuracy remains only 1.46%; relative improvement should not obscure the unresolved task difficulty.
Ablation Study¶
Table 12 changes the Qwen3-4B weighting scheme sequentially. The table below retains the authors' reported incremental-gain convention.
| Config | Four-benchmark Mean (%) | Gain over preceding config (percentage points) | Note |
|---|---|---|---|
| Global entropy, binary selection | 55.74 | — | 80/20 |
| Proximal entropy, binary selection | 56.87 | 1.14 | Local-reference substitution |
| Proximal entropy, continuous weighting on all trajectories | 57.53 | 0.66 | Intermediate config lacks run-level uncertainty |
| Proximal entropy, continuous weighting on successful trajectories only | 58.01 | 0.48 | Full PEPO |
The reported 1.14 averages the four benchmark-level differences in Table 3 before rounding; subtracting the already-rounded Mean values gives 1.13. These conventions should not be mixed. In the Qwen3-1.7B substitution experiment in Table 3, Mean rises from 43.67 to 44.58, but AIME2025 falls from 20.83 to 20.69. Its AIME2024 score of 22.01 also exceeds PEPO's 21.74. Thus, “substitution improves every benchmark” and “PEPO is best on every metric in every comparison” are too strong.
In the window ablation in Table 6, \(W=51,101,151\) produces AIME2024 scores of 36.25, 39.24, and 36.81, and MATH500 scores of 90.67, 91.40, and 92.27. The default 101 has the better four-benchmark average, but 151 performs better on MATH500. Stronger locality is not always better, and window preferences differ across benchmarks.
Key Findings¶
- Single-stream transfer improves the aggregate, but not every metric. In Table 4, SPO/PESPO achieve Mean scores of 57.13 ± 0.42/58.93 ± 0.48, while the global-entropy extensions score 54.65 and 55.99. PESPO's MATH500 score of 90.67 is below S-80/20's 91.00.
- Intervention supports selection efficiency within a limited setting. Table 5 fixes Qwen3-4B Base, the replacement GRPO checkpoint, AIME2024, and the replacement budget, changing only token ranking. At 5% replacement, global/proximal entropy achieves avg@16 of 27.34/31.62. Proximal entropy reaches 33.33 at 10%, whereas global entropy reaches it at 30%. This supports ranking effectiveness under that intervention protocol, not universal causal attribution.
- Coding results require the training boundary to be retained. Table 10 evaluates 1055 LiveCodeBench problems, with GRPO/80/20/PEPO pass@1 of 31.75/41.04/46.35. These are point estimates for one Qwen3-4B held-out split. That evaluation was not used in RL training, but the appendix states that the ROLL training mixture included a separate verifiable coding source, KodCode; the model was not trained exclusively without coding data.
Highlights & Insights¶
- Changing the reference from absolute uncertainty to neighborhood-relative uncertainty addresses a specific credit-assignment confounder more directly than simply increasing entropy rewards. The reusable idea is the reference frame, not an equation between high entropy and correct reasoning.
- Local softmax and sequence-level renormalization serve distinct roles: the former chooses the comparison context, and the latter controls the scalar trajectory budget. Transferring the method across algorithms requires preserving that distinction and rechecking advantage estimation and failure handling.
Limitations & Future Work¶
- There is no comparison with value-model methods such as PPO, so the paper does not establish that its lightweight proxy matches more expensive credit estimators. Evidence primarily covers mathematical reasoning in small models, with limited coding generalization.
- Constant shifts are exactly removed, but positional trends receive only approximate robustness under small slopes and controlled curvature. Reflect padding and short sequences particularly warrant separate evaluation.
- Appendix E claims that a simple window average also removes positional trends, without specifying an additional detrending transformation. Ordinary moving averages generally retain slowly varying trends, so its text and caption do not establish a general mathematical result.
- The main text describes DeepMath-103K training, whereas Appendix F states that the ROLL mixture includes KodCode. The training-data composition is incompletely described. Clearer mixture proportions, boundary analyses, and multi-seed coding results would improve reproducibility.
- Appendix B's window examples are 31/51/101, corresponding to the bias analysis in Figure 4. The performance ablation in Table 6 uses 51/101/151; these should not be merged into one experiment.
Related Work & Insights¶
- vs GRPO: PEPO retains group-relative reward advantages and the clipped objective, redistributing token weights only within successful responses; it does not supply a true step-level value function.
- vs 80/20 and Entropy Adv.: The former uses global-entropy hard selection, while the latter modulates advantages using global entropy. PEPO uses local-relative comparison and preserves continuous learning signals. Table 3 isolates a contribution from the local reference frame, but gains are not monotonic on every benchmark.
- vs ARES: The paper describes ARES as using forward-window arithmetic averages with a batch threshold, versus PEPO's centered window and relative softmax. The key distinction is the reference frame; symmetry should not be presented as proof of exact removal of arbitrary positional trends.
Rating¶
- Novelty: 4/5 — Local reference frames address a specific entropy-credit confounder with a small algorithmic change.
- Experimental Thoroughness: 4/5 — Three models, multiple runs, substitution, and single-stream analyses; value baselines and broad coding validation remain absent.
- Writing Quality: 3/5 — The core mechanism is clear, but trend claims, some generalizations, and data descriptions require careful cross-checking.
- Value: 4/5 — A low-cost RLVR modification, provided entropy weights are treated as an empirical proxy rather than true credit labels.