Action Chunking Proximal Policy Optimization with Feedback Correction¶
Conference: NeurIPS 2026 (author-reported acceptance in the arXiv list; acceptance type unverified)
arXiv: 2609.36250
Code: https://github.com/hshhahn/ACPPO.git
Area: Robotics & Embodied AI
Keywords: action chunking, proximal policy optimization, closed-loop feedback, chunked advantage, residual regularization
TL;DR¶
ACPPO-Corr extends PPO with a low-frequency action-chunk planner and a stepwise feedback corrector while retaining a state-only value network, improving final normalized interquartile mean over PPO by 30.4% across 25 simulated robotics tasksโnot by 30.4 percentage points of success rate.
Background & Motivation¶
Action chunking predicts several future actions at once instead of making a completely new action decision at every step. In robot imitation learning, it helps represent coherent motion; in reinforcement learning (RL), it also organizes exploration into short sequences and assigns multi-step consequences to a planning decision. However, many chunked RL methods also learn an action-chunk value function: the critic must evaluate both the state and a sequence whose dimensionality grows with action degrees of freedom and chunk length. This is harder than standard PPO state-value regression in high-dimensional bimanual manipulation, and several existing approaches depend on offline data or heavier sequence architectures.
A separate problem arises during execution. If a chunk runs entirely open-loop after observing the state at its beginning, object slip, changing finger contacts, or a misplaced foot cannot immediately affect the current action. This is not merely a reduction in control calls: actions within a chunk can differ, but they are still selected from an old observation. The authors therefore address the two problems separately. ACPPO first introduces actor-side chunking into PPO without an action-chunk Q-network; ACPPO-Corr then restores within-chunk feedback. Open-loop ACPPO actually underperforms PPO on the full main benchmark, showing that temporal abstraction alone is insufficient for these contact-rich tasks.
The setting is fully online, on-policy training from scratch, not a deployment patch attached to a frozen imitation-learning controller. Core Idea: assign temporal action structure to a low-frequency planner and local responses to a corrector that reads fresh observations at every step, then jointly train both with a joint chunk likelihood ratio and a chunk-level advantage so that chunk planning need not sacrifice closed-loop reactivity.
Method¶
Overall Architecture¶
At a chunk boundary, the planner reads the current state and outputs a sequence of \(h\) action means. Plain ACPPO subsequently takes these means in order and samples executed actions without changing the means using new within-chunk states. ACPPO-Corr instead runs a feedback corrector at every step, adds a current-state correction to the corresponding plan entry, and supplies that step's Gaussian exploration scale.
Training still collects stepwise states, rewards, termination flags, and action probabilities. A state-value network predicts values at every step, and standard stepwise generalized advantage estimation (GAE) supplies value-regression targets. The actor signal is aggregated into a chunk-level advantage and paired with the joint importance ratio of the entire valid chunk. Inference requires only the planner and corrector, not the critic or rewards.
Solid arrows below represent sampling and execution data flow; dashed arrows represent trajectory-derived training targets and parameter updates. The planner's within-chunk outputs are open-loop plans, whereas the corrector receives fresh feedback. Training nodes are not additional online control modules.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
S["Chunk-start state"] --> P["Action-chunk planning"]
P --> C["Stepwise feedback correction"]
E["Environment: actions, states, rewards"] -->|Fresh state each step| C
C -->|Sampled action| E
E -.-> A["Projected value and chunked advantage"]
A -.-> J["Joint chunk objective and decoupled updates"]
J -.->|Training update| P
J -.->|Training update| C
Key Designs¶
1. Action-chunk planning: place temporal abstraction only in the actor, keeping critic input independent of chunk length
The planner outputs \(h\) action means at each chunk boundary, each retaining the environment's original action dimensionality. The main setting uses \(h=4\): one planning call produces four steps, but those steps do not repeat an identical action. This distinguishes ACPPO from PPO-Repeat, which only reduces decision frequency. The implementation uses multilayer perceptrons rather than introducing a Transformer for chunking. ACPPO primarily expands the actor's output layer while retaining the original state-value input interface.
Predicting a sequence of means is different from executing a deterministic chunk. Open-loop ACPPO still samples each executed action from a Gaussian; open-loop means that its means do not incorporate fresh within-chunk states, not that exploration disappears. ACPPO-Corr retains this planning backbone and delegates subsequent responses to the corrector. The planner therefore need not regenerate an entire chunk at every step or explicitly learn environment dynamics, but it may still adhere to an outdated plan after a fundamental contact change.
2. Stepwise feedback correction: adjust the current plan entry from fresh state instead of replanning the whole chunk
At every step, the corrector reads the current state and outputs an additive action-dimensional correction and an exploration scale. It does not explicitly take the action chunk as input. The current plan entry and the corrector output are added at the policy mean to form the executed Gaussian policy. Here \(u_{t_0,k}\) is the plan for offset \(k\) produced at the chunk boundary, \(c_t\) and \(\sigma_t\) come from the current state, and \(t_0\) is the current chunk start:
This restores within-chunk closed-loop responses rather than refreshing the planner every step. If an object slips slightly, the corrector can adjust the current finger action from its new position, while the next step still starts from the next entry of the original plan. Exploration also comes from this stepwise branch, so correction is not a purely deterministic post-processing module. Both branches learn from online interaction; the corrector is not layered onto a frozen pretrained policy.
Addition alone does not enforce the intended division of labor. An unrestricted corrector could generate almost the entire action, leaving the planner as a negligible background component. The authors penalize the squared norm of the correction vector rather than imposing a fixed hard correction limit. This encourages the planner to carry persistent action structure and reserves corrections for immediate deviations. Excessive regularization can also suppress necessary feedback, so its strength is selected per task; smaller corrections are not inherently better.
3. Projected value and chunked advantage: retain stepwise value regression while assigning actor credit to chunk consequences
The same physical state encountered midway through different chunks can have different future returns because the remaining plan, chunk-start state, and within-chunk offset differ. An exact state representation would include these active chunk variables; a critic observing only the physical state cannot recover this augmented-state value instance by instance. The paper does not claim that it is an exact Markov value. Instead, it interprets the state-only predictor as the conditional-expectation projection under squared loss:
Here \(z_t\) collects active chunk information, with its averaging distribution induced by the current policy. Avoiding a chunked Q-function does not eliminate every value approximation error introduced by temporal abstraction; it selects a predictor with fixed input dimensionality and lower training complexity. The critic still regresses against stepwise GAE plus the current value, rather than directly regressing an action-chunk Q-value.
The actor forms a single advantage at the chunk start. It first accumulates discounted stepwise TD residuals within the chunk, then adds a standard stepwise GAE tail at the next chunk boundary. With \(\delta_t=r_t+\gamma V_\phi(s_{t+1})-V_\phi(s_t)\), the paper's Eq. (1) is:
This is neither an average of within-chunk stepwise GAE estimates nor a reduction of critic training to once per chunk. The chunk head uses \(\gamma^j\), whereas standard GAE uses \((\gamma\lambda)^j\). The former telescopes intermediate value terms and evaluates the chunk as a coherent decision. A GAE tail remains, so the estimator should not be reduced to a pure \(h\)-step return without subsequent information.
Appendix C analyzes approximation error relative to the projected value: intermediate value errors cancel, and the remaining recursive bootstrap-error term gains an additional factor of \(\gamma^h\) relative to stepwise GAE. This does not remove all error or establish uniform superiority. For \(\lambda<1\), larger within-chunk weights can increase variance under noisy rewards or transitions. Sharing one advantage among all chunk actions also reduces stepwise credit resolution. Short chunks are a concrete choice between less dependence on intermediate values and finer credit assignment.
4. Joint chunk objective and decoupled updates: clip the complete closed-loop trajectory ratio while regulating branch roles
Actor updates must not clip each within-chunk step independently and then treat the result as a chunk objective. For the same sampled trajectory, ACPPO-Corr evaluates each action-density ratio between the current and behavior policies and multiplies the factors over the valid prefix:
\(\mathcal T_{t_0}\) contains the executed offsets before termination. Appendix B does not justify this by assuming that within-chunk states remain fixed. It factorizes the joint density along the realized closed-loop trajectory: environment transition probabilities are independent of policy parameters and cancel in the new-to-old trajectory ratio. Fresh state observations inside the chunk therefore do not invalidate the product-ratio interpretation. PPO clipping applies to the final product paired with one chunk-level advantage, not separately to each factor with a different stepwise advantage.
When an episode ends midway through a chunk, its unexecuted suffix contributes neither likelihood factors nor regularizer terms. In vectorized sampling, that environment remains inactive padding until the next global chunk boundary triggers a new planning call. This avoids mixing actions from a new episode into an old chunk. Implementation must therefore handle both termination-aware returns and valid-action masks rather than assuming that every chunk has its full nominal length.
The shared objective includes the negative chunk-level PPO surrogate, standard value regression, entropy and action-bound regularizers, and a correction-magnitude penalty. Each minibatch uses two Adam updates: the planner pass stops gradients through correction and exploration scale, while the corrector/value pass stops gradients through the plan. Both branches still coordinate through the same executed policy and chunk advantage. They do not learn separate chunk and stepwise PPO objectives.
A Worked Example¶
Consider simulated in-hand manipulation with \(h=4\). At the chunk start, the planner outputs four finger-action means. The first-step corrector observes the current object position, adds a local correction, and samples an action. If the object slips at the second step, the planner is not rerun, but the corrector observes the new state and changes the executed mean for the second plan entry. The third and fourth steps work the same way; only the next chunk boundary produces a new four-step plan.
During training, the four rewards and the GAE tail after the boundary form one chunk-level advantage, and the four action-density ratios form one joint ratio. The critic still receives value-regression supervision at all four stepwise states. If the third step terminates the episode, only the first three likelihood factors remain; the fourth step is padding rather than an executed action. This illustrates the algorithm, not a reported real-hardware trajectory.
Loss & Training¶
Let \(J_{\mathrm{chunk}}\) denote the clipped surrogate built from the joint chunk ratio and chunk-level advantage. The paper's Eq. (3) can be written as:
The value loss is mean squared error against a stepwise GAE return target, entropy preserves exploration, and action-bound regularization follows the robotics PPO implementation. This objective does not introduce a separate reward for each branch; stop gradients determine which variables each pass changes. The planner is frozen for the first \(N/20\) training epochs while the corrector/value branch continues learning, after which joint training begins. This reduces the influence of unstable early advantages on chunk planning.
The main experiments inherit task-specific PPO hyperparameters and select correction strength from \(\lambda_r\in\{0.03,0.1,0.3\}\). QC-FQL likewise uses three candidate behavior-regularization coefficients, \(\{1,3,10\}\). Preliminary tuning runs select the coefficients, which are then fixed for the reported five-seed evaluations. Multiple actor branches use reduced first-layer MLP widths to control parameter count; the gains should not simply be attributed to duplicating a full PPO network.
To control accumulated changes in the joint ratio, learning rates adapt to measured average KL: divide by 1.5 above twice the threshold and multiply by 1.5 below half the threshold. Appendix D's fixed-covariance Gaussian analysis provides intuition, not a strict independence guarantee for arbitrary closed-loop trajectories. Longer chunks can still accumulate greater KL and clipping pressure.
Key Experimental Results¶
Main Results¶
Evaluation covers 9 IsaacGym tasks and 16 Bi-DexHands tasks, with 5 random seeds for every method-task pair. Goal-completion tasks use success rate, while progress-based control tasks use episodic return. Scores are normalized using task-specific bounds, then aggregated over task-seed scores using IQM, the mean of the middle 50% after sorting. Final normalized IQM across all 25 tasks improves over PPO by 30.4%.
Frequency subsets use the relative absolute difference between PPO and \(h=4\) action repetition in final normalized per-task scores, with a threshold of 0.3. Four tasks are excluded only from this analysis because their denominator is near zero. The remaining 21 tasks comprise 11 sensitive and 10 neutral tasks; the main benchmark is not reduced to 21 tasks. The table reports each subset's IQM relative to PPO's IQM on the same subset, not the mean of per-task ratios.
| Method | Frequency-sensitive: relative IQM, 95% interval | Frequency-neutral: relative IQM, 95% interval |
|---|---|---|
| PPO | 1.00 [1.00, 1.00] | 1.00 [1.00, 1.00] |
| PPO-Repeat | 0.61 [0.55, 0.68] | 0.97 [0.92, 1.02] |
| SAC | 0.33 [0.21, 0.44] | 0.59 [0.49, 0.68] |
| QC-FQL | 0.76 [0.64, 0.87] | 0.77 [0.66, 0.87] |
| ACPPO | 0.92 [0.87, 0.97] | 0.97 [0.85, 1.10] |
| ACPPO-Corr | 1.33 [1.24, 1.43] | 1.24 [1.15, 1.34] |
Source: Figure 2(b). Intervals use seed bootstrap with the task set fixed, not confidence intervals for generalization to unseen robotics tasks. Appendix Table 3 defines wins and losses using a 0.03 threshold on normalized score differences against the strongest competitor. Across all tasks, ACPPO-Corr records 13 wins / 8 ties / 4 losses; the sensitive subset records 7 / 2 / 2 and the neutral subset 5 / 5 / 0.
Ablation Study¶
| Configuration or analysis | Reported result | Interpretation and scope |
|---|---|---|
| Replace corrector input with chunk-start state only at evaluation | Retains 84.0% of full-model performance | Architecture unchanged, but fresh within-chunk observations removed; ยง5.2 |
| Retrain from scratch with chunk-start state input | Retains 88.1% of full-model performance | Retraining partly recovers performance, still below fresh feedback; ยง5.2 |
| Replace chunked advantage with stepwise GAE | Below full model; no exact scalar tabulated | Figure 4(a) supports the direction; no numbers guessed from curves |
| Remove correction regularization | Aggregate performance below full model; no exact scalar tabulated | Figure 4(a); the corrector can absorb most of the action signal |
| Chunk lengths \(h\in\{1,2,4,8\}\) | \(h=4\) has the strongest aggregate performance | Figure 4(b); \(h=1\) improves only marginally over PPO |
| Chunk-level clip fraction at \(h=2/4/8\) | 0.364 / 0.369 / 0.408 | Appendix Table 9; PPO is 0.365, with more clipping at longer chunks |
The regularizer's mechanism is also visible in the ratio of correction norm to plan norm. The following table preserves final correction ratios from source Table 1. This ratio is not a success rate and does not identify the best performance across tasks. The main text cites these results as โTable 2,โ but the cache labels correction ratios as Table 1 and longer-horizon evaluation as Table 2. This note distinguishes them by caption and content without changing the source numbers.
| Task | \(\lambda_r=0.03\) | \(\lambda_r=0.1\) | \(\lambda_r=0.3\) |
|---|---|---|---|
| AllegroHand | 0.222 | 0.053 | 0.007 |
| Humanoid | 0.021 | 0.009 | 0.002 |
| ShadowHand | 0.163 | 0.029 | 0.009 |
Key Findings¶
- Closed-loop feedback and temporal abstraction contribute together: open-loop ACPPO does not outperform PPO, and the two-branch \(h=1\) configuration offers only a small improvement. The contribution cannot be reduced to lower decision frequency or an additional correction network.
- Stronger regularization consistently lowers the correction ratio, but the preferred strength is task-dependent. AllegroHand favors stronger regularization, whereas Humanoid and ShadowHand favor weaker regularization.
- Longer-chunk results have a separate scope: across five tasks with sufficiently long rollouts, mean per-task score ratios at \(h=16\) are 1.21, 0.68, and 0.29 for ACPPO-Corr, ACPPO, and QC-FQL, respectively (Table 2), not all-task IQM. Only FrankaCubeStack and Humanoid are evaluated at \(h=32\), where ACPPO-Corr ratios are 1.09 and 0.65 (Table 8); it does not outperform PPO everywhere.
- Performance has a computational cost: representative training throughput and parameter count are 41,714 steps/sec and 1,557,258 for ACPPO-Corr versus 47,894 steps/sec and 1,463,785 for PPO (Appendix Table 5). Throughput includes rollout and optimization, not real-hardware inference control frequency.
Highlights & Insights¶
- Keeping critic inputs unchunked and extending actor credit across multiple steps are separate choices. Retaining the familiar state-value interface while acknowledging its projected nature is more precise than claiming an exact stepwise value after chunking.
- A closed-loop chunk likelihood ratio is evaluated along the realized trajectory, allowing environment transition factors to cancel. This makes fresh-state correction a trained policy component rather than a heuristic execution patch outside training.
- Residual regularization controls the division of labor between planning and feedback, not simply action magnitude. Task-specific correction-ratio analysis exposes this division without establishing a universal optimal threshold.
Limitations & Future Work¶
- Experiments use only state observations and continuous actions in simulated robotics. They provide no real-hardware, visual-policy, or tactile-policy validation and do not directly extend to discrete token actions.
- Joint likelihood changes accumulate with chunk length, and clipping already increases at \(h=8\). Short-chunk stability does not guarantee arbitrary temporal horizons; adaptive chunk length and dedicated stabilization are potential directions.
- Additive correction around an existing plan may struggle to abandon the current manipulation mode. ShadowHandUpsideDown and ShadowHandBottleCap achieve only 0.706 and 0.829 times PPO's score. The authors propose regrasping or global finger reconfiguration as possible explanations, not established causal findings.
- Every method has zero success on ShadowHandPushBlock, so this does not identify a failed component. QC-FQL's nonzero performance on ShadowHandDoorCloseOutward is also confounded by policy class, replay, and Q-optimization differences and cannot be attributed to one design alone.
Related Work & Insights¶
- vs PPO / PPO-Repeat: PPO regenerates actions each step, whereas Repeat executes the same action repeatedly. ACPPO-Corr generates distinct sequence entries at low frequency while retaining fresh-state feedback each step, separating planning frequency from reaction frequency.
- vs QC-FQL and action-chunk Q-learning: These comparators learn action-sequence values; this paper avoids expanding critic inputs through projected state values and a chunked actor surrogate. The conclusion applies to the matched online simulation protocol, not a rejection of chunked Q-methods in offline or pretrained settings.
- vs residual RL and pretrained chunk-policy correction: Prior approaches often correct frozen controllers or depend on pretrained visual/tactile policies; both branches here are jointly trained online from scratch. A useful extension is learning when to abandon a plan rather than only adjusting its current action mean.
Rating¶
- Novelty: 4/5. Integrates chunk-level PPO, projected state values, and online feedback correction into one training formulation; the main novelty is their combination and credit assignment.
- Experimental Thoroughness: 4/5. Includes 25 tasks, five seeds, and multiple ablations, but remains limited to state-based simulation.
- Writing Quality: 4/5. Appendices clarify likelihood ratios and the limits of value-error cancellation; the main text contains a mismatched correction-ratio table reference.
- Value: 4/5. Offers reusable ideas for online high-dimensional control, while real-hardware benefits and longer-chunk stability remain unverified.