PARL-VLA: Pruning-Aware On-Policy Reinforcement Learning for Vision-Language-Action Model¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Robotics & Embodied AI
Keywords: Vision-Language-Action Model / Visual Token Pruning / On-Policy Reinforcement Learning / GRPO / Embodied AI Robustness
TL;DR¶
Addressing the severe exploration-exploitation mismatch and training collapse caused by naive token pruning in on-policy VLA post-training, PARL-VLA records rollout pruning configurations and replays them during policy updates to align likelihood ratios, coupled with a spectrum pruning curriculum that discards 55% of visual tokens while accelerating training and boosting out-of-distribution robustness.
Background & Motivation¶
Vision-Language-Action (VLA) models have emerged as the foundational generalist paradigm for robotic manipulation, translating multimodal visual inputs and natural language instructions into low-level motor actions. To transcend the intrinsic constraints of offline behavioral cloning, such as covariate shift and narrow trajectory support, online reinforcement learning—particularly via Group Relative Policy Optimization (GRPO)—has been increasingly adopted during post-training. Nevertheless, scaling online RL over large multimodal transformer architectures incurs prohibitive computational overhead due to repeated forward evaluations across tens of thousands of environment rollout transitions.
Visual token pruning represents an appealing path toward alleviating sequence length complexity, as extensive literature demonstrates that a substantial fraction of visual tokens in manipulation observations encode task-irrelevant background distractors. However, directly grafting existing inference-time pruning or offline fine-tuning pruning methods into online RL triggers severe optimization failure. If pruning is applied purely at evaluation time, the policy encounters an abrupt train-test distribution shift without adaptation. Conversely, if token pruning is performed during trajectory rollouts but policy updates recompute action log-probabilities using full tokens or mismatched pruning budgets, the importance sampling likelihood ratios \(r_t(\theta)\) deviate entirely from the rollout generation regime. In practice, this exploration-exploitation mismatch drives the empirical KL divergence above 1.0, nullifies the relative advantage signal, and destabilizes or collapses policy optimization.
The foundational insight of this work is that visual token pruning and policy optimization must co-adapt within a closed reinforcement learning loop. Core Idea: Introduce PARL-VLA, a pruning-aware on-policy RL framework that enforces exploration-exploitation pruning consistency by logging per-sample rollout pruning configurations and replaying them during policy updates, combined with a spectrum token curriculum to train a single robust policy across varying information budgets.
Method¶
Overall Architecture¶
PARL-VLA incorporates a differentiable visual token pruning module into an OpenVLA-OFT backbone optimized via on-policy GRPO. At each execution step, raw visual observations and task instructions are mapped to visual patch tokens and instruction tokens. Before entering the language transformer backbone, visual tokens are scored by a text-conditioned gated projection layer and pruned via hard Top-\(k\) selection under a budget \(k\) and perturbation \(\epsilon_t\). Crucially, the rollout phase caches these pruning parameters alongside trajectory steps, allowing the subsequent policy gradient update to replay identical pruned sequences, recompute matching log-probabilities, and propagate straight-through gradients back into both policy and scoring networks.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multimodal Observation<br/>Image It + Instruction x"] --> B["Text-Conditioned Scoring & Straight-Through Selection<br/>Gated instruction query pooled with visual key similarity"]
B --> C["Spectrum Pruning Curriculum<br/>Dynamic token budget k and exploration perturbation εt"]
C --> D["Pruned Sequence Input<br/>Truncated visual tokens + instruction tokens"]
D --> E["VLA Backbone Rollout Execution<br/>Record trajectory and pruning metadata"]
E --> F["Exploration-Exploitation Consistency Replay<br/>Replay rollout pruning mask and align likelihood ratio"]
F --> G["GRPO Policy Update & End-to-End Co-Adaptation<br/>Clipped surrogate loss and straight-through gradient backprop"]
Key Designs¶
1. Exploration-Exploitation Pruning Consistency: Eliminating Distribution Shift in Policy Gradients In on-policy actor-critic and policy gradient methods, optimization integrity depends on valid importance sampling ratios \(r_{i,t}(\theta) = \exp(\log \pi_\theta(a_{i,t}|s_{i,t}) - \log \pi_{\theta_{\text{old}}}(a_{i,t}|s_{i,t}))\). When an agent explores the environment conditioned on a pruned token sequence \(\tilde{V}_t\), computing update log-probabilities over unpruned full sequences or under a divergent token budget fundamentally violates the underlying Markov decision process observation dynamics. This discrepancy causes empirical KL divergence to surge past 1.0 within early training steps, leading to erratic ratio swings and near-zero effective advantage signals. By caching the exact selection configuration \((k, \epsilon_t)\) as lightweight transition metadata and replaying the identical token mask during actor updates, PARL-VLA guarantees that optimization operates on the exact observation distribution generated during rollout, stabilizing approximate KL divergence near 0.01 throughout training.
2. Text-Conditioned Scoring and Straight-Through Selection: Dynamic Relevance Prioritization To dynamically preserve task-critical semantics—such as the robot gripper and manipulation targets—while filtering background clutter, PARL-VLA implements a lightweight cross-modal projection layer. Given visual tokens \(V_t \in \mathbb{R}^{N \times d}\) and instruction tokens \(X_t \in \mathbb{R}^{L \times d}\), representations are projected into a shared metric space and \(L_2\)-normalized: \(q_{t,\ell} = \text{norm}(W_q x_{t,\ell})\) and \(\kappa_{t,n} = \text{norm}(W_k v_{t,n})\). To attenuate uninformative padding tokens, a learned gating scalar \(g_{t,\ell} = \sigma(f_{\text{gate}}(x_{t,\ell}))\) weights the instruction tokens into an aggregated query vector \(\bar{q}_t\):
The relevance score of visual token \(n\) is computed via cosine similarity \(s_{t,n} = \bar{q}_t^\top \kappa_{t,n}\). During forward inference, a uniform perturbation \(\epsilon_{t,n} \sim \mathcal{U}(-\sigma_t/2, \sigma_t/2)\) provides stochastic exploration, followed by a hard Top-\(k\) operator to physically truncate the sequence. In the backward pass, a straight-through estimator (STE) utilizing temperature-controlled softmax relaxation routes policy gradient backpropagation into the scoring parameters \(\phi\), enabling task-reward-driven token selection.
3. Spectrum Pruning Curriculum: Budget Generalization and Robustness Regularization Optimizing an embodied policy exclusively at a single static pruning budget risks overfitting to a narrow information density, precipitating severe performance drops under test-time budget adjustments or visual domain shifts. PARL-VLA overcomes this by treating the token keep ratio \(\rho = k/N\) as a sampled distribution variable governed by a loose-to-aggressive curriculum: early training (0% to 10% progress) samples \(\rho \in [0.75, 1.0]\) to facilitate initial exploration in dense regimes; mid-stage training (10% to 20%) shifts to \(\rho \in [0.5, 0.75]\); and late-stage training operates aggressively within \(\rho \in [0.25, 0.5]\). This spectrum training regularizes the policy to extract invariant spatial and geometric cues under sparse visual evidence, preventing brittle reliance on transient lighting and texture artifacts while producing a single model deployable across arbitrary compute constraints.
Loss & Training¶
PARL-VLA optimizes the policy parameters \(\theta\) and pruning module \(\phi\) via Group Relative Policy Optimization (GRPO). For each task instruction \(x\), a group of \(G\) trajectories \(\{\tau_i\}_{i=1}^G\) is gathered, receiving environment returns \(R_i\). The group-relative advantage is normalized as:
The optimization objective optimizes the standard clipped surrogate loss:
where gradients are computed strictly over valid action tokens. Training runs for 100 epochs on the RLinf distributed infrastructure with FSDP; evaluations are conducted using checkpoint epoch 75, which resides stably within the target sparse spectrum phase (\(\rho \in [0.25, 0.5]\)).
Key Experimental Results¶
Main Results¶
On the LIBERO-Plus out-of-distribution benchmark, PARL-VLA is evaluated against full-token RLinf and inference-only pruning baselines (FastV and VLA-Pruner) under three visual perturbation regimes: lighting variations, background textures, and object layout shifts. PARL-VLA is tested under Low (30% keep ratio), Medium (50% keep ratio), and High (70% keep ratio) evaluation budgets.
| Perturbation | Method | Spatial (%) | Object (%) | Goal (%) | Long (%) | Avg. (%) |
|---|---|---|---|---|---|---|
| Light | RLinf (Full tokens) | 83.9 | 44.1 | 63.8 | 59.1 | 62.7 |
| Light | FastV | 71.9 | 38.7 | 74.6 | 56.9 | 60.5 |
| Light | VLA-Pruner | 77.7 | 47.5 | 62.4 | 57.3 | 61.2 |
| Light | PARL-VLA (Low, 30% kept) | 86.6 | 44.1 | 74.9 | 61.7 | 66.8 |
| Light | PARL-VLA (Medium, 50% kept) | 84.9 | 49.8 | 76.3 | 57.3 | 67.1 |
| Light | PARL-VLA (High, 70% kept) | 82.9 | 51.2 | 75.6 | 65.0 | 68.7 (+6.0) |
| Background | RLinf (Full tokens) | 91.9 | 81.0 | 79.0 | 69.6 | 80.4 |
| Background | FastV | 82.2 | 69.8 | 78.3 | 70.2 | 75.1 |
| Background | VLA-Pruner | 84.5 | 79.4 | 75.8 | 67.1 | 76.7 |
| Background | PARL-VLA (Low, 30% kept) | 95.3 | 82.3 | 83.3 | 73.0 | 83.5 |
| Background | PARL-VLA (Medium, 50% kept) | 96.5 | 82.3 | 85.8 | 71.6 | 84.1 (+3.7) |
| Background | PARL-VLA (High, 70% kept) | 93.0 | 81.0 | 86.1 | 70.2 | 82.6 |
| Layout | RLinf (Full tokens) | 87.0 | 62.0 | 57.2 | 63.5 | 67.4 |
| Layout | FastV | 80.8 | 61.8 | 58.4 | 60.3 | 65.3 |
| Layout | VLA-Pruner | 73.5 | 58.6 | 56.2 | 57.7 | 61.5 |
| Layout | PARL-VLA (Low, 30% kept) | 89.4 | 60.3 | 59.1 | 62.8 | 67.9 |
| Layout | PARL-VLA (Medium, 50% kept) | 90.1 | 64.5 | 58.6 | 64.1 | 69.3 |
| Layout | PARL-VLA (High, 70% kept) | 93.0 | 63.5 | 58.8 | 67.3 | 70.6 (+3.2) |
| Mean | RLinf (Full tokens) | 87.6 | 62.4 | 66.7 | 64.1 | 70.2 |
| Mean | FastV | 78.3 | 56.8 | 70.4 | 62.5 | 67.0 |
| Mean | VLA-Pruner | 78.6 | 61.8 | 64.8 | 60.7 | 66.5 |
| Mean | PARL-VLA (Low, 30% kept) | 90.4 | 62.2 | 72.4 | 65.8 | 72.7 |
| Mean | PARL-VLA (Medium, 50% kept) | 90.5 | 65.5 | 73.6 | 64.3 | 73.5 |
| Mean | PARL-VLA (High, 70% kept) | 89.6 | 65.2 | 73.5 | 67.5 | 74.0 (+3.8) |
In physical hardware experiments on an AgileX Piper manipulator across 405 trials under lighting, background, and layout shifts (RoboTwin2.0 benchmark), PARL-VLA achieves 46.9% overall real-robot success (95% Wilson CI: 42.1%–51.8%), surpassing RLinf (37.0%), VLA-Pruner (26.7%), and FastV (21.2%).
Ablation Study¶
Spectrum Training vs. Fixed-Ratio Training Ablation on LIBERO-Plus: The ablation isolates the impact of the token keep-ratio distribution during on-policy training when evaluated across diverse test-time budgets:
| Training Configuration | Eval \(\rho=0.30\) (%) | Eval \(\rho=0.40\) (%) | Eval \(\rho=0.50\) (%) | Avg. Across Budgets (%) | Min. Across Budgets (%) | Note |
|---|---|---|---|---|---|---|
| Fixed \(\rho = 0.30\) | 63.9 | 71.1 | 72.7 | 69.2 | 63.9 | Overfits sparse budget; poor generalizability |
| Fixed \(\rho = 0.40\) | 69.2 | 71.9 | 72.1 | 71.1 | 69.2 | Strongest fixed baseline, yet drops on edge budgets |
| Spectrum (Ours) | 72.7 | 74.4 | 73.5 | 73.5 | 72.7 | Consistently outperforms fixed setups; min gain +3.5% |
Wall-Clock Acceleration and In-Distribution Performance: PARL-VLA eliminates 55% of visual tokens on average, yielding substantial speedups across distributed rollout and optimization stages:
| Metric | RLinf (Full Tokens) | PARL-VLA (Ours) | Efficiency Gain / Speedup |
|---|---|---|---|
| Visual Token Reduction | 0.00% | 55.0% | 55% reduction in attention sequence |
| Rollout Policy Forward Latency | ~298 ms | ~229 ms | 1.30× speedup |
| Actor Optimization Step Time | 432–453 ms | 256–264 ms | 1.66×–1.74× speedup |
| End-to-End Training Step Time | 0.81–0.84 s | 0.67–0.70 s | 1.21×–1.23× end-to-end acceleration |
| LIBERO In-Distribution Success Rate | 94.4% | 94.7%–94.9% | Preserves full in-distribution efficacy |
Key Findings¶
- Consistency replay is indispensable for on-policy stability: Deviations between rollout and update token budgets cause the approximate KL divergence to escalate beyond 1.0, destroying the group advantage signal and triggering policy collapse; replaying recorded configurations stabilizes KL around 0.01.
- Spectrum training drives robust out-of-distribution transfer: Exposing the policy to variable token budgets prevents over-specialization, boosting worst-case evaluation success by 3.5 points on LIBERO-Plus and yielding an 8.8-point improvement at \(\rho=0.30\) over fixed \(\rho=0.30\) training.
- Token pruning acts as visual feature regularization: Removing non-essential background tokens inherently attenuates environmental lighting and texture perturbations, elevating out-of-distribution success (e.g., from 80.4% to 84.1% under background shifts) while maintaining identical in-distribution accuracy.
Highlights & Insights¶
- Pruning parameters as trajectory transition metadata: Storing \((k, \epsilon_t)\) per step resolves the foundational exploration-exploitation mismatch in on-policy reinforcement learning with virtually zero memory overhead.
- Spectrum curriculum bridges exploration and efficiency: Progressing from dense to aggressive keep ratios allows stable early policy exploration while progressively unlocking 1.7× backpropagation speedups.
- Pruning improves embodied robustness without accuracy trade-offs: Demonstrates that reducing visual token volume filters nuisance background shifts, transforming efficiency pruning into a structural regularizer for physical generalization.
Limitations & Future Work¶
- Bottlenecked by simulator physics and vision encoding: While transformer inference and training update times drop by 1.3× and 1.7×, end-to-end training acceleration is capped at ~1.23× due to simulation physics and image preprocessing overhead.
- Temporal and multiview token budgeting remain static: Token pruning operates independently on individual frames without cross-view dynamic allocation across primary and wrist cameras.
- Open-loop action chunking sensitivity: For fine-grained manipulation requiring continuous visual feedback, excessive token pruning at contact boundaries may introduce action jitter.
Related Work & Insights¶
- vs. RLinf / SimpleVLA-RL: While RLinf and SimpleVLA-RL established scalable on-policy RL for VLAs, they compute full visual sequences across all rollout and update phases. PARL-VLA maintains identical in-distribution success while accelerating updates by 1.7× and improving out-of-distribution generalization.
- vs. FastV / VLA-Pruner: FastV and VLA-Pruner operate primarily at test-time inference or offline adaptation; applying them directly to RL produces severe distribution mismatch, trailing PARL-VLA by 6–8 percentage points in robustness.
- vs. DTP (Distracting Token Pruning): DTP highlighted the benefits of pruning background tokens within supervised fine-tuning. PARL-VLA establishes the theoretical and empirical foundation for joint token pruning and on-policy policy gradient optimization.
Rating¶
- Novelty: ⭐⭐⭐⭐ [Identifies the rollout-learn pruning mismatch and introduces consistency replay with spectrum curriculum training]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive verification across LIBERO, thousands of LIBERO-Plus perturbations, and 405 real-robot trials]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear articulation of reinforcement learning failure modes and mathematical alignment mechanics]
- Value: ⭐⭐⭐⭐⭐ [Provides a practical blueprint for resource-efficient, robust foundation model reinforcement post-training in robotics]