Skip to content

Grounding Sim-to-Real Generalization in Dexterous Manipulation: An Empirical Study with Vision-Language-Action Models

Conference: ECCV 2026
Paper: ECCV Official
Area: Robotics & Embodied AI
Keywords: Sim-to-Real Generalization, Vision-Language-Action (VLA) Models, Domain Randomization, Embodied Manipulation, Reinforcement Fine-Tuning (RFT)

TL;DR

Through over 10,000 physical robot trials across five bimanual manipulation tasks, this work systematically decouples and quantifies the critical factors governing zero-shot Sim-to-Real transfer for Vision-Language-Action (VLA) models, demonstrating that spatial variations outshine appearance factors and that frame-wise randomization combined with reinforcement fine-tuning dramatically boosts real-world robustness.

Background & Motivation

Building generalist embodied manipulation policies is a foundational milestone toward Artificial General Intelligence (AGI). Vision-Language-Action (VLA) models, which blend multimodal foundation backbones with continuous motor action outputs, have demonstrated remarkable capabilities in semantic comprehension, open-ended task planning, and zero-shot reasoning. However, VLA architectures are heavily data-driven and require massive volumes of interaction trajectories. Collecting demonstrations on physical robots is prohibitively expensive, prone to hardware wear-and-tear, and strictly constrained in environmental diversity. Generating synthetic data within physics-based simulators (such as Isaac Lab, MuJoCo, and RoboTwin 2.0) offers scalable, zero-marginal-cost data generation, yet significant domain discrepancies in lighting, textures, camera poses, frictional dynamics, and contact restitution cause simulation-trained policies to suffer severe performance degradation upon real-world deployment.

While classical robotics has developed numerous Sim-to-Real paradigmsโ€”such as Domain Randomization (DR), domain adaptation, high-fidelity rendering, and reinforcement fine-tuning (RFT)โ€”the vast majority of studies evaluate these techniques monolithically on low-dimensional reinforcement learning policies. There is a marked lack of rigorous empirical studies grounding these methods in high-capacity VLA policies deployed on physical bimanual robots. It remains largely unknown which specific simulation factors dominate Sim-to-Real transfer, whether appearance shifts matter as much as 3D spatial perturbations, how temporal randomization granularity impacts policy attention, and where visual/physical rendering fidelity reaches diminishing returns. Without such mechanistic insights, optimizing simulation pipelines and diagnosing physical deployment failures remain largely trial-and-error.

To bridge this fundamental gap, this work establishes a standardized, factorized evaluation benchmark to probe the zero-shot Sim-to-Real transferability of VLA models. By isolating individual domain randomization factors, temporal sampling granularities, photorealistic and physics-realistic fidelity tiers, and policy optimization regimes across more than 10,000 real-world trials, the authors provide clear empirical principles. The core idea is to factorize and decouple multi-level domain randomization, temporal granularity, and policy fine-tuning paradigms, revealing that spatial geometric perturbations (camera pose and table height) dominate visual appearance factors for VLA transfer, while high-frequency frame-wise randomization and group-relative reinforcement learning synergistically maximize real-world execution robustness.

Method

Overall Architecture

The framework establishes a factorized evaluation methodology for zero-shot Sim-to-Real transfer of VLA models, utilizing the RoboTwin 2.0 simulation suite and physical bimanual robotic platforms (Cobot Magic / Piper). The pipeline decomposes the investigation across four complementary axes: structured domain randomization (orthogonally disentangling appearance vs. spatial factors), temporal sampling granularity (episode-wise vs. frame-wise), simulation realism tiers (ray-tracing visual presets and perturbed contact dynamics), and policy optimization paradigms (supervised behavioral cloning vs. group-relative reinforcement learning fine-tuning).

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Simulation Setup<br/>RoboTwin 2.0 + 5 Bimanual Tasks"] --> B["Spatially-Dominant Structured DR<br/>Orthogonal factorized visual vs spatial variations"]
    B --> C["High-Frequency Frame-wise DR<br/>Disrupting static background spurious correlations"]
    C --> D["Tiered Simulation Fidelity Tuning<br/>Ray-traced photorealism + Contact physics modeling"]
    D --> E["Group-Relative Reinforcement Fine-Tuning<br/>GRPO relative advantage optimization without KL"]
    E --> F["Zero-Shot Physical Deployment<br/>Physical bimanual robot >10k evaluation trials"]

Key Designs

1. Spatially-Dominant Structured DR: Decoupling visual appearance from 3D geometric perturbations Conventional domain randomization routinely scrambles textures, illumination, and physical attributes concurrently, obscuring the primary drivers of generalization. The authors factorize the simulation parameter space into structured components \(\xi = \{\xi_{\text{BG}}, \xi_{\text{TD}}, \xi_{\text{CP}}, \xi_{\text{LT}}, \xi_{\text{TH}}\}\), distinguishing visual appearance factors (background and table textures \(\xi_{\text{BG}}\), tabletop distractor configurations \(\xi_{\text{TD}}\), and directional lighting conditions \(\xi_{\text{LT}}\)) from spatial geometric factors (camera pose translational offsets \(\xi_{\text{CP}}\) and table height shifts \(\xi_{\text{TH}}\)). The policy optimizes the multi-step action regression objective across the broadened parameter distribution: $\(\mathcal{L}_{\text{DR}}(\phi) = \mathbb{E}_{\xi \sim p(\xi)} \left[ \mathbb{E}_{x \sim p(x \mid \xi)} \ell(f_\phi(x), y) \right]\)$ Empirical evidence demonstrates that spatial geometric perturbationsโ€”specifically table height \(\xi_{\text{TH}}\) and camera position \(\xi_{\text{CP}}\)โ€”serve as the primary drivers of transfer, as they compel the model to internalize robust hand-eye coordination and spatial relative positioning. Purely visual variations like tabletop color or background swapping provide secondary, complementary gains only when paired with spatial perturbations.

2. High-Frequency Frame-wise DR: Disrupting static background spurious correlations In standard imitation learning, domain randomization is typically applied on an episode-wise basis, sampling parameters once at rollout initialization and keeping them static throughout the trajectory. However, the expressiveness of modern VLA backbones leads them to latch onto static background features, lighting gradients, or surface patterns as spurious causal predictors for action execution. By transitioning to frame-wise randomization, background textures \(\xi_{\text{BG}}\), lighting \(\xi_{\text{LT}}\), and slight camera translations \(\xi_{\text{CP}}\) are resampled at every single control step. This high-frequency temporal perturbation eliminates spurious temporal correlations between background pixels and task progression, forcing the visual encoder's spatial attention to concentrate tightly on the target objects and the robot's end-effectors.

3. Tiered Simulation Fidelity Tuning: Identifying saturation thresholds in visual and physical realism To determine how visual realism and physical dynamics impact the Sim-to-Real gap, the benchmark establishes three visual tiers: Low (ray-tracing disabled, minimal sampling and path depth), Medium (ray-tracing enabled with global illumination and shadow transport), and High (increased samples per pixel and path depth), alongside a physically degraded baseline with distorted gravity, surface friction, and restitution. The study finds that visual photorealism dominates contact dynamics in tabletop manipulation, yet its benefits saturate once moderate photorealism (accurate ambient lighting and soft contact shadows) is attained, indicating that hyper-realistic visual rendering offers diminishing returns for zero-shot VLA transfer.

4. Group-Relative Reinforcement Fine-Tuning: Expanding beyond supervised behavioral cloning Supervised fine-tuning (SFT) restricts the model to mimicking demonstration manifolds, making it susceptible to compounding error drift in physical settings. To overcome this limitation, the authors introduce Group Relative Policy Optimization (GRPO) directly in simulation without reference KL penalties. At each training step, \(G\) trajectories \(\{\tau_i\}_{i=1}^G\) are sampled from the current policy, and normalized group advantage estimates \(\hat{A}_i = \frac{R_i - \text{mean}(\{R_j\})}{\text{std}(\{R_j\})}\) are calculated from sparse completion rewards \(R_i\): $\(\mathcal{J}_{\text{RFT}}(\theta) = \mathbb{E} \left[ \frac{1}{G} \sum_{i=1}^G \frac{1}{|\tau_i|} \sum_{t=1}^{|\tau_i|} \min \left( r_{i,t}(\theta)\hat{A}_i, \, \text{clip}(r_{i,t}(\theta), 1-\epsilon, 1+\epsilon)\hat{A}_i \right) \right]\)$ where \(r_{i,t}(\theta)\) denotes the policy probability ratio. By updating exclusively on trajectory groups exhibiting mixed success outcomes, the policy actively explores self-correction trajectories and recovers from trajectory drift, equipping the VLA with dynamic resilience when facing real-world physical discrepancies.

Loss & Training

The base policy uses OpenVLA-OFT, predicting continuous multi-step joint targets and gripper states via an L1 regression loss: $\(\mathcal{L}_{\text{SFT}}(\theta) = \mathbb{E}_{(o_t, a_{t:t+M}, l) \sim \mathcal{D}} \left[ \left\| \pi_\theta(o_t, l) - a_{t:t+M} \right\|_1 \right]\)$ For the reinforcement fine-tuning stage, the action head is converted to discrete action tokens optimized with autoregressive cross-entropy, followed by the aforementioned GRPO advantage optimization across simulated rollouts.

Key Experimental Results

Main Results

The benchmark evaluates five dual-arm manipulation tasks: Click Bell, Place Empty Cup, Beat Block Hammer, Stack Bowls Two, and Pick Dual Bottles, spanning both simulation OOD setups and physical robot deployments.

The following table presents the zero-shot physical deployment and simulation OOD success rates across domain randomization factor configurations (from Table 1 of the paper):

Data Setting Click Bell (Sim / Real) Place Cup (Sim / Real) Beat Block (Sim / Real) Stack Bowls (Sim / Real) Dual Bottles (Sim / Real)
Clean (Baseline) 14% / 2.7% 11% / 5.4% 17% / 0.0% 36% / 26.2% 19% / 1.5%
BG (Background) 23% / 11.5% 15% / 10.2% 30% / 2.7% 42% / 40.4% 27% / 12.3%
LT (Lighting) 22% / 12.3% 12% / 8.6% 29% / 4.6% 43% / 31.5% 24% / 7.3%
TD (Distractor) 15% / 3.1% 12% / 6.9% 23% / 0.0% 38% / 27.3% 21% / 4.6%
CP (Camera Pose) 34% / 23.5% 20% / 17.5% 32% / 4.6% 47% / 42.3% 30% / 16.2%
TH (Table Height) 40% / 36.9% 26% / 24.8% 37% / 6.5% 58% / 49.6% 32% / 15.4%
TH + CP + BG 50% / 47.7% 34% / 33.5% 43% / 8.5% 62% / 60.0% 42% / 21.2%
TH + CP + LT 48% / 44.2% 31% / 27.9% 46% / 7.3% 60% / 54.6% 39% / 20.0%
TH + CP + TD 43% / 40.0% 29% / 26.3% 38% / 6.5% 56% / 52.7% 32% / 16.5%
All Factors 54% / 49.7% 44% / 41.0% 49% / 11.5% 65% / 63.1% 52% / 23.8%

Ablation Study

The table below shows the orthogonal synergy between reinforcement fine-tuning (RL) and domain randomization (DR) on discretized action policies (from Table 3 of the paper):

Training Variant Click Bell (Sim / Real) Place Cup (Sim / Real) Beat Block (Sim / Real) Stack Bowls (Sim / Real) Dual Bottles (Sim / Real) Real Average
SFT (Clean Simulation) 11% / 0.0% 23% / 6.8% 15% / 0.0% 33% / 21.2% 10% / 0.0% 5.6%
SFT + RL (Clean Simulation RL) 42% / 36.9% 65% / 34.8% 52% / 9.6% 60% / 59.2% 48% / 26.5% 33.4%
SFT + RL + DR (Randomized RL) 60% / 50.8% 87% / 51.4% 72% / 16.2% 68% / 64.6% 67% / 30.8% 42.8%

Furthermore, in temporal granularity ablations (Table 2), frame-wise randomization outpaced episode-wise sampling consistently, yielding real-world improvements of +3.4% to +8.4% for camera pose perturbations and up to +12.9% for background texture randomization (e.g., boosting Place Empty Cup from 10.2% to 23.1%).

Key Findings

  • Spatial geometry dominates visual appearance: Single spatial factors like table height (TH) and camera pose (CP) deliver the most decisive zero-shot improvements, vastly outperforming appearance perturbations like background textures or lighting.
  • Frame-level granularity prevents static overfitting: Continual per-step randomizations eliminate spurious static visual cues, channeling transformer attention to critical object interaction zones and boosting transfer accuracy.
  • RL fine-tuning instills closed-loop resilience: Simulation-based RL fine-tuning without real demonstrations lifts physical deployment success from 5.6% to 33.4% even in clean environments, reaching 42.8% when united with domain randomization.

Highlights & Insights

  • First rigorous 10k-trial physical benchmarking: Decouples the intertwined factors of Sim-to-Real transfer on physical bimanual robots at an unprecedented empirical scale, establishing reproducible guidelines for embodied foundation models.
  • Debunking visual-centric assumptions: Demonstrates that physical VLA policies care far more about 3D geometric and viewpoint alignments than precise photometric rendering match.
  • Exploration vs. imitation dynamics: While imitation memorizes trajectory correlations, reinforcement learning equips the policy with dynamic recovery capabilities necessary to overcome physical deployment disturbances.

Limitations & Future Work

  • High-contact dynamic tasks remain a bottleneck: In contact-rich, impulsive manipulation tasks like Beat Block Hammer, real-world success remains low (11.5% under full DR), highlighting persistent simulation inaccuracies in micro-contact dynamics and impact restitution.
  • Single forward-camera limitation: The setup relies exclusively on a single front-facing RealSense D435 camera, omitting wrist-mounted eye-in-hand perspectives or tactile sensing streams.
  • Future directions: Integrating tactile sensing to resolve contact ambiguities in stiff manipulation and investigating real-to-sim closed-loop exploration for continuous deployment adaptation.
  • vs. RoboTwin 2.0 [4] / MimicGen [28]: While prior systems emphasize data generation pipelines, this work provides a rigorous causal decomposition of which synthetic generation factors genuinely transfer to physical VLA models.
  • vs. SimpleVLA-RL [19]: While SimpleVLA-RL introduced reinforcement fine-tuning for VLA scaling, this work illuminates how RL inherently promotes zero-shot out-of-distribution transfer and synergizes orthogonally with environmental domain randomization.

Rating

  • Novelty: โญโญโญโญ [Pioneering, large-scale factorized empirical investigation into zero-shot Sim-to-Real VLA transfer]
  • Experimental Thoroughness: โญโญโญโญโญ [Exceeds 10,000 real-world robot trials across 5 dual-arm manipulation tasks]
  • Writing Quality: โญโญโญโญโญ [Clear problem formulation, elegant empirical structure, and insightful causal takeaways]
  • Value: โญโญโญโญโญ [Provides essential design principles for synthetic data curation and sim-trained VLA deployment]