Skip to content

One Demonstration Is Enough for Real-World Robotic Reinforcement Learning

Conference: ECCV 2026
Paper: ECCV Official
Code: https://autoserl.github.io/
Area: Robotics & Embodied AI
Keywords: robotic reinforcement learning, one-shot learning, automated intervention, safety recovery, real-world manipulation

TL;DR

AutoSERL replaces expensive human teleoperation and multi-demonstration dependence in real-world robotic reinforcement learning by deriving automated sliding-window guidance, stagnation recovery replays, and adaptive termination criteria from a single demonstration, achieving 100% success across 6 contact-intensive physical tasks while matching HIL-SERL.

Background & Motivation

Deploying reinforcement learning on physical robotic hardware faces two persistent fundamental barriers. First, physical real-world exploration incurs non-trivial hardware risks: unlike in simulation where execution failures carry no cost, unconstrained exploratory motions can cause dangerous hardware collisions or motor overload. Second, contact-intensive manipulation tasksโ€”such as precision insertion, hanging, and drawer manipulationโ€”feature exceptionally sparse reward landscapes where task-completing states occupy only a negligible fraction of the configuration space, causing model-free RL agents to struggle through thousands of unproductive iterations without encountering learning signals.

To bypass sample inefficiency and dangerous exploratory drift, recent state-of-the-art frameworks have combined demonstrations with online experience. SERL initializes policy learning with a pool of offline demonstrations to bootstrap value estimation, while HIL-SERL introduces human-in-the-loop intervention during training, where human operators actively teleoperate the robot away from deadlocks and unsafe states. However, continuous human supervision imposes a prohibitive operational bottleneck: human supervisors suffer from cognitive fatigue, inconsistent response latency, and substantial labor costs per training hour, rendering unassisted continuous training infeasible.

A rigorous breakdown of intervention episodes in HIL-SERL reveals four primary triggers: convergence to local optima with low-magnitude oscillations, Q-value overestimation driving the end-effector away from targets, environmental obstacles intercepting free-space exploration, and physical stagnation caused by unaligned contacts on object surfaces. The core insight of this paper is that the corrective bias and safety constraints provided by human operators can be fully automated using the geometric and temporal structure of a single reference demonstration. Core idea: by deriving forward sliding window guidance, stagnation-triggered safety recovery replays, and an intervention termination threshold entirely from a single expert demonstration, AutoSERL enables fully automated, safe, and sample-efficient real-world robotic reinforcement learning without continuous human supervision.

Method

Overall Architecture

AutoSERL formulates an autonomous intervention loop around a single expert demonstration trajectory collected prior to training. On this trajectory, two critical geometric anchors are annotated: a safe recovery target \(\text{recover\_point}_0\) positioned in free space away from surrounding clutter, and a stable contact point \(\text{recover\_point}_1\) where the robot establishes steady physical engagement with the target object. Both anchor points are identified by replaying the demonstration on the physical setup.

The automated supervision system coordinates three complementary components: a sliding window intervention that actively pulls the robot back toward the demonstration path during trajectory deviations; a passive safety recovery mechanism that detects physical deadlocks and restores progress via replay; and an intervention termination criterion that permanently disengages all guidance once the policy achieves autonomous proficiency. Transitions recorded during automated interventions and the single expert demonstration are continuously piped into the Demo Buffer and Replay Buffer to stabilize online policy optimization.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Single Expert Demonstration<br/>Annotate safe target and stable contact points"] --> B["Sliding Window Intervention<br/>Forward tangent angle and distance gating"]
    B --> C["Safety Recovery Mechanism<br/>Temporal stagnation check and demo segment replay"]
    C --> D["Intervention Termination<br/>Threshold gating on intervention count per episode"]
    D --> E["Autonomous Policy Exploration and Trajectory Optimization"]

Key Designs

1. Sliding Window Intervention: forward-constrained active geometric guidance To prevent policies from vibrating in local optima, drifting away from the task object due to inaccurate Q-value estimation, or colliding with ambient obstacles, AutoSERL introduces a sliding window along the demonstration trajectory. The window length corresponds to the trajectory index distance between \(\text{recover\_point}_0\) and \(\text{recover\_point}_1\). When uninhibited by interventions, the window advances along the demonstration path by one step at each environment time step. At each step, the Euclidean distance between the current end-effector pose and all poses inside the sliding window is calculated, and the minimum-distance point is selected as a candidate interference point \(p_{\text{ipoint}}\). Intervention is conditionally considered only when this distance exceeds the trigger threshold \(th_2 = 0.02\text{ m}\).

To strictly prevent the guidance controller from pulling the robot backward into previously traversed trajectory segments, AutoSERL performs a directional sanity check. Let \(v_1\) denote the forward tangent vector of the demonstration trajectory at the closest point \(c_{\text{point}}\), and \(v_2\) denote the displacement vector from the current end-effector pose to \(p_{\text{ipoint}}\). The angle between the two vectors is evaluated: $\(\theta = \arccos\left(\frac{v_1 \cdot v_2}{\|v_1\| \|v_2\|}\right) \le 90^\circ\)$ If \(\theta > 90^\circ\), the candidate point is recognized as lying on a historically visited path segment and the intervention is discarded. Otherwise, when \(\theta \le 90^\circ\), an operational space motion planner guides the end-effector toward \(p_{\text{ipoint}}\) until the remaining tracking error falls below the convergence threshold \(th_1 = 0.005\text{ m}\). Because the reference demonstration is inherently collision-free, this windowed attraction implicitly enforces obstacle avoidance throughout free-space navigation.

2. Safety Recovery Mechanism: temporal distance stagnation detection and replay recovery During intricate contact phasesโ€”such as connector pins misaligned against socket margins or hanger hooks binding against support barsโ€”active sliding-window attractions cannot resolve physical interlocks and can lead to persistent stalls. AutoSERL addresses this contact-induced deadlocking through a passive safety recovery mechanism based on temporal projection statistics. The system continuously tracks the closest point \(c_{\text{point}}\) across the entire demonstration trajectory, maintaining a history buffer over the preceding \(l_{\text{stag}} = 20\) time steps.

Stagnation is detected when the end-effector's closest projection remains virtually stationary over the past \(l_{\text{stag}}\) steps, satisfying \(\|c_{\text{point}, i} - c_{\text{point}, i - l_{\text{stag}}}\| < th_1 = 0.005\text{ m}\), while \(c_{\text{point}, i}\) lies within the contact-critical segment \([\text{recover\_point}_0, \text{recover\_point}_1]\). Once this condition is met, policy execution is halted, and a two-stage recovery sequence executes: motion planning first disengages the robot back to the safe clearance pose \(\text{recover\_point}_0\), after which the controller replays the open-loop demonstration trajectory segment between \(\text{recover\_point}_0\) and \(\text{recover\_point}_1\). This deterministic replay resets the end-effector into stable physical contact, allowing exploration to resume safely.

3. Intervention Termination: adaptive transition to autonomous reinforcement learning Because human teleoperated demonstrations are frequently sub-optimal, maintaining continuous external intervention throughout training risks over-constraining the agent, causing it to merely mimic demonstrated motions rather than discovering more efficient trajectories. To preserve the exploratory advantages of reinforcement learning, AutoSERL introduces an episode-level termination criterion.

In training setups without initial state randomization, the system counts the cumulative number of triggered interventions within each episode. Once an episode achieves task success while logging fewer than \(l_{\text{term}} = 10\) intervention steps, the policy is judged to have achieved sufficient autonomous competence. From that episode forward, both sliding-window interventions and safety recovery mechanisms are permanently disabled. The policy is thereafter trained exclusively via unconstrained online reinforcement learning, enabling it to optimize beyond the demonstration trajectory.

Loss & Training

AutoSERL builds upon the SERL software suite, employing an off-policy Actor-Critic backbone (Soft Actor-Critic with asymmetric inputs). Observations consist of dual-view RGB images from RealSense cameras and proprioceptive states (end-effector poses, linear/angular velocities, forces/torques, and gripper status). The action space is a 6D delta end-effector Cartesian pose \(\Delta \text{pose} \in \mathbb{R}^6\). Rewards are binary and sparse (\(r \in \{0, 1\}\)), assigned upon task completion. Maximum episode length is 300 steps.

The framework maintains two replay pools: a Demo Buffer storing transitions from the initial single demonstration alongside intervention-guided trajectories, and an online Replay Buffer storing autonomous exploratory rollouts. Training batches maintain a balanced sampling ratio between demonstration and online transitions, ensuring accurate Q-value propagation in sparse reward environments while entropy regularization drives fine-grained force-motion coordination.

Key Experimental Results

Main Results

AutoSERL was evaluated on 6 real-world contact-intensive tasks across two robot arms (Franka Emika Panda and Universal Robots UR5): USB insertion, electrical plug insertion, hanger suspension, correction tape suspension, spoon suspension, and drawer opening with a hook. During evaluation, all automated intervention mechanisms were deactivated, and each task was evaluated across 50 independent episodes. Success rate and minimum real-robot training time required to reach 100% success were recorded.

Task Category Specific Task Training Budget SERL (20 demos) AutoSERL (1 demo) Gain
Insertion USB Insertion (Franka) 8 min 20/50 (40%) 50/50 (100%) +60%
Insertion Plug Insertion (Franka) 8 min 0/50 (0%) 50/50 (100%) +100%
Hanging Hanger Suspension (UR5) 33 min 0/50 (0%) 50/50 (100%) +100%
Hanging Correction Tape Suspension (UR5) 25 min 6/50 (12%) 50/50 (100%) +88%
Hanging Spoon Suspension (UR5) 35 min 0/50 (0%) 50/50 (100%) +100%
Hinge-based Drawer Opening (UR5) 45 min 0/50 (0%) 50/50 (100%) +100%

In comparisons measuring the minimum training time required to achieve a 50/50 (100%) success rate, AutoSERL matched or outperformed HIL-SERL (which requires continuous human teleoperation): for USB insertion, AutoSERL required 8 min (vs. 6 min for HIL-SERL); for plug insertion, both required 8 min; for hanger suspension, AutoSERL took 33 min (vs. 48 min for HIL-SERL); for correction tape suspension, AutoSERL took 25 min (vs. 60 min for HIL-SERL); for spoon suspension, AutoSERL took 35 min (vs. 70 min for HIL-SERL); and for drawer opening, both took 45 min.

Compared against imitation learning baselines, MILES (a dedicated one-shot imitation method) achieved only 0/50 on USB insertion, 33/50 on plug insertion, 1/50 on correction tape, 42/50 on hanger, 2/50 on spoon, and 0/50 on drawer opening. Standard behavior cloning (BC) achieved 5/50 on USB, 2/50 on plug, 38/50 on correction tape, 37/50 on hanger, 0/50 on spoon, and 35/50 on drawer opening. AutoSERL attained 50/50 on every single task, proving the necessity of reinforcement learning fine-tuning.

Ablation Study

To quantify the contributions of individual mechanisms, ablation studies were conducted on the plug insertion, USB insertion, and drawer opening tasks under controlled configurations.

Configuration Task Evaluated Steps to 100% / Performance Outcome Note
Full AutoSERL Plug / USB Insertion Reaches 100% success earliest Seamless synergy between sliding guidance and stagnation replay
No sliding window intervention Plug / USB Insertion Requires substantially more training steps Lacks active directional pull; agent struggles with local minima
No recovery mechanism Plug / USB Insertion Performs worse than SERL baseline; fails to converge Sliding window pulls arm into contact traps with no self-healing
No intervention termination Drawer Opening Peaks at ~19.5k steps then degrades Over-relies on corrections; small actions misclassified as stalls
Termination threshold \(l_{\text{term}}=50\) Drawer Opening Slower and lower convergence than \(l_{\text{term}}=10\) Oversized threshold terminates guidance prematurely before mastery

Key Findings

  • The recovery mechanism is essential for sliding window safety: Removing the recovery mechanism while keeping sliding window interventions caused performance to drop below the SERL baseline. Without recovery replays, the window attraction frequently forces the gripper into unrecoverable contact pinches, trapping the policy in deadlocks that sparse-reward exploration cannot resolve.
  • Intervention termination unlocks policy optimization beyond the demonstration: In the plug insertion task, the initial demonstration trajectory spanned 99 time steps due to human teleoperation noise. The policy trained by AutoSERL completed the same insertion in only 54 steps with a visibly smoother 3D spatial path, demonstrating trajectory-level optimization enabled by disabling external interventions.
  • Hyperparameter sensitivity favors moderate values: For the tracking threshold \(th_1\), a value that is too small (0.001 m) introduces excessive planning delays, while a value that is too large (0.02 m) leaves substantial residual offset. For the intervention threshold \(th_2\), small values (0.005 m) trigger hyper-frequent interventions that suppress RL exploration, while large values (0.04 m) fail to arrest dangerous drift. The default settings (\(th_1=0.005\text{ m}, th_2=0.02\text{ m}\)) achieve optimal stability.
  • High repeatability across seeds and spatial perturbations: AutoSERL maintained near 100% success across 5 random seeds (40โ€“44) in plug insertion and converged reliably under \(\pm 3\text{ cm}\) initial 2D planar position randomizations, exhibiting strong spatial generalization.

Highlights & Insights

  • Tangent-vector angular gating: Enforcing the \(\theta \le 90^\circ\) angular constraint between trajectory tangents and intervention vectors provides an elegant, parameter-free filter that guarantees the robot is never pulled backward along previously visited paths.
  • Dual-anchor contact reset protocol: Structuring recovery around \(\text{recover\_point}_0\) (free space) and \(\text{recover\_point}_1\) (contact engagement) automates the intuitive human habit of retracting, realigning, and re-inserting, eliminating physical interlocks without complex force-torque modeling.
  • Decoupled bootstrapping and exploration: Transitioning from guided correction to unconstrained policy optimization via an intervention-count threshold resolves the dilemma of demonstration-guided RL, where policies often overfit to sub-optimal demonstrations.

Limitations & Future Work

  • Single-trajectory failure diversity limit: When a policy encounters anomalous failure modes unrepresented in the single demonstration trajectory, the rigid two-point recovery protocol may fail to dislodge the end-effector. Integrating multi-trajectory recovery policies (e.g., UniIntervene or FARL) represents an important future step.
  • Restriction to 6D Cartesian action spaces: The current implementation operates over 6D delta Cartesian poses of the end-effector. Extending automated geometric guidance to high-dimensional joint spaces or multi-fingered dexterous hands requires more sophisticated topological constraints.
  • Large-scale visual and semantic domain shifts: The reference trajectory assumes the object remains within moderate positional offsets (\(\pm 3\text{ cm}\)). Tasks with random object rotations, flips, or visual clutter will require coupling with visual foundation models for dynamic trajectory warping.
  • vs HIL-SERL: HIL-SERL relies on human supervisors continuously monitoring the setup to provide corrective actions via teleoperation, limiting scalability; AutoSERL automates all four primary human intervention triggers, matching or reducing physical training time without human involvement.
  • vs SERL: SERL relies on 20+ offline demonstrations and struggles to solve highly constrained contact tasks with sparse rewards; AutoSERL requires only 1 demonstration and achieves 100% success across all tasks where SERL achieved 0%โ€“40% under identical training budgets.
  • vs MILES / BC: One-shot and few-shot imitation learning methods suffer from compounding errors and distribution shift in high-precision contact tasks; AutoSERL uses the single demonstration strictly as an exploration scaffolding while online RL learns resilient closed-loop feedback policies.

Rating

  • Novelty: โญโญโญโญโญ [Extracts complete automated intervention and recovery loops from a single demonstration using clean geometric principles, removing the need for human teleoperation]
  • Experimental Thoroughness: โญโญโญโญโญ [Evaluated on 6 challenging real-world tasks across two distinct robot platforms, featuring seed robustness, spatial variation, hyperparameter sweeps, and ablations]
  • Writing Quality: โญโญโญโญโญ [Clear structural narrative, well-formulated technical mechanisms, and transparent motivation]
  • Value: โญโญโญโญโญ [Significantly lowers the barrier to practical real-world robotic reinforcement learning, offering a blueprint for autonomous physical policy learning]