Related Work
We do not claim a fundamentally new curriculum principle. Instead, we show that a simple variant of success-based sampling, combined with diverse resets and large-scale reinforcement learning, can solve challenging locomotion and assembly tasks that prior sampling methods struggle to learn.
Algorithmic Decisions for Reinforcement Learning Training
Beyond the sampler itself, SGS makes two deliberate algorithmic choices that work together to support training across a wide task configuration space with PPO.
Shared Rewards Within Each Domain
We use one simple reward function across all locomotion tasks and another across all manipulation tasks. Each combines a boolean task-success reward obtained at the end of the episode with relatively small robot-specific, per-step regularization terms.
The shared-reward design avoids specifying a different reward function for each terrain or assembly task. Appendix A.3 gives the terms and weights for each domain.
Experimental Results
Through a series of large-scale scaling studies, we ask:
- Scaling: Does SGS scale better than uniform sampling and PLR as parallel training environments grow?
- Capability: Does SGS unlock robotics tasks beyond what prior methods could solve, even at smaller scales?
- Transfer: Can dexterous manipulation behaviors learned with SGS transfer zero-shot to real hardware?
Manipulation
We use the same simple reward function and diverse simulator resets across all manipulation tasks. Our main scaling manipulation experiments use a gravity curriculum that we later found to be unnecessary; Appendix B.4 describes its implementation and ablates its effect.

Locomotion
Terrains from previous works typically have a linear structure, where difficulty is determined by parameters such as slope or step height, linearly interpolated from easy to hard. For our procedurally generated terrains, a clear ordering in terms of difficulty may not exist, so the difficulty axis can be interpreted as different seeds.
The terrains used in this paper are shown in the middle and bottom rows of Figure 3, with the exception of Inverted Slope, which we omit due to space constraints.
Observations, Actions, and Rewards
We describe the observation and action spaces for each embodiment, followed by the reward definitions and weights.
The locomotion policy observes base linear and angular velocity, projected gravity, joint positions and velocities, the previous action, the goal command, and a local terrain height scan. Its 12-dimensional action specifies joint-position offsets from the robot’s nominal pose, which are tracked by the joint controllers.
Comparison with Sampling for Learnability
Table 1 in the main text compares SFL with SGS on locomotion. We use SFL’s default hyperparameters and report means and 95% confidence intervals over three training seeds.
SFL’s final performance is lower than SGS at both scales. At 32K, its best checkpoint comes close, but still underperforms SGS. Notably, performance significantly degrades as training continues, most noticeably at larger scales, whereas SGS monotonically improves during training.
SGS Hyperparameters and Configuration Density
For each task configuration, we store a window of the latest Boolean outcomes in a fixed-length history in a circular buffer. Our sampling floors the kernel weight before forming logits and enforces a minimum temperature of one.
Hardware and Control Setup
We use a UR5e arm with a Robotiq 2F-85 gripper and follow OmniReset’s relative Cartesian operational-space control interface. Our camera setup uses one third-person camera and one wrist-mounted camera, rather than OmniReset’s three cameras. We describe the point-cloud teacher and RGB student in Appendix C.3.
The NIST board is made of polished acrylic. When the wrist camera faces the board directly, the robot observes its own reflection, which is out of distribution. The low surface friction also impairs policies that use the board to reorient the object. For these reasons, we covered the board with tape during evaluation.

Simulation Modeling and Domain Randomization
We follow OmniReset’s domain-randomization setup, adding small Gaussian noise to the point-cloud observations and randomizing object masses around their measured real-world values.
Teacher–Student Distillation
Each iteration collects one step across all environments, updates the student toward the teacher’s actions, and discards the data. This limits GPU memory use, a bottleneck when simulation and rendering run alongside training.
However, training success can be misleading: updates from the immediately preceding step can guide the student’s next action toward the teacher, helping it solve a task it cannot yet perform without these updates. We therefore reserve 6.67% of environments for evaluation and exclude their data from gradient updates. Success in these held-out environments lags training success, revealing the need for longer distillation.
Transfer Observations and Policy Behaviors
A consistent lesson across tasks was that policies which commit to a clean grasp and insert directly transfer better than ones that exploit simulator contact dynamics.
For rod-in-hole insertion, adding 1 mm per-step Gaussian noise to the point-cloud observation further helped by encouraging the policy to re-aim after a miss and wiggle the rod when it was not fully seated.
For nut-and-bolt assembly, the policy repeatedly dropped the nut onto the bolt instead of seating it, flipping the nut each time. We speculate this was due to the teacher reward, which only counted one upright nut orientation for success despite the nut being symmetric. Removing this constraint and using appropriately tuned action scaling on the 6-DOF end-effector control space produced a policy that completed the task.
Real-World Evaluation Protocol
Figure 4 shows hardware trajectories, and Table 2 reports task-level results. We follow the evaluation protocol and metric definitions of OmniReset, reporting overall success rate, first-try success rate, and throughput, with the differences noted below.
We sample initial object configurations by uniformly randomizing object positions over the NIST board. The board is fixed to the table with command strips. In simulation, the receptive objects are static, and the policy learns to reorient the insertive object relative to a fixed reference. If the board is not secured, the receptive objects can shift during contact, which causes failures.
Frequently Asked Questions
What does the paper show about SGS? A simple variant of success-based sampling, combined with diverse resets and large-scale reinforcement learning, can solve challenging locomotion and assembly tasks that prior sampling methods struggle to learn.
How are rewards structured across tasks? The method uses one reward function for locomotion and another for manipulation. Each combines an end-of-episode boolean task-success reward with small robot-specific, per-step regularization terms.
Why are some environments held out during teacher–student distillation? The held-out environments reveal whether the student can succeed without guidance from updates based on immediately preceding steps. Their success lagged training success, indicating a need for longer distillation.
What helped policies transfer to real-world insertion tasks? Policies that commit to a clean grasp and insert directly transferred better than policies that exploit simulator contact dynamics. For rod-in-hole insertion, point-cloud noise also encouraged re-aiming after a miss.
