Autonomous Robot Manipulation Skills Through Language-Driven Quality-Diversity

Autonomous Robot Manipulation Skills Through Language-Driven Quality-Diversity

Émiland Garrabé, Mahdi Khoramshahi, Stéphane Doncieux

4 min branja1. sep. 2026

Overall, most methods for autonomously acquiring robot skills focus on obtaining single, high-performing solutions. This can lead to brittleness when deployed to the real world, due to the high dimensionality of manipulation task spaces, and curricula based on these techniques need to be carefully curated. Conversely, while motion primitives obtained with quality-diversity pipelines are well-suited to out-of-the-box integration due to their diversity, the design of useful, task-specific diversity metrics and fitness signals still requires expert efforts, and this is incompatible with autonomy requirements. Language-model-based reward shaping techniques can be adapted for quality-diversity algorithms in the context of manipulation tasks, leading to archives of diverse solutions while being compatible with free-form instructions.

Trajectory Datasets in Robotics

Overall, most methods for autonomously acquiring robot skills focus on obtaining single, high-performing solutions. This can lead to brittleness when deployed to the real world, due to the high dimensionality of manipulation task spaces, and curricula based on these techniques need to be carefully curated.

Conversely, while motion primitives obtained with quality-diversity pipelines are well-suited to out-of-the-box integration due to their diversity, the design of useful, task-specific diversity metrics and fitness signals still requires expert efforts, and this is incompatible with autonomy requirements.

Language-model-based reward shaping techniques can be adapted for quality-diversity algorithms in the context of manipulation tasks, leading to archives of diverse solutions while being compatible with free-form instructions.

Method

This section begins by formalizing the robotic skill archive design problem and the parametrization of a MAP-Elites-Success algorithm for robotic skill acquisition. It then proposes a sequential technique for exploring the behavior descriptor and fitness spaces using language models, without expert input or task-based adaptation. Finally, the generated mesh of behavior-descriptor samples can be used with a multi-behavior-descriptor MAP-Elites-Success algorithm to generate an archive of diverse motion primitives for a given task.

Trajectory Dataset Generation for Robotics

Given a known task and unknown task instance distribution, the robotic quality-diversity design problem is cast as follows:

Find an archive of motion primitives satisfying the task requirements across the unknown task instance distribution.

Here, the archive of motion primitives is a collection of candidate robot behaviors, and the motion-primitive space is the space of possible motion primitives.

Language-driven quality-diversity exploration process

LLM Inference for Multi-BD MAP-Elites-Success

This section introduces a strategy for parameterizing quality-diversity algorithms with language models.

While obtaining a correct success criterion is critically important for the resulting archive’s quality, it is relatively easy to design. Accordingly, only one success condition is inferred for a given task, reducing computational load.

Experiments

This section details the implementation of the proposed algorithm used in the experiments and describes the experimental setting used to validate the method.

Robot grasping motion primitive

Implementation: Perception API

The following information is available to the language model:

  1. The object base’s position and orientation. Euler angles are chosen because they are easier to interpret for a non-expert system such as a language model.
  2. The contacts between the table and the robot, respectively the object.
  3. The contacts between the robot’s gripper and each link of the object, or the object base for non-articulated objects.
  4. The robot gripper’s position and orientation.
  5. For articulated objects, the position of each joint, represented by its angle or prismatic displacement.

Environment and Tasks

Expert success condition: The object must touch the gripper and not the table.

Ground-truth behavior descriptor: The gripper’s XYZ position at first contact with the object.

Robot interaction with an articulated drawer

Metrics

Success Archive Size

The number of successful individuals, measured with an expert-written success condition, output by each run.

Ground-Truth Behavior Descriptor Coverage

For each task, the output archive from each baseline is projected into the ground-truth behavior-descriptor space used for the hand-crafted baseline. This ground-truth behavior-descriptor space is hand-crafted to capture meaningful diversity in task approaches, providing a measure of actual solution diversity.

Frequently Asked Questions

What problem does language-driven quality-diversity address? It addresses the autonomous acquisition of diverse robot manipulation skills without requiring expert-designed task-specific diversity metrics and fitness signals.

Why are diverse motion primitives useful for robot manipulation? Their diversity makes them well-suited for out-of-the-box integration and can reduce brittleness in high-dimensional manipulation task spaces.

What information is provided to the language model? The language model receives object pose, contact information, gripper pose, and joint positions for articulated objects.

How is ground-truth behavior-descriptor coverage measured? Output archives are projected into a hand-crafted ground-truth behavior-descriptor space that captures meaningful diversity in task approaches.

🍪 Nastavitve piškotkov

Uporabljamo piškotke za merjenje zmogljivosti. Politika zasebnosti