Related Experiment Video
Updated: Jan 14, 2026

07:52
Investigating Motor Skill Learning Processes with a Robotic Manipulandum
Published on: February 12, 2017
9.2K
LLMs augmented hierarchical reinforcement learning with action primitives for long-horizon manipulation tasks.
Ning Zhang1, Yongjia Zhao2,3, Minghao Yang4
1China's State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing, 100191, China.
Scientific Reports
|October 21, 2025
Summary
This study introduces LARAP, a novel hierarchical reinforcement learning agent that uses large language models (LLMs) to guide policy learning for complex manipulation tasks, improving sample efficiency and performance.
Area of Science:
- Robotics
- Artificial Intelligence
- Machine Learning
Background:
- Deep reinforcement learning (RL) faces challenges with long-horizon manipulation tasks due to large state spaces and sparse rewards.
- Hierarchical RL improves skill learning but struggles with training efficiency and transferability.
- Large language models (LLMs) offer world knowledge and reasoning but lack real-world task grounding.
Purpose of the Study:
- To develop a novel approach that combines the planning capabilities of LLMs with RL for long-horizon manipulation tasks.
- To improve sample efficiency and performance in complex robotic tasks by integrating LLMs into a hierarchical RL framework.
Main Methods:
- Proposed a hierarchical agent, LARAP (LLM-guided hierarchical agent with parameterized action primitives).
- Utilized LLMs to guide a high-level policy, enhancing sample efficiency during training.
- Combined LLMs with parameterized action primitives for long-horizon manipulation.
Main Results:
- LARAP significantly outperformed baseline methods in various simulated manipulation tasks.
- The approach demonstrated improved sample efficiency compared to traditional RL methods.
- The integration of LLMs effectively guided the hierarchical agent's learning process.
Conclusions:
- LARAP offers a promising solution for addressing long-horizon manipulation challenges in robotics.
- Combining LLMs with RL provides a powerful framework for complex task learning.
- The LARAP agent shows potential for real-world applications requiring sophisticated manipulation skills.
Related Concept Videos
Observational Learning
824
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
824
Long-term Potentiation
58.3K
Long-term potentiation, or LTP, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTP is the process of synaptic strengthening that occurs over time between pre- and postsynaptic neuronal connections. The synaptic strengthening of LTP works in opposition to the synaptic weakening of long-term depression (LTD) and together are the main mechanisms that underlie learning and memory.
58.3K
Long-term Potentiation
3.4K
Long-term potentiation, or LTP, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTP is the process of synaptic strengthening that occurs over time between pre and postsynaptic neuronal connections. The synaptic strengthening of LTP works in opposition to the synaptic weakening of long-term depression (LTD) and together are the main mechanisms that underlie learning and memory.
Hebbian LTP
LTP can occur when...
Hebbian LTP
LTP can occur when...
3.4K
Reinforcement Schedules
453
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
453
Cognitive Learning
1.0K
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
1.0K

