在LLM中,对长视野操纵任务的动作原始体进行了增强层次的强化学习
Ning Zhang1, Yongjia Zhao2,3, Minghao Yang4
1China's State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing, 100191, China.
Scientific reports
|October 21, 2025
概括
本研究介绍了LARAP,这是一种新的层次增强学习代理,它使用大型语言模型 (LLM) 来指导复杂的操纵任务的政策学习,提高样本效率和性能.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 深度强化学习 (deep reinforcement learning,RL) 面临的挑战是长时间的操纵任务,因为状态空间大,奖励稀少.
- 层次的RL可以提高技能学习,但在培训效率和可转移性方面存在困难.
- 大型语言模型 (LLM) 提供世界知识和推理,但缺乏现实世界的任务接地.
研究的目的:
- 开发一种新的方法,将LLMs的规划能力与RL结合起来,用于长时间的操纵任务.
- 通过将LLMs集成到一个层次化的RL框架中,提高复杂的机器人任务中的样本效率和性能.
主要方法:
- 提出了一个等级代理,LARAP (具有参数化的动作原始的LLM引导等级代理).
- 利用LLM指导高层政策,提高培训期间的样本效率.
- 结合LLM与参数化的动作原始体进行长视野操纵.
主要成果:
- 在各种模拟操纵任务中,LARAP显著超过了基线方法.
- 与传统的RL方法相比,这种方法证明了样本效率的提高.
- 整合LLMs有效指导等级代理人的学习过程.
结论:
- 拉拉普为解决机器人技术中长期操纵挑战提供了一个有希望的解决方案.
- 将LLM与RL结合起来,为复杂任务学习提供了一个强大的框架.
- 拉拉普代理显示了潜在的现实世界的应用程序,需要复杂的操纵技能.
相关概念视频
Observational Learning
824
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
824
Long-term Potentiation
58.3K
Long-term potentiation, or LTP, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTP is the process of synaptic strengthening that occurs over time between pre- and postsynaptic neuronal connections. The synaptic strengthening of LTP works in opposition to the synaptic weakening of long-term depression (LTD) and together are the main mechanisms that underlie learning and memory.
58.3K
Long-term Potentiation
3.4K
Long-term potentiation, or LTP, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTP is the process of synaptic strengthening that occurs over time between pre and postsynaptic neuronal connections. The synaptic strengthening of LTP works in opposition to the synaptic weakening of long-term depression (LTD) and together are the main mechanisms that underlie learning and memory.
Hebbian LTP
LTP can occur when...
Hebbian LTP
LTP can occur when...
3.4K
Reinforcement Schedules
453
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
453
Cognitive Learning
1.0K
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
1.0K


