一个两阶段的强化学习框架,用于人形机器人坐着和站着
Xisheng Jiang1,2,3,4, Shihai Zhao1,2, Yudi Zhu1,2,3,4
1School of Optoelectronic Information and Computer Engineering, University of Shanghai for Science and Technology, Shanghai 200093, China.
Biomimetics (Basel, Switzerland)
|November 26, 2025
概括
人形机器人现在可以自主地坐和站,使用一种新的两阶段强化学习 (RL) 框架. 这种方法提高了机器人的稳定性和流性,用于日常生活中的实际应用.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
背景情况:
- 人形机器人需要自主坐和站起来进行实际应用,但传统的控制器却在复杂的动态和各种场景中扎.
- 强化学习 (RL) 提供了运动控制的潜力,但往往导致不稳定或突然的运动.
研究的目的:
- 开发一个强大的两阶段强化学习框架,用于在人形机器人中实现平稳和稳定的自主坐和站.
- 克服直接RL训练在实现平衡运动控制方面的局限性.
主要方法:
- 一个两阶段的强化学习方法:初始的政策培训与宽松的约束,其次是轨迹跟踪改进.
- 适应性课程学习的双层优化,根据错误动态调整跟踪精度.
主要成果:
- 在一个1.7米长的成人规模的人形机器人上,稳定地执行自主坐和站动作.
- 在现实世界的场景中成功演示,包括从椅子上坐下来和站起来.
结论:
- 拟议的两阶段RL框架有效地产生平滑和稳定的人形机器人运动,用于日常任务.
- 这种方法为增强人形机器人的自主性和实用性在现实环境中提供了一个有希望的解决方案.
相关概念视频
Reinforcement Schedules
438
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
438
Reinforcement
797
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
797


