通过协作式交互式反向增强学习方法为社会机器人提供以目标为导向的自主决策
Mingyue Luo1, Hui Li2, Wanbo Luo3,4
1School of Mechatronic Engineering, Changchun University of Technology, Yan'an St., Changchun, 130012, Jilin Province, China.
Scientific reports
|July 30, 2025
概括
本研究介绍了社会机器人的目标导向自主决策 (GO-ADM) 方法,通过从专家演示中学习而改善导航,而不是预测行人路径. 该GO-ADM方法确保安全和高效的导航,即使有干扰.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
- 人与机器人的交互
背景情况:
- 对于社交机器人而言,现有的反向增强学习 (IRL) 通常依赖于轨迹规划,这对于具有明确目标和长距离预测的机器人来说是不切实际的.
- 目前的方法面临的局限性是由于未知的行人目标信息和预测轨迹的复杂性,特别是在更长的距离.
- 社会机器人需要强大的导航策略,考虑到社会规范和安全.
研究的目的:
- 为社会机器人提出一种新的以目标为导向的自主决策 (GO-ADM) 方法.
- 通过使机器人能够在没有预测行人轨迹的情况下做出连续决策来解决社会兼容性导航问题.
- 通过社会距离考虑,加强人机交互安全.
主要方法:
- 开发了一个以目标为导向的自主决策 (GO-ADM) 框架,使用在离散时间步骤中的顺序行动.
- 定义并收集了针对培训的以目标为导向的专家演示.
- 提出了一个协作式的交互式反向强化学习 (IRL) 框架,将明确和隐性行人运动规范集成到目标导向的奖励功能和社会安全距离处罚中.
主要成果:
- GO-ADM方法在纵向和横向主导导航任务中实现了合理的自主决策,平均目的地偏差分别低于0.13m和0.23m.
- 与其他社交导航和决策算法相比,其成功率明显更高.
- 对未知的干扰表现出强大的强度,在恶劣条件下 (0.5m噪声) 达到75%以上的成功率.
结论:
- 拟议的GO-ADM方法有效地解决了具有明确目标的机器人的社会兼容性导航问题.
- 这种方法消除了复杂的行人轨迹预测的需要,提供了更实用的解决方案.
- 在人机交互场景中,GO-ADM确保了安全,高效和强大的导航.
相关概念视频
Observational Learning
314
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
314
Reinforcement
343
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
343
Avoidance Learning and Learned Helplessness
1.9K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.9K
Decision Making
233
Decision-making is a fundamental cognitive process that involves evaluating alternatives and selecting among them. This process can range from simple choices, such as deciding what to wear, to complex decisions, like choosing a major in college or a career path. The complexity of the decision often dictates the approach we use, which can be broadly categorized into two types: automatic and controlled decision-making.
Automatic decision-making is fast, intuitive, and relies on gut feelings...
Automatic decision-making is fast, intuitive, and relies on gut feelings...
233
Decision Making: Traditional Method
4.2K
The process of hypothesis testing based on the traditional method includes calculating the critical value, testing the value of the test statistic using the sample data, and interpreting these values.
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
4.2K
Purposive Learning
207
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
207


