在混合人群中进行多目的地导航的模块化层次增强学习
Wen Ou1, Biao Luo1, Bingchuan Wang1
1School of Automation, Central South University, Changsha 410083, China.
概括
本研究介绍了一个模块化层次强化学习 (MHRL) 方法,用于在拥挤的环境中进行机器人导航. MHRL有效地处理动态和静态人群,优于现有方法.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 现实世界的机器人导航通常涉及多个目的地.
- 拥挤的环境会因为动态和静态障碍而带来挑战.
- 现有的方法与复杂的群众互动作斗争.
研究的目的:
- 开发一种用于复杂,拥挤的环境中的机器人导航的新方法.
- 为了应对动态和静态人群带来的挑战.
- 为了提高导航效率和性能.
主要方法:
- 开发了一种模块化层次强化学习 (MHRL) 方法.
- MHRL包括目的地评估,政策切换和运动网络模块.
- 模块旨在解决导航问题的不同阶段.
主要成果:
- MHRL有效地处理混合人群 (动态和静态).
- 与最先进的技术相比,该方法显示出更高的性能.
- 广泛的模拟验证了MHRL的有效性.
结论:
- MHRL为机器人在具有挑战性的环境中进行导航提供了强大的解决方案.
- 模块化的层次结构增强了适应群众动态的能力.
- 这种方法为更复杂的自主导航系统铺平了道路.
相关概念视频
Multi-input and Multi-variable systems
106
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
106
Observational Learning
179
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
179
Reinforcement Schedules
148
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
148
Reinforcement
212
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
212
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K
Design Example: Identifying the Locations of Monuments in the Field Using Global Positioning System Device
36
Surveyors use Global Positioning System (GPS) technology to measure the precise location and elevation of points on Earth. In a recent survey, GPS receivers were used to determine the coordinates and elevations of two park monuments. The process involved careful mission planning, data collection, and correction to ensure accuracy. The survey began with mission planning to identify optimal satellite visibility and minimize Position Dilution of Precision (PDOP). A geodetic control point...
36


