基于强化学习的智能路径规划,以在动态环境中实现最佳导航
Anil Kumar Yadav1, Purushottam Sharma2, Xiaochun Cheng3
1VIT Bhopal University, Bhopal-Indore Highway, Bhopal, India.
概括
在强化学习 (RL) 中优化奖励功能显著改善了自主移动机器人的导航. 这种增强的RL方法减少了动态环境中的路径距离和学习时间.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 路径选择和规划对于自主移动机器人 (AMR) 进行高效导航和避免障碍物至关重要.
- 传统方法通常使用分析搜索最短路径,但强化学习 (RL) 通过动作序列优化提供了增强的性能.
- 一个常见的RL算法Q-learning由于依赖累积奖励而在动态系统中与环境概括作斗争.
研究的目的:
- 在基于RL的AMR路径规划中,优化奖励功能以实现高效的导航和避开障碍.
- 在动态环境中增强RL算法的概括能力.
- 评估优化奖励机制对路径规划效率和学习绩效的影响.
主要方法:
- 该研究提出了基于RL的路径规划的优化奖励函数,考虑动态环境中的总步骤,计数步骤和折扣率.
- 通过使用优化奖励机制在不同环境中实现和分析状态奖励值.
- 评估了对Q-Learning和深度Q-Learning算法的影响,比较了基于状态-动作对的性能.
主要成果:
- 优化的奖励函数显著减少了学习所需的代和情节的数量.
- 与传统方法相比,总体轨迹距离减少了30%至70%.
- 证明改善了路径优化,学习速率,情节完成和决策准确性.
结论:
- 优化的奖励功能提高了RL在动态环境中的AMR路径规划的有效性.
- 拟议的方法在导航效率和避开障碍方面取得了显著的改进.
- 结合多个代理和先进的技术,如联合和转移学习,可以进一步提高在更大的地图上的融合和性能.
相关概念视频
Mean free path and Mean free time
5.0K
Consider the gas molecules in a cylinder. They move in a random motion as they collide with each other and change speed and direction. The average of all the path lengths between collisions is known as the "mean free path."
5.0K
Path Between Thermodynamics States
3.9K
Consider the two thermodynamic processes involving an ideal gas that are represented by paths AC and ABC in Figure 1:
3.9K
Interference: Path Lengths
1.9K
Consider two sources of sound, that may or may not be in phase, emitting waves at a single frequency, and consider the frequencies to be the same.
Two special sources may be considered when they are in phase. This can be easily achieved by feeding the two sources from the same source. An example would be synchronizing the two speakers by feeding them with the same source, such as the sound waves produced by a tuning fork. This setup ensures that the two sources have the same frequency and are...
Two special sources may be considered when they are in phase. This can be easily achieved by feeding the two sources from the same source. An example would be synchronizing the two speakers by feeding them with the same source, such as the sound waves produced by a tuning fork. This setup ensures that the two sources have the same frequency and are...
1.9K
Reinforcement
872
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
872
Intelligence
8.5K
The term "intelligence" is complex because it refers to both behavior and individuals, and its interpretation varies across cultures. European Americans tend to link intelligence with reasoning and cognitive skills, while in Kenya, it is tied to responsible participation in family and social life. In Uganda, intelligence is seen as the ability to know the right actions and carry them out effectively, while the Iatmul people of Papua New Guinea associate it with the capacity to remember...
8.5K
Behavior of Gas Molecules: Molecular Diffusion, Mean Free Path, and Effusion
31.2K
Although gaseous molecules travel at tremendous speeds (hundreds of meters per second), they collide with other gaseous molecules and travel in many different directions before reaching the desired target. At room temperature, a gaseous molecule will experience billions of collisions per second. The mean free path is the average distance a molecule travels between collisions. The mean free path increases with decreasing pressure; in general, the mean free path for a gaseous molecule will be...
31.2K


