相关实验视频
Updated: Jul 19, 2025

16:14
Trajectory Data Analyses for Pedestrian Space-time Activity Study
Published on: February 25, 2013
13.6K
使用强化学习进行信息性轨迹规划,用于空间时间场的最小时间探索
IEEE transactions on neural networks and learning systems
|August 15, 2023
概括
本研究介绍了有效的自动驾驶汽车轨迹规划,用于探索时空领域. 强化学习方法将探索时间最小化,同时确保满足累积信息约束.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
- 自主系统 自主系统
背景情况:
- 现有的研究重点是最大限度地提高空间领域的信息获取.
- 对于自主系统来说,有效地探索未知分布的时空场是非常重要的.
- 在信息限制下,当前的方法可能无法优化最小时间.
研究的目的:
- 为探索时空领域的自动驾驶车辆开发一个最小时间轨迹规划方法.
- 为了应对在累积信息限制下高效勘探的挑战.
- 为持续的政策学习提出基于强化学习的方法.
主要方法:
- 将问题建模为马尔科夫决策过程 (MDP).
- 提出一种强化学习 (RL) 算法来学习持续规划政策.
- 设计一种新的奖励函数,使用现场近似来加速政策学习.
主要成果:
- 证明在温和条件下存在最小时间轨迹.
- 证明已学习的RL政策可以实现有效的探索.
- 在勘探时间方面,与覆盖计划相比,表现优越.
结论:
- 拟议的基于RL的轨迹规划方法可以实现高效的时空实地探索.
- 这种方法有效地平衡了尽量减少勘探时间和满足信息约束.
- 这项工作在复杂的环境调查中提高了自动驾驶汽车的能力.
相关概念视频
Reinforcement Schedules
204
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
204
Relative Motion Analysis using Rotating Axes-Problem Solving
421
Consider a crane whose telescopic boom rotates with an angular velocity of 0.04 rad/s and angular acceleration of 0.02 rad/s2. Along with the rotation, the boom also extends linearly with a uniform speed of 5 m/s. The extension of the boom is measured at point D, which is measured with respect to the fixed point C on the other end of the boom. For the given instant, the distance between points C and D is 60 meters.
Here, in order to determine the magnitude of velocity and acceleration for point...
Here, in order to determine the magnitude of velocity and acceleration for point...
421

