基于意图推断和深度强化学习的逃避性城市目标的长期跟踪
IEEE transactions on neural networks and learning systems
|August 11, 2023
概括
本研究介绍了一种混合方法,用于无人机 (UAV) 目标跟踪,使用目标意图推断和深度强化学习 (DRL). 这种方法提高了复杂的城市环境中的跟踪性能和稳定性.
科学领域:
- 机器人技术和自主系统
- 人工智能的人工智能
- 计算机视觉 计算机视觉
背景情况:
- 无人驾驶飞行器 (UAV) 追踪城市目标对于公共安全至关重要,但因逃避性目标和复杂环境而面临挑战.
- 由于不可预测的移动和非结构化的城市环境,现有的方法在目标损失方面扎.
研究的目的:
- 为无人机在城市环境中开发一种强大的混合目标追踪方法.
- 通过整合意图推断和深度强化学习来改善逃避目标的长期跟踪.
主要方法:
- 基于卷积神经网络 (CNN) 的模型通过融合环境数据和轨迹观测来推断目标意图.
- 深度强化学习 (DRL) 框架开发了一个目标搜索策略,以深度神经网络 (DNN) 为模型.
- 通过与任务环境的交互来训练DRL政策,以优化搜索策略.
主要成果:
- 目标意图推断有效指导无人机搜索操作,显著提高目标追踪性能.
- 拟议的基于DRL的搜索策略显示出对不确定的目标行为具有很高的稳定性.
- 模拟结果验证了混合追踪方法的提高效率和可靠性.
结论:
- 目标意图推断和DRL的融合为挑战无人机目标追踪场景提供了一个有希望的解决方案.
- 这种混合方法提高了无人机在动态的城市景观中追踪逃避目标的能力.
- 该方法显示了通过更可靠的监控和跟踪能力来提高公共安全的巨大潜力.
相关概念视频
Observational Learning
210
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
210
Avoidance Learning and Learned Helplessness
1.8K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.8K
Instinctive Drift
250
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
250
Reinforcement
277
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
277
Purposive Learning
142
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
142


