实时在线目标识别在连续领域通过深度强化学习通过深度强化学习
Zihao Fang1, Dejun Chen1, Yunxiu Zeng1
1College of Systems Engineering, National University of Defense Technology, Changsha 410000, China.
Entropy (Basel, Switzerland)
|October 28, 2023
概括
本研究介绍了一种新的深度强化学习算法,用于在连续环境中实时在线目标识别. 该方法有效地推断了代理目标,克服了先前离散和离线方法的局限性.
科学领域:
- 人工智能的人工智能
- 机器人技术 机器人技术 机器人技术
- 机器学习 机器学习
背景情况:
- 目标识别从观察到的行为中推断出代理任务.
- 现有的方法在连续环境和实时处理方面遇到了困难.
- 离线建模需要大量的计算资源.
研究的目的:
- 为连续域开发一个高效的实时在线目标识别算法.
- 解决当前离线和离散环境方法的局限性.
- 在复杂,动态的环境中实现准确的目标推断.
主要方法:
- 提出了一个使用深度强化学习 (DRL) 的高级算法.
- 杆离线DRL建模用于观察到的代理行为.
- 在连续模拟环境中实现在线目标识别.
主要成果:
- 实现了实时目标识别能力.
- 在连续环境中证明了算法的准确性和稳定性.
- 在通信限制下评估的性能.
结论:
- 拟议的基于DRL的算法为实时目标识别提供了有效的解决方案.
- 成功克服连续域和计算需求所带来的挑战.
- 提供了一种强大的方法来推断复杂环境中的代理目标.
相关概念视频
Reinforcement
221
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
221
Observational Learning
188
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
188
Reinforcement Schedules
160
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
160
Associative Learning
412
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
412
Introduction to Learning
446
Learning is the process of acquiring knowledge or skills through practice or experience, leading to long-lasting behavioral changes. This acquisition occurs through interaction with the environment and requires practice or experience. For instance, mastering a skill such as surfing requires considerable practice and experience, highlighting the essential role of repeated interactions with the environment in learning.
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
446
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K


