基于memristor的尖端神经网络与在线强化学习
Danila Vlasov1, Anton Minnekhanov1, Roman Rybka2
1NRC "Kurchatov Institute", Akademika Kurchatova sq., 1 Moscow, Russian Federation.
概括
这项研究介绍了一种在线强化学习算法,用于使用基于memristor的突触来增强神经网络 (SNN). 这种新的算法可以在神经形态系统中实现高效的实时学习,用于控制任务.
科学领域:
- 神经形态工程的神经形态工程
- 人工智能的人工智能
- 材料科学 材料科学 材料科学
背景情况:
- 基于memristor的硬件为神经网络提供高效的内存计算.
- 传统的学习方法,如反向传播,对memristor硬件具有挑战性.
- 尖端神经网络 (SNN) 能够实现本地,自我组织的重量变化,适合这样的硬件.
研究的目的:
- 为在memristor硬件上实现的SNN开发一个在线强化学习算法.
- 整合来自真实memristor设备的STDP类学习规则.
- 用神经形态系统在连续时间环境中展示实时代理学习的可行性.
主要方法:
- 开发了一个在线强化学习算法,在每个环境状态交互后,连接权重更新.
- 该算法应用于SNN,使用基于memristor的STDP类学习规则.
- 塑性功能是从实验组装的聚-p-xylylene和CoFeB-LiNbO3纳米复合物记忆器中得出的.
主要成果:
- 由泄漏的整合和发射神经元组成的SNN成功地解决了卡特-波尔基准任务.
- 环境状态由输入尖峰时间编码,控制行动由第一个尖峰解码.
- 该算法显示了基于本地学习规则的有效权重变化.
结论:
- 拟议的在线强化学习算法对具有记忆性突触的SNN有效.
- 这项工作代表了在神经形态系统的连续时间环境中实现实时代理学习的重要一步.
- 使用来自memristor的STDP类规则促进了高效和本地化的学习.
相关概念视频
Reinforcement
277
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
277
Reinforcement Schedules
204
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
204
Neural Regulation
39.5K
Digestion begins with a cephalic phase that prepares the digestive system to receive food. When our brain processes visual or olfactory information about food, it triggers impulses in the cranial nerves innervating the salivary glands and stomach to prepare for food.
39.5K
Associative Learning
444
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
444
Observational Learning
210
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
210


