完全尖端的演员网络与内层连接用于强化学习
IEEE transactions on neural networks and learning systems
|February 6, 2024
概括
这项研究介绍了一个完全可扩展的参与者网络 (SAN),用于节能的人工智能 (AI) 控制任务. 这种新的方法使用膜电压来表示动作,使其能够在无浮点运算的情况下在神经形态硬件上部署.
科学领域:
- 神经形态计算是一种神经形态计算.
- 人工智能的人工智能
- 计算神经科学是一种神经科学.
背景情况:
- 尖端神经网络 (SNN) 与深度强化学习 (DRL) 结合,提供节能的人工智能解决方案,特别是用于现实的控制任务.
- 现有的基于尖的增强学习方法通常使用发射率,需要浮点运算,这阻碍了直接部署在神经形态硬件上.
- 多维决定性政策对于许多现实世界的控制场景至关重要.
研究的目的:
- 开发一个完全尖端的演员网络 (SAN),避免浮点矩阵操作,以与神经形态硬件无集成.
- 通过使用非刺刺的神经元的膜电压来表示连续的动作空间,灵感来自昆虫的神经机制.
- 通过神经元群中的内层连接来增强输出层的表示能力.
主要方法:
- 引入了非尖端内部神经元启发的机制,其中膜电压直接编码动作值.
- 利用人口神经元来解码不同的动作维度,每个人口中的神经元在时间和空间领域都连接在一起.
- 实现了输出群体内的内层连接,以提高表示权力,创建了内层连接SAN (ILC-SAN).
主要成果:
- 拟议的ILC-SAN在OpenAI体育馆的持续控制任务上取得了最先进的性能.
- 与现有的基于尖峰的强化学习方法相比,表现优越.
- 预计神经形态芯片部署的理论能源消耗,突出显著的能源效率增长.
结论:
- ILC-SAN成功地实现了全面的深度强化学习,用于在没有浮点操作的情况下进行连续控制.
- 拟议的架构非常适合在节能神经形态硬件上部署.
- 这项工作推进了节能AI和神经形态计算领域,用于复杂的控制任务.
相关概念视频
Reinforcement
209
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
209
Observational Learning
175
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
175
Propagation of Action Potentials
5.7K
The propagation of an action potential refers to the process by which a nerve impulse, or "action potential," travels along a neuron.
Neurons (nerve cells) have a resting membrane potential, with a slightly negative charge inside compared to outside. This is maintained by ion channels, such as sodium (Na+) and potassium (K+) channels, which control the flow of ions. When a stimulus, like a touch or a signal from another neuron, triggers the neuron, sodium channels open, allowing sodium ions to...
Neurons (nerve cells) have a resting membrane potential, with a slightly negative charge inside compared to outside. This is maintained by ion channels, such as sodium (Na+) and potassium (K+) channels, which control the flow of ions. When a stimulus, like a touch or a signal from another neuron, triggers the neuron, sodium channels open, allowing sodium ions to...
5.7K
Neural Circuits
1.2K
Neural circuits and neuronal pools are two of the main structures found in the nervous system. Neural circuits are networks of neurons that work together to carry out a specific task or process. They consist of interconnected neurons and glial cells, which provide structural and metabolic support.
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
1.2K
Introduction to Learning
404
Learning is the process of acquiring knowledge or skills through practice or experience, leading to long-lasting behavioral changes. This acquisition occurs through interaction with the environment and requires practice or experience. For instance, mastering a skill such as surfing requires considerable practice and experience, highlighting the essential role of repeated interactions with the environment in learning.
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
404
Cognitive Learning
243
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
243


