Related Experiment Video
Updated: Jul 4, 2025

08:08
Real-time Electrophysiology: Using Closed-loop Protocols to Probe Neuronal Dynamics and Beyond
Published on: June 24, 2015
11.5K
Fully Spiking Actor Network With Intralayer Connections for Reinforcement Learning
IEEE Transactions on Neural Networks and Learning Systems
|February 6, 2024
Summary
This study introduces a fully spiking actor network (SAN) for energy-efficient artificial intelligence (AI) control tasks. The novel approach uses membrane voltage to represent actions, enabling deployment on neuromorphic hardware without floating-point operations.
Area of Science:
- Neuromorphic computing
- Artificial Intelligence
- Computational Neuroscience
Background:
- Spiking neural networks (SNNs) combined with deep reinforcement learning (DRL) offer energy-efficient AI solutions, particularly for realistic control tasks.
- Existing spike-based reinforcement learning methods often use firing rates, requiring floating-point operations that hinder direct deployment on neuromorphic hardware.
- Multidimensional deterministic policies are crucial for many real-world control scenarios.
Purpose of the Study:
- To develop a fully spiking actor network (SAN) that avoids floating-point matrix operations for seamless integration with neuromorphic hardware.
- To represent continuous action spaces using the membrane voltage of nonspiking neurons, inspired by insect neural mechanisms.
- To enhance the representation capacity of the output layer through intralayer connections in neuron populations.
Main Methods:
- Introduced nonspiking interneuron-inspired mechanisms where membrane voltage directly encodes action values.
- Utilized population neurons to decode different action dimensions, with neurons within each population connected in both time and space domains.
- Implemented intralayer connections within output populations to boost representational power, creating the intralayer connection SAN (ILC-SAN).
Main Results:
- The proposed ILC-SAN achieved state-of-the-art performance on continuous control tasks from OpenAI gym.
- Demonstrated superior performance compared to existing spike-based reinforcement learning methods.
- Estimated theoretical energy consumption for neuromorphic chip deployment, highlighting significant energy efficiency gains.
Conclusions:
- The ILC-SAN successfully enables fully spiking deep reinforcement learning for continuous control without floating-point operations.
- The proposed architecture is highly suitable for deployment on energy-efficient neuromorphic hardware.
- This work advances the field of energy-efficient AI and neuromorphic computing for complex control tasks.
Related Concept Videos
Reinforcement
209
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
209
Observational Learning
175
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
175
Propagation of Action Potentials
5.7K
The propagation of an action potential refers to the process by which a nerve impulse, or "action potential," travels along a neuron.
Neurons (nerve cells) have a resting membrane potential, with a slightly negative charge inside compared to outside. This is maintained by ion channels, such as sodium (Na+) and potassium (K+) channels, which control the flow of ions. When a stimulus, like a touch or a signal from another neuron, triggers the neuron, sodium channels open, allowing sodium ions to...
Neurons (nerve cells) have a resting membrane potential, with a slightly negative charge inside compared to outside. This is maintained by ion channels, such as sodium (Na+) and potassium (K+) channels, which control the flow of ions. When a stimulus, like a touch or a signal from another neuron, triggers the neuron, sodium channels open, allowing sodium ions to...
5.7K
Neural Circuits
1.2K
Neural circuits and neuronal pools are two of the main structures found in the nervous system. Neural circuits are networks of neurons that work together to carry out a specific task or process. They consist of interconnected neurons and glial cells, which provide structural and metabolic support.
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
1.2K
Introduction to Learning
404
Learning is the process of acquiring knowledge or skills through practice or experience, leading to long-lasting behavioral changes. This acquisition occurs through interaction with the environment and requires practice or experience. For instance, mastering a skill such as surfing requires considerable practice and experience, highlighting the essential role of repeated interactions with the environment in learning.
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
404
Cognitive Learning
243
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
243

