Related Experiment Video
Updated: Oct 11, 2025

08:18
WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
5.1K
QC_SANE: Robust Control in DRL Using Quantile Critic With Spiking Actor and Normalized Ensemble
IEEE Transactions on Neural Networks and Learning Systems
|December 7, 2021
Summary
Quantile Critic with Spiking Actor and Normalized Ensemble (QC_SANE) improves continuous control by using quantile loss and a spiking neural network for actor training. This deep reinforcement learning approach shows better results than existing methods in complex robotic simulations.
Area of Science:
- Artificial Intelligence
- Robotics
- Computational Neuroscience
Background:
- Deep reinforcement learning (DRL) has advanced fields like online gaming and robotics.
- Continuous control problems present unique challenges for DRL agents.
- Existing methods like population coded spiking actor networks (PopSAN) have limitations.
Purpose of the Study:
- To introduce a novel DRL approach, Quantile Critic with Spiking Actor and Normalized Ensemble (QC_SANE), for continuous control problems.
- To leverage quantile loss for critic training and spiking neural networks for actor ensembles.
- To enhance robustness and performance in complex robotic tasks.
Main Methods:
- Developed QC_SANE, integrating quantile loss for critic and a spiking neural network (NN) for actor ensembles.
- Utilized scaled exponential linear unit (SELU) activation within the NN for internal normalization and robustness.
- Conducted empirical studies on MuJoCo-based environments featuring multijoint dynamics with contact.
Main Results:
- QC_SANE demonstrated superior training and testing performance compared to the state-of-the-art PopSAN.
- The proposed method achieved improved results in complex, contact-rich robotic simulations.
- Internal normalization via SELU contributed to the robustness of the spiking NN actors.
Conclusions:
- QC_SANE offers a significant advancement for deep reinforcement learning in continuous control.
- The integration of quantile critics and spiking actor ensembles enhances performance and robustness.
- This approach shows promise for tackling complex robotic manipulation and control tasks.
Related Concept Videos
Reinforcement
423
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
423
Randomized Experiments
8.2K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
8.2K
Neural Regulation
40.5K
Digestion begins with a cephalic phase that prepares the digestive system to receive food. When our brain processes visual or olfactory information about food, it triggers impulses in the cranial nerves innervating the salivary glands and stomach to prepare for food.
40.5K
Observational Learning
372
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
372
Actor-Observer Effect
29
The actor-observer effect, a cognitive bias closely linked to the fundamental attribution error, refers to the tendency for individuals to attribute their behavior to external, situational factors while explaining others’ behavior in terms of internal, dispositional traits. This asymmetry in attribution significantly influences social perception and judgment.Cognitive Mechanisms Behind the EffectTwo primary psychological mechanisms contribute to the actor-observer effect: differences in...
29
Confidence Coefficient
8.5K
The confidence coefficient is also known as the confidence level or degree of confidence. It is the percent expression for the probability, 1-α, that the confidence interval contains the true population parameter assuming that the confidence interval is obtained after sufficient unbiased sampling; for example, if the CL = 90%, then in 90 out of 100 samples the interval estimate will enclose the true population parameter. Here α is the area under the curve, distributed equally under...
8.5K

