Related Experiment Video
Updated: Nov 21, 2025

11:18
Closed-loop Neuro-robotic Experiments to Test Computational Properties of Neuronal Networks
Published on: March 2, 2015
10.6K
t-soft update of target network for deep reinforcement learning
Taisuke Kobayashi1, Wendyam Eric Lionel Ilboudo1
1Nara Institute of Science and Technology, Nara, Japan.
Summary
A new t-soft update rule for deep reinforcement learning (DRL) target networks improves robustness. This method, inspired by Student-t distribution, prevents wrong updates while maintaining learning speed, outperforming conventional approaches in robotics simulations.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Robotics
Background:
- Deep reinforcement learning (DRL) uses target networks to stabilize training by generating reference signals.
- Conventional exponential moving average update rules can propagate incorrect parameter updates, increasing learning variance.
- Existing methods to mitigate this risk often slow down the learning process.
Purpose of the Study:
- To introduce a novel, robust update rule for DRL target networks.
- To enhance learning stability and speed by addressing the limitations of conventional update methods.
- To develop a method that prevents erroneous parameter updates without sacrificing learning efficiency.
Main Methods:
- A new t-soft update rule is proposed, drawing inspiration from the Student-t distribution.
- The method leverages an analogy between exponential moving average and normal distribution properties.
- The t-soft update rule is analyzed to demonstrate its unique properties, including a heavy-tailed characteristic.
Main Results:
- The t-soft update rule exhibits properties of the Student-t distribution, notably a heavy-tailed characteristic.
- It automatically excludes extreme parameter updates that diverge from past experiences.
- The method accelerates learning when updates align with past data, outperforming conventional methods in PyBullet robotics simulations.
Conclusions:
- The t-soft update rule offers a robust alternative to conventional target network updates in DRL.
- It effectively balances the need for stability with the requirement for rapid learning.
- Empirical results in robotics simulations demonstrate superior performance in terms of return and reduced variance.
More Related Videos
Related Concept Videos
Reinforcement
611
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
611
Reinforcement Schedules
329
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
329
Improving Translational Accuracy
12.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
12.7K
Improving Translational Accuracy
3.3K
3.3K
Observational Learning
638
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
638
Neural Regulation
41.9K
Digestion begins with a cephalic phase that prepares the digestive system to receive food. When our brain processes visual or olfactory information about food, it triggers impulses in the cranial nerves innervating the salivary glands and stomach to prepare for food.
41.9K

