Related Experiment Video
Updated: Oct 21, 2025

08:18
WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
5.1K
A Maximum Divergence Approach to Optimal Policy in Deep Reinforcement Learning
IEEE Transactions on Cybernetics
|September 3, 2021
Summary
This study introduces divergence Markov decision processes (MDPs) to enhance model-free reinforcement learning by learning intrinsic state transition information. This approach yields more robust and performant stochastic policies for complex control tasks.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Control Theory
Background:
- Model-free reinforcement learning (RL) often uses entropy regularization for policy learning.
- Existing methods face challenges in balancing exploration and exploitation, particularly in high-dimensional continuous spaces.
Purpose of the Study:
- To propose a novel perspective for RL by explicitly learning intrinsic information in state transitions.
- To develop a new class of Markov decision processes (MDPs) called divergence MDPs for improved policy learning.
Main Methods:
- Introduced divergence MDPs that maximize expected rewards plus a divergence term representing state transition information.
- Developed a divergence actor-critic (DivAC) algorithm using divergence policy iteration for large-scale continuous problems.
Main Results:
- The proposed DivAC algorithm demonstrated superior performance and robustness compared to existing methods in complex environments.
- The divergence term effectively captures implicit information in state transitions, leading to better stochastic policies.
Conclusions:
- Divergence MDPs offer a new framework for reinforcement learning, enhancing policy robustness and performance.
- The developed DivAC algorithm is effective for high-dimensional continuous control tasks.
Related Concept Videos
Reinforcement
474
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
474
Decision Making: P-value Method
5.9K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.9K
Observational Learning
398
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
398
Collisions in Multiple Dimensions: Problem Solving
4.6K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
4.6K
Divergence and Stokes' Theorems
2.6K
The divergence and Stokes' theorems are a variation of Green's theorem in a higher dimension. They are also a generalization of the fundamental theorem of calculus. The divergence theorem and Stokes' theorem are in a way similar to each other; The divergence theorem relates to the dot product of a vector, while Stokes' theorem relates to the curl of a vector. Many applications in physics and engineering make use of the divergence and Stokes' theorems, enabling us to write...
2.6K
Reinforcement Schedules
275
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
275
