Related Experiment Video
Updated: Sep 6, 2025

05:41
A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
9.5K
Understanding via Exploration: Discovery of Interpretable Features With Deep Reinforcement Learning
Summary
This study introduces dual-world-based attentive feature selection (D-AFS) for deep reinforcement learning (DRL). D-AFS effectively identifies crucial input features for mastering complex systems by using a real and virtual world approach.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Control Systems
Background:
- Understanding complex systems often requires analyzing environmental interactions.
- Deep reinforcement learning (DRL) excels at control but its opaque nature (DNNs) hinders feature relevance identification.
- Mastering unknown systems necessitates understanding which inputs are critical for control.
Purpose of the Study:
- To propose a novel online feature selection framework, dual-world-based attentive feature selection (D-AFS).
- To quantitatively identify the contribution of inputs to the control process in DRL.
- To enhance the interpretability of deep neural networks in control applications.
Main Methods:
- Introduced a dual-world framework with a real and a virtual world possessing distinct features.
- Developed an attention-based evaluation (AR) module for dynamic mapping between the real and virtual worlds.
- Modified existing DRL algorithms to learn within this dual-world environment.
Main Results:
- D-AFS quantitatively identifies feature importance by analyzing DRL responses across both worlds.
- Experiments on four OpenAI Gym classical control systems demonstrated D-AFS's effectiveness.
- D-AFS generated feature combinations comparable or superior to human experts and seven baseline methods.
Conclusions:
- The proposed D-AFS framework effectively addresses the challenge of feature relevance in DRL.
- Selected feature representations by D-AFS correlate strongly with underlying system dynamics.
- D-AFS offers a powerful tool for understanding and improving DRL-based control systems.
Related Concept Videos
Reinforcement
328
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
328
Observational Learning
297
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
297
Associative Learning
557
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
557
Introduction to Learning
524
Learning is the process of acquiring knowledge or skills through practice or experience, leading to long-lasting behavioral changes. This acquisition occurs through interaction with the environment and requires practice or experience. For instance, mastering a skill such as surfing requires considerable practice and experience, highlighting the essential role of repeated interactions with the environment in learning.
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
524
Reinforcement Schedules
237
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
237
Generalization, Discrimination, and Extinction
779
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
779
