ATA:一个抽象-列车-抽象的方法,以解释友好的深度强化学习学习
Shi Peng1, Si Liu2, Dapeng Zhi1
1Shanghai Key Laboratory of Trustworthy Computing, East China Normal University, Shanghai, China.
概括
我们开发了Abstract-Train-Abstract (ATA),一种新的方法,用于在深度强化学习 (DRL) 中创建更准确,更易于理解的抽象政策图 (APG). ATA显著提高了模型的可解释性和预测准确性.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 深度强化学习 (deep reinforcement learning) 是一种深度强化学习的方法.
背景情况:
- 解释深度强化学习 (DRL) 神经网络中的决策是困难的.
- 抽象政策图 (APG) 对于模型解释是有效的,但在实现高准确性和可解释性方面面临挑战.
- 在APG中较大的集群大小与更高的保真度相关.
研究的目的:
- 引入一种新的方法,抽象-列车-抽象 (ATA),用于构建高保真性和可解释的APG.
- 提高DRL模型中与抽象状态相关的预测行为的准确性.
- 提高用户对DRL决策过程的理解和预测准确度.
主要方法:
- 开发了抽象-列车-抽象 (ATA) 方法,集成基于抽象的培训和以抽象为导向的集群.
- 基于抽象的培训扩大了抽象状态集群的范围.
- 面向抽象的集群确保集群地图中的状态与相同的动作有关.
主要成果:
- 与最先进的方法相比,ATA实现了高达26.63%的更高保真度.
- 这种方法保持了竞争力的奖励水平.
- 一项用户研究显示,ATA平均提高了35.7%的用户预测准确度.
结论:
- ATA在为DRL创建可解释和高保真APG方面取得了重大进展.
- 该方法通过有效集群抽象状态来提高动作预测的准确性.
- ATA显然提高了模型性能和人类对DRL系统的理解.
相关概念视频
Observational Learning
321
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
321
Reinforcement
353
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
353
Generalization, Discrimination, and Extinction
823
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
823
Associative Learning
605
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
605
Introduction to Learning
551
Learning is the process of acquiring knowledge or skills through practice or experience, leading to long-lasting behavioral changes. This acquisition occurs through interaction with the environment and requires practice or experience. For instance, mastering a skill such as surfing requires considerable practice and experience, highlighting the essential role of repeated interactions with the environment in learning.
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
551
Role of Shaping in Operant Conditioning
511
Shaping is a technique used in operant conditioning to train complex behaviors by rewarding successive approximations toward the target behavior. This method is necessary because organisms are unlikely to perform complex behaviors spontaneously. Instead, shaping breaks down the desired behavior into small, manageable steps.
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
511


