一个统一的体验重复框架,用于深度强化学习的深度强化学习
IEEE transactions on pattern analysis and machine intelligence
|December 11, 2025
概括
我们引入了弹性体验重复方法,以平衡能源效率和性能,在深度强化学习 (DRL) 中. 这种方法可以动态地管理重播缓冲器,提高样本质量并提高跨任务的DRL模型性能.
科学领域:
- 人工智能的人工智能
- 计算神经科学是一种神经科学.
- 节能计算 节能计算 节能计算 节能计算
背景情况:
- 深度增强学习 (DRL) 提供了显著的功能,但受到高能耗的影响.
- 尖端神经网络 (SNN) 减少了DRL中的能源消耗,但在短时间模拟中面临样本质量和性能方面的挑战.
- 现有的Spiking DRL方法由于固定的重复缓冲器而扎着能源性能权衡.
研究的目的:
- 为Spiking DRL.开发一种通用的弹性体验重播方法.
- 为了解决Spiking DRL的能源消耗和模型性能之间的权衡问题.
- 提高样品质量和模型性能,而不损害能源效率.
主要方法:
- 实现了一个与训练样本扩展的动态重复缓冲区.
- 引入了自适应式缓冲器管理,以缩小缓冲器并删除多余的样本.
- 将该方法无地集成到现有的最先进的Spiking DRL算法中.
主要成果:
- 显著提高了五种SOTA Spiking DRL方法的性能 (回报).
- 在16个不同的任务和各种模拟持续时间中证明了改进.
- 在不影响性能的情况下保持能源效率.
结论:
- 弹性体验重播方法有效地解决了斯派金DRL的能源性能权衡问题.
- 动态和自适应性缓冲器管理对于提高基于SNN的RL的样本效率至关重要.
- 这种方法为在现实世界应用中部署节能AI提供了实际解决方案.
更多相关视频
11:20Recording Single Neurons' Action Potentials from Freely Moving Pigeons Across Three Stages of Learning
Published on: June 2, 2014
12.4K
09:13A Fully Automated and Highly Versatile System for Testing Multi-cognitive Functions and Recording Neuronal Activities in Rodents
Published on: May 3, 2012
14.8K
相关概念视频
Reinforcement
786
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
786
Observational Learning
791
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
791
Reinforcement Schedules
436
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
436
Associative Learning
1.2K
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
1.2K
Elaborative Rehearsals
306
Elaborative rehearsal is a crucial cognitive strategy that strengthens information encoding in long-term memory by making meaningful connections between new data and pre-existing knowledge. This approach contrasts with maintenance rehearsal, which involves simple repetition without delving into the significance of the information. While maintenance rehearsal might temporarily keep information active in short-term memory, it is less effective for long-term retention.
The effectiveness of...
The effectiveness of...
306
Avoidance Learning and Learned Helplessness
2.5K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.5K
