竞争性群体强化学习提高了深度强化学习的稳定性和性能
Xindong Huang1, Weihang Luo2, Ling Li1
1School of Opto-Electronic and Communication Engineering, Xiamen University of Technology, Xiamen, 361021, China.
Scientific reports
|December 11, 2025
概括
竞争性群体强化学习 (CSRL) 通过使用各种代理策略来增强深度强化学习 (RL). 这种方法提高了稳定性和性能,在复杂的任务中优于现有方法.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 深度强化学习学习 (deep reinforcement learning) 是一种深度强化学习的方法.
背景情况:
- 深度强化学习 (RL) 将深度学习与RL整合在一起,扩展能力,但引入稳定性问题.
- 关键的挑战包括不一致的样本采集和对超参数的高度敏感性.
研究的目的:
- 引入竞争群强化学习 (CSRL) 以解决深度RL中的稳定性和超参数灵敏性.
- 通过基于人口的优化原则来提高代理的性能,稳定性和适应性.
主要方法:
- CSRL雇佣了一群多元化的代理人,他们用各种策略探索环境.
- 为了高效的数据使用和体验重复,创建了一个多样化的样本共享池.
- 同时训练具有不同超参数的多个策略,以减少灵敏度和提高强度.
主要成果:
- 与三个基线算法相比,CSRL显示出更高的收率和稳定性.
- 在10个任务中,在9个任务中获得了2%以上的平均回报,而在HumanoidBulletEnv-v0.0上获得了31.09%的改进.
- 保持样本效率和平衡的勘探开发,在奖励方面超过基线,特别是在复杂的场景中.
结论:
- 竞争激烈的群体动态显著提高了RL系统的性能,稳定性和适应性.
- CSRL为深度强化学习提供了一个强大的框架,克服了常见的局限性.
- 该方法对在复杂和动态环境中推进RL应用有前途.
相关概念视频
Reinforcement
786
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
786
Observational Learning
791
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
791
Survival Tree
369
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
369
Avoidance Learning and Learned Helplessness
2.5K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.5K
Reinforcement Schedules
436
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
436
Pole and System Stability
864
The transfer function is a fundamental concept representing the ratio of two polynomials. The numerator and denominator encapsulate the system's dynamics. The zeros and poles of this transfer function are critical in determining the system's behavior and stability.
Simple poles are unique roots of the denominator polynomial. Each simple pole corresponds to a distinct solution to the system's characteristic equation, typically resulting in exponential decay terms in the system's...
Simple poles are unique roots of the denominator polynomial. Each simple pole corresponds to a distinct solution to the system's characteristic equation, typically resulting in exponential decay terms in the system's...
864
