一个改进的差异进化算法,基于强化学习及其应用.
Guangwei Yang1,2, Peng Sun1, Jieyong Zhang3
1Information and Navigation College, Air Force Engineering University, Xi'an 710077, China.
Scientific reports
|November 6, 2025
概括
这项研究引入了一种基于强化学习的新差异进化 (RLDE) 算法,以克服参数灵敏度和在群集智能优化中的过早融合. 在复杂的,高维度的问题上,RLDE表现出卓越的全球优化性能.
科学领域:
- 计算智能是一种计算智能.
- 优化算法 优化算法
- 群集情报 群集情报 群集情报
背景情况:
- 差异进化 (DE) 是一种强大的群体智能方法,用于解决高维问题.
- 由于DE具有参数敏感性和过早的趋同,这限制了其实际应用.
- 现有的优化方法需要提高效率和适应性.
研究的目的:
- 提出一个改进的差异进化算法,RLDE,利用强化学习.
- 为了提高全球优化性能,并解决标准DE算法的局限性.
- 为了验证算法的有效性在基准函数和现实世界的工程问题.
主要方法:
- 使用哈尔顿序列进行种群初始化,以提高厄尔戈迪性.
- 通过强化学习政策梯度网络进行动态参数调整,以适应性缩放因子和交叉概率.
- 基于人口适应性分类的差异化突变策略.
主要成果:
- 在26个标准测试函数中,RLDE显著提高了全球优化性能.
- 该算法在10个,30个和50个维度中超过了多个启发式优化算法.
- 对无人机任务分配问题的成功应用证明了实际的工程价值.
结论:
- 拟议的RLDE算法有效地解决了DE的参数灵敏度和过早的融合.
- 对于复杂的,高维度的问题,RLDE提供了增强的全球优化功能.
- 该算法显示了对现实世界工程应用的巨大潜力,例如无人机任务分配.
相关概念视频
Reinforcement
804
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
804
Reinforcement Schedules
440
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
440
Generalization, Discrimination, and Extinction
1.3K
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
1.3K
Observational Learning
804
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
804
Differential Leveling
647
Differential leveling is a precise method in surveying used to determine the elevation difference between two points. Its primary goal is to establish accurate vertical measurements to create level surfaces or grade lines critical for designing and constructing infrastructures such as roads, bridges, and buildings.The procedure for differential leveling begins with setting up and leveling the instrument at a point where the benchmark can be seen. The level rod is held on the benchmark (BM), and...
647
Law of Effect
2.4K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
2.4K

