通过使用交叉融合注意力网络的深度强化学习算法优化多目标旅行销售员问题.
Xiaoyu Fu1, Shenshen Gu1, Chee-Meng Chew2
1School of Mechatronic Engineering and Automation, Shanghai University, 99 Shangda Road, Shanghai, 200444, China.
概括
本研究介绍了一种新的深度强化学习算法,即交叉融合注意网络 (CFAN),以有效地解决复杂的多目标旅行销售员问题. 在各种问题实例中,CFAN表现出卓越的性能和概括性.
科学领域:
- 人工智能的人工智能
- 运营研究 运营研究
- 计算机科学 计算机科学
背景情况:
- 多目标旅行销售员问题 (MOTSP) 是一个具有广泛应用的关键组合优化挑战.
- 传统的算法由于大量的搜索空间和相互矛盾的目标而与MOTSP作斗争.
研究的目的:
- 开发一个高效的深度强化学习 (DRL) 算法来解决MOTSP.
- 增强DRL模型处理不同重量偏好的能力,并探索边界解决方案.
主要方法:
- 开发了一个新的交叉融合注意网络 (CFAN) 架构.
- CFAN的交叉融合关注编码器捕捉了实例-问题关系和重量偏好,用于统一的上下文特征.
- 使用重量分布调整来改善边界解决方案的探索.
主要成果:
- 与经典进化和先进的DRL算法相比,CFAN表现出优越的性能.
- 在超量 (HV) 度量方面观察到显著的改善:KroAB的1.43%,三目标的3.12%,大规模实例的2.17%.
- 该CFAN模型显示了强大的概括能力在不同的MOTSP实例.
结论:
- 拟议的CFAN算法有效地解决了MOTSP的挑战.
- 对于多目标的组合优化问题,CFAN提供了一种强大而通用的方法.
相关概念视频
Collisions in Multiple Dimensions: Problem Solving
4.4K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
4.4K
Reinforcement
343
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
343
Reinforcement Schedules
242
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
242
Multi-input and Multi-variable systems
150
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
150
Reducing Line Loss
194
In a three-phase circuit, line loss is an indicator of energy dissipated as heat due to the resistance of transmission lines. To address this, incorporating transformers into the system—a step-up transformer at the source and a step-down transformer at the load—is a strategic solution. Two three-phase transformers are introduced to improve this.
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
194
Associative Learning
579
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
579

