对于基于航母的飞机飞行甲板操作的深度增强学习,调度飞行甲板操作问题.
Changjiu Li1, Wei Han1, Haixu Li2
1Naval Aviation University, Yantai 264001, Shandong, PR China.
概括
一个新的深度强化学习框架优化了飞行甲板操作计划,优于传统方法. 这种人工智能方法显著减少了决策时间,同时保持了对复杂的实时需求的高质量时间表.
科学领域:
- 运营研究 运营研究
- 人工智能的人工智能
- 计算机科学 计算机科学
背景情况:
- 飞行甲板操作调度是一个复杂的,NP难题.
- 传统方法与效率-解决方案质量权衡作斗争.
研究的目的:
- 开发一个先进的AI框架,以优化飞行甲板调度.
- 克服现有的计算方法的局限性.
主要方法:
- 一个与图形神经网络集成的深度强化学习框架.
- 作为基于代理的调度的马尔科夫决策过程的表述.
- 使用软max勘探策略,其折扣因子为1.0.0.
主要成果:
- 人工智能代理在优先调度规则上显著提高了解决方案质量.
- 在小规模问题上实现了与元启发学相比的竞争性表现.
- 在大规模实例上展示了卓越的搜索能力.
- 将调度决策时间从几分钟缩短到几秒.
结论:
- 拟议的深度强化学习框架为飞行甲板调度提供了一种优越的方法.
- 这种人工智能解决方案以高质量,高效的时间表满足实时的运营需求.
相关概念视频
Reinforcement Schedules
651
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
651
Reinforcement
1.1K
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
1.1K
Controller Configurations
418
Controller configurations are crucial in a car's cruise control system because they manage speed over time to maintain a consistent pace regardless of road conditions, thereby meeting design goals. In traditional control systems, fixed-configuration design involves predetermined controller placement. System performance modifications are known as compensation.
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
418
Observational Learning
1.1K
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.1K
Multi-input and Multi-variable systems
453
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...
453
