相关实验视频
Updated: Sep 15, 2025

08:30
Operant Procedures for Assessing Behavioral Flexibility in Rats
Published on: February 15, 2015
21.0K
离线强化学习学习学习调度工作车间调度
Jesse van Remmerden1, Zaharah Bukhsh1, Yingqian Zhang1
1Information Systems IE&IS, Eindhoven University of Technology, De Zaale, Eindhoven, 5600 MB Netherlands.
概括
离线学习调度 (Offline-LD) 通过从历史数据中学习来改进工作室安排,克服在线强化学习 (RL) 的局限性. 这种方法有效地找到有效的解决方案,而不需要进行广泛的模拟或从头开始进行新的培训.
科学领域:
- 运营研究 运营研究
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 工作坊调度问题 (JSSP) 是一个重要的组合优化挑战.
- 在线强化学习 (RL) 为JSSP提供了快速解决方案,但受到样本效率低下的影响,需要不切实际的模拟环境.
- 像约束编程 (CP) 这样的现有方法提供了高质量的解决方案,但不容易与RL集成.
研究的目的:
- 引入一个离线强化学习方法,离线学习调度 (Offline-LD),用于工作坊调度问题 (JSSP).
- 通过利用历史调度数据,解决在线RL的局限性,包括样本效率低下和模拟环境的需求.
- 允许在RL框架内利用CP等方法的先前存在的高质量解决方案.
主要方法:
- 开发了Q学习方法的可掩盖的变体:可掩盖的量子回归DQN (mQRDQN) 和离散的可掩盖的软演员-批判性 (d-mSAC).
- 采用保守的Q学习 (CQL) 来实现从历史数据中学习.
- 引入了一个新的奖金修改用于d-mSAC在可掩饰的动作空间和一个新的奖励规范化技术在JSSP下线RL.
主要成果:
- 与在线RL相比,离线LD在生成和基准JSSP实例上表现出更好的表现.
- 这种方法取得了强的结果,即使在有限的100个解决方案由CP生成集训练.
- 将噪音引入专家数据集的结果与使用清洁专家数据集相比或优于使用清洁专家数据集的性能,表明了稳定性.
结论:
- 离线LD通过从历史数据中学习,有效地解决了JSSP在线RL的关键局限性.
- 该方法对现实世界的应用有希望,在现实世界中,历史数据是可用的,并且可能会有噪音.
- 离线LD为复杂的安排环境提供了可行的替代方案,在复杂的环境中,在线培训是不切实际的.
相关概念视频
Reinforcement Schedules
243
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
243
Reinforcement
345
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
345
Distributed Loads: Problem Solving
738
Beams are structural elements commonly employed in engineering applications requiring different load-carrying capacities. The first step in analyzing a beam under a distributed load is to simplify the problem by dividing the load into smaller regions, which allows one to consider each region separately and calculate the magnitude of the equivalent resultant load acting on each portion of the beam. The magnitude of the equivalent resultant load for each region can be determined by calculating...
738
Avoidance Learning and Learned Helplessness
1.9K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.9K
Observational Learning
317
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
317
Operant Conditioning
1.8K
Operant conditioning, a key concept in behavioral psychology, involves using reinforcement and punishment to alter the likelihood of a behavior being repeated. B.F. introduced this type of conditioning. Skinner focused on voluntary behaviors and the consequences that follow them, influencing whether these behaviors will be strengthened or diminished.
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
1.8K

